AI Infrastructure & Agent Observability engineer working on telemetry correctness across inference engines and agent runtimes.
My current focus:
- distributed tracing and low-overhead metrics for LLM serving;
- KV-cache and prefill/decode disaggregation observability;
- streaming, cancellation, context propagation, and span-finalization correctness;
- OpenTelemetry instrumentation and semantic conventions for agent systems.
- SGLang: speculative decoding tracing and chunked-prefill observability work.
- vLLM: OpenTelemetry request-attribute propagation.
- LoongSuite: Agent telemetry context-isolation fixes for Hermes and AgentScope ReAct.
- OpenDerisk: observability improvements in #128 and #130.
I am contributing and reviewing around SGLang observability, OpenTelemetry Python GenAI lifecycle correctness, and vLLM metrics/KV telemetry. I prefer scoped changes with reproducible tests, bounded metric cardinality, and clear success/error/cancellation semantics.
You can reach me through GitHub issues and pull requests related to these areas.




