Building AI systems that act reliably.
I'm a master's student at HKUST, with engineering experience in backend systems, retrieval-augmented generation, and AI agents. I focus on the systems around models: how agents use tools, how we verify outcomes, and how execution recovers when something goes wrong.
Website · X · LinkedIn · Email
- Bounded actions — tool permissions, approval boundaries, and safe execution.
- Verifiable outcomes — task evaluation, execution traces, and reproducible failure analysis.
- Recoverable execution — state management, retries, and the trade-offs between reliability, latency, and cost.
I explore these questions through open-source tools and small, reproducible experiments, with post-training and efficient inference as supporting interests.
A local-first tracing proxy for MCP tool calls. It helps inspect the tools an agent called, their timing, results, and errors—making execution easier to investigate.
Start with the README, try it on your own workflow, and share a reproducible issue or integration request.
An educational PyTorch implementation of SFT and DPO, with LoRA and evaluation utilities. A compact project for understanding how post-training components fit together.
回路之外 · Beyond the Loop is my Chinese publication about AI systems in practice: builds, experiments, failure cases, and engineering decisions.
- GitHub: code, tests, and reproducible artifacts.
- 回路之外 on WeChat: Chinese long-form explanations and experiment reports. Search for 回路之外.
- X: short English findings, demos, and technical discussion.
- Substack: an early-stage home for English field notes.
- Website: selected work and long-term context.
- LinkedIn: professional milestones and collaboration.
I'm interested in collaborating with teams building agents that interact with real tools and applications—especially on tool-call observability, evaluation, permission boundaries, and failure recovery.
Have a workflow that fails in a hard-to-explain way, or an integration idea for LoopPrism? Get in touch with the task, expected outcome, and a sanitized failure example. Please don't send credentials, private logs, or employer-confidential material.


