I work on LLM serving at Google, and on the open-source inference stack everywhere else. Mostly that means KV-cache management and layout, prefill-decode disaggregation, and the test and transfer infrastructure underneath them.
- Triage collaborator on kvcached, GPU KV-cache virtualization for vLLM and SGLang
- Workgroup member (Evaluation & Quality) in vllm-project/semantic-router
- In Mooncake I wrote the MasterScenario DSL its store tests now run on. In LMCache I'm proposing a declarative KV-layout descriptor.
- Contributions across SGLang, vLLM, llm-d, TileLang, LLM Compressor, NIXL, and inference-perf
Recent papers: Latent-State Auditing for Tool-Using Language Agents · SuperRed: AI Red-Teaming & Security Benchmarking


