About me

I lead the GPU Communication team at Meta. We build the communication systems that power the full lifecycle of modern AI—from large-scale LLM pre-training and post-training to latency-sensitive, high-throughput inference. Our work spans collective communication libraries, GPU host networking software, and high-performance RDMA fabrics. By co-designing these layers across GPU platforms, we help researchers train increasingly capable models, iterate efficiently through fine-tuning and reinforcement-learning workloads, and serve them reliably at production scale. This infrastructure supports Meta's frontier language, multimodal, and coding models, including Muse Spark, Muse Video, Muse Image, and Muse Code. Since joining Meta in 2014, I have worked across AI infrastructure, network analytics, traffic engineering, routing, and switch software. We build production infrastructure: understanding a problem rigorously, designing practical mechanisms across hardware and software boundaries, and operating them at scale.

I received my PhD in Computer Science from Stanford University, advised by Professor Nick McKeown and Professor George Varghese. My doctoral research focused on automated network testing and verification. I also hold a Master's Degree in Electrical Engineering from Stanford and a Bachelor's Degree in Electronic Engineering from Tsinghua University. My interests span computer networks, distributed systems, and AI infrastructure—particularly collective communication, AI training and inference systems, data-center and backbone networks, software-defined networking, network verification, and programmable hardware. I enjoy building systems that connect research ideas with the realities of large-scale production.

Selected Publications

Professional Services

More