Chenchen Hong AI Infrastructure · MLSys · Compiler Systems Large-model serving performance across kernels, schedulers, and distributed inference. Inference systems SGLang-omni, SGLang, and vLLM serving; scheduling and memory efficiency. RL systems Rollout orchestration, training/inference infrastructure, and scaling workflows. Compiler kernels Triton/CUDA kernels, H100/B200 tuning, codegen, and graph optimization. Email · LinkedIn · Blog · X / Twitter · WeChat: hayden-gai GitHub Stats