I work at the intersection of RL and ML systems — making LLM post-training both smarter and faster.
You can reach me by email: linkai0508@gmail.com
- RL for LLMs — GRPO/PPO post-training, test-time scaling, self-rewarding and intrinsic reward signals
- ML systems — train–inference consistency, distributed training (Megatron-LM), high-performance CUDA/Triton kernels
- RL & Bayesian optimization — non-Markovian RL for multi-objective BO
- Core maintainer of RL-Align/RL-Kernel — CUDA/Triton kernels for bitwise-consistent RL training across training/inference engines.
- RL.cu — from-scratch RLVR training + inference engine for LLMs in pure CUDA/C++
| Languages | Python, C++, CUDA, Triton |
| ML / RL | PyTorch, Megatron-LM, vLLM, GRPO/PPO, LLM post-training |
| Systems | Kernel optimization, numerical consistency, TP/CP parallelism, Slurm |



