personal page / cv · scholar · email · X · rednote
I am currently a Ph.D. student at Peking University.
I build models:
- Post-train: Data-efficient and stable reinforcement learning across domains, including RL and OPD for agentic AI and reasoning.
- Data & Eval: Data synthesis research and reliable evaluation.
Open for any intern opportunities.


