I am a master’s student at Peking University. My research focuses on reinforcement learning, agentic RL, and multimodal learning. I am interested in how models learn to reason, use tools, and understand and generate multimodal content.
M.S. Student in Computer Science · Peking University B.S. in Computer Science · Sichuan University
Research Intern · Mar 2026 – Present Foundation Model Department
As a main contributor to UniRL, a reinforcement learning framework for unified multimodal models, I primarily developed the prompt enhancement (PE) trainer and agentic RL components, supporting image-feedback-driven prompt optimization and multi-turn tool-use training.
Research Intern · Aug 2024 – Feb 2025 Data Intelligence Platform & Services Department
I contributed to DLRover, an automatic distributed deep learning system for efficient and reliable training. My work focused on designing and implementing Flash-checkpoint, along with Auto Accelerate experiments and analysis for large-scale recommendation model training.
For AlignedGen, I co-led the work as an equal-contribution first author, designing the core method, implementing the pipeline, and conducting the main experiments.
In collaboration with THU IVG Lab, I led the core method design, end-to-end implementation, and major experiments for DDAVS on audio-visual segmentation.
A FOUNDATION IN PROBLEM SOLVING
Honors & awards
Gold Medal, ICPC Jinan Regional Contest, 2022
Bronze Medal, ICPC East Asia Regional Final (ICPC-EC Final), 2023
Gold Medal, CCF Collegiate Computer Systems and Programming Contest (CCSP), 2023
Gold Medal (2nd place), Sichuan Collegiate Programming Contest, 2024
Bronze Medal, National Olympiad in Informatics (NOI), 2021