YD / Yiheng Du

REINFORCEMENT LEARNING & MULTIMODAL AI

Yiheng Du杜毅衡

Exploring how machines
perceive and create.

I am a master’s student at Peking University. My research focuses on reinforcement learning, agentic RL, and multimodal learning. I am interested in how models learn to reason, use tools, and understand and generate multimodal content.

M.S. Student in Computer Science · Peking University
B.S. in Computer Science · Sichuan University

Portrait of Yiheng Du
YIHENG DURESEARCH / ENGINEERING
01 Reinforcement Learning
02 Agentic RL
03 Multimodal Learning

BUILDING IN THE OPEN

Open source & industry

Contributions through internships
at Tencent Hunyuan and Ant Group.

TencentTencent Hunyuan
Tencent · Hunyuan · INTERNSHIP☆ ~1k

UniRL

Research Intern · Mar 2026 – Present
Foundation Model Department

As a main contributor to UniRL, a reinforcement learning framework for unified multimodal models, I primarily developed the prompt enhancement (PE) trainer and agentic RL components, supporting image-feedback-driven prompt optimization and multi-turn tool-use training.

Explore repository ↗
Ant Group
Ant Group · INTERNSHIP☆ 1k+

DLRover

Research Intern · Aug 2024 – Feb 2025
Data Intelligence Platform & Services Department

I contributed to DLRover, an automatic distributed deep learning system for efficient and reliable training. My work focused on designing and implementing Flash-checkpoint, along with Auto Accelerate experiments and analysis for large-scale recommendation model training.

Explore repository ↗

IDEAS INTO PRACTICE

Research & publications03

Generative models, perception,
and the systems behind them.

* Equal contribution.

THE PATH SO FAR

Education

SEPT 2026 — JUN 2029 (EXPECTED)

Peking University

M.S. in Computer Science and Technology

SEPT 2022 — JUN 2026

Sichuan University

B.S. in Computer Science and Technology

RESEARCH EXPERIENCE

Peking University · VILLA Lab

RESEARCH INTERN · FROM MAR 2025

For AlignedGen, I co-led the work as an equal-contribution first author, designing the core method, implementing the pipeline, and conducting the main experiments.

In collaboration with THU IVG Lab, I led the core method design, end-to-end implementation, and major experiments for DDAVS on audio-visual segmentation.

A FOUNDATION IN PROBLEM SOLVING

Honors & awards

  • Gold Medal, ICPC Jinan Regional Contest, 2022
  • Bronze Medal, ICPC East Asia Regional Final (ICPC-EC Final), 2023
  • Gold Medal, CCF Collegiate Computer Systems and Programming Contest (CCSP), 2023
  • Gold Medal (2nd place), Sichuan Collegiate Programming Contest, 2024
  • Bronze Medal, National Olympiad in Informatics (NOI), 2021
  • Silver Medal, Asia-Pacific Informatics Olympiad (APIO), 2021
More awards
  • Silver Medal, ICPC Nanjing Regional Contest, 2022
  • Silver Medal, China Collegiate Programming Contest (CCPC), 2022
  • First Prize, The 15th Lanqiao Cup National Software Competition, 2023
  • Gold Medal (3rd place), Sichuan Collegiate Programming Contest, 2023
  • Gold Medal (3rd place), Sichuan Collegiate Programming Contest, 2022
  • First Prize, National Olympiad in Informatics in Provinces (NOIP), 2021
  • First Prize, National Olympiad in Informatics in Provinces (NOIP), 2020
  • Bronze Medal, Asia-Pacific Informatics Olympiad (APIO), 2020

LET’S CONNECT

Good research starts
with a conversation.

yihengdu42@gmail.com