I am a third-year undergraduate in Physics at Shanghai Jiao Tong University. I study how scalable feedback can help AI systems acquire, verify, and improve generalizable reasoning.

I am currently a research intern at THU C3I, supervised by Ning Ding. Previously, I worked with Jie Fu at Shanghai AI Lab and with Junchi Yan and Renqiu Xia at Shanghai Jiao Tong University.

CV / Publications / GitHub / OpenReview

Research Agenda

Feedback-based learning for reasoning

I am interested in when and why reinforcement learning improves reasoning beyond imitation, especially under weak supervision, sparse feedback, and long-horizon exploration. I care about practical failure modes such as instability, reward hacking, and train-inference mismatch.

Verifiable and self-evolving AI systems

I study how models can interact with formal or programmatic verifiers, generate tasks for themselves, and use curriculum or self-play mechanisms to create scalable supervision. The broader goal is to move from externally curated data toward systems that can propose, solve, and verify increasingly challenging problems.

Hierarchical evaluation of multimodal reasoning

I am interested in decomposing reasoning failures into interpretable stages such as perception, planning, theorem application, and self-reflection. This motivates my work on GeoBench, a structured testbed for multimodal geometry reasoning.

Selected Work

ICLR 2026Benchmark

GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation

Question. When a multimodal model fails on geometry, is the bottleneck perception, planning, theorem application, or backtracking?

We introduce a hierarchical benchmark that separates these stages and makes multimodal reasoning failures easier to diagnose.

News

  • Co4ICF was accepted to SIGKDD 2026.
  • GeoBench was accepted to ICLR 2026.
  • Received the Xiaomi Scholarship.
  • Joined THU C3I as a research intern.

Experience

Jul 2025 - Dec 2025
Research Intern, Big AI Dream Lab, Shanghai AI Lab, supervised by Jie Fu.

Mar 2025 - May 2025
Research Intern, SAI, Shanghai Jiao Tong University, supervised by Junchi Yan and Renqiu Xia.

Education

BSc in Physics, Zhiyuan Honor College, Shanghai Jiao Tong University, expected 2027.

Contact

I am happy to discuss interesting problems in reinforcement learning, verifiable AI, multimodal reasoning, and AI for science. The best way to reach me is email. You can also find my work on GitHub and OpenReview.