My research focuses on training embodied foundation models that can perform a wide range of tasks in the physical world. I believe that building on top of pretrained foundation models is important for achieving generalization and scalability.
To this end, I have explored VLM-based VLA Policy and Video-model-based Policy. Previously, I also do research on LLM agent and RL.
Besides research, I enjoy playing soccer ⚽. I also love physics 👾 —the simplicity and elegance of physics law strongly appeal to me.
Awards: [2024.07] Best Paper Award Finalists in RSS 2024.
[2022.06] Outstanding Graduates Award (Top 10% Tsinghua undergraduate students).
[2017.11] Silver Medal in 34th National Physics Olympiad (CPhO).
Selected Research (* indicates equal contribution)
Video Prediction Policy 2: Predict Better, Act Better Yanjiang Guo*, Haodong Yan*, Zhide Zhong*, Zhongru Zhang*, Qingyuan Yang*, Qingzhou Lu, Xiaoyu Chen, Yen-Jen Wang, Shuying Deng, Chenghan Yang, Puzhen Yuan, Chenxin Liu, Tun Ban, Xiang Zhu, Yichen Liu, Kun Feng, Haoang Li, Jianyu Chen
Preprint, 2026 project page
/
code
/
arXiv
/
Hugging Face
/
量子位
We introduce VPP2, a world action model with strong zero-shot generalization in video prediction and action generation, combining event-level video pretraining, single-step visual planning, and a mixture-of-transformers action module.
We incoperate both multi-modal understanding (MMU) and future prediction into VLA model, enhancing both high-level semantic knowledge and low-level visual dynamics.
We make some initial exploration on leveraging online RL to improve the VLA model! We notice that online RL for VLA can be extremely unstable and thus we adopted a iterative approach.