Video Prediction Policy 2: Predict Better, Act Better

Yanjiang Guo1,2,*, Haodong Yan1,3,*, Zhide Zhong1,3,*, Zhongru Zhang2,*, Qingyuan Yang1,2,*,
Qingzhou Lu1,2, Xiaoyu Chen1, Yen-Jen Wang4, Shuying Deng1,2,4, Chenghan Yang2, Puzhen Yuan1,2,
Chenxin Liu1,2, Tun Ban1,5, Xiang Zhu1,2, Yichen Liu1,2, Kun Feng1,2, Haoang Li3, Jianyu Chen1,2
*Equal Contribution 1Robotera 2Tsinghua University 3HKUST (GZ)
4University of California, Berkeley 5Shanghai Jiaotong University
Overview of Video Prediction Policy 2 (VPP2) VPP2 training architecture: video pretraining, post-training and distillation, and action training

We introduce Video Prediction Policy 2 (VPP2), a WAM that enables strong zero-shot generalization in both video prediction and action generation.

The whole pipeline:

(1) First, we curate a large-scale, diverse dataset of manipulation videos to continue pretraining the base video foundation model. We annotate video clips with detailed captions and perform event-level video pretraining to promote generalization across open-ended manipulation tasks.

(2) Second, we post-train and distill the video model into a single-step visual planner with fixed prediction horizon.

(3) Finally, we introduce action module via a mixture-of-transformers (MoT) architecture to learn implicit inverse dynamics model.

Open-World Manipulation Video Prediction

Video predictions of human-hand and robot-arm manipulation in open environments.

Human-hand manipulation

Approach the AirPod in front of the bottle.
Place the leftmost bottle in front of the cup.
Use the scissors to cut the paper.
Close the screen of the MacBook.
Grasp a tissue and wipe the screen.
Place the spoon on the iPad.
Pull out the gas nozzle.
Open the oven.
Remove the lid and place it on the table.
Pick up the left slipper and place it on the right side.

Robot-arm manipulation

Grasp the sponge and clean the whiteboard.
Place the electronic device into the box.
Place the green block on the controller.
Pour water into the blue bowl.
Fold the towel from right to left.
Place the lid back on the bottle.
Open the upper drawer.
Sweep the block into the dustbin.
Grasp the blue towel and wipe the plate.
Flip the package.

Faithful Instruction Following

Video predictions from Wan2.1-I2V-14B, Cosmos3-Super-64B, and VPP2 given the same initial frame and different instructions.

Model comparison

Initial frame
Instruction
Wan2.1-I2V-14B
Cosmos3-Super-64B
VPP2 (Ours)
Shared initial frame for fig5
Same initial frame
Place the second bottle from the left in front of the leftmost bottle.
Place the leftmost bottle in front of the cup.
Place the cup in front of the leftmost cup.
Initial frame
Instruction
Wan2.1-I2V-14B
Cosmos3-Super-64B
VPP2 (Ours)
Shared initial frame for fig5-2
Same initial frame
Stack the red paper cup on the leftmost black cup.
Stack the red paper cup on the second black cup from the left.

Citation

If you use VPP2 in your research, please cite our paper. Download BibTeX

@misc{guo2026videopredictionpolicy2,
  title = {Video Prediction Policy 2: Predict Better, Act Better},
  author = {Yanjiang Guo and Haodong Yan and Zhide Zhong and Zhongru Zhang and
            Qingyuan Yang and Qingzhou Lu and Xiaoyu Chen and Yen-Jen Wang and
            Shuying Deng and Chenghan Yang and Puzhen Yuan and Chenxin Liu and
            Tun Ban and Xiang Zhu and Yichen Liu and Kun Feng and Haoang Li and
            Jianyu Chen},
  year = {2026},
  eprint = {2610.10270},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url = {https://arxiv.org/abs/2610.10270}
}