Next Forcing: Causal World Modeling with Multi-Chunk Prediction

Published
Source
arXiv
Paper number
383
Field
Computer Vision
arXiv ID
2606.11187

Key points

  • We solve the myopic supervision problem in autoregressive video generation with multi-chunk prediction (MCP).
  • A causal MCP chain lets near-future predictions inform far-future predictions and provides dense supervision to the main model.
  • We achieve SOTA on RoboTwin, with 94.1% on Clean and 93.5% on Random.
  • At 50 fps, it delivers a 93.1% relative improvement over LingBot-VA and a 2.3x training speedup.
  • General video pretraining also reduces FVD by more than 50%, demonstrating generality beyond robot-specific data.
  • Keeping the MCP module at inference time doubles inference speed while preserving accuracy.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)