Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

Published
Source
arXiv
Paper number
357
Field
Computer Vision
arXiv ID
2606.05769

Key points

  • We propose an interleaved latent visual reasoning architecture that alternates between language tokens and continuous latent visual tokens.
  • Using visual-gain data curation, we built 50K high-quality training examples in which future visual hints genuinely help prediction.
  • LA-DAPO combines outcome-contrastive reward and temporal-diversity reward as a latent-awareness RL objective.
  • On FutureBench, Qwen3-VL-8B improves from 61.0 to 85.4, a 10.4-point gain over the previous best, Video-CoE.
  • On TwiFF-Bench, the average score rises from 2.44 to 3.04, and latent span usage increases adaptively as reasoning depth grows.
  • With the same data, text-only SFT reaches only 65.0, whereas interleaved latent SFT reaches 73.2, showing that the gain does not come from extra supervision.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)