LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Published
Source
arXiv
Paper number
247
Field
Computer Vision
arXiv ID
2605.25979

Key points

  • In stage 3, which focuses on long videos, the model is exposed to a length-hierarchical video caption corpus of about 8 million clips and 104 billion tokens in total, with videos ranging from 30 seconds to 15 minutes, so that it learns to maintain context over long durations.
  • In stage 4, which focuses on codec and spatial reasoning, codec-native training is introduced and 10 to 15 minute videos are re-encoded with a variable-length GOP pipeline, while the model is also trained on the LLaVA-OneVision-2-Spatial-4M corpus containing 4 million samples focused on 2D and 3D spatial reasoning such as identifying objects in indoor scans or tracking motion in embodied simulators.
  • The codec-stream approach performs substantially better on temporal grounding and visual perception because it reallocates tokens to event-bearing moments and captures motion dynamics that uniform sampling misses.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)