LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
- Published
- Source
- arXiv
- Paper number
- 247
- Field
- Computer Vision
- arXiv ID
- 2605.25979
Key points
- In stage 3, which focuses on long videos, the model is exposed to a length-hierarchical video caption corpus of about 8 million clips and 104 billion tokens in total, with videos ranging from 30 seconds to 15 minutes, so that it learns to maintain context over long durations.
- In stage 4, which focuses on codec and spatial reasoning, codec-native training is introduced and 10 to 15 minute videos are re-encoded with a variable-length GOP pipeline, while the model is also trained on the LLaVA-OneVision-2-Spatial-4M corpus containing 4 million samples focused on 2D and 3D spatial reasoning such as identifying objects in indoor scans or tracking motion in embodied simulators.
- The codec-stream approach performs substantially better on temporal grounding and visual perception because it reallocates tokens to event-bearing moments and captures motion dynamics that uniform sampling misses.
Paper links
External research summaries. These are not HDATF publications or measured product results.