Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio
- Published
- Source
- arXiv
- Paper number
- 591
- Field
- Computer Vision
- arXiv ID
- 2607.08127
Key points
- The VAG gap is the phenomenon in which the compositional prior of video foundation models is lost after robot behavior fine-tuning, and we define and analyze it systematically.
- Temporal Ratio (TR) measures how much the action head's attention focuses on future latent rollouts compared with the current frame, and it predicts compositional generalization ability.
- TR changes dynamically across task phases: it rises during planning and falls during precise manipulation, which is the natural attention pattern.
- TR-Adaptive Guidance is an inference-time method that adaptively amplifies compositional conditioning signals during planning phases.
- On the LIBERO benchmark, it improves average OOD success by more than 5x over the previous VAM baseline, from about 10% to 59.4% with guidance.
- It also achieves an average of 83.3% on real bimanual robot tasks in YAM, far above pi0 at 36.7% and pi0.5 at 55.0%.
Paper links
External research summaries. These are not HDATF publications or measured product results.