Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

Published
Source
arXiv
Paper number
591
Field
Computer Vision
arXiv ID
2607.08127

Key points

  • The VAG gap is the phenomenon in which the compositional prior of video foundation models is lost after robot behavior fine-tuning, and we define and analyze it systematically.
  • Temporal Ratio (TR) measures how much the action head's attention focuses on future latent rollouts compared with the current frame, and it predicts compositional generalization ability.
  • TR changes dynamically across task phases: it rises during planning and falls during precise manipulation, which is the natural attention pattern.
  • TR-Adaptive Guidance is an inference-time method that adaptively amplifies compositional conditioning signals during planning phases.
  • On the LIBERO benchmark, it improves average OOD success by more than 5x over the previous VAM baseline, from about 10% to 59.4% with guidance.
  • It also achieves an average of 83.3% on real bimanual robot tasks in YAM, far above pi0 at 36.7% and pi0.5 at 55.0%.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)