Native Video-Action Pretraining for Generalizable Robot Control

Published
Source
arXiv
Paper number
585
Field
Robotics
arXiv ID
2607.08639

Key points

  • Semantic visual-action tokenizer SemVAE aligns visual representations and actions in a shared latent space.
  • Causal pretraining avoids catastrophic forgetting that appears when applying a bidirectional architecture, and it trains from scratch.
  • A Sparse MoE backbone expands model capacity while keeping the active compute per step unchanged.
  • Foresight Reasoning asynchronous inference achieves 225 Hz real-time closed-loop control, a 6.5x speedup over the 35 Hz baseline.
  • It achieves few-shot generalization with 10 to 15 demos and outperforms π0.5 and LingBot-VA on both simulation and real evaluation.
  • Multi-chunk prediction, or MCP, encourages trajectory-level dynamics learning instead of short-term visual continuity.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)