Native Video-Action Pretraining for Generalizable Robot Control
- Published
- Source
- arXiv
- Paper number
- 585
- Field
- Robotics
- arXiv ID
- 2607.08639
Key points
- Semantic visual-action tokenizer SemVAE aligns visual representations and actions in a shared latent space.
- Causal pretraining avoids catastrophic forgetting that appears when applying a bidirectional architecture, and it trains from scratch.
- A Sparse MoE backbone expands model capacity while keeping the active compute per step unchanged.
- Foresight Reasoning asynchronous inference achieves 225 Hz real-time closed-loop control, a 6.5x speedup over the 35 Hz baseline.
- It achieves few-shot generalization with 10 to 15 demos and outperforms π0.5 and LingBot-VA on both simulation and real evaluation.
- Multi-chunk prediction, or MCP, encourages trajectory-level dynamics learning instead of short-term visual continuity.
Paper links
External research summaries. These are not HDATF publications or measured product results.