ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
- Published
- Source
- arXiv
- Paper number
- 538
- Field
- Computer Vision
- arXiv ID
- 2607.00678
Key points
- It identifies three structural bottlenecks: temporal resolution mismatch between coarse video and fine-grained control, action-space mismatch between navigation and manipulation, and context mismatch between training and inference.
- It introduces an intermediate latent action that bridges video latent representations and embodiment-specific control while capturing local visual state transitions.
- A dual-level Mixture-of-Transformers separates modality representation from heterogeneous action subspaces, such as base movement and arm manipulation, so that each can be optimized separately.
- Dream forcing trains dynamics on model-predicted dream video, which aligns autoregressive inference with training conditions and reduces long-horizon rollout error accumulation.
- On RoboCasa365 mobile manipulation, it reaches 90.8 percent, 75.4 percent, and 58.4 percent across difficulty levels, and on real-robot peg insertion it achieves a 70 percent success rate and a 96 percent process score, compared with 50 percent and 90 percent for pi0.5 and 30 percent and 77 percent for Fast-WAM.
- On real robot long-horizon tasks such as plate sorting, fruit sorting, and flower arrangement, it achieves 60 to 80 percent success, which is a large improvement over Fast-WAM at 20 to 40 percent.
Paper links
External research summaries. These are not HDATF publications or measured product results.