ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Published
Source
arXiv
Paper number
538
Field
Computer Vision
arXiv ID
2607.00678

Key points

  • It identifies three structural bottlenecks: temporal resolution mismatch between coarse video and fine-grained control, action-space mismatch between navigation and manipulation, and context mismatch between training and inference.
  • It introduces an intermediate latent action that bridges video latent representations and embodiment-specific control while capturing local visual state transitions.
  • A dual-level Mixture-of-Transformers separates modality representation from heterogeneous action subspaces, such as base movement and arm manipulation, so that each can be optimized separately.
  • Dream forcing trains dynamics on model-predicted dream video, which aligns autoregressive inference with training conditions and reduces long-horizon rollout error accumulation.
  • On RoboCasa365 mobile manipulation, it reaches 90.8 percent, 75.4 percent, and 58.4 percent across difficulty levels, and on real-robot peg insertion it achieves a 70 percent success rate and a 96 percent process score, compared with 50 percent and 90 percent for pi0.5 and 30 percent and 77 percent for Fast-WAM.
  • On real robot long-horizon tasks such as plate sorting, fruit sorting, and flower arrangement, it achieves 60 to 80 percent success, which is a large improvement over Fast-WAM at 20 to 40 percent.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)