$μ_0$: A Scalable 3D Interaction-Trace World Model

Published
Source
arXiv
Paper number
418
Field
Robotics
arXiv ID
2606.13769

Key points

  • TraceExtract automates 3D trace extraction by selecting semantic keypoints through DINOv2 entity clustering, expanding the scale by about 8x over prior work.
  • It represents 3D traces with B-spline control points and trains a VLM backbone plus a permutation-equivariant trace expert with flow matching.
  • It pretrains from video only, without action labels, then freezes the model and trains a separate action expert for cross-embodiment transfer.
  • It reaches a 30.25% average on RoboCasa 8 tasks, beating π0's 25.25% by 5.0 points, and achieves 91.7% on three real-world UR3 tasks.
  • A +18.4-point gap over a VLM plus action expert without trace features shows that trace representations provide real motion information.
  • Event-centric captioning aligns language with trace segments, and depth input is not needed at inference time.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)