$μ_0$: A Scalable 3D Interaction-Trace World Model
- Published
- Source
- arXiv
- Paper number
- 418
- Field
- Robotics
- arXiv ID
- 2606.13769
Key points
- TraceExtract automates 3D trace extraction by selecting semantic keypoints through DINOv2 entity clustering, expanding the scale by about 8x over prior work.
- It represents 3D traces with B-spline control points and trains a VLM backbone plus a permutation-equivariant trace expert with flow matching.
- It pretrains from video only, without action labels, then freezes the model and trains a separate action expert for cross-embodiment transfer.
- It reaches a 30.25% average on RoboCasa 8 tasks, beating π0's 25.25% by 5.0 points, and achieves 91.7% on three real-world UR3 tasks.
- A +18.4-point gap over a VLM plus action expert without trace features shows that trace representations provide real motion information.
- Event-centric captioning aligns language with trace segments, and depth input is not needed at inference time.
Paper links
External research summaries. These are not HDATF publications or measured product results.