Robots Need More than VLA and World Models
- Published
- Source
- arXiv
- Paper number
- 370
- Field
- Robotics
- arXiv ID
- 2606.06556
Key points
- It redefines VLA models as one layer in the physical intelligence stack and argues that upstream and downstream grounding mechanisms are essential.
- It advocates a shift to grounding-centered pipelines that convert unstructured physical experience into robot-usable supervision.
- It proposes four essential interfaces: a physical data engine, embodiment retargeting, physics-based world models, and task-conditioned reward grounding.
- It presents a vision for self-improving systems that use success, failure, and human correction as structured supervision in the deployment loop.
- It revisits state-of-the-art research in robot-native data, video-based methods, simulation, and world models from the perspective of grounding bottlenecks.
Paper links
External research summaries. These are not HDATF publications or measured product results.