Robots Need More than VLA and World Models

Published
Source
arXiv
Paper number
370
Field
Robotics
arXiv ID
2606.06556

Key points

  • It redefines VLA models as one layer in the physical intelligence stack and argues that upstream and downstream grounding mechanisms are essential.
  • It advocates a shift to grounding-centered pipelines that convert unstructured physical experience into robot-usable supervision.
  • It proposes four essential interfaces: a physical data engine, embodiment retargeting, physics-based world models, and task-conditioned reward grounding.
  • It presents a vision for self-improving systems that use success, failure, and human correction as structured supervision in the deployment loop.
  • It revisits state-of-the-art research in robot-native data, video-based methods, simulation, and world models from the perspective of grounding bottlenecks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)