Show-Harness: Just a VLM Agent Can Play Robots

Published
Source
arXiv
Paper number
1077
Field
Robotics
arXiv ID
2609.10522

Key points

  • Without robot-specific retraining as in vision-language-action models, it taps a foundation VLM's intelligence for robot control through a meaningful motion-primitive interface alone.
  • Real robot-arm experiments confirm that state-of-the-art closed VLMs can be used directly as robot agents zero-shot, with no training.
  • Small open-source VLMs (2B class) handle robots in the same action space after just a few GPU-hours of light fine-tuning, opening a low-cost deployment path.
  • It generalizes better across tasks, robot morphologies, and environments than prior agent approaches and VLA models.
  • A GUI-based control interface called GUMI lets humans and agents collect robot demonstration data together without expensive teleoperation rigs.
  • Component ablations show planning, failure recovery, and motion logging lift success rates by up to 24–35 percentage points.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)