In-Context Robot Learning with VLM Agents

Published
Source
arXiv
Paper number
1090
Field
Computer Vision
arXiv ID
2609.19138

Key points

  • It is one of the first systematic measurements of how far commercial VLMs alone, without robot training, can adapt in the field.
  • With only a human video demonstration, task progress rose from 55% to 100% while run time and token usage decreased.
  • It compared five context types, including goal images, robot video with actions, and interaction history, within a single execution framework.
  • On human-robot interaction tasks like tic-tac-toe, GPT-6 Astra succeeded 3/3 while handling turn-taking and strategic reasoning together.
  • On the other hand, it clearly exposed the limitation that precise contact, outcome verification, and safety issues cannot be solved by general-purpose models alone.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)