In-Context Robot Learning with VLM Agents
- Published
- Source
- arXiv
- Paper number
- 1090
- Field
- Computer Vision
- arXiv ID
- 2609.19138
Key points
- It is one of the first systematic measurements of how far commercial VLMs alone, without robot training, can adapt in the field.
- With only a human video demonstration, task progress rose from 55% to 100% while run time and token usage decreased.
- It compared five context types, including goal images, robot video with actions, and interaction history, within a single execution framework.
- On human-robot interaction tasks like tic-tac-toe, GPT-6 Astra succeeded 3/3 while handling turn-taking and strategic reasoning together.
- On the other hand, it clearly exposed the limitation that precise contact, outcome verification, and safety issues cannot be solved by general-purpose models alone.
Paper links
External research summaries. These are not HDATF publications or measured product results.