VISTA: A Visual Harness for Reasoning in an Interactive World

Published
Source
arXiv
Paper number
1147
Field
AI / Agents
arXiv ID
2610.02200

Key points

  • The system was designed to preserve past screens in their original form and let the model enlarge and revisit relevant views.
  • The authors reported solving all 25 public ARC-AGI-3 games while using 57.4% fewer actions than first-time human participants.
  • The authors reported improving accuracy from 41.0% to 63.2% on 39 BabyVision visual tracking questions.
  • The authors stated that they could not rule out the possibility that the public games were included in model training.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)