VISTA: A Visual Harness for Reasoning in an Interactive World
- Published
- Source
- arXiv
- Paper number
- 1147
- Field
- AI / Agents
- arXiv ID
- 2610.02200
Key points
- The system was designed to preserve past screens in their original form and let the model enlarge and revisit relevant views.
- The authors reported solving all 25 public ARC-AGI-3 games while using 57.4% fewer actions than first-time human participants.
- The authors reported improving accuracy from 41.0% to 63.2% on 39 BabyVision visual tracking questions.
- The authors stated that they could not rule out the possibility that the public games were included in model training.
Paper links
External research summaries. These are not HDATF publications or measured product results.