Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
- Published
- Source
- arXiv
- Paper number
- 710
- Field
- Computer Vision
- arXiv ID
- 2607.21072
Key points
- It proposes the ProVisE framework, which makes image generation models express spatial answers directly in pixel space.
- It builds the SpatialGen-Bench diagnostic benchmark with 470 samples and 14 subtasks.
- An agent builder automatically constructs and validates an evaluation protocol for a new benchmark.
- Image generation models are competitive on intuitive visual-expression tasks.
- Eighty-eight percent of failures come from errors in spatial reasoning itself, not from the visual communication channel.
Paper links
External research summaries. These are not HDATF publications or measured product results.