Visual General Intelligence: A White Paper
- Published
- Source
- arXiv
- Paper number
- 1019
- Field
- Computer Vision
- arXiv ID
- 2608.25924
Key points
- Bringing together multiple researchers' perspectives, it examined whether visual experience, which predates language by far, could provide an independent starting point for general intelligence.
- It proposed large-scale generative learning, future prediction, spatial memory, continual learning, active perception, and embodied interaction as candidate paths to visual intelligence.
- It argued that evaluation should move beyond fixed-task scores to assess transfer to new situations, continual adaptation, creativity, physical consistency, and active information seeking.
- It concluded that multiple paths, including vision-only, vision-first, and language-integrated approaches, should remain open, rather than prematurely fixing visual general intelligence to a single model or definition.
- This work is a white paper organizing a research agenda rather than a paper presenting a new model or integrated experiments, and current visual capabilities remain fragmented, with limits in computational cost and temporal scope.
Paper links
External research summaries. These are not HDATF publications or measured product results.