Visual General Intelligence: A White Paper

Published
Source
arXiv
Paper number
1019
Field
Computer Vision
arXiv ID
2608.25924

Key points

  • Bringing together multiple researchers' perspectives, it examined whether visual experience, which predates language by far, could provide an independent starting point for general intelligence.
  • It proposed large-scale generative learning, future prediction, spatial memory, continual learning, active perception, and embodied interaction as candidate paths to visual intelligence.
  • It argued that evaluation should move beyond fixed-task scores to assess transfer to new situations, continual adaptation, creativity, physical consistency, and active information seeking.
  • It concluded that multiple paths, including vision-only, vision-first, and language-integrated approaches, should remain open, rather than prematurely fixing visual general intelligence to a single model or definition.
  • This work is a white paper organizing a research agenda rather than a paper presenting a new model or integrated experiments, and current visual capabilities remain fragmented, with limits in computational cost and temporal scope.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)