Orca: The World is in Your Mind

Published
Source
arXiv
Paper number
520
Field
Computer Vision
arXiv ID
2606.30534

Key points

  • It proposes Next-State Prediction as a unified paradigm instead of Next-Token, Next-Frame, or Next-Action Prediction.
  • It designs two complementary paradigms, unconscious learning through dense state transitions in continuous video and conscious learning through language-conditioned sparse event transitions.
  • It builds a large-scale world-learning dataset with 125,000 hours of video and 160 million event annotations.
  • It freezes the encoder and trains only a lightweight decoder to produce text, image, and action readouts.
  • It validates scalability at the 4B and 0.8B scales, showing that larger world latents translate directly into stronger downstream performance.
  • It also openly states clear limitations, including the current restriction to vision and language modalities, short-horizon transitions, and limited model scale.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)