Orca: The World is in Your Mind
- Published
- Source
- arXiv
- Paper number
- 520
- Field
- Computer Vision
- arXiv ID
- 2606.30534
Key points
- It proposes Next-State Prediction as a unified paradigm instead of Next-Token, Next-Frame, or Next-Action Prediction.
- It designs two complementary paradigms, unconscious learning through dense state transitions in continuous video and conscious learning through language-conditioned sparse event transitions.
- It builds a large-scale world-learning dataset with 125,000 hours of video and 160 million event annotations.
- It freezes the encoder and trains only a lightweight decoder to produce text, image, and action readouts.
- It validates scalability at the 4B and 0.8B scales, showing that larger world latents translate directly into stronger downstream performance.
- It also openly states clear limitations, including the current restriction to vision and language modalities, short-horizon transitions, and limited model scale.
Paper links
External research summaries. These are not HDATF publications or measured product results.