Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
- Published
- Source
- arXiv
- Paper number
- 903
- Field
- Computer Vision
- arXiv ID
- 2608.13546
Key points
- It stores persistent state in an external geometric bank indexed by camera pose, which keeps the context cost fixed regardless of session length.
- The teacher combines chunk-level sparse attention with a linear-attention global state so that supervision cost grows only linearly with length.
- It distills a three-stage, CFG-free student on a 30-second long-horizon distribution-matching objective, which gives both low-latency interaction and long-horizon consistency.
- It demonstrates two hours of continuous generation and achieves top results on WBench and a score of 66.77 on VBench-2.0, which ranks first among 10 compared systems.
- During a session, text instruction replacement can inject 67 percent of new content, but only 4 percent of content already fixed in geometry is overwritten.
Paper links
External research summaries. These are not HDATF publications or measured product results.