Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Published
Source
arXiv
Paper number
903
Field
Computer Vision
arXiv ID
2608.13546

Key points

  • It stores persistent state in an external geometric bank indexed by camera pose, which keeps the context cost fixed regardless of session length.
  • The teacher combines chunk-level sparse attention with a linear-attention global state so that supervision cost grows only linearly with length.
  • It distills a three-stage, CFG-free student on a 30-second long-horizon distribution-matching objective, which gives both low-latency interaction and long-horizon consistency.
  • It demonstrates two hours of continuous generation and achieves top results on WBench and a score of 66.77 on VBench-2.0, which ranks first among 10 compared systems.
  • During a session, text instruction replacement can inject 67 percent of new content, but only 4 percent of content already fixed in geometry is overwritten.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)