Video = World + Event Stream

Published
Source
arXiv
Paper number
642
Field
Computer Vision
arXiv ID
2607.15038

Key points

  • It proposes a decomposition in which video equals world plus event stream, and uses it to define a general-purpose pretraining task.
  • It adds free-form actions, expressed as natural-language instructions in parentheses, to the event stream so that the agent can perform open-vocabulary actions rather than a fixed action set.
  • It maintains real-time duplex audio-visual dialogue with 200 ms model response latency and 550 ms end-to-end interaction latency.
  • It demonstrates flexible transfer from pretraining to downstream tasks, suggesting that the same streaming logic can adapt to a range of real-time problems.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)