Video = World + Event Stream
- Published
- Source
- arXiv
- Paper number
- 642
- Field
- Computer Vision
- arXiv ID
- 2607.15038
Key points
- It proposes a decomposition in which video equals world plus event stream, and uses it to define a general-purpose pretraining task.
- It adds free-form actions, expressed as natural-language instructions in parentheses, to the event stream so that the agent can perform open-vocabulary actions rather than a fixed action set.
- It maintains real-time duplex audio-visual dialogue with 200 ms model response latency and 550 ms end-to-end interaction latency.
- It demonstrates flexible transfer from pretraining to downstream tasks, suggesting that the same streaming logic can adapt to a range of real-time problems.
Paper links
External research summaries. These are not HDATF publications or measured product results.