Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
- Published
- Source
- arXiv
- Paper number
- 492
- Field
- Computer Vision
- arXiv ID
- 2606.25041
Key points
- Wan-Streamer models language, audio, and video as both inputs and outputs with a single transformer, and uses block-causal attention to support incremental streaming.
- It jointly learns cognition, reasoning, generation, response timing, turn management, and cross-modal synchronization in one model instead of relying on separate VAD, ASR, language, TTS, avatar, and video generation modules.
Paper links
External research summaries. These are not HDATF publications or measured product results.