Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Published
Source
arXiv
Paper number
492
Field
Computer Vision
arXiv ID
2606.25041

Key points

  • Wan-Streamer models language, audio, and video as both inputs and outputs with a single transformer, and uses block-causal attention to support incremental streaming.
  • It jointly learns cognition, reasoning, generation, response timing, turn management, and cross-modal synchronization in one model instead of relying on separate VAD, ASR, language, TTS, avatar, and video generation modules.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)