HelloWorld: Enabling Socially Interactive Characters in Video World Models

Published
Source
arXiv
Paper number
828
Field
Computer Vision
arXiv ID
2608.05070

Key points

  • This is the first video world model to realize social interaction between a character and a user, such as waving, greeting, and speaking to the user.
  • The paper proposes a self-distillation pipeline that reuses interaction videos generated by the model itself as training data.
  • Without additional training, it controls the timing of interactions at the frame level using only cross-attention masks.
  • It builds the first social-interaction benchmark, HelloWorldBench, with 400 samples and 3 interaction metrics.
  • It substantially outperforms existing world models on interaction quality while maintaining SOTA-level video aesthetics and camera following.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)