EgoCS-400K: An Egocentric Gameplay Dataset for World Models

Published
Source
arXiv
Paper number
443
Field
Computer Vision
arXiv ID
2606.18180

Key points

  • The dataset contains more than 400K first-person videos and more than 10,000 hours of gameplay data built from professional CS and CS2 demos.
  • It uses tick-level precise alignment so that keyboard and mouse input, camera motion, weapon and movement state, and game events are all synchronized.
  • The hierarchy goes from player-view sequence to segment to action chain to atomic action to per-tick state trace.
  • It covers more than 1,000 matches, more than 40,000 rounds, 13 maps, and up to 10 player viewpoints per round.
  • Its replay-grounded structure makes parsing, replaying, rendering, and temporal alignment all possible, which distinguishes it from simple video capture.
  • It supports action-conditioned future prediction, state-aware scene rollout, and egocentric action understanding.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)