PhiZero: A World Model Built Around Physical Language

Published
Source
arXiv
Paper number
772
Field
Computer Vision
arXiv ID
2607.28624

Key points

  • We discovered a discrete compressed representation called a physics language in natural videos through self-supervised learning. It symbolizes patterns of state change rather than pixels.
  • A reasoning-then-rendering architecture explicitly predicts physical evolution and then synthesizes video with a diffusion model.
  • It substantially improves physical consistency over prior world models on video generation and physics-understanding benchmarks.
  • It demonstrates zero-shot action transfer by transferring human motions to humanoid robots or robotic hands.
  • It learns physical common sense from video data alone, without physics laws or simulators.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)