PhiZero: A World Model Built Around Physical Language
- Published
- Source
- arXiv
- Paper number
- 772
- Field
- Computer Vision
- arXiv ID
- 2607.28624
Key points
- We discovered a discrete compressed representation called a physics language in natural videos through self-supervised learning. It symbolizes patterns of state change rather than pixels.
- A reasoning-then-rendering architecture explicitly predicts physical evolution and then synthesizes video with a diffusion model.
- It substantially improves physical consistency over prior world models on video generation and physics-understanding benchmarks.
- It demonstrates zero-shot action transfer by transferring human motions to humanoid robots or robotic hands.
- It learns physical common sense from video data alone, without physics laws or simulators.
Paper links
External research summaries. These are not HDATF publications or measured product results.