4DAnyone: Create Anyone in 4D from a Casual Monocular Video

Published
Source
arXiv
Paper number
961
Field
Computer Vision
arXiv ID
2608.20335

Key points

  • It built a pipeline that reconstructs 4D humans, moving 3D people, by consistently generating videos from 16 viewpoints using only video from a single camera.
  • Reference-view compression (RCP) and viewpoint-group cycling (TCR) resolved the loss of consistency caused by increasing the number of viewpoints.
  • On the DNA-Rendering and DyMVHumans benchmarks, its 4D-reconstruction PSNR of 24.15/23.28 substantially exceeded the previous best method's 20.55/19.86.
  • It uses only a 3D skeleton as conditioning instead of unreliable depth information, allowing it to work on videos with unknown camera parameters.
  • It built the MVGameHuman training dataset using its own game engine, achieving generalization robust to varied outdoor videos.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)