Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Published
Source
arXiv
Paper number
914
Field
Computer Vision
arXiv ID
2608.14530

Key points

  • ActionGPT selects actions and PoseGPT creates poses, while a fixed graphics pipeline computes skeletons and occlusion relationships and then a video generation model paints in the appearance.
  • The action model was trained on 1,395 segments at 20 frames per second, and the current action range is limited to one kind of monster.
  • On 48 test segments, forcing incorrect actions increased joint error from 0.272 meters to 0.357 meters, a 31 percent rise, which shows that action tokens control pose.
  • Applying a ground-collision rule reduced the share of frames where joints passed through the ground from 0.337 to 0.114, and a distance-limit rule lowered the distance between characters from 21.2 meters to 5.1 meters.
  • The video quality metric FVD was 831, lower than the pixel-based model's 975, but the confidence intervals overlapped, and long-horizon appearance changes and unseen-pose inputs remain unsolved.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)