Marionette: Predicting World States, Rendering Geometry, Painting Appearance
- Published
- Source
- arXiv
- Paper number
- 914
- Field
- Computer Vision
- arXiv ID
- 2608.14530
Key points
- ActionGPT selects actions and PoseGPT creates poses, while a fixed graphics pipeline computes skeletons and occlusion relationships and then a video generation model paints in the appearance.
- The action model was trained on 1,395 segments at 20 frames per second, and the current action range is limited to one kind of monster.
- On 48 test segments, forcing incorrect actions increased joint error from 0.272 meters to 0.357 meters, a 31 percent rise, which shows that action tokens control pose.
- Applying a ground-collision rule reduced the share of frames where joints passed through the ground from 0.337 to 0.114, and a distance-limit rule lowered the distance between characters from 21.2 meters to 5.1 meters.
- The video quality metric FVD was 831, lower than the pixel-based model's 975, but the confidence intervals overlapped, and long-horizon appearance changes and unseen-pose inputs remain unsolved.
Paper links
External research summaries. These are not HDATF publications or measured product results.