GraphVid: Interactive Graph-Controllable Video Generation

Published
Source
arXiv
Paper number
718
Field
Computer Vision
arXiv ID
2607.21580

Key points

  • It proposes the first framework for generating video conditioned on a scene graph, where objects are nodes and interactions are edges.
  • The Edge-Aware Graph Reasoning module performs message passing while reflecting the directional properties of edges.
  • It builds a new GraphVid-Bench dataset with about 27K interaction-centered videos.
  • It reduces FID by 39.9%, reduces FVD by 37.6%, and improves PSNR from 9.87 to 15.98 compared with Motion-I2V.
  • With only 0.6B trainable parameters, it is far more efficient than existing methods that use billions of parameters.
  • Because the graph representation is backbone-independent, it works with both LTX (2B) and Wan 2.2 (5B).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)