GraphVid: Interactive Graph-Controllable Video Generation
- Published
- Source
- arXiv
- Paper number
- 718
- Field
- Computer Vision
- arXiv ID
- 2607.21580
Key points
- It proposes the first framework for generating video conditioned on a scene graph, where objects are nodes and interactions are edges.
- The Edge-Aware Graph Reasoning module performs message passing while reflecting the directional properties of edges.
- It builds a new GraphVid-Bench dataset with about 27K interaction-centered videos.
- It reduces FID by 39.9%, reduces FVD by 37.6%, and improves PSNR from 9.87 to 15.98 compared with Motion-I2V.
- With only 0.6B trainable parameters, it is far more efficient than existing methods that use billions of parameters.
- Because the graph representation is backbone-independent, it works with both LTX (2B) and Wan 2.2 (5B).
Paper links
External research summaries. These are not HDATF publications or measured product results.