VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
- Published
- Source
- arXiv
- Paper number
- 777
- Field
- Computer Vision
- arXiv ID
- 2607.27380
Key points
- It uses executable code, Blender, as a process-level thought chain to preserve physical consistency.
- The code agent programs the scene and a simulator renders a draft video, and then the video editor composes the final result realistically.
- It builds the VideoCoCo-3K dataset, which contains draft, instruction, and target pairs, for editor adaptation training.
- It improves PhyGenBench from 0.475 to 0.558 and VBench-2.0 from 52.18% to 77.88%, which are both state of the art.
- Without tuning, the draft alone improves from 0.475 to 0.506, and LoRA tuning raises it further to 0.558, which separates the contribution of each stage.
Paper links
External research summaries. These are not HDATF publications or measured product results.