VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Published
Source
arXiv
Paper number
777
Field
Computer Vision
arXiv ID
2607.27380

Key points

  • It uses executable code, Blender, as a process-level thought chain to preserve physical consistency.
  • The code agent programs the scene and a simulator renders a draft video, and then the video editor composes the final result realistically.
  • It builds the VideoCoCo-3K dataset, which contains draft, instruction, and target pairs, for editor adaptation training.
  • It improves PhyGenBench from 0.475 to 0.558 and VBench-2.0 from 52.18% to 77.88%, which are both state of the art.
  • Without tuning, the draft alone improves from 0.475 to 0.506, and LoRA tuning raises it further to 0.558, which separates the contribution of each stage.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)