Geo-Align: Video Generation Alignment via Metric Geometry Reward

Published
Source
arXiv
Paper number
229
Field
Computer Vision
arXiv ID
2605.23903

Key points

  • Standard training losses compare pixels or abstract features, and they do not understand physical units such as meters or degrees.
  • By using a 3D evaluator, the model learns to interpret camera motion in physical units, which resolves the scale drift that has long been a problem in video synthesis.
  • This RL framework removes the need for hard-to-obtain synchronized multiview video, allowing the model to learn from a wide range of real-world distributions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)