Geo-Align: Video Generation Alignment via Metric Geometry Reward
- Published
- Source
- arXiv
- Paper number
- 229
- Field
- Computer Vision
- arXiv ID
- 2605.23903
Key points
- Standard training losses compare pixels or abstract features, and they do not understand physical units such as meters or degrees.
- By using a 3D evaluator, the model learns to interpret camera motion in physical units, which resolves the scale drift that has long been a problem in video synthesis.
- This RL framework removes the need for hard-to-obtain synchronized multiview video, allowing the model to learn from a wide range of real-world distributions.
Paper links
External research summaries. These are not HDATF publications or measured product results.