Visual prompt engineering for video models
- Published
- Source
- arXiv
- Paper number
- 755
- Field
- Computer Vision
- arXiv ID
- 2607.25537
Key points
- It systematically proposes and validates visual prompt engineering, or VIPE, for video models.
- On Veo 3.1, converting sketches into photorealistic scenes raises physical reasoning accuracy from 41.3 percent to 59.3 percent, a gain of 18 points.
- It uncovers a realism bias, meaning video models reason much better on realistic scenes, so abstract benchmarks underestimate their ability.
- VIPE is often more effective than test-time scaling methods such as text-prompt optimization or self-consistency.
- Stepwise experiments confirm that scene consistency improves whenever realism is increased by one level.
Paper links
External research summaries. These are not HDATF publications or measured product results.