Visual prompt engineering for video models

Published
Source
arXiv
Paper number
755
Field
Computer Vision
arXiv ID
2607.25537

Key points

  • It systematically proposes and validates visual prompt engineering, or VIPE, for video models.
  • On Veo 3.1, converting sketches into photorealistic scenes raises physical reasoning accuracy from 41.3 percent to 59.3 percent, a gain of 18 points.
  • It uncovers a realism bias, meaning video models reason much better on realistic scenes, so abstract benchmarks underestimate their ability.
  • VIPE is often more effective than test-time scaling methods such as text-prompt optimization or self-consistency.
  • Stepwise experiments confirm that scene consistency improves whenever realism is increased by one level.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)