GeoWAM: Visual Geometry World Action Models for Autonomous Driving

Published
Source
arXiv
Paper number
995
Field
Computer Vision
arXiv ID
2608.23486

Key points

  • It proposed world-model pretraining that predicts future scene geometry, point maps, rather than future observations in pixels.
  • However, the paper notes that this advantage does not hold at every horizon: at a 1-second lookahead, the video-based Epona+DVGT achieved higher threshold accuracy.
  • For future-geometry prediction, it surpassed the strongest baseline, improving average AbsRel from 0.274 to 0.257 and δ<1.25 from 0.655 to 0.754.
  • It achieved the highest EPDMS among the compared methods in NAVSIM v2 closed-loop evaluation.
  • Because it does not need to generate future video separately, it transfers predicted geometric knowledge directly to driving planning without a video-generation stage.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)