GeoWAM: Visual Geometry World Action Models for Autonomous Driving
- Published
- Source
- arXiv
- Paper number
- 995
- Field
- Computer Vision
- arXiv ID
- 2608.23486
Key points
- It proposed world-model pretraining that predicts future scene geometry, point maps, rather than future observations in pixels.
- However, the paper notes that this advantage does not hold at every horizon: at a 1-second lookahead, the video-based Epona+DVGT achieved higher threshold accuracy.
- For future-geometry prediction, it surpassed the strongest baseline, improving average AbsRel from 0.274 to 0.257 and δ<1.25 from 0.655 to 0.754.
- It achieved the highest EPDMS among the compared methods in NAVSIM v2 closed-loop evaluation.
- Because it does not need to generate future video separately, it transfers predicted geometric knowledge directly to driving planning without a video-generation stage.
Paper links
External research summaries. These are not HDATF publications or measured product results.