SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
- Published
- Source
- arXiv
- Paper number
- 853
- Field
- Computer Vision
- arXiv ID
- 2608.07468
Key points
- We propose SimWAM, a structure that uses video generation only as a training signal and removes it at inference time, so real-time video generation is unnecessary and inference is faster.
- Because the video expert and action expert do not share parameters, it is easy to replace the video model with a better one or scale the action model.
- On NAVSIM, it achieves 91.5 PDMS. It performs better than prior world-action-model-based planners while keeping inference latency much lower.
- We additionally apply reinforcement learning (GRPO) to optimize complex driving rewards beyond trajectory imitation.
- Zero-shot transfer to the nuScenes dataset also confirms generality.
Paper links
External research summaries. These are not HDATF publications or measured product results.