SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

Published
Source
arXiv
Paper number
853
Field
Computer Vision
arXiv ID
2608.07468

Key points

  • We propose SimWAM, a structure that uses video generation only as a training signal and removes it at inference time, so real-time video generation is unnecessary and inference is faster.
  • Because the video expert and action expert do not share parameters, it is easy to replace the video model with a better one or scale the action model.
  • On NAVSIM, it achieves 91.5 PDMS. It performs better than prior world-action-model-based planners while keeping inference latency much lower.
  • We additionally apply reinforcement learning (GRPO) to optimize complex driving rewards beyond trajectory imitation.
  • Zero-shot transfer to the nuScenes dataset also confirms generality.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)