Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
- Published
- Source
- arXiv
- Paper number
- 1058
- Field
- Computer Vision
- arXiv ID
- 2609.00111
Key points
- Without changing the architecture of the pretrained VLM (Qwen3.5-4B), it added just two external modules, a BEV perception head and a Planning Expert, to integrate 3D perception, question answering, and path planning.
- It used the BEV head as a '3D probe' to make it possible to test how well the model understands actual 3D scene structure, beyond simply producing fluent text answers.
- Staged training that mixes driving data with general-purpose vision-language data preserved both driving and general capabilities, mitigating catastrophic forgetting (losing old capabilities while learning new ones).
- It achieved 43.95 mAP for 3D perception and 60.99 mIoU for mapping on nuScenes, a NAVSIM driving score (PDMS) of 90.7, and a Waymo E2E evaluator score of 7.91.
- In line with the trend toward vehicle cockpits (multimedia) and driving functions sharing a single computing platform, it aims for a model that handles them together rather than a model dedicated only to driving.
Paper links
External research summaries. These are not HDATF publications or measured product results.