WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving
- Published
- Source
- arXiv
- Paper number
- 983
- Field
- Computer Vision
- arXiv ID
- 2608.20974
Key points
- It replaced random-mask reconstruction with pretraining that masks future segments, aligning the video-representation learning objective with the temporal direction of autonomous-driving planning.
- Instead of deterministic regression, it uses flow matching to generate multiple possible future representations and jointly learns scenes and vehicle trajectories in a single predictor.
- It achieved 91.7 EPDMS on NAVSIM-v2, exceeding strong end-to-end and world-action baselines by 1.6 and 1.3 points, respectively, and also achieved the highest score, 0.4462, in HUGSIM closed-loop evaluation.
- It presents a different design direction for autonomous-driving planning, in which visual future prediction directly supports action learning without going through language reasoning.
- Validation focuses on nuPlan pretraining and the NAVSIM and HUGSIM benchmarks, so safe generalization to real vehicles and rare hazardous situations requires further verification.
Paper links
External research summaries. These are not HDATF publications or measured product results.