WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

Published
Source
arXiv
Paper number
983
Field
Computer Vision
arXiv ID
2608.20974

Key points

  • It replaced random-mask reconstruction with pretraining that masks future segments, aligning the video-representation learning objective with the temporal direction of autonomous-driving planning.
  • Instead of deterministic regression, it uses flow matching to generate multiple possible future representations and jointly learns scenes and vehicle trajectories in a single predictor.
  • It achieved 91.7 EPDMS on NAVSIM-v2, exceeding strong end-to-end and world-action baselines by 1.6 and 1.3 points, respectively, and also achieved the highest score, 0.4462, in HUGSIM closed-loop evaluation.
  • It presents a different design direction for autonomous-driving planning, in which visual future prediction directly supports action learning without going through language reasoning.
  • Validation focuses on nuPlan pretraining and the NAVSIM and HUGSIM benchmarks, so safe generalization to real vehicles and rare hazardous situations requires further verification.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)