AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

Published
Source
arXiv
Paper number
1137
Field
Robotics
arXiv ID
2609.34085

Key points

  • AD-E2E-JEPA is an autonomous-driving world model that predicts future latent visual representations and selects trajectories close to a target image.
  • A learnable projector with SIGReg regularization reduces the number of planning patches to one-sixteenth and the feature dimension to one-quarter.
  • In an A100 experiment evaluating 256 candidate trajectories, planning takes 0.8 seconds per scene, compared with 91.8 seconds for DINO-WM and 101.0 seconds for JEPA-WM.
  • Using the pretrained projector in a separate imitation-learning setup raises EPDMS from 80.2 with random initialization to 85.4, demonstrating the value of reusing its representations.
  • Zero-shot evaluation supplies ground-truth future images as goals and does not directly optimize safety, so the results do not establish safety on real roads.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)