What Matters for Latent Actions in Robot Learning

Published
Source
arXiv
Paper number
964
Field
Robotics
arXiv ID
2608.19613

Key points

  • It compared 41 latent action model (LAM) design choices under unified conditions across 3 simulation benchmarks + a real Franka robot.
  • A latent-action dimension of 32 was optimal for both single-arm and dual-arm robots, and additional regularization was unnecessary when pretraining regularization was done properly.
  • Fine-tuning the VLM backbone with latent actions beforehand raised success on 4 real-robot tasks from 64.75%→79.25% (+14.5 percentage points).
  • The LAM-tuned model reached 85% success after just 10k steps, surpassing the baseline's performance at 40k steps (76.25%) and demonstrating data efficiency.
  • The simple reconstruction (FDM) metric was a more reliable predictor of latent-action quality than the conventionally used MLP-probe metric.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)