LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

Published
Source
arXiv
Paper number
814
Field
Robotics
arXiv ID
2608.03701

Key points

  • It trains a lightweight world-action model that unifies future-state prediction and action generation in a single stream on one 24 GB GPU.
  • It proposes the Visual Transition Token, or VTT, which specifies a task only through visual feature direction without language.
  • On 50 RoboTwin 2.0 tasks, it achieves a 90.48 percent success rate and shows strong efficiency relative to parameter count.
  • It also performs competitively with much larger models in LIBERO and real-robot experiments.
  • Using only a DINOV3 visual backbone, it shows that effective robot control is possible without dependence on a language model.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)