LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
- Published
- Source
- arXiv
- Paper number
- 814
- Field
- Robotics
- arXiv ID
- 2608.03701
Key points
- It trains a lightweight world-action model that unifies future-state prediction and action generation in a single stream on one 24 GB GPU.
- It proposes the Visual Transition Token, or VTT, which specifies a task only through visual feature direction without language.
- On 50 RoboTwin 2.0 tasks, it achieves a 90.48 percent success rate and shows strong efficiency relative to parameter count.
- It also performs competitively with much larger models in LIBERO and real-robot experiments.
- Using only a DINOV3 visual backbone, it shows that effective robot control is possible without dependence on a language model.
Paper links
External research summaries. These are not HDATF publications or measured product results.