MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models

Published
Source
arXiv
Paper number
385
Field
Robotics
arXiv ID
2606.09827

Key points

  • We introduce a cognition-inspired full-temporal model that combines working memory, episodic memory, and future imagination through a world model.
  • The Perceptual-Cognitive Memory Bank stores low-level details and high-level semantics together, and updates memory when redundancy is detected.
  • A latent-space world model imagines future states through partial denoising only, without requiring pixel-level prediction.
  • We achieve 98.4% on Libero and 74.0% on SimplerEnv, with real-robot gains of +9% for general capability, +26% for memory, and +28% for imagination.
  • It is especially strong on long-horizon tasks, with 4.29 on Calvin and 44.4% on Mikasa-Robo.
  • It is ready for real-time deployment, with 66.4 Hz inference on an RTX 4090.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)