UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models

Published
Source
arXiv
Paper number
999
Field
Robotics
arXiv ID
2608.22869

Key points

  • An event classifier detects task-critical moments in the backbone's latent space and updates text and keyframe memory.
  • It integrates memory and low-level control into a single backbone, π0.5, without a separate VLM for memory management.
  • Keyframe caching preserves single-frame-level speed at approximately 90ms per frame, making it 6× faster than hierarchical memory methods and allowing scaling to a dual-arm configuration with 4 cameras on a standard workstation GPU without a speed penalty.
  • It achieved an average of 93.4% across 5 simulated tasks, versus 68.2% for the fixed-interval-frame baseline, and 80.0% across 4 real-robot tasks, versus 43.5% for hierarchical MemER.
  • However, training-compute limits restrict it to retaining only three past milestones, and it discards the oldest frame when this limit is exceeded.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)