T$^2$Mem: Learning Test-Time Memory for Robotics

Published
Source
arXiv
Paper number
1140
Field
Robotics
arXiv ID
2609.36720

Key points

  • T²Mem adds internal memory to a pretrained vision-language-action model so that actions can use information no longer visible in the current observation.
  • An observation-grounded interface stores history in fast weights, while the memory module and action policy are trained in alternating stages.
  • Across 16 RoboMME tasks, average success increases from 17.93% for the memory-free baseline to 56.83%.
  • The method uses memory without an external reasoning model, and controlled RTX A5000 profiling shows approximately three times faster foreground computation than FrameSamp+Modul.
  • Reasoning about swapped object locations and precise insertion remain limitations, and the timing measurements exclude communication and environment execution.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)