FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation
- Published
- Source
- arXiv
- Paper number
- 667
- Field
- Robotics
- arXiv ID
- 2607.18231
Key points
- It identifies a problem with contact-rich repetitive tasks: because the screen changes little, visual memory alone cannot track progress.
- It first trains a variational autoencoder on force and torque time series with reconstruction as the objective, then freezes the encoder and uses the latent representation as memory tokens.
- Using force memory alone makes pre-contact motion unstable, so it supplements it with a short record of joint states.
- It achieves an average success rate of 83.3% over three tasks, outperforming pi-zero without memory at 27.8%, TA-VLA with a short force window at 22.2%, and visual-memory pi-MEM at 53.7%.
- Inference latency increases by only 3.3 ms, from 60.7 ms to 64.0 ms, whereas visual memory grows from 100 ms to 190 ms depending on the number of frames.
- Eight memory tokens were optimal, while increasing the count to 32 exceeded the token budget of the pretrained action module and hurt performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.