TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Published
Source
arXiv
Paper number
757
Field
Computer Vision
arXiv ID
2607.27205

Key points

  • It proposes a direct vision-plus-language-to-action approach by removing the LLM-centered architecture.
  • With 0.2B parameters, the model is only 6% the size of π0.5, which has about 3.4B parameters, yet matches or exceeds its performance with a 97.7% success rate on LIBERO.
  • It achieves 32 Hz inference, 31.2 ms, on a consumer RTX 4090 and uses less than 1 GB of VRAM, making edge deployment possible.
  • It outperforms π0.5 consistently on four real-robot tasks with an AgileX Piper.
  • Removing language drops the success rate to 70.8%, proving that language conditioning is still important even at the execution level.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)