TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
- Published
- Source
- arXiv
- Paper number
- 757
- Field
- Computer Vision
- arXiv ID
- 2607.27205
Key points
- It proposes a direct vision-plus-language-to-action approach by removing the LLM-centered architecture.
- With 0.2B parameters, the model is only 6% the size of π0.5, which has about 3.4B parameters, yet matches or exceeds its performance with a 97.7% success rate on LIBERO.
- It achieves 32 Hz inference, 31.2 ms, on a consumer RTX 4090 and uses less than 1 GB of VRAM, making edge deployment possible.
- It outperforms π0.5 consistently on four real-robot tasks with an AgileX Piper.
- Removing language drops the success rate to 70.8%, proving that language conditioning is still important even at the execution level.
Paper links
External research summaries. These are not HDATF publications or measured product results.