FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference

Published
Source
arXiv
Paper number
1032
Field
Robotics
arXiv ID
2608.27384

Key points

  • It began by profiling π0.5 and confirming that action decoding consumes 75% of per-step inference time.
  • It created a streaming architecture with a buffer of action chunks at staggered noise levels, denoising all of them by one step in a single forward pass.
  • Per-step action-decoding latency was reduced by a factor of up to 20, and the full pipeline became up to 2.43 times faster.
  • Causal attention between chunks eliminated temporal misalignment in asynchronous execution without a separate future-state predictor.
  • On LIBERO, RoboTwin 2.0, and real Franka experiments, it maintained or improved π0.5 performance while achieving control at 30Hz or higher on a single GPU.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)