FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference
- Published
- Source
- arXiv
- Paper number
- 1032
- Field
- Robotics
- arXiv ID
- 2608.27384
Key points
- It began by profiling π0.5 and confirming that action decoding consumes 75% of per-step inference time.
- It created a streaming architecture with a buffer of action chunks at staggered noise levels, denoising all of them by one step in a single forward pass.
- Per-step action-decoding latency was reduced by a factor of up to 20, and the full pipeline became up to 2.43 times faster.
- Causal attention between chunks eliminated temporal misalignment in asynchronous execution without a separate future-state predictor.
- On LIBERO, RoboTwin 2.0, and real Franka experiments, it maintained or improved π0.5 performance while achieving control at 30Hz or higher on a single GPU.
Paper links
External research summaries. These are not HDATF publications or measured product results.