ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation

Published
Source
arXiv
Paper number
802
Field
Robotics
arXiv ID
2608.02326

Key points

  • It divides execution state into two parts. The backward-looking part is progress, combining a recurrent task state and sparsely logged event memory. The forward-looking part is the unexecuted tail of the most recent prediction. This state is used only as revisable prior information, not as a fixed plan.
  • On RMBench, the 1.2B model scores 62.8 on average. Among methods that cover all tasks, the strongest Mem-0 uses 8B plus an extra 2B and gets 52.8.
  • The largest gain is on Put Back Block. This task requires putting back a block whose original location is no longer visible, and it scores 96, compared with 90 for Mem-0 and 50 for MemoryVLA.
  • The ablation results are asymmetric. Keeping only progress gives 11.2, keeping only the action tail gives 3.0, and removing both gives 1.6, while the full model gets 62.8. A half-state is not very different from no state.
  • A control that mimicked prior methods was insufficient. A fixed-length observation history combined with temporal ensembling, which smooths overlapping predictions from the back, reached 35.6, 27.2 points below the full model.
  • It shows that smoothness alone is not the cause. The linear extrapolation control has lower error than the full model on all four boundary metrics for Put Back Block and on three of four for Swap Blocks. Yet its success rates are 0% and 2%, while the full model reaches 96% and 74%.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)