SPIRAL: Learning to Search and Aggregate

Published
Source
arXiv
Paper number
485
Field
AI / General
arXiv ID
2606.23595

Key points

  • It is the first framework to jointly optimize sequential, parallel, and aggregation compute within a single reinforcement learning pipeline.
  • Set reinforcement learning trains parallel traces so that they are collectively useful for aggregation.
  • It achieves up to 11x scaling efficiency in pass@k versus GRPO and 15 percent better performance when parallel and aggregation scale together.
  • It points out that existing training paradigms optimize only sequential reasoning, leaving a gap with test-time scaffolds.
  • Experiments with Qwen3-4B-Instruct validate the method, and the authors identify 8B-plus parameter scaling as future work.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)