SPIRAL: Learning to Search and Aggregate
- Published
- Source
- arXiv
- Paper number
- 485
- Field
- AI / General
- arXiv ID
- 2606.23595
Key points
- It is the first framework to jointly optimize sequential, parallel, and aggregation compute within a single reinforcement learning pipeline.
- Set reinforcement learning trains parallel traces so that they are collectively useful for aggregation.
- It achieves up to 11x scaling efficiency in pass@k versus GRPO and 15 percent better performance when parallel and aggregation scale together.
- It points out that existing training paradigms optimize only sequential reasoning, leaving a gap with test-time scaffolds.
- Experiments with Qwen3-4B-Instruct validate the method, and the authors identify 8B-plus parameter scaling as future work.
Paper links
External research summaries. These are not HDATF publications or measured product results.