SPADE: Self-Play in Adaptive Synthetic Executable Environments
- Published
- Source
- arXiv
- Paper number
- 953
- Field
- LLMs / NLP
- arXiv ID
- 2608.19197
Key points
- SPADE makes environment design itself a learnable component. The distribution of environments produced by the designer coevolves with the agent's capabilities.
- A hint-based regret signal is the key mechanism. A purely adversarial designer proposes unsolvable tasks, while a purely cooperative designer proposes tasks with nothing to teach. The performance gap with and without hints automatically targets environments that are both feasible and worth learning from.
- Because each environment is written as executable code with a Gym-style interface, one interface can cover both single-turn reasoning problems and multi-turn tool-use tasks.
- Grounding the designer in documents from a pretraining corpus and maintaining an accumulated memory of previously created environments were both critical to performance.
- At the 30B scale, SPADE improves over the strongest fixed-environment baseline by an average of 5.3 points across eight mathematics, science, coding, and reasoning benchmarks, by 5.7 points on BFCL v4 multi-turn, and by 13.9 points on ACEBench-Agent.
Paper links
External research summaries. These are not HDATF publications or measured product results.