SPADE: Self-Play in Adaptive Synthetic Executable Environments

Published
Source
arXiv
Paper number
953
Field
LLMs / NLP
arXiv ID
2608.19197

Key points

  • SPADE makes environment design itself a learnable component. The distribution of environments produced by the designer coevolves with the agent's capabilities.
  • A hint-based regret signal is the key mechanism. A purely adversarial designer proposes unsolvable tasks, while a purely cooperative designer proposes tasks with nothing to teach. The performance gap with and without hints automatically targets environments that are both feasible and worth learning from.
  • Because each environment is written as executable code with a Gym-style interface, one interface can cover both single-turn reasoning problems and multi-turn tool-use tasks.
  • Grounding the designer in documents from a pretraining corpus and maintaining an accumulated memory of previously created environments were both critical to performance.
  • At the 30B scale, SPADE improves over the strongest fixed-environment baseline by an average of 5.3 points across eight mathematics, science, coding, and reasoning benchmarks, by 5.7 points on BFCL v4 multi-turn, and by 13.9 points on ACEBench-Agent.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)