LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
- Published
- Source
- arXiv
- Paper number
- 184
- Field
- Reasoning / Test-Time Scaling
- arXiv ID
- 2605.08083
Key points
- Test-time scaling strategies for large language models are usually developed through manual design and heuristic tuning.
- This reliance on human intuition limits the search over complex computation-allocation policies and makes them vulnerable to human cognitive bias.
- Manual methods often produce strategies that are specific to a task or model, which leads to suboptimal accuracy-cost trade-offs across broader settings.
- AutoTTS formulates test-time scaling as a controller synthesis problem, where a searcher LLM repeatedly proposes and refines policies defined in code.
- The method uses an offline replay environment to collect reasoning trajectories in advance, which allows candidate controllers to be evaluated cheaply and deterministically without live LLM calls.
- The framework introduces a beta parameterization to simplify the hyperparameter search space and provides fine-grained execution traces as feedback for robust discovery.
Paper links
External research summaries. These are not HDATF publications or measured product results.