LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

Published
Source
arXiv
Paper number
184
Field
Reasoning / Test-Time Scaling
arXiv ID
2605.08083

Key points

  • Test-time scaling strategies for large language models are usually developed through manual design and heuristic tuning.
  • This reliance on human intuition limits the search over complex computation-allocation policies and makes them vulnerable to human cognitive bias.
  • Manual methods often produce strategies that are specific to a task or model, which leads to suboptimal accuracy-cost trade-offs across broader settings.
  • AutoTTS formulates test-time scaling as a controller synthesis problem, where a searcher LLM repeatedly proposes and refines policies defined in code.
  • The method uses an offline replay environment to collect reasoning trajectories in advance, which allows candidate controllers to be evaluated cheaply and deterministically without live LLM calls.
  • The framework introduces a beta parameterization to simplify the hyperparameter search space and provides fine-grained execution traces as feedback for robust discovery.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)