Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Published
Source
arXiv
Paper number
825
Field
Machine Learning
arXiv ID
2608.04001

Key points

  • It formalizes test-time scaling into three structural regimes: single-path, leaf-level, and prefix-level.
  • It emphasizes that the target of evaluation is the entire inference system, including prompts, decoders, verifiers, and stopping rules, not just the model weights.
  • It shows the risk of a poor selection criterion, with mean log probability accuracy dropping from 75.6% to 65.8% as the sample count increases.
  • It releases more than 2 billion inference traces, covering mathematics, science, and competitions, to provide a foundation for reproducible research.
  • It organizes the open-weight reasoning-model ecosystem by training method and interface control criteria.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)