Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
- Published
- Source
- arXiv
- Paper number
- 825
- Field
- Machine Learning
- arXiv ID
- 2608.04001
Key points
- It formalizes test-time scaling into three structural regimes: single-path, leaf-level, and prefix-level.
- It emphasizes that the target of evaluation is the entire inference system, including prompts, decoders, verifiers, and stopping rules, not just the model weights.
- It shows the risk of a poor selection criterion, with mean log probability accuracy dropping from 75.6% to 65.8% as the sample count increases.
- It releases more than 2 billion inference traces, covering mathematics, science, and competitions, to provide a foundation for reproducible research.
- It organizes the open-weight reasoning-model ecosystem by training method and interface control criteria.
Paper links
External research summaries. These are not HDATF publications or measured product results.