QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling
- Published
- Source
- arXiv
- Paper number
- 557
- Field
- Machine Learning
- arXiv ID
- 2607.01179
Key points
- It reparameterizes autoregressive sampling as inverse-CDF sampling, converting QMC uniform numbers into LM samples.
- Each sample is marginally exact, meaning it comes from the same LM distribution, while negative correlation across the batch dimension increases coverage.
- On pass@k, it achieves the same performance with 25% to 47% fewer samples across four reasoning benchmarks.
- On GRPO RL, it reaches the same pass@1 with 50% fewer training steps and reduces the zero-variance group ratio.
- Coverage increases in the order lattice, stratified, Sobol, and i.i.d., quantifying the freedom-versus-coverage tradeoff.
- The limitation is that it was validated only on 1B to 2B models and short outputs, so extension to long CoT remains future work.
Paper links
External research summaries. These are not HDATF publications or measured product results.