QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling

Published
Source
arXiv
Paper number
557
Field
Machine Learning
arXiv ID
2607.01179

Key points

  • It reparameterizes autoregressive sampling as inverse-CDF sampling, converting QMC uniform numbers into LM samples.
  • Each sample is marginally exact, meaning it comes from the same LM distribution, while negative correlation across the batch dimension increases coverage.
  • On pass@k, it achieves the same performance with 25% to 47% fewer samples across four reasoning benchmarks.
  • On GRPO RL, it reaches the same pass@1 with 50% fewer training steps and reduces the zero-variance group ratio.
  • Coverage increases in the order lattice, stratified, Sobol, and i.i.d., quantifying the freedom-versus-coverage tradeoff.
  • The limitation is that it was validated only on 1B to 2B models and short outputs, so extension to long CoT remains future work.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)