Reasoning with Sampling: Your Base Model is Smarter Than You Think

Published
Source
arXiv
Paper number
089
Field
Reasoning / Inference
arXiv ID
2510.14901

Key points

  • It is unclear whether RL fundamentally teaches new reasoning capabilities to LLMs or simply sharpens already latent capabilities.
  • Current RL-based LLM post-training methods inherently require high compute cost, curated datasets, and external verifiers.
  • RL-post-trained models often suffer from diversity collapse, sacrificing multi-sample reasoning performance, pass@k, in favor of single-sample gains.
  • The base LLM's power distribution is defined by sampling from the original distribution raised to exponent α (α >= 1), which mathematically reweights the original distribution to favor higher-likelihood sequences.
  • The Metropolis-Hastings (MH) algorithm, a Markov chain Monte Carlo (MCMC) method, is used to enable approximate sampling from the unnormalized power distribution.
  • An iterative, blockwise autoregressive MCMC approach is introduced that progressively refines sequence segments, managing complexity and improving mixing time for long-sequence generation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)