Reasoning with Sampling: Your Base Model is Smarter Than You Think
- Published
- Source
- arXiv
- Paper number
- 089
- Field
- Reasoning / Inference
- arXiv ID
- 2510.14901
Key points
- It is unclear whether RL fundamentally teaches new reasoning capabilities to LLMs or simply sharpens already latent capabilities.
- Current RL-based LLM post-training methods inherently require high compute cost, curated datasets, and external verifiers.
- RL-post-trained models often suffer from diversity collapse, sacrificing multi-sample reasoning performance, pass@k, in favor of single-sample gains.
- The base LLM's power distribution is defined by sampling from the original distribution raised to exponent α (α >= 1), which mathematically reweights the original distribution to favor higher-likelihood sequences.
- The Metropolis-Hastings (MH) algorithm, a Markov chain Monte Carlo (MCMC) method, is used to enable approximate sampling from the unnormalized power distribution.
- An iterative, blockwise autoregressive MCMC approach is introduced that progressively refines sequence segments, managing complexity and improving mixing time for long-sequence generation.
Paper links
External research summaries. These are not HDATF publications or measured product results.