Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
- Published
- Source
- arXiv
- Paper number
- 122
- Field
- Reasoning / Probabilistic
- arXiv ID
- 2503.17523
Key points
- Off-the-shelf large language models, or LLMs, lack the ability to form and update probabilistic beliefs effectively in uncertain, interactive environments.
- Current LLMs perform poorly at implicitly inferring latent states, such as user preferences, from observed behavior across multiple interactions.
- Their recommendation accuracy often plateaus early, which suggests a fundamental limitation in dynamically updating beliefs from new information.
- The authors developed Bayesian teaching, a supervised fine-tuning strategy that trains LLMs on data generated by a normative Bayesian assistant.
- The Bayesian assistant maintains and updates a probability distribution over user preferences with Bayes' rule and produces probabilistically optimal recommendations.
- This approach was compared with oracle teaching, which fine-tunes the LLM on consistently correct answers from an assistant with full knowledge.
Paper links
External research summaries. These are not HDATF publications or measured product results.