Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models

Published
Source
arXiv
Paper number
122
Field
Reasoning / Probabilistic
arXiv ID
2503.17523

Key points

  • Off-the-shelf large language models, or LLMs, lack the ability to form and update probabilistic beliefs effectively in uncertain, interactive environments.
  • Current LLMs perform poorly at implicitly inferring latent states, such as user preferences, from observed behavior across multiple interactions.
  • Their recommendation accuracy often plateaus early, which suggests a fundamental limitation in dynamically updating beliefs from new information.
  • The authors developed Bayesian teaching, a supervised fine-tuning strategy that trains LLMs on data generated by a normative Bayesian assistant.
  • The Bayesian assistant maintains and updates a probability distribution over user preferences with Bayes' rule and produces probabilistically optimal recommendations.
  • This approach was compared with oracle teaching, which fine-tunes the LLM on consistently correct answers from an assistant with full knowledge.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)