Forecasting Rare Language Model Behaviors
- Published
- Source
- arXiv
- Paper number
- 042
- Field
- Alignment / Safety
- arXiv ID
- 2502.16797
Key points
- Standard LLM evaluations cover thousands of queries, but deployed models process billions, leaving a significant blind spot for rare but potentially harmful behaviors.
- Behavior that appears safe in limited tests can create substantial risk when exposed to far more queries in deployment.
- Traditional evaluation methods do not account for the fundamental scale gap or the stochastic nature of language models, which can generate different answers to the same query.
- Researchers from Anthropic, Mila, and UC Berkeley developed a Gumbel-tail method to predict the risk of undesirable LLM behavior at deployment scale using limited evaluation data.
- The method estimates each query's trigger probability by repeatedly sampling the model and measuring the fraction of responses that exhibit the target behavior.
- It uses extreme value theory, specifically the Gumbel distribution, to model and extrapolate the tail behavior of these trigger probabilities from a small set of evaluation queries.
Paper links
External research summaries. These are not HDATF publications or measured product results.