Leadership
Can changing AI benchmarks support a single progress forecast?
- Published
- Source
- arXiv
- Source type
- Research
Summary
An audit examines 62 selected systems and 12 benchmarks, identifying changing tests and concentrated sources. It is a limited sample, not a refutation of all forecasting.
Why it matters for our work
Our model comparisons need consistent tasks, budgets and criteria. A higher score and better workplace outcomes are separate claims.
Translated from the Korean original. Summaries may be translated and edited. Commentary reflects our perspective; forecasts remain the source’s views.