Leadership

Can changing AI benchmarks support a single progress forecast?

Published
Source
arXiv
Source type
Research

Summary

An audit examines 62 selected systems and 12 benchmarks, identifying changing tests and concentrated sources. It is a limited sample, not a refutation of all forecasting.

Why it matters for our work

Our model comparisons need consistent tasks, budgets and criteria. A higher score and better workplace outcomes are separate claims.

Translated from the Korean original. Summaries may be translated and edited. Commentary reflects our perspective; forecasts remain the source’s views.

Read original (opens in a new tab)