You Don't Need to Run Every Eval

Published
Source
arXiv
Paper number
493
Field
Machine Learning
arXiv ID
2606.24020

Key points

  • The authors build a public score matrix over 84 frontier models and 133 benchmarks, with 2,604 cells and 23.3% coverage, and show that it is effectively rank-2.
  • They design BENCHPRESS, a rank-2 ALS matrix completion method in logit space, to recover hidden scores within 4.6 points.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)