PACE: A Proxy for Agentic Capability Evaluation

Published
Source
arXiv
Paper number
563
Field
AI / General
arXiv ID
2607.02032

Key points

  • The core idea is to predict expensive agentic evaluation with cheap single-turn benchmark scores, and 100 proxy instances are enough.
  • It combines SVD leverage, which measures geometric information content, with rank correlation, which measures target relevance, using target-specific weights.
  • Leave-one-out cross-validation gives a MAE of 3.80 percent, Spearman of 0.81, and pairwise accuracy of 84.4 percent across 14 models and 4 targets.
  • It is about 100 times cheaper than random sampling with the same quality, which makes agentic evaluation practical in everyday model development.
  • The selected instances are interpretable because their capability distribution clearly reflects the abilities required by each benchmark, such as IF plus verification for GAIA and planning plus testing for SWT.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)