PACE: A Proxy for Agentic Capability Evaluation
- Published
- Source
- arXiv
- Paper number
- 563
- Field
- AI / General
- arXiv ID
- 2607.02032
Key points
- The core idea is to predict expensive agentic evaluation with cheap single-turn benchmark scores, and 100 proxy instances are enough.
- It combines SVD leverage, which measures geometric information content, with rank correlation, which measures target relevance, using target-specific weights.
- Leave-one-out cross-validation gives a MAE of 3.80 percent, Spearman of 0.81, and pairwise accuracy of 84.4 percent across 14 models and 4 targets.
- It is about 100 times cheaper than random sampling with the same quality, which makes agentic evaluation practical in everyday model development.
- The selected instances are interpretable because their capability distribution clearly reflects the abilities required by each benchmark, such as IF plus verification for GAIA and planning plus testing for SWT.
Paper links
External research summaries. These are not HDATF publications or measured product results.