Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

Published
Source
arXiv
Paper number
978
Field
LLMs / NLP
arXiv ID
2608.20169

Key points

  • Existing harness optimization is expensive because it evaluates all validation tasks at every iteration. This work was the first to formalize the problem of evolving validation-task selection itself together with the harness.
  • Based on the observation that tasks on which candidate harnesses differ in success or failure (tasks with high success-rate variance) are most informative for distinguishing candidates, it used variance-weighted sampling to focus evaluation near the boundary of agent capabilities.
  • It corrects for sampling probabilities to estimate the full score from partial evaluation, allowing fair comparisons across iterations even when a different subset is evaluated each time.
  • On Terminal-Bench 2.1, it achieved final performance similar to full search with 20% of the evaluation budget, substantially reducing token use (for example, from 741M to 246M) and search time (from 38.0 hours to 20.5 hours).
  • At the same 20% budget, final performance was 3.3 points higher than with random sampling, confirming that 'what is evaluated' determines efficiency.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)