Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

Published
Source
arXiv
Paper number
1187
Field
LLMs / NLP
arXiv ID
2610.10478

Key points

  • They built a framework that replays successful trajectories to find the decisive step T* where tests first flip from failing to passing, and probes the base model directly at that point.
  • The three probes (Decisive-Action BPB, Patch MCQ, prefix-conditioned pass@K) showed strong rank agreement with post-trained SWE-bench Verified pass@1, at Spearman rho 0.964, 0.903, and 0.988 respectively.
  • Existing non-agentic coding benchmarks (HumanEval and others) were unstable, with correlations from -0.39 to 0.83, confirming they are ill-suited for base model selection.
  • These probes require only a benchmark's successful trajectories and its verifier, so future agentic benchmarks can be repurposed as base-model evaluations.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)