Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
- Published
- Source
- arXiv
- Paper number
- 1187
- Field
- LLMs / NLP
- arXiv ID
- 2610.10478
Key points
- They built a framework that replays successful trajectories to find the decisive step T* where tests first flip from failing to passing, and probes the base model directly at that point.
- The three probes (Decisive-Action BPB, Patch MCQ, prefix-conditioned pass@K) showed strong rank agreement with post-trained SWE-bench Verified pass@1, at Spearman rho 0.964, 0.903, and 0.988 respectively.
- Existing non-agentic coding benchmarks (HumanEval and others) were unstable, with correlations from -0.39 to 0.83, confirming they are ill-suited for base model selection.
- These probes require only a benchmark's successful trajectories and its verifier, so future agentic benchmarks can be repurposed as base-model evaluations.
Paper links
External research summaries. These are not HDATF publications or measured product results.