SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

Published
Source
arXiv
Paper number
670
Field
LLMs / NLP
arXiv ID
2607.18213

Key points

  • The core claim is that there is no need to recreate pruning signals outside the model, because the backbone already encodes that judgment in its representations when it reads tool outputs.
  • As evidence, the authors froze Qwen3-Coder-Next and ran a logistic-regression probe over line-averaged hidden states, reaching an AUC of 0.83 and a best F1 of 0.63 on held-out data.
  • The method uses the final-layer hidden states from the prefill stage, and a small head reads them to score tokens and aggregate them into line-level keep or delete decisions.
  • Two design choices improve performance: a length-aware embedding that includes the number of output lines, and a balanced focal loss that keeps the keep and delete ratios aligned per sample.
  • The result is 39% token savings on SWE-QA-Pro and 30% on Oolong across two open-weight backbones and four multi-turn benchmarks, and with MiMo-V2-Flash it also improves Oolong accuracy by 2.2 points and SWE-Bench Verified solve rate by 3.8 percentage points.
  • The extra cost is limited. In a replay setup with the head attached inside the inference engine, it adds only 15.0% wall-clock time over total generation time and requires no separate model call.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)