FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelines

Published
Source
arXiv
Paper number
464
Field
Software Engineering
arXiv ID
2606.19605

Key points

  • The Claude Code-based optimization loop follows a pipeline of evaluation, intermediate-step analysis for failure attribution, scoped change proposal, safety review, and verification.
  • The prompt-first strategy prioritizes prompt edits and escalates to chain-structure or parameter changes only when attribution identifies a structural bottleneck.
  • It represents the pipeline as a stateful graph built with LangGraph, and it ensures reproducibility through tenant-isolated workspaces.
  • Across 6 benchmarks and 3 models, meaning 18 comparisons, GEPA wins 15 times with an average gain of 14.1 percentage points, and it wins all 6 structure-change comparisons on HoVer and IFBench with an average gain of 33.8 percentage points.
  • On the security task CTIBench-RCM, it improves GPT-5 by 4.0 points, Foundation-Sec-8B-Instruct by 7.1 points, and Foundation-Sec-8B-Reasoning by 2.0 points.
  • A step-attribution subagent separates prompt-editable issues from structural bottlenecks, which prevents unnecessary structural changes.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)