What is Missing from AI Post-Training AI: An Empirical Analysis
- Published
- Source
- arXiv
- Paper number
- 957
- Field
- AI / General
- arXiv ID
- 2608.19072
Key points
- The paper separates execution-level capability, meaning local refinement within a chosen strategy, from strategy-level capability, meaning revision of high-level decisions as evidence accumulates, and identifies the latter as the bottleneck.
- The training strategy becomes fixed at the outset, and the entire remaining budget is spent on local adjustments. An agent can spend more than 10 hours iterating efficiently within the wrong strategy.
- Strategy fixation reflects the agent's prior knowledge rather than the task itself. The same agent chooses similar strategies across different tasks, while different agents choose different strategies for the same task.
- An experience scaffold combining an experiment log, a skill library, and an evaluator agent improves execution by 12.6 points on GSM8K and 40.8 points on HumanEval, but does not change the selected strategy.
- Human corrections can redirect the strategy only during a brief window before training begins; once training starts, the agent returns to local adjustment loops. Increasing inference compute by two to eight times yields almost no benefit on the hardest task.
Paper links
External research summaries. These are not HDATF publications or measured product results.