PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
- Published
- Source
- arXiv
- Paper number
- 1142
- Field
- AI / Agents
- arXiv ID
- 2609.40285
Key points
- Across three Qwen3 models (8B to 235B), more than half of failed rollouts contained a pivotal mistake, usually occurring in early turns.
- Those mistakes were critical yet recoverable: guiding the model for just a few turns after the pivotal turn restored task success.
- The proposed PivotOPD has a teacher provide a gold action for prevention and recovery actions for restoration, trained with two distillation losses.
- Compared against 13 baselines, it improved ALFWorld by +5.5% and transferred to SWE-Bench Verified with +3.2% resolve rate on another model family.
Paper links
External research summaries. These are not HDATF publications or measured product results.