PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

Published
Source
arXiv
Paper number
1142
Field
AI / Agents
arXiv ID
2609.40285

Key points

  • Across three Qwen3 models (8B to 235B), more than half of failed rollouts contained a pivotal mistake, usually occurring in early turns.
  • Those mistakes were critical yet recoverable: guiding the model for just a few turns after the pivotal turn restored task success.
  • The proposed PivotOPD has a teacher provide a gold action for prevention and recovery actions for restoration, trained with two distillation losses.
  • Compared against 13 baselines, it improved ALFWorld by +5.5% and transferred to SWE-Bench Verified with +3.2% resolve rate on another model family.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)