ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
- Published
- Source
- arXiv
- Paper number
- 831
- Field
- AI / General
- arXiv ID
- 2608.05102
Key points
- The paper proposes the ABC framework, which backtracks from the answer to reconstruct intermediate clues and evaluates each search step against those clues.
- The key idea is precise credit assignment: useful steps in failed trajectories are rewarded, while unnecessary steps in successful trajectories are penalized.
- ABC-SFT uses step-wise loss weighting, and ABC-GRPO uses step-wise rewards to improve both conventional SFT and GRPO.
- With a 4B model, it achieves state-of-the-art performance at the same scale, improving from 37.3 to 55.3 percent on BrowseComp and from 39.1 to 52.9 percent on BrowseComp-ZH.
- Trained on only 8.5K examples, it also generalizes well to other benchmarks such as xbench and GAIA.
Paper links
External research summaries. These are not HDATF publications or measured product results.