SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents

Published
Source
arXiv
Paper number
078
Field
Research Agents / RL
arXiv ID
2509.06283

Key points

  • It equips large language models, or LLMs, with advanced reasoning and tool-use capabilities so they can carry out complex, long-horizon tasks while mitigating issues such as hallucination.
  • It develops an effective autonomous single-agent system for deep research that works without rigid multi-agent orchestration and generalizes better.
  • It stabilizes reinforcement learning, or RL, for multi-turn agentic tasks, where existing methods struggle with stability and can degrade the base model's reasoning ability.
  • It develops an agentic reasoning scaffold for a specific LLM family with the minimal set of three tools, Internet search, page viewing, and a code interpreter, and an iterative single-turn workflow.
  • It builds a new synthetic data generation pipeline that produces difficult multi-hop QA and long-form report-writing tasks for agentic training.
  • It implements an end-to-end RL recipe with a length-normalized REINFORCE objective, trajectory filtering, and partial rollouts to ensure stable policy optimization over multi-turn interactions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)