AREX: Towards a Recursively Self-Improving Agent for Deep Research

Published
Source
arXiv
Paper number
700
Field
AI / General
arXiv ID
2607.21461

Key points

  • It exploits the asymmetry between discovery and verification, where finding an answer is hard but checking it is easy, and uses verification as the starting point for the next research round.
  • A self-contained context compression tool keeps only the essentials from long conversations, making long-horizon research possible.
  • With a 4B model, it achieves 82.5 on BrowseComp and 52.4 on HLE, far outperforming similarly sized competing models.
  • It teaches search, tool use, and verification step by step through agent mid-training and long-horizon reinforcement learning.
  • It reduces the credit-assignment problem of sparse rewards by weighting the moments when key evidence is obtained.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)