AREX: Towards a Recursively Self-Improving Agent for Deep Research
- Published
- Source
- arXiv
- Paper number
- 700
- Field
- AI / General
- arXiv ID
- 2607.21461
Key points
- It exploits the asymmetry between discovery and verification, where finding an answer is hard but checking it is easy, and uses verification as the starting point for the next research round.
- A self-contained context compression tool keeps only the essentials from long conversations, making long-horizon research possible.
- With a 4B model, it achieves 82.5 on BrowseComp and 52.4 on HLE, far outperforming similarly sized competing models.
- It teaches search, tool use, and verification step by step through agent mid-training and long-horizon reinforcement learning.
- It reduces the credit-assignment problem of sparse rewards by weighting the moments when key evidence is obtained.
Paper links
External research summaries. These are not HDATF publications or measured product results.