From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
- Published
- Source
- arXiv
- Paper number
- 803
- Field
- AI / General
- arXiv ID
- 2608.02163
Key points
- The benchmark consists of 500 deep-research tasks spanning 31 topics and 10 broad categories, and it separates different capabilities with three query types.
- The construction loop works as follows: the Explorer solves tasks in a real environment, the Formalizer organizes the results into a DAG and scoring checkpoints, and the Challenger creates the next-round query using clues that have not yet been used.
- As the rounds progress, the average DAG depth rises quickly from 2.5 in round 1 to 5.2 in round 5 and then levels off at about 7.0, while the number of nodes grows from 4.6 to 25.2, checkpoints from 13.6 to 123.2, and Explorer tool calls from 3.9 to 23.3. The average evolution rounds is 14.19.
- The difficulty genuinely increases. For the two query types other than the one that gives ample information, the average round score falls from about 0.93 and 0.95 to about 0.29 and 0.40. Even so, the relative ranking of models is preserved in most rounds.
- Removing tools lowers the score for all ten models. The overall score drops by an average of 0.14 and the evidence-based and analytical checkpoints drop by an average of 0.07 each, which means the tasks require long-tail information that is not well covered in pretraining.
- The scoring is stable as well. On 100 manually checked tasks, repeated scoring by the same model or by different models both show high Pearson correlation with human evaluation, between 0.884 and 0.916.
Paper links
External research summaries. These are not HDATF publications or measured product results.