From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

Published
Source
arXiv
Paper number
803
Field
AI / General
arXiv ID
2608.02163

Key points

  • The benchmark consists of 500 deep-research tasks spanning 31 topics and 10 broad categories, and it separates different capabilities with three query types.
  • The construction loop works as follows: the Explorer solves tasks in a real environment, the Formalizer organizes the results into a DAG and scoring checkpoints, and the Challenger creates the next-round query using clues that have not yet been used.
  • As the rounds progress, the average DAG depth rises quickly from 2.5 in round 1 to 5.2 in round 5 and then levels off at about 7.0, while the number of nodes grows from 4.6 to 25.2, checkpoints from 13.6 to 123.2, and Explorer tool calls from 3.9 to 23.3. The average evolution rounds is 14.19.
  • The difficulty genuinely increases. For the two query types other than the one that gives ample information, the average round score falls from about 0.93 and 0.95 to about 0.29 and 0.40. Even so, the relative ranking of models is preserved in most rounds.
  • Removing tools lowers the score for all ten models. The overall score drops by an average of 0.14 and the evidence-based and analytical checkpoints drop by an average of 0.07 each, which means the tasks require long-tail information that is not well covered in pretraining.
  • The scoring is stable as well. On 100 manually checked tasks, repeated scoring by the same model or by different models both show high Pearson correlation with human evaluation, between 0.884 and 0.916.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)