From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
Ranking
Overall
68
Content
80
Popularity
41
Observed public metrics from 1 member.
Merged summary
TL;DR - An arXiv paper introducing a 500-task verifiable benchmark for deep-research agents, built fully automatically by iteratively evolving simple QA into expert-level tasks. It matters because it removes the expert-authoring bottleneck while keeping evaluation traceable and reproducible.
- Benchmark covers 500 deep research tasks across 31 topics and 10 major categories, with three query forms probing complementary deep-research capabilities.
- Construction uses an iterative Explorer–Formalizer–Challenger pipeline that progressively transforms simple questions into harder research tasks.
- Each task is encoded as a DAG of atomic steps plus checkpoints, so query, DAG, and rubrics co-evolve in a controlled, traceable way.
- Reported experiments show the benchmark discriminates among models and query types, with fact-grounded pointwise rubrics enabling fine-grained, human-aligned, stable scoring; data, code, and results are public.
Sources (1)
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
Public signals
Hugging Face upvotes 0
TL;DR - An arXiv paper introducing a 500-task verifiable benchmark for deep-research agents, built fully automatically by iteratively evolving simple QA into expert-level tasks. It matters because it removes the expert-authoring bottleneck while keeping evaluation traceable and reproducible.
- Benchmark covers 500 deep research tasks across 31 topics and 10 major categories, with three query forms probing complementary deep-research capabilities.
- Construction uses an iterative Explorer–Formalizer–Challenger pipeline that progressively transforms simple questions into harder research tasks.
- Each task is encoded as a DAG of atomic steps plus checkpoints, so query, DAG, and rubrics co-evolve in a controlled, traceable way.
- Reported experiments show the benchmark discriminates among models and query types, with fact-grounded pointwise rubrics enabling fine-grained, human-aligned, stable scoring; data, code, and results are public.