WANDR Benchmark Exposes Major Gaps in AI Research Agents' Ability to Collect Wide-and-Deep Evidence
Summary
- • WANDR is a new open benchmark of 500 tasks testing AI agents on broad entity discovery plus per-record evidence gathering
- • Best AI system achieves only 0.363 soft F1 and 0.133 hard F1—wide-and-deep research is far from solved
- • Tasks mirror real knowledge work: due diligence, competitive mapping, talent sourcing, and literature review
- • WANDR pairs with the DRACO benchmark to cover both long-form report writing and large-scale structured data collection
Details
Best AI system scores 0.363 soft F1 and 0.133 hard F1 on WANDR
Even the strongest system evaluated falls far short of professional research quality, revealing substantial capability gaps in wide-and-deep data collection tasks
500-task open benchmark built from real knowledge-work patterns
Tasks derived from de-identified production usage covering competitive mapping, due diligence, talent sourcing, literature review, and product comparison at professional scale
Hierarchical qualification keys make every evidence path independently verifiable
Tasks use compound hierarchies (e.g., company→employee→URL) so each node can be scored independently, enabling precise measurement of where agents fail
WANDR is the wide counterpart to the DRACO deep research benchmark
DRACO tests deep long-form report writing; WANDR tests wide structured data collection—together they cover the full spectrum of AI research agent evaluation
Sustaining breadth without sacrificing per-record accuracy is the core unsolved challenge
Wide discovery and deep evidence verification compound each other's difficulty—agents that find many entities tend to sacrifice quality on each one
Example task required 70+ company CEO/CFO appointments with multi-source evidence each
Representative WANDR task: find 70+ qualifying US companies with executive appointments in a 2-month window, each requiring both an appointment record and a listing-authority page
Benchmark released by the WANDR research team, surfaced via TLDR AI newsletter. Open benchmark and evaluation harness available publicly.
What This Means
WANDR fills a critical gap in AI evaluation by measuring how well research agents handle large-scale, evidence-backed data collection—the kind of structured work increasingly delegated to AI in business settings. The benchmark's dismal top scores (sub-0.37 soft F1) confirm that AI agents are still far from being reliable research partners for professional-scale tasks requiring both breadth and accuracy. This open benchmark gives researchers a reproducible way to track progress on one of the most practically important unsolved challenges in agentic AI.
