← Back to feed
6

WANDR Benchmark Exposes Major Gaps in AI Research Agents' Ability to Collect Wide-and-Deep Evidence

Research1 source·Jul 15

Summary

  • • WANDR is a new open benchmark of 500 tasks testing AI agents on broad entity discovery plus per-record evidence gathering
  • • Best AI system achieves only 0.363 soft F1 and 0.133 hard F1—wide-and-deep research is far from solved
  • • Tasks mirror real knowledge work: due diligence, competitive mapping, talent sourcing, and literature review
  • • WANDR pairs with the DRACO benchmark to cover both long-form report writing and large-scale structured data collection
Adjust signal

Details

Stat

Best AI system scores 0.363 soft F1 and 0.133 hard F1 on WANDR

Even the strongest system evaluated falls far short of professional research quality, revealing substantial capability gaps in wide-and-deep data collection tasks

Research

500-task open benchmark built from real knowledge-work patterns

Tasks derived from de-identified production usage covering competitive mapping, due diligence, talent sourcing, literature review, and product comparison at professional scale

Tech Info

Hierarchical qualification keys make every evidence path independently verifiable

Tasks use compound hierarchies (e.g., company→employee→URL) so each node can be scored independently, enabling precise measurement of where agents fail

Industry Update

WANDR is the wide counterpart to the DRACO deep research benchmark

DRACO tests deep long-form report writing; WANDR tests wide structured data collection—together they cover the full spectrum of AI research agent evaluation

Insight

Sustaining breadth without sacrificing per-record accuracy is the core unsolved challenge

Wide discovery and deep evidence verification compound each other's difficulty—agents that find many entities tend to sacrifice quality on each one

Context

Example task required 70+ company CEO/CFO appointments with multi-source evidence each

Representative WANDR task: find 70+ qualifying US companies with executive appointments in a 2-month window, each requiring both an appointment record and a listing-authority page

Benchmark released by the WANDR research team, surfaced via TLDR AI newsletter. Open benchmark and evaluation harness available publicly.

What This Means

WANDR fills a critical gap in AI evaluation by measuring how well research agents handle large-scale, evidence-backed data collection—the kind of structured work increasingly delegated to AI in business settings. The benchmark's dismal top scores (sub-0.37 soft F1) confirm that AI agents are still far from being reliable research partners for professional-scale tasks requiring both breadth and accuracy. This open benchmark gives researchers a reproducible way to track progress on one of the most practically important unsolved challenges in agentic AI.

Sources

Similar Events