AARRI-Bench: Best AI Agents Score Only 68.3% on Research Intern Tasks, Revealing Key Judgment Gaps
Summary
- • A new benchmark called AARRI-Bench tests whether AI agents can perform research intern-level tasks with real professionalism and judgment.
- • The best-performing setup — Mini-SWE-Agent with Claude Opus 4.7 — achieved only 68.3% success, missing subtle critical details.
- • AI agents struggle with field sensitivity, research ethics, and nuanced scientific judgment that human interns handle naturally.
- • Researchers conclude that better scaffolding alone won't close the gap — models themselves need to learn research-specific behavior.
Details
AARRI-Bench evaluates AI agents on granular research lifecycle tasks
First benchmark in the AARR series; focuses on professionalism, thoroughness, and nuanced reasoning rather than macro execution capability.
Mini-SWE-Agent with Claude Opus 4.7 achieves 68.3% — the best result of all tested configurations
No configuration crossed 70%, indicating a meaningful and consistent capability ceiling across frontier models.
AI agents frequently overlook subtle critical details that human researchers catch easily
The failure mode is not length or context — agents miss nuanced cues that are obvious to trained human researchers, pointing to a training gap.
Benchmark tests field sensitivity, research ethics, and nuanced scientific judgment
Distinct from existing benchmarks that measure task completion speed or code correctness; targets the soft skills of research work.
Full benchmark dataset released publicly for community evaluation
Open release enables researchers to test other frontier models and agent scaffolding combinations against the same standardized tasks.
AARRI-Bench submitted to arXiv June 5, 2026; evaluates frontier AI agents on research-intern-level tasks across multiple frontier models.
What This Means
AARRI-Bench gives one of the clearest measurements yet of where frontier AI stands on real research work — capable enough to handle roughly two-thirds of research intern tasks, but not reliably enough to operate independently. The finding that scaffolding complexity doesn't explain the gap is significant: it points the finger at core model training rather than tooling, suggesting that closing the remaining 31.7% will require teaching AI systems research-specific norms and judgment. For research institutions considering AI automation of intern-level tasks, the benchmark provides a useful and public calibration tool.
