← Back to feed
7

Multi-Institution Study Finds AI Agents Excel at Research Engineering But Fail at Open-Ended Scientific Creativity

Research2 sources·Aug 3

Summary

  • • Frontier AI agents given 6 days and thousands in compute completed all engineering tasks on two unpublished NeurIPS 2026 papers, but failed to produce publishable research contributions
  • • Both papers were 'unambiguously rejected' by original authors citing poor research bar judgment, uncreative responses to design shortcomings, and instruction drift
  • • Study introduces 'shadow evaluations' — a new benchmark where agents tackle real open-ended paper questions graded by original authors — as a third way beyond narrow tasks and blind peer review
  • • Robustness check with a second model and scaffold reproduced all five failure modes, suggesting fundamental architectural gaps rather than fixable prompt engineering issues
Adjust signal

Details

Research

Multi-institution team tests frontier agents on real NeurIPS 2026 papers

Team from Princeton, UC Berkeley, Georgetown, Johns Hopkins, and other institutions tested frontier AI agents on two unpublished NeurIPS 2026 paper questions using the shadow evaluation methodology (submitted to arXiv July 29, 2026)

New Tech

Shadow evaluations: new benchmark for open-ended research capability

'Shadow evaluations' methodology: agent takes on the central open-ended research question of a real high-quality unpublished paper; original authors serve as expert graders — a third way between narrow verifiable benchmarks and overstretched blind peer review

Insight

Full engineering done; zero publishable research progress

Agents completed 100% of required engineering tasks without human assistance but made no substantial progress on core research questions; original authors rejected both papers 'unambiguously' as falling below publishable standards

Insight

Five failure modes identified across both evaluations

Five recurring failure modes: (1) poor judgment about publishable research bar, (2) uncreative responses to design shortcomings, (3) ineffective backtracking from dead ends, (4) poor resource awareness, (5) instruction drift

Research

Robustness check reproduces all failures across second model and scaffold

A separate robustness check using a different model and agent scaffold reproduced all five failure modes, ruling out model-specific or scaffold-specific explanations and strengthening generalizability of findings

Stat

6 days and thousands of dollars of compute per evaluation

Each shadow evaluation provided frontier agents with 6 days of runtime and thousands of dollars of compute budget — substantial resource allocation that still failed to bridge the creativity gap

Context

Direct challenge to explosive-AI-progress forecasts

Forecasts of explosive AI capability growth often hinge on AI agents automating AI research itself; this study provides early evidence that the creative and judgment components of that research loop remain unsolved today

Source: arXiv preprint submitted July 29, 2026, multi-institution team (Princeton, UC Berkeley, Georgetown, Johns Hopkins). Reported by Import AI newsletter. Curator summary credits institutions including Princeton, UC Berkeley, Georgetown, Johns Hopkins.

What This Means

This multi-institution study provides the first systematic shadow-evaluation evidence that frontier AI agents can fully execute the engineering component of AI research autonomously — but cannot replicate the creative scientific judgment that makes research publishable. Both evaluated papers were rejected by original authors despite agents completing all required engineering without human help. The finding is particularly significant for forecasters who assume AI will soon automate its own R&D loop, driving explosive capability gains: the five identified failure modes appear robust across models and scaffolds, pointing to fundamental architectural gaps in research creativity rather than solvable implementation issues. The shadow evaluation methodology itself is a methodological contribution — a more rigorous alternative to narrow benchmarks and overstretched blind peer review for assessing genuine open-ended AI research capability.

Sources

Similar Events