Stanford Launches Terminal-Bench-Science: AI Agent Benchmark for Real Scientific Research Workflows
Summary
- • Stanford-led Terminal-Bench-Science evaluates AI agents on 70 expert-curated scientific research tasks.
- • Best model tested, Claude Opus 5, achieves only 30% task resolution — a significant capability gap.
- • Tasks span life, physical, Earth, mathematical, and engineering sciences from real researcher workflows.
- • Unlike static benchmarks, Terminal-Bench-Science evolves continuously as scientists contribute new tasks.
Details
Terminal-Bench-Science 0.1 released by Stanford-led team
Led by Stanford University researchers and built by the Terminal-Bench team, this benchmark evaluates AI agents on authentic scientific research workflows contributed by practicing scientists across disciplines worldwide.
70 tasks across 5 scientific domains in version 0.1
Initial release covers life sciences, physical sciences, Earth sciences, mathematical sciences, and engineering sciences. Tasks are drawn from real researcher workflows, not textbook exercises or standardized tests.
Claude Opus 5: 30% task resolution — best evaluated model
The strongest model tested on Terminal-Bench-Science 0.1 achieves a 30% resolution rate, revealing a significant gap between current frontier AI capabilities and authentic scientific research competency.
12 scientific workflow types evaluated across tasks
Tasks span scientific data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, signal processing, inverse problems, sensor calibration, model fitting, classification, and scientific machine learning.
Continuous benchmark: evolves via open GitHub contributions
Scientists contribute tasks through an open GitHub process with structured review and Discord discussion. Regular releases add new tasks and improve existing ones, keeping the benchmark relevant as models improve.
Built on Terminal-Bench framework for software engineering AI
Terminal-Bench originally drove progress in AI agents for software engineering; Terminal-Bench-Science applies the same rigorous real-workflow methodology to scientific research domains.
Goal: free scientists from time-consuming technical workflows
Designed to accelerate AI agents useful as research assistants, handling demanding scientific workflows so researchers can focus on defining questions, forming hypotheses, interpreting results, and communicating findings.
Source: Terminal-Bench-Science official page (via Hacker News)
What This Means
Terminal-Bench-Science is a significant methodological contribution to AI evaluation in science — it measures what AI can actually do when given the kinds of real, messy workflows that practicing scientists face, not idealized problems. The 30% ceiling for Claude Opus 5 is a sobering data point: even the best current models succeed on less than a third of authentic scientific tasks, confirming that AI as a useful autonomous research assistant remains aspirational across most scientific disciplines. The benchmark's continuous evolution and community-contributed design could make it a durable reference standard, avoiding the saturation problem that undermines most one-time academic benchmarks. If widely adopted, Terminal-Bench-Science could become a key signal for tracking AI progress toward genuine scientific usefulness.
Sources
- Terminal-Bench-Science: Evaluating AI agents on scientific research workflowsTerminal-bench-science
- Terminal-Bench-Science 0.1Terminal-bench-science
