← Back to feed
6

Stanford Launches Terminal-Bench-Science: AI Agent Benchmark for Real Scientific Research Workflows

Research2 sources·5d ago

Summary

  • • Stanford-led Terminal-Bench-Science evaluates AI agents on 70 expert-curated scientific research tasks.
  • • Best model tested, Claude Opus 5, achieves only 30% task resolution — a significant capability gap.
  • • Tasks span life, physical, Earth, mathematical, and engineering sciences from real researcher workflows.
  • • Unlike static benchmarks, Terminal-Bench-Science evolves continuously as scientists contribute new tasks.
Adjust signal

Details

Research

Terminal-Bench-Science 0.1 released by Stanford-led team

Led by Stanford University researchers and built by the Terminal-Bench team, this benchmark evaluates AI agents on authentic scientific research workflows contributed by practicing scientists across disciplines worldwide.

Stat

70 tasks across 5 scientific domains in version 0.1

Initial release covers life sciences, physical sciences, Earth sciences, mathematical sciences, and engineering sciences. Tasks are drawn from real researcher workflows, not textbook exercises or standardized tests.

Stat

Claude Opus 5: 30% task resolution — best evaluated model

The strongest model tested on Terminal-Bench-Science 0.1 achieves a 30% resolution rate, revealing a significant gap between current frontier AI capabilities and authentic scientific research competency.

Tech Info

12 scientific workflow types evaluated across tasks

Tasks span scientific data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, signal processing, inverse problems, sensor calibration, model fitting, classification, and scientific machine learning.

Research

Continuous benchmark: evolves via open GitHub contributions

Scientists contribute tasks through an open GitHub process with structured review and Discord discussion. Regular releases add new tasks and improve existing ones, keeping the benchmark relevant as models improve.

Context

Built on Terminal-Bench framework for software engineering AI

Terminal-Bench originally drove progress in AI agents for software engineering; Terminal-Bench-Science applies the same rigorous real-workflow methodology to scientific research domains.

Strategy

Goal: free scientists from time-consuming technical workflows

Designed to accelerate AI agents useful as research assistants, handling demanding scientific workflows so researchers can focus on defining questions, forming hypotheses, interpreting results, and communicating findings.

Source: Terminal-Bench-Science official page (via Hacker News)

What This Means

Terminal-Bench-Science is a significant methodological contribution to AI evaluation in science — it measures what AI can actually do when given the kinds of real, messy workflows that practicing scientists face, not idealized problems. The 30% ceiling for Claude Opus 5 is a sobering data point: even the best current models succeed on less than a third of authentic scientific tasks, confirming that AI as a useful autonomous research assistant remains aspirational across most scientific disciplines. The benchmark's continuous evolution and community-contributed design could make it a durable reference standard, avoiding the saturation problem that undermines most one-time academic benchmarks. If widely adopted, Terminal-Bench-Science could become a key signal for tracking AI progress toward genuine scientific usefulness.

Sources

Update history (1)
4d agoAdded TLDR AI newsletter as a corroborating source for the Terminal-Bench-Science 0.1 announcement; no new information.

Similar Events