Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
Terminal-Bench-Science is a new, continuous benchmark led by Stanford researchers, designed to rigorously evaluate AI agents on real-world scientific workflows. It sets the bar for AI capabilities in scientific research by having scientists, not model developers, define and contribute challenging tasks. This initiative aims to accelerate discovery by developing AI agents that act as powerful research assistants, freeing human scientists for higher-level judgment.
The Lowdown
Terminal-Bench-Science, a new benchmark spearheaded by researchers at Stanford University and the team behind Terminal-Bench, aims to rigorously evaluate and advance AI agents for scientific research. It distinguishes itself by involving scientists directly in defining and contributing challenging, expert-curated workflows, ensuring the evaluations reflect real-world scientific practice rather than theoretical exercises. The goal is to develop AI agents capable of executing technically demanding and time-consuming tasks, thereby allowing human researchers to focus on critical aspects like hypothesis formation, interpretation, and communication, ultimately accelerating scientific discovery.
- Real-world Scientific Workflows: The benchmark focuses on practical research tasks contributed by practicing scientists, covering diverse domains like life, physical, Earth, mathematical, and engineering sciences.
- Verifiable Capability: Evaluations are conducted in realistic environments, grading concrete artifacts such as analyses, simulations, and data products with reproducible, task-specific tests.
- Continuous Evolution: Designed as a continuous benchmark, Terminal-Bench-Science evolves alongside frontier AI. It features regular releases, an open contribution process for new tasks, and mechanisms to improve existing ones, fostering a feedback loop between scientific needs and AI development.
- Initial Findings (0.1 Release): The first release includes 70 tasks, with Claude Opus 5 achieving the highest resolution rate at 30%, followed by GPT-5.6 Sol (22.4%) and Claude Fable 5 (21.4%). This highlights substantial room for AI agent improvement in scientific domains.
- Open Contribution Process: Researchers can propose, build, and review tasks through GitHub and Discord, ensuring a community-driven, transparent, and highly selective process for task inclusion.
By grounding AI agent evaluation in authentic scientific practice and fostering a continuous feedback loop, Terminal-Bench-Science seeks to build AI tools that genuinely extend researchers' capabilities and push the boundaries of scientific understanding.