Terminal-Bench-Science tests scientific workflows with verifiable artifacts
Terminal-Bench-Science 0.1 contains 70 expert-curated tasks across five scientific domains and grades analyses, simulations, proofs, code, and data products with reproducible task-specific tests. Why it matters: The benchmark makes a more useful claim than a research demo: whether an agent can complete a versioned workflow and leave an artifact that a verifier can inspect. Its 30% leading resolution rate is still a result for this release, harness, and configuration—not a measure of general scientific autonomy.
Try this: Pick one task close to an internal analytical workflow, inspect or run its verification script, and record resolution rate, elapsed time, human intervention, and artifact quality before automating a scientific step.