← Back to issue3 / 13 · Week of Aug 31, 2026

Terminal-Bench-Science tests scientific workflows with verifiable artifacts

Terminal-Bench-Science 0.1 contains 70 expert-curated tasks across five scientific domains and grades analyses, simulations, proofs, code, and data products with reproducible task-specific tests. Why it matters: The benchmark makes a more useful claim than a research demo: whether an agent can complete a versioned workflow and leave an artifact that a verifier can inspect. Its 30% leading resolution rate is still a result for this release, harness, and configuration—not a measure of general scientific autonomy.

Try this: Pick one task close to an internal analytical workflow, inspect or run its verification script, and record resolution rate, elapsed time, human intervention, and artifact quality before automating a scientific step.

Hacker News 117 pts · Aug 31verify ↗
Source
Terminal-Bench-Science — 0.1 announcement
View source →

Get the field brief every week.

Important AI developments, useful explanations, and practical resources in one weekly read. Context to understand what matters, with links to the original sources and deeper reading.

Subscribe free →
Free weekly·No spam·Unsubscribe anytime