← Back to issue8 / 22 · Week of Jul 27, 2026

smevals keeps harness failures out of model scores

smevals is an MIT-licensed CLI framework that records immutable runs for task and model-or-harness pairs, then grades them with ordered checks and reusable checkers. Why it matters: A network drop or crashed runner is a harness failure, not a model result. Keeping those records without grading them makes a routing comparison easier to audit.

Try this: Build a ten-task eval around a recurring review job, add a deterministic citation or schema checker, and compare two model or harness configs before changing the default.

GitHub 120 stars · Aug 2verify ↗
Source
GitHub — prime-radiant-inc/smevals
View source →

Get the field brief every week.

Important AI developments, useful explanations, and practical resources in one weekly read. Context to understand what matters, with links to the original sources and deeper reading.

Subscribe free →
Free weekly·No spam·Unsubscribe anytime