← Back to issue8 / 22 · Week of Jul 27, 2026

smevals keeps harness failures out of model scores

smevals is an MIT-licensed CLI framework that records immutable runs for task and model-or-harness pairs, then grades them with ordered checks and reusable checkers. Why it matters: A network drop or crashed runner is a harness failure, not a model result. Keeping those records without grading them makes a routing comparison easier to audit.

Try this: Build a ten-task eval around a recurring review job, add a deterministic citation or schema checker, and compare two model or harness configs before changing the default.

GitHub 120 stars · Aug 2verify ↗
Source
GitHub — prime-radiant-inc/smevals
View source →

Get the field brief every week.

One lead signal, three quick hits, one thing to try, one concept decoded - and the rest of the week on the wire. For people who want to know what matters and what to do next.

Subscribe free →
Free weekly·No spam·Unsubscribe anytime