← Back to issue11 / 22 · Week of Jul 20, 2026

Opus 5 makes agentic work a routing problem

Artificial Analysis reports Claude Opus 5 as the new leader on AA-Briefcase, its proprietary benchmark for agentic knowledge work across private-file tasks such as reports, presentations, and spreadsheets. At max effort, Opus 5 scores 1720 Elo, 146 points ahead of Claude Fable 5, while reducing cost per task from $22.30 to $17.79. Why it matters: The useful signal is not just a leaderboard win. Opus 5 exposes the operating dial teams now need for knowledge-work agents: effort setting, rubric pass rate, analytical quality, presentation quality, cost per task, turns, and wall-clock time. The top settings lead, but they also average more than 25 minutes per task.

Try this: Before routing all knowledge-work agents to the top model, run the same task at two effort settings and compare receipts: rubric pass, cost per task, minutes per task, turns, and how much presentation cleanup a human still needs to do.

Source
Artificial Analysis on Claude Opus 5 AA-Briefcase
View source →

Get the field brief every week.

Important AI developments, useful explanations, and practical resources in one weekly read. Context to understand what matters, with links to the original sources and deeper reading.

Subscribe free →
Free weekly·No spam·Unsubscribe anytime