Opus 5 makes agentic work a routing problem
Artificial Analysis reports Claude Opus 5 as the new leader on AA-Briefcase, its proprietary benchmark for agentic knowledge work across private-file tasks such as reports, presentations, and spreadsheets. At max effort, Opus 5 scores 1720 Elo, 146 points ahead of Claude Fable 5, while reducing cost per task from $22.30 to $17.79. Why it matters: The useful signal is not just a leaderboard win. Opus 5 exposes the operating dial teams now need for knowledge-work agents: effort setting, rubric pass rate, analytical quality, presentation quality, cost per task, turns, and wall-clock time. The top settings lead, but they also average more than 25 minutes per task.
Try this: Before routing all knowledge-work agents to the top model, run the same task at two effort settings and compare receipts: rubric pass, cost per task, minutes per task, turns, and how much presentation cleanup a human still needs to do.