← Back to issue18 / 22 · Week of Jul 27, 2026

HANDBOOK.md tests policy-following under real tool use

HANDBOOK.md presents 65 agent tasks in mock workplace environments, where agents must follow 20- to 124-page operating procedures while using MCP-connected email, chat, calendar, issue, and commerce services. Why it matters: Long policy files look like control layers until an agent has to carry their rules across a tool sequence. The benchmark grades required and prohibited actions with 824 programmatic criteria, so a plausible final answer cannot hide a broken process.

Try this: Read one task rubric before trusting a policy-file approach. Check whether the evaluation tests prohibited actions and rule conflicts, not only task completion.

Hacker News 325 pts · Aug 2verify ↗
Source
arXiv — HANDBOOK.md
View source →

Get the field brief every week.

Important AI developments, useful explanations, and practical resources in one weekly read. Context to understand what matters, with links to the original sources and deeper reading.

Subscribe free →
Free weekly·No spam·Unsubscribe anytime