HANDBOOK.md tests policy-following under real tool use
HANDBOOK.md presents 65 agent tasks in mock workplace environments, where agents must follow 20- to 124-page operating procedures while using MCP-connected email, chat, calendar, issue, and commerce services. Why it matters: Long policy files look like control layers until an agent has to carry their rules across a tool sequence. The benchmark grades required and prohibited actions with 824 programmatic criteria, so a plausible final answer cannot hide a broken process.
Try this: Read one task rubric before trusting a policy-file approach. Check whether the evaluation tests prohibited actions and rule conflicts, not only task completion.