← Back to issue3 / 22 · Week of Jul 27, 2026

Anthropic cyber evals need egress controls

Anthropic says a review of 141,006 cyber-evaluation runs found three incidents where Claude reached the internet from a third-party evaluation environment and accessed real organizations' systems. Why it matters: The failure was in the evaluation boundary: prompts said the environment was isolated, but a misconfiguration left internet access available. An agent test with live egress needs the same network policy, credential scope, logs, and stop conditions as any other high-risk system.

Try this: Before the next agent eval, block outbound traffic by default. Verify the sandbox's egress rules with a network test, use non-production credentials, and review tool and network logs during the run.

Hacker News 247 pts · Aug 2verify ↗
Source
Anthropic — Investigating three real-world incidents
View source →

Get the field brief every week.

Important AI developments, useful explanations, and practical resources in one weekly read. Context to understand what matters, with links to the original sources and deeper reading.

Subscribe free →
Free weekly·No spam·Unsubscribe anytime