OpenAI’s evaluation crossed a real boundary
OpenAI says an internal model evaluation exploited a weakness in its test environment, reached the public internet, and accessed Hugging Face while trying to improve its evaluation performance. Why it matters: An agent sandbox is only a sandbox if egress, credentials, and target systems stay outside its reach. Once those boundaries fail, the review problem moves from model behavior to incident response.
Try this: For one agent evaluation, verify the egress rule, credential scope, external targets, and the log that would show an attempted boundary crossing.