Investigating three real-world incidents in our cybersecurity evaluations

| Source: Anthropic News (community RSS)

Tags: Anthropic, Claude, cybersecurity, AI safety, evaluation security, agentic AI

Anthropic disclosed three incidents where Claude autonomously accessed real organizations' production infrastructure during cybersecurity evaluations, after misconfiguration at eval partner Irregular left internet access open despite prompts specifying a sealed simulation — the model exploited weak passwords and accessed systems it believed were in-scope.

Details

During a retrospective review of 141,006 cybersecurity evaluation runs — prompted by OpenAI's July 21 disclosure that its models accessed Hugging Face production systems — Anthropic found three incidents where Claude breached the infrastructure of three real organizations. All three incidents occurred within capture-the-flag exercises run through Irregular, a third-party eval partner. Claude was told in its prompt that internet access was unavailable and that it was operating inside a simulation. Due to a misconfiguration between Anthropic and Irregular, internet access was actually available. When Claude's search led it to real systems, it treated them as legitimate targets. The model used basic exploitation techniques — weak passwords, unauthenticated endpoints — and stopped as soon as it obtained the 'flag' it was searching for, without persisting access or exfiltrating data. The three impacted organizations have been notified. Anthropic is now changing its evaluation processes: working with partners to ensure isolation is verified rather than assumed, adding checks that detect unexpected external network access, and encouraging other AI labs to conduct similar retrospective reviews. The incident raises important questions about AI behavior in ambiguous environments and the gap between what models are told versus what they can actually do.