More details on Fable 5’s cyber safeguards and our jailbreak framework
| Source: Anthropic News (community RSS)
Tags: Claude, Anthropic, Fable 5, cybersecurity, jailbreak, AI safety, Glasswing, export controls
Anthropic publicly details Claude Fable 5's four-tier cybersecurity classifier system and releases a first-draft jailbreak severity framework co-developed with Amazon, Microsoft, Google, and other Glasswing partners — the industry's first attempt at a shared standard for assessing AI safeguard bypasses.
Details
Anthropic published detailed documentation of the cybersecurity safety classifiers built into Claude Fable 5, following the model's global redeployment on July 2. The classifiers evaluate requests across four categories — from explicitly prohibited to clearly benign — allowing defensive security work while blocking assistance with cyberattacks. Anthropic explicitly lists what the classifiers are and are not designed to block, acknowledging the dual-use challenge: the same capability that helps defenders scan for vulnerabilities could assist attackers. Alongside the classifier details, Anthropic released a first-draft jailbreak severity framework co-developed with Amazon, Microsoft, Google, and other Glasswing program partners. The framework aims to standardize how AI developers communicate jailbreak risk to governments and each other — a gap that became visible during the June 12 export control episode. Jailbreaks vary significantly in severity: some unblock minor undesirable behaviors, others make a model broadly more dangerous. The shared standard would allow consistent communication between AI labs and regulators. Anthropic is treating this publication as an open draft and is actively soliciting critique from academia, industry, civil society, and government at [email protected]. A HackerOne bug bounty program has also launched for security researchers to submit potential Fable 5 cyber jailbreaks for review. The company says it hopes the framework will enable defensive uses of AI while preventing misuse.