Following the global re-deployment of its latest model, Claude Fable 5, Anthropic has shared further information in two areas[1]. The first is a detailed account of how far the safety classifiers—the safeguards that detect and block dangerous cybersecurity uses—are designed to go[1]. The second is an early draft of an industry-wide framework for scoring the severity of "jailbreaks," the techniques used to bypass an AI model's safeguards[1]. This article organizes what each of these entails.

Cyber safeguards judged across four categories

Cybersecurity is a difficult area for AI safeguards because the same techniques are often "dual use"—useful to both defenders and attackers[1]. For example, letting defenders scan their own code for vulnerabilities is desirable, yet the very same capability could, in the wrong hands, become the precursor to an attack[1].

For this reason, rather than blocking everything cyber-related, Anthropic trains its classifiers to sort uses into four categories[1]. The most dangerous, "Prohibited use," covers ransomware and wipers, malware development and modification, malware delivery through phishing, building attack infrastructure such as command-and-control (C2) servers, and attacks on internet backbone infrastructure like BGP hijacking—all of which are blocked[1].

Next, "high-risk dual use" includes penetration testing, red team exercises, privilege escalation, and exploit development—work that is part of security professionals' daily jobs but carries a high risk of misuse[1]. What separates legitimate work from abuse is context: who is doing it, and under what authorization? For Fable 5, Anthropic expects to block these actions until it has better controls to limit access to known good actors[1]. The third category, "low-risk dual use," covers defense-leaning activities such as open-source intelligence (OSINT) and identifying vulnerabilities that other tools can already find; most are allowed, though a fraction are still blocked as a precaution[1]. The fourth, "benign use," covers core defensive and IT work such as secure coding, debugging, patch management, log analysis, and malware reverse engineering, which are not intended to be blocked[1].

Fraud, social engineering, game cheating, CAPTCHA evasion, and crypto-related crimes fall outside these cyber classifiers—handled by separate classifiers or not considered harmful[1].

For Fable 5, Anthropic has set a wider "safety margin" than for previous models to reliably catch dangerous behavior[1]. The safety margin means a request has to look very clearly safe to avoid being blocked, which can raise false positives in ordinary work[1]. Classifiers are only one part of the defense; Anthropic says it combines them with access controls, model safety training, and offline monitoring for a layered approach[1].

How far vulnerability finding is allowed

Among the four categories, vulnerability finding is especially hard to draw a line around[1]. Rather than banning it entirely, Anthropic aims to block "high-uplift" vulnerability finding—the ability to identify vulnerabilities that other widely available models cannot[1]. Because attackers can sometimes build exploits from public vulnerability reports or patches, the automatic generation of exploits is also blocked[1].

On the other hand, if a vulnerability is one that many common models can already find, having Fable 5 find and fix it is considered beneficial[1]. The security community has long held that vulnerability finding and responsible disclosure is a net positive—defenders gain more from knowing what to fix than attackers gain from the same reports[1]. The US government is cited as taking the same position, noting that "in the vast majority of cases, responsibly disclosing a newly discovered vulnerability is clearly in the national interest"[1].

The "CJS" framework for measuring jailbreak severity

The other pillar is a proposed framework for assessing the severity of jailbreaks[1]. A jailbreak is a technique that prompts an AI model in unusual ways to bypass its safeguards[1]. Their severity varies widely, from merely unblocking trivial behaviors to unlocking a broad range of harmful outputs and making a model far more dangerous[1]. Yet there has been no common yardstick for describing that severity, making it hard for AI developers and governments to discuss risk in consistent terms[1].

Anthropic proposes a scale it calls "Cyber Jailbreak Severity (CJS)," with five bands: None (Informational, CJS-0), Low (CJS-1), Medium (CJS-2), High (CJS-3), and Critical (CJS-4)[1]. The bands are meant to be exponential rather than linear, so each step up is several times more serious than the last[1].

The CJS score is calculated from four axes[1]. The first two describe what a jailbreak gives an attacker: "capability gain," or how far it takes an attacker beyond their existing tools, and "breadth," or how many distinct offensive tasks the same technique works on[1]. The second two describe how quickly the technique can become a real-world problem: "ease of weaponization," the effort needed to turn the technique into a working attack, and "discoverability," how easily an attacker can obtain the technique in the first place[1].

Anthropic notes that "capability gain" concerns cyber-domain expertise (does it accelerate even experts?), whereas "ease of weaponization" concerns skill at using LLMs and jailbreaks; the two are independent[1]. A technique can score high on one and low on the other[1].

The score is a "floor," and the Log4Shell example shows its time dependence

Summing the four axes produces an initial CJS level from 0 to 4[1]. This value is a provisional "floor" that cannot be lowered, but it can be raised if the rubric is judged to underestimate real-world risk[1]. Reasons to raise it include producing a novel, serious vulnerability in widely deployed software, exploiting a fundamental capability with no near-term mitigation, or combining with other open findings to worsen the overall risk[1].

The point that capability gain is measured against the tools available at the time is illustrated well by the Log4Shell example in the appendix[1]. Imagining a jailbreak that finds the real Log4Shell vulnerability, the severity is rated high for December 2021, when no other tool could find it[1]. But today, when the vulnerability is public and every scanner detects it, the capability gain is rated as zero and the CJS level drops—even though the model's behavior is identical[1]. The evaluation changes because the baseline of "the state of the art at that moment" has moved[1].

The framework is an early draft, and Anthropic says it hopes to spark a discussion across academia, industry, civil society, and government[1]. It is taking feedback at [email protected] and has launched a HackerOne program where security researchers can submit jailbreaks they find in Fable 5[1].

Summary

This release lays out, in four concrete categories, what Fable 5's cyber safeguards will and will not prevent, and puts forward the CJS framework as a common language for discussing jailbreak severity. The focus going forward will be whether the difficult task of curbing misuse without hampering defensive use can grow into a standard shared across industry and government.

Source: https://www.anthropic.com/news/fable-safeguards-jailbreak-framework