Dev.to
7/10/2026

The original headline is: "Anthropic publishes Cyber Jailbreak Severity (CJS) framework for grading AI jailbreaks"
Original: Anthropic wants to grade AI jailbreaks like CVEs. Here's the framework.
Short summary
Anthropic published a Cyber Jailbreak Severity (CJS) scale — a CVE-like framework for grading AI jailbreaks from CJS-0 (informational) to CJS-4 (critical). The framework evaluates jailbreaks on capability gain, breadth, ease of weaponization, and discoverability, with exponential severity bands. Anthropic also launched a HackerOne program for Fable 5 jailbreak submissions and outlined a four-tier classifier for cybersecurity use cases.
- •CJS scale grades jailbreaks 0-4 on capability gain, breadth, weaponization ease, and discoverability
- •Four-tier cyber use-case classifier: prohibited, high-risk dual use, low-risk dual use, benign
- •Fable 5 ships with a deliberately larger safety margin, accepting more false positives at launch
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



