Back to feed
Dev.to
Dev.to
7/10/2026
The original headline is: "Anthropic publishes Cyber Jailbreak Severity (CJS) framework for grading AI jailbreaks"

The original headline is: "Anthropic publishes Cyber Jailbreak Severity (CJS) framework for grading AI jailbreaks"

Original: Anthropic wants to grade AI jailbreaks like CVEs. Here's the framework.

Short summary

Anthropic published a Cyber Jailbreak Severity (CJS) scale — a CVE-like framework for grading AI jailbreaks from CJS-0 (informational) to CJS-4 (critical). The framework evaluates jailbreaks on capability gain, breadth, ease of weaponization, and discoverability, with exponential severity bands. Anthropic also launched a HackerOne program for Fable 5 jailbreak submissions and outlined a four-tier classifier for cybersecurity use cases.

  • CJS scale grades jailbreaks 0-4 on capability gain, breadth, weaponization ease, and discoverability
  • Four-tier cyber use-case classifier: prohibited, high-risk dual use, low-risk dual use, benign
  • Fable 5 ships with a deliberately larger safety margin, accepting more false positives at launch

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more