A severity scale for AI jailbreaks, proposed by Anthropic

Anthropic and partners like Google propose Cyber Jailbreak Severity, a five-level scale to rate AI guardrail bypasses from informational to critical.

Back on a global scale and for all users, Fable 5 comes with two Anthropic publications dedicated to safety. The first details the model's cyber guardrails, these classifiers that identify and block potentially dangerous uses. The publisher explains that it does not seek to prohibit all cybersecurity activity, as many uses are double-edged, defensive as well as offensive. Its classifiers therefore sort requests into four risk levels, from prohibited use to benign use, with a deliberately expanded safety margin for Fable 5, at the cost of a higher false positive rate.

The second publication proposes, with partners from the Glasswing program including Amazon, Microsoft, and Google, a framework for rating the severity of jailbreaks, these methods of bypassing guardrails. Named Cyber Jailbreak Severity, it classifies a bypass on five levels, from "informational" to "critical," according to several axes including the capability gain provided to an attacker and the ease of implementation.

The stated objective is to provide AI developers and public authorities with a common language to discuss these risks, in the absence of an existing standard. Anthropic specifies that this is a draft submitted for discussion, and has opened a reporting program for security researchers.