Anthropic Claude 3.5 Sonnet Jailbroken to Generate Stack Exploits
Key Takeaways Anthropic’s newly launched Fable 5 AI model was jailbroken within days of its release. The jailbreak allowed the generation of detailed instructions for creating stack buffer...
Key Takeaways
- Anthropic’s newly launched Fable 5 AI model was jailbroken within days of its release.
- The jailbreak allowed the generation of detailed instructions for creating stack buffer overflow exploits and synthesizing controlled substances.
- The researcher, “Pliny the Liberator,” leveraged multi-agent attacks, Unicode tricks, and narrative framing to bypass safety classifiers.
- Fable 5’s 120,000-character system prompt was also leaked, revealing Anthropic’s internal safety mechanisms.
Anthropic’s latest flagship AI model, Fable 5, introduced on June 9, 2026, has been successfully jailbroken just days after its public debut. The model, the first in the company’s new Mythos class, was touted as Anthropic’s most advanced AI, demonstrating superior capabilities in areas like software engineering, knowledge processing, and vision. However, a prominent AI red-teamer has circumvented its safety mechanisms, leading to the generation of highly sensitive information and the exposure of its core system prompt.
Table Of Content
Researcher “Pliny the Liberator” exploited a combination of multi-agent decomposition, Unicode manipulation, and narrative framing techniques to bypass Fable 5’s safety classifiers. This breach not only enabled the model to produce detailed instructions for stack exploits but also resulted in the leak of its extensive 120,000-character system prompt.
A notable design choice accompanied Fable 5’s release: it shares its underlying model with the more restricted Claude Mythos 5, with safety classifiers acting as the primary differentiator. When Fable 5 detects a query falling into high-risk categories such as cybersecurity, biology, chemistry, or model distillation, it is designed to silently reroute the request to the less capable Claude Opus 4.8, informing the user of the fallback. Anthropic had previously stated that extensive pre-launch bug bounty testing, spanning over 1,000 hours, yielded no universal jailbreaks. This claim was quickly put to the test.
Multi-Agent Bypass Within Days
Within a mere few days of Fable 5’s release, the prolific AI red-teamer Pliny the Liberator announced a successful bypass of the model’s safety layers. Pliny described his coordinated multi-agent attack strategy as “a pack hunt.”
Screenshots shared by Pliny showcased the model’s ability to generate highly technical and potentially dangerous outputs. These included step-by-step guidance for exploiting stack buffer overflows on x86 Linux systems. The instructions covered critical details such as disabling Address Space Layout Randomization (ASLR), crafting vulnerable C server code utilizing strcpy overflows, and compiling code without essential protections. Furthermore, the model was coaxed into detailing the Birch reduction mechanism, a well-known method for synthesizing methamphetamine.
Attack Vectors Utilized
Pliny meticulously documented the various attack vectors employed to achieve these bypasses:
- Unicode and Homoglyph Evasion: The use of Unicode characters, homoglyphs, and Cyrillic character substitutions to circumvent keyword-based classifiers.
- Long-Context Reference Tracking: Smuggling malicious intent within lengthy conversations, leveraging the model’s ability to track references across extended contexts.
- Taxonomy and Document-Structure Framing: Embedding harmful queries within what appeared to be legitimate study guides or academic references.
- Fiction and Narrative Framing: Masking offensive intent by framing requests as creative content or fictional scenarios.
- Decomposition and Recomposition: Extracting sensitive technical information in seemingly benign, isolated fragments, and then reassembling them into actionable offensive knowledge.
The decomposition and recomposition technique proved particularly effective. Pliny noted that “getting uplift on the process itself, like Birch reduction method or reductive amination, is much more doable” than directly requesting a specific harmful compound. The difficulty was further reduced by using an already jailbroken Opus instance to assist in the backend operations.
Beyond the technical bypasses, Pliny also made public Fable 5’s approximately 120,000-character system prompt, posting it to GitHub. This leak exposes the internal framing and safety instructions that Anthropic uses to govern the model’s foundational behavior.
Implications for AI Safety and Security
This incident reignites the ongoing debate concerning the balance between AI capabilities and safety containment. Anthropic’s design, which routes flagged requests to a weaker fallback model instead of outright refusal, aimed to reduce friction for legitimate users. However, Pliny contends that this approach creates a false sense of security while simultaneously hindering legitimate security researchers who require access to offensive techniques for defensive purposes. As of this report, Anthropic has not publicly addressed the jailbreak claims or the leaked system prompt.
The episode also highlights the complex challenges associated with securing agentic, multi-model AI pipelines. The fact that one jailbroken model (Opus) could facilitate the evasion of controls in another (Fable 5) suggests that safety evaluations focused solely on individual models may be fundamentally inadequate for complex AI systems.
What You Should Do
- Organizations deploying or integrating advanced AI models should actively red-team these systems for potential jailbreaks and bypasses, especially in multi-model architectures.
- Developers of AI models should re-evaluate fallback mechanisms, ensuring they do not inadvertently create new attack surfaces or a false sense of security.
- Security researchers and AI safety experts should continue to explore novel jailbreaking techniques to improve defensive measures and understand evolving threats.
- Users of AI models, particularly those in sensitive domains, should be aware that even models designed with safety features can be exploited to generate harmful content.
Disclaimer: HackersRadar reports on cybersecurity threats and incidents for informational and awareness purposes only. We do not engage in hacking activities, data exfiltration, or the hosting or distribution of stolen or leaked information. All content is based on publicly available sources.



JAILBREAK ALERT 


…
󠅫󠄼󠄿󠅆󠄵󠄐󠅀󠄼󠄹󠄾󠅉󠅭 (@elder_plinius)
No Comment! Be the first one.