Anthropic Debuts Claude Fable 5, First Mythos-Class AI Model
Key Takeaways Anthropic has launched Claude Fable 5, the first AI model in its new “Mythos” capability tier. Fable 5 incorporates built-in cybersecurity safeguards, automatically...
Key Takeaways
- Anthropic has launched Claude Fable 5, the first AI model in its new “Mythos” capability tier.
- Fable 5 incorporates built-in cybersecurity safeguards, automatically rerouting high-risk prompts to a less capable model.
- A specialized version, Claude Mythos 5, with lifted safeguards, is available to a restricted group of cybersecurity defenders and infrastructure providers.
- The new models set a high bar for AI capabilities in both offensive and defensive cybersecurity applications.
Anthropic Unveils Claude Fable 5: A New Era of AI with Integrated Cybersecurity Defenses
Anthropic has officially released Claude Fable 5, marking the debut of its “Mythos” capability tier, a new class of AI models. The company asserts that this advanced generation is so potent that robust cybersecurity safeguards have been integrated directly into its foundational release.
Table Of Content
Positioned above the existing Claude Opus line, Fable 5 demonstrates leading performance across a broad spectrum of capability benchmarks. Its most significant advancements are observed in handling intricate, multi-step tasks requiring extensive processing.
Dual-Use Capabilities and Integrated Safeguards
For cybersecurity professionals, a key aspect of the Mythos-class models is their profound offensive prowess. These models exhibit significant skill in identifying and exploiting software vulnerabilities, as well as executing “agentic hacking” — a sophisticated approach that chains together reconnaissance, discovery, lateral movement, and exploit development throughout a complete attack lifecycle. Recognizing this dual-use capability, Anthropic has centered the Fable 5 launch around comprehensive containment strategies.
Instead of outright rejecting potentially risky prompts, Fable 5 employs an intelligent routing mechanism. A distinct layer of classifiers is designed to detect requests related to cybersecurity, biology, chemistry, or model distillation. When such prompts are identified, the session is seamlessly redirected to a less capable model, specifically Claude Opus 4.8, preventing Fable 5 from processing the request directly. Users are transparently informed when this fallback mechanism is activated.
Conservative Tuning and Robust Defenses
Anthropic acknowledges that its classifiers have been conservatively tuned, which may lead to some benign requests being flagged. However, the company reports that these fallback triggers occur in less than 5% of sessions, ensuring that over 95% of user interactions leverage Fable 5’s full capabilities.
On the cybersecurity front, internal assessments confirm that the classifiers effectively prevent Fable 5 from making substantial progress on offensive tasks. An extensive external bug bounty program, spanning over 1,000 hours of testing, revealed no universal jailbreaks. Similarly, independent red-teaming organizations reported no universal jailbreaks in long-form agentic tasks.
Anthropic did note one specific instance where the UK AI Safety Institute made initial headway toward a jailbreak within a limited testing period. Despite this, an external partner reportedly found Fable 5’s defenses to be superior to any other model tested. It demonstrated zero compliance on harmful single-turn requests involving attack planning, exploit development, or defense evasion, even when confronted with 30 publicly known jailbreak techniques.
Mythos 5 for Defenders
In parallel with Fable 5, Anthropic is also making Claude Mythos 5 available. This version utilizes the same underlying model but with its cyber safeguards selectively lifted, offered exclusively to a restricted cohort of cyber defenders and critical infrastructure providers.
This specialized model is initially being deployed through Project Glasswing, a collaborative initiative with the U.S. government. Claude Mythos 5 is heralded as possessing the most advanced cybersecurity capabilities of any model globally. Anthropic anticipates expanding access through a structured trusted-access program in the future.
Both Fable 5 and Mythos 5 are priced at $10 per million input tokens and $50 per million output tokens. A new policy mandates a 30-day data retention period for all Mythos-class traffic. This data will be used strictly for safety monitoring — to identify novel jailbreaks, multi-request attacks, and false positives — and will never be employed for model training purposes.
Developers can integrate claude-fable-5 via the Claude API starting today.
Disclaimer: HackersRadar reports on cybersecurity threats and incidents for informational and awareness purposes only. We do not engage in hacking activities, data exfiltration, or the hosting or distribution of stolen or leaked information. All content is based on publicly available sources.



No Comment! Be the first one.