
Security researchers push back on the guardrails in Anthropic's Fable
Anthropic's new model Fable blocks prompts that look even loosely related to cybersecurity. Researchers say the filters catch ordinary engineering work and push them back to an older model.
Anthropic released Fable on Tuesday, presenting it as a public and limited version of Mythos, its cybersecurity model. The restrictions that come with it have drawn complaints from security researchers and professionals, who say the filters trigger on work that has nothing to do with attacking systems.
Filters that flag the topic, not the request
Valentina “Chompie” Palmiotti, a security researcher at IBM X-Force, wrote that Fable “rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post.” When a prompt trips the guardrails, the model pauses the chat and tells the user that its “safety measures flagged this message for cybersecurity or biology topics.” Another researcher reported on X that even asking for a code review was enough to set the filters off.
Why the limits exist
Anthropic put the restrictions in place to reduce the risk that Fable could be used to write malware or compromise software, a long-standing concern for the company. The biology limits follow a similar worry about biological weapons. The model is also programmed to fall back to Claude Opus 4.8 when it hits a guardrail. Separately, Anthropic runs a Cyber Verification Program: security professionals who are approved get fewer limitations when using Claude for security work. OpenAI offers a comparable programme called Trusted Access for Cyber.
Mythos and Project Glasswing
When Anthropic released Mythos in April, it limited the model to a small number of companies and organisations under an effort called Project Glasswing, aimed at securing critical software and infrastructure. Last week the company expanded access to hundreds of organisations in 15 countries.
Keyword matching, not judgement
Matt Suiche, a cybersecurity veteran who now works at the AI security startup Tolmo, told TechCrunch that “if you ask it to write secure code, it assumes it is cybersecurity related work instead of software engineering best practices, and you get downgraded.” In his view the triggers look keyword based, so anything in the lexical field of cybersecurity sets them off. He added that the approach is understandable at this stage and that the guardrails will probably be relaxed as frontier labs work more closely with security companies: “It's better to catch more people than not enough when you do such a release and to relax the guardrails over time.” Anthropic did not immediately respond to a request for comment.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.