Back
Security researchers push back on the guardrails in Anthropic's Fable
SiTech AI Team2 წთ. საკითხავი

Security researchers push back on the guardrails in Anthropic's Fable

Anthropic's new model Fable blocks prompts that look even loosely related to cybersecurity. Researchers say the filters catch ordinary engineering work and push them back to an older model.

Anthropic released Fable on Tuesday, presenting it as a public and limited version of Mythos, its cybersecurity model. The restrictions that come with it have drawn complaints from security researchers and professionals, who say the filters trigger on work that has nothing to do with attacking systems.

Filters that flag the topic, not the request

Valentina “Chompie” Palmiotti, a security researcher at IBM X-Force, wrote that Fable “rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post.” When a prompt trips the guardrails, the model pauses the chat and tells the user that its “safety measures flagged this message for cybersecurity or biology topics.” Another researcher reported on X that even asking for a code review was enough to set the filters off.

Why the limits exist

Anthropic put the restrictions in place to reduce the risk that Fable could be used to write malware or compromise software, a long-standing concern for the company. The biology limits follow a similar worry about biological weapons. The model is also programmed to fall back to Claude Opus 4.8 when it hits a guardrail. Separately, Anthropic runs a Cyber Verification Program: security professionals who are approved get fewer limitations when using Claude for security work. OpenAI offers a comparable programme called Trusted Access for Cyber.

Mythos and Project Glasswing

When Anthropic released Mythos in April, it limited the model to a small number of companies and organisations under an effort called Project Glasswing, aimed at securing critical software and infrastructure. Last week the company expanded access to hundreds of organisations in 15 countries.

Keyword matching, not judgement

Matt Suiche, a cybersecurity veteran who now works at the AI security startup Tolmo, told TechCrunch that “if you ask it to write secure code, it assumes it is cybersecurity related work instead of software engineering best practices, and you get downgraded.” In his view the triggers look keyword based, so anything in the lexical field of cybersecurity sets them off. He added that the approach is understandable at this stage and that the guardrails will probably be relaxed as frontier labs work more closely with security companies: “It's better to catch more people than not enough when you do such a release and to relax the guardrails over time.” Anthropic did not immediately respond to a request for comment.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.