Back
Anthropic apologizes for invisible Claude Fable guardrails
SiTech AI Team3 წთ. საკითხავი

Anthropic apologizes for invisible Claude Fable guardrails

Anthropic admits that silently degrading answers to suspected distillation attempts was the wrong tradeoff and says users will now see when Claude Fable 5 hands a query to Opus 4.8.

Anthropic has apologized for quietly throttling its new model, Claude Fable 5, with hidden guardrails that also affected researchers and rivals trying to use it to build competing systems. The company says it is reversing course and will be more transparent about when the restrictions take effect — even if that means Fable refuses more queries.

Hidden safeguards, then an apology

Fable is the first widely available model in Anthropic's Mythos class, a family the company spent months warning was too dangerous for public release. Anthropic said it addressed some of those risks by shipping Fable with safeguards that block certain "high-risk" queries. One restricted area is distillation, the technique of training smaller models on the outputs of larger ones. In Fable's system card, the company said it would handle suspected distillation attempts by altering and degrading the model's answers directly, without telling users that they had triggered the measure or that their responses had been changed.

Routing to Opus 4.8 instead

That approach is now changing. Queries that look like distillation attempts will fall back to Claude Opus 4.8, Anthropic's previous flagship model, the company said in a post on X. Users will be told when it happens: "You will see this every time it happens." The mechanism mirrors how Fable handles other high-risk areas such as biology, chemistry and cybersecurity, where queries are routed through Opus 4.8 unless they are blocked outright under broader rules covering drugs, weapons or other prohibited content. In biology, the safeguards were calibrated so broadly that Fable is practically unusable for even basic questions — something Anthropic spokesperson Paruul Maheshwary acknowledged to The Verge.

Backlash from researchers

The reversal follows intense criticism from the AI research community, which warned that silently limiting users suspected of distilling Fable could also hit third parties trying to evaluate the frontier model. In the system card, Anthropic argued that newer models' ability to accelerate AI development justified targeting such requests, noting that "using Claude to develop competing models already violates our Terms of Service". The company has previously accused Chinese rivals such as DeepSeek of distilling its models on an "industrial" scale.

In its post on X, Anthropic explained the tradeoff it had made: "Visible safeguards can be probed, so they have to be robust, which takes time to get right. Invisible safeguards can be targeted more narrowly, allowing us to ship quickly with very few false positives. We went with invisible safeguards for this reason — and that was the wrong tradeoff. You should have visibility into the safeguards we have in place, and why. We're sorry for not getting the balance right."

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.