
Anthropic may reroute Claude Sonnet 5.5 requests to Sonnet 5 under cyber safeguards
Anthropic says Claude Sonnet 5.5 is the first Sonnet model to carry its three-stage cyber classifiers and model fallbacks. Blocked cyber requests are routed to Sonnet 5, which changes the rules for developers working through the API.
Routing reaches the cheaper tier
Anthropic released Claude Sonnet 5.5 this week. It is the first Sonnet model to carry the three-stage cyber classifiers and model fallbacks the company built for its most capable systems, putting classifier-driven routing in the tier many teams use for production.
Anthropic says Sonnet 5.5 does not advance the frontier of its models' capabilities, but it rates its cybersecurity skills as comparable to Opus 5. On Terminal-Bench 4.0, Sonnet 5.5 scores 70.6% against 66.4% for Opus 5.5. With cyber safeguards off, it reached full arbitrary code execution in 178 of 410 ExploitBench runs and completed 46.1% of challenges on Irregular's CyScenarioBench, up from 0.7% for Sonnet 5.
Three-stage enforcement
Filtering runs in three stages: a probe that reads the model's internal activations, a lightweight classifier running on Sonnet 5.5 itself, and a trained LLM classifier that weighs the probe's verdict before blocking a conversation. Anthropic says the classifiers catch harmful cyber requests at a rate comparable to Opus 5, but it chose less aggressive jailbreak protections because Sonnet 5.5 is weaker at cybersecurity than Opus 5 or Fable 5.1. Users should expect more refusals than with Sonnet 5, including on legitimate security work.
Blocked cyber requests, plus a narrow set tied to frontier LLM development such as kernel work on certain ML accelerators, are routed to Sonnet 5. Blocks for biology, conventional weapons and anti-distillation end the request with no fallback model at all, and Anthropic says they never covertly change responses.
API fallback is opt-in
Anthropic's own apps send blocked cyber requests to Sonnet 5 automatically, but API developers must enable the fallback themselves, and other platforms may handle blocked requests differently. Without fallback, a blocked request simply stops, so moving from Sonnet 5 to Sonnet 5.5 is not a straight model swap. The policy still allows vulnerability discovery in source code, keeping secure-coding workflows intact, but blocks it in compiled binaries. Support documentation says the checks review everything the model reads, including memory, connector content, web search results and files.
Fallback and prompt injection
In Anthropic's coding-environment tests, 25% of requests sent to Sonnet 5.5 were handled by Sonnet 5 after triggering a cyber block, often because injected instructions to wipe disks or delete files set off the classifier. Of those rerouted requests, 12.01% were compromised, while Sonnet 5.5 itself was compromised in only four of the 5,901 requests it handled.
Anthropic says it is still tuning Sonnet 5.5's classifiers to reduce false positives and plans to give verified defenders fewer restrictions through an expanded Cyber Verification Program. Sonnet 5.5 is not part of that programme at launch.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.