
Claude Opus 5.5 wants to finish your coding tasks, not just start them
Anthropic's Claude Opus 5.5 is pitched as an agent that carries a coding task across the development lifecycle, from spec to verified fix. Practitioners warn that “completed” and “correct” are not the same thing.
Claude Opus 5.5, the first release in Anthropic's new Claude 5.5 family, is positioned less as a faster code autocomplete than as an agent that carries a task through the whole development lifecycle: from a design specification and debugging to code generation and testing. The New Stack reports the pitch is not about starting work faster, but about finishing it.
An agent that closes the loop
The shift developers describe is architectural. Sergey Ermakovich, co-founder of HasData, argues the terminal “is no longer just a place where a developer pastes generated code — the terminal now becomes part of the model's workspace.” Because a model can run commands, inspect failures, modify files and verify the result, it closes the loop instead of handing unfinished work back. That, he says, makes small complete tasks more valuable: fix one failing test, update one dependency, verify it, then move on.
Migrations, audits and price
Anthropic claims Opus 5.5 is particularly strong on long, sprawling jobs such as codebase-wide migrations and audits. An early tester reportedly audited and fixed a 200,000-line codebase in under three hours, where Opus 5 needed over 20 hours and 2.5 times as many tokens. In an internal test, Opus 5.5 and Fable 5.1 rewrote HAProxy from C into Rust and passed nearly all of its regression tests, with Opus 5.5 finishing in 9.5 hours against 12, at 51% lower cost. Pricing is $4 per million input tokens and $20 per million output tokens, 20% below Opus 5, with cache reads cut 60%.
“Completed” is not “correct”
Practitioners The New Stack spoke to welcome the capability but warn that verification, not generation, is now the bottleneck. Independent SRE and AI reliability architect Akash Thakur says models at this level are good at breaking a project into finishable pieces, yet “completed” and “correct” are not the same thing — the task that looks done often costs a team later, so the win is moving people from writing code to verifying it. Maxime Vermeir, vice president of AI strategy at Abbyy, says the autocomplete era is over and the expectation is now a whole Jira ticket.
Safeguards and model routing
Anthropic says Opus 5.5 was tested before release by outside organisations including Frontier Design and METR, and is the strongest model on its most comprehensive alignment test to date, with gains in behaviours linked to recent cybersecurity incidents. When safeguards intervene, requests fall back to another model transparently: most cybersecurity tasks route to Opus 4.8, and biology or frontier-LLM-development requests route to Opus 5.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.