
Claude Opus 5.5 vs. Opus 5 on reasoning tasks: cheaper and faster, but not better
Anthropic says Claude Opus 5.5 is 40% cheaper and 30% faster than Opus 5. The New Stack ran three reasoning problems against both models: the savings held up, the speed claim did not, and the answers were identical.
The New Stack compared Anthropic's two models, Claude Opus 5.5 and Claude Opus 5, on three hard reasoning problems. The result is simple: the newer model is cheaper and quicker, but not smarter. Both solved two of the tasks outright, and both failed the third.
What Anthropic claims
Anthropic says Opus 5.5 costs 40% less than Opus 5 on typical workloads and generates output more than 30% faster, while performing at the level of Claude Fable 5.1, ahead of the previous model. The API price also fell: $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5. That cut alone accounts for a 20% saving, and the rest of the claimed 40% has to come from the model using fewer tokens.
The logic grid: identical answers, lower cost
Both models were called through the Anthropic API with identical prompts, running adaptive thinking at the default effort level. Opus 5.5 does not allow thinking to be turned off. On a logic grid with seven engineers and 22 clues, both filled all 28 cells correctly. Opus 5.5 needed 65 seconds, 7,573 output tokens and $0.16; Opus 5 needed 108 seconds, 10,621 output tokens and $0.27, making the new model 43% cheaper for the same answer.
The stone game: same result, far fewer tokens
The stone game with memory ended 3 out of 3 correct for both models. The first player loses from 200 stones, sizes from 120 to 500 are losing ones, and the smallest losing size above 340 is 344. The gap was in thinking volume: Opus 5.5 finished in 215 seconds with 28,740 output tokens for $0.58, while Opus 5 took 624 seconds, 74,981 tokens and $1.88. That is 62% fewer tokens and 69% less cost for the same answer.
The problem neither model could solve
Neither model answered the constrained-ordering problem, where the correct counts are 27, 1,695 and 159,019. With a 48,000-token output limit both spent the whole budget thinking and never replied. After the limit was raised to 128,000, Opus 5 burned every token, ran for more than 25 minutes and stopped with no answer. Opus 5.5 ended at 112,733 tokens after 19 minutes, but the API returned a refusal stop reason with no text. The prompt contains nothing sensitive, so the writer treats it as a mistake by Anthropic's safety filter.
What it means in practice
Across the whole test Opus 5.5 used 1,701 input and 197,046 output tokens; Opus 5 used 1,693 and 261,602. Cost came to $3.95 against $6.55. Output speed was 103.4 tokens per second versus 93.1, about 11% faster and short of the 30% Anthropic claims. The recommendation is straightforward: if you run Opus 5 today, moving to Opus 5.5 gives the same results for less money and less waiting. Set a hard output limit, since both models can think for 20 minutes and return nothing, and give the model a code execution tool for counting problems.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.