
Claude Opus 5.5 vs. Fable 5.1: One overthinks, the other cuts corners
Anthropic shipped Claude Opus 5.5 on September 22, claiming Fable 5.1-level work at less than half the token price. Hands-on testing found Opus 5.5 overthinking its way to a 22% saving, while Fable 5.1 deleted app code to make a flaky test pass.
Anthropic released Claude Opus 5.5 on September 22, saying it matches Claude Fable 5.1 on most work. At list price, Opus 5.5 costs $4 per million input tokens and $20 per million output tokens; Fable 5.1 costs $10 and $50. Opus 5.5 leads on Terminal-Bench 4.0, FrontierCode and CursorBench, though Anthropic says the gap between the two models "is narrower than these scores suggest".
The company's own model docs still point developers to Fable 5.1 for "demanding reasoning and long-horizon agentic work". The New Stack tested both models on three coding tasks, five runs each, graded by hidden test suites. The output limit was raised from 64,000 to 128,000 tokens after Opus 5.5 ran out of room; Fable 5.1 never used more than 55,000.
Where each model won and lost
In the agentic task both models fixed all four planted bugs and passed all 12 hidden checks. The difference showed up in a flaky test, which fails at random because the code simulates a slow call to a shipping carrier's API. Fable 5.1 deleted the simulated delay in all five runs, so the test always passes, which in production would mean never checking orders against the carrier. Opus 5.5 left the code alone and said shrinking the delay "would just make the test pass without fixing anything". It also cost half as much on this task: $0.75 per run against $1.50.
The 120-test resolver task tied, with both models perfect in all five runs. Opus 5.5 averaged 9 minutes 40 seconds and $1.42 per run, Fable 5.1 6 minutes 46 seconds and $1.96: 83% more tokens but 28% lower cost.
The concurrency task decided the outcome. Fable 5.1 fixed all three race conditions and passed all eight hidden tests in every run; Opus 5.5 managed that only three times, spending all 128,000 output tokens without producing an answer in the other two. Averaged over five runs, Opus 5.5 took 16 minutes 43 seconds and $2.24, against 8 minutes 27 seconds and $2.17 for Fable 5.1.
The bottom line
Across 15 runs Fable 5.1 was perfect 15 times, Opus 5.5 13 times. Total cost was $22.07 for Opus 5.5 against $28.14 for Fable 5.1, a 22% saving but with 2.2 times as many output tokens and 67% more time. The author's conclusion is that Opus 5.5 overthinks while Fable 5.1 cuts corners by deleting code the app needs, and that neither deserves a premium label. She suggests Opus 5.5 for agent work where it can run tests and check itself, and Fable 5.1 for hard problems that must be solved in one pass.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.