
GPT-6 Sol vs. Claude Opus 5.5: cheaper only when the results repeat
The New Stack ran OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5 through three hard developer tests, five times each. Sol came out about six times cheaper, but missed three of 15 runs while Opus 5.5 was perfect every time.
OpenAI released GPT-6 Sol on September 22, 2026, claiming it beats Claude Opus 5 on business workflow tasks at 9% of the per-task cost: 33.2% on Zapier's AutomationBench against 26.9%. Anthropic shipped Claude Opus 5.5 the same day, so The New Stack's Jessica Wachtel tested that model instead.
List pricing puts Sol at $2 per million input tokens and $10 per million output tokens, half of GPT-5.6 Sol; Opus 5.5 costs $4 and $20. Even with a heavier token bill, Sol could still come out cheaper.
Three tests, five runs each
The first round ended in a tie: both models fixed three planted bugs in a Python billing service, triaged 40 failed CI jobs and answered 20 questions about a fake SDK without inventing a method. Harder variants, a 3,664-line outage postmortem and a spec graded by 120 hidden tests, also ended level. Only repeats separated them.
Each model was called through its API with identical prompts at its highest effort setting, five times per test. CI triage decides retry, block or page for 40 failed jobs; incident logs asks seven questions about 3,664 lines from five services during a two-hour outage; the resolver spec asks for a dependency resolver written from a two-page spec without running code.
What the numbers show
CI triage, the closest match to OpenAI's claim, went 5 for 5 for both. Opus 5.5 averaged 1 minute 27 seconds, 11,127 output tokens and $0.24 per run; Sol averaged 18 seconds, 1,143 tokens and $0.02, about 8% of the cost and in line with the 9% OpenAI advertises.
Incident logs changed the picture: Opus 5.5 was perfect on all five runs, Sol on two. It twice missed the same customer, whose charge was confirmed 52 seconds after a successful retry, and once counted 27 failed checkouts instead of 28. Opus 5.5 averaged 6 minutes 44 seconds, 53,308 output tokens and $1.68 per run; Sol averaged 1 minute 41 seconds, 7,073 tokens and $0.30, reading the same log in 25% fewer input tokens, 113,966 against 152,345.
On the resolver spec Opus 5.5 passed every run and Sol four of five: a stray closing parenthesis on line 78 crashed the module on import and failed all 120 tests. Opus 5.5 averaged 9 minutes 40 seconds, 70,687 output tokens and $1.42; Sol averaged 6 minutes 28 seconds, 21,435 tokens and $0.22.
Cheaper, not consistently right
Across 15 runs Opus 5.5 was perfect every time and Sol 12 times. Sol's 15 runs cost $2.68, about 16% of Opus 5.5's $16.72. Wachtel splits by stakes: run Sol on high-volume work where a person or a test suite checks the output, and at those prices you can run it twice and compare; use Opus 5.5 where a wrong answer is expensive and nobody is checking. Sol's misses slip past review, a customer left off a refund list or a file that will not import.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.