Back
GPT-5.6 Sol Pricing Cut by 50% on OpenRouter
SiTech AI Team2 წთ. საკითხავი

GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

OpenRouter now lists GPT-5.6 Sol with a 50% discount: $2 per million input tokens and $10 per million output tokens. The cut covers OpenAI's standard, Fast and Flex endpoints, while Azure and Bedrock keep their rates.

OpenRouter's model page for openai/gpt-5.6-sol now carries a 50% discount across OpenAI's own endpoints. The headline pricing on the page reads $2 per million input tokens and $10 per million output tokens — the standard OpenAI endpoint's rate with the discount applied.

Which endpoints are discounted

According to the provider table, OpenAI's standard endpoint listed $4 for input and $20 for output per million tokens before the cut, dropping to $2 and $10. OpenAI Fast, the lowest-latency tier, moves from $8/$40 to $4/$20, and OpenAI Flex falls from $2/$10 to $1/$5. The Azure and Amazon Bedrock rows show no discount: Azure charges $5/$30, Azure in the US and EU charge $5.50/$33, and Amazon Bedrock (US) charges $4.40/$22 per million tokens.

What customers actually pay

OpenRouter's weighted average tracks the price customers really pay rather than the posted rate. It currently stands at $0.6271 per million input tokens and $14.96 per million output tokens. Caching explains most of that gap: the OpenAI endpoint carries 83% of traffic on the page with a 91.5% cache hit rate, which pulls its effective input price down to $0.4625 per million tokens. After the discount, cache reads cost $0.20 per million tokens on the standard endpoint, $0.40 on Fast and $0.10 on Flex.

What the model offers

GPT-5.6 Sol is the flagship of OpenAI's GPT-5.6 series. OpenRouter describes it as suited to complex reasoning, coding and agentic workflows, and particularly strong at command-line and multi-step coding tasks and long-horizon problem solving. It has a 1.1 million token context window and was released on 9 July 2026. Requests can be routed in three modes: Balanced, which weighs price and speed; Nitro, which prioritises the fastest provider; and Exacto, which targets the highest tool-calling accuracy.

Why it matters

A 50% cut lands at a moment when long-running agent workloads make token spend the dominant line item for teams building on hosted models. The gap between the fast and standard endpoints is now a factor of two, though Fast also offers the lowest latency on the page at 2.62 seconds and the highest throughput at 70 tokens per second. Flex is the cheapest option but comes with 13.14-second latency and 88.21% uptime, while the standard OpenAI endpoint reports 99.43% uptime.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.