Back
Cloudflare's Auto Router picks a capable model for every AI request
SiTech AI Team3 წთ. საკითხავი

Cloudflare's Auto Router picks a capable model for every AI request

Cloudflare opened AI Gateway's Auto Router in public beta, reporting up to 30% internal cost savings, and added a model-overkill view to User Insights so teams can see where costly models are used.

Cloudflare has released its Auto Router in public beta through AI Gateway. The routing layer picks a model for every request instead of leaving the choice to the user, and setting the model to cloudflare/auto turns it on. Cloudflare says early internal use through its OpenCode harness cut spending by up to 30% versus relying only on frontier models such as OpenAI Sol and Anthropic Claude Opus.

What the benchmark shows

Cloudflare evaluated cloudflare/auto against Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Sol on an internal general knowledge-work benchmark built around workspace tools: email, calendars, Slack, files, travel and finance. Across 97 tasks with three samples each, the router solved 252 of 291 trials, an 86.6% success rate, at a total cost of $2.10, or $0.0084 per success. Claude Opus 5.5 reached 96.6% for $5.91, and GPT-6 Sol 84.2% for $2.64. In short, the router landed at about 80% of Sol's cost and 35% of Opus's.

How the routing decision is made

The gateway first narrows the pool of models that can actually serve a request, filtering by format, execution mode, credentials, billing, access control and spend limits, and dropping unhealthy providers during outages. A classifier running on Workers AI at the edge then reads the most recent turns of the conversation, sorts the task into one of 14 categories and scores it on complexity, ambiguity, stakes and dependence on earlier context. A scoring matrix combines those signals with benchmark results, and the router selects the model with the highest utility, defined as expected quality minus an adaptive cost penalty. On long agentic sessions it also weighs cache reads and writes, because switching models throws away a warm context that must then be paid for again.

Finding the spend you did not notice

Alongside the router, Cloudflare expanded User Insights, its dashboard for AI traffic. The new model-overkill view flags conversations where the selected model looks more capable than the task requires, and shows which users, agents or applications drive it. A Potential Savings view and task analysis covering categories such as coding, research, writing, summarization and data analysis help teams see whether a default model is applied too broadly. Classification runs asynchronously after the gateway answers, so it adds no latency, but analysis trails incoming traffic by about a day. The insights are free for AI Gateway users, and Cloudflare Access, which links traffic to authenticated users, is free for teams of up to 50.

Why it matters

Most organizations have settled on the tools they use daily, yet their AI bills grow faster than their output. Budgets and hard limits only go so far, because someone still has to decide, request by request, whether a summary really needs a frontier model. Cloudflare's bet is that the gateway sitting in the path of every request can make that call itself. The Auto Router is free while in beta, and Cloudflare says a cloudflare/auto-best variant that ignores the cost tradeoff is planned.

Sources: Cloudflare Blog (Auto Router) · Cloudflare Blog (User Insights)

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.