
Cloudflare releases Clef: open-source decision models and an RL fine-tuning platform
Cloudflare has published Clef and Clef-flash, decision models that return typed probabilities instead of free text. The weights are on Hugging Face under Apache 2.0, and a new RL service fine-tunes them for customers.
Cloudflare has released Clef and Clef-flash, its first in-house trained decision models, hosted on Workers AI and published on Hugging Face under Apache 2.0. A new RL service lets customers fine-tune Clef for their own workloads, at first with the company's engineers and later self-serve.
Cloudflare says Clef currently leads the Jev Decision Index and that both models are fully API-compatible with Jev, the decision model from Typesafe AI that popularized the concept.
What a decision model does
A decision model turns inputs into typed, bounded outputs with probabilities attached. A support message, for instance, can be scored as urgent or not and routed to the right team, and the code around it can escalate or defer to a human. Unlike large language models, decision models aim to be cheap, fast and consistent, so agents can act without a human in the loop.
Cloudflare tested Clef inside its Threat Intelligence team, classifying website domains with Browser Run. The model fetched, rendered and classified a domain in 2.2 seconds, assigning probabilities such as 95% fashion site, 85% ecommerce and under 1% phishing. Its fastest general-purpose model, gpt-oss-120b, took 4.7 seconds on the same workflow and produced only two classifications.
How Clef compares
Clef has two features Jev lacks: a vision encoder for images as well as text, and a 64k-token context window against Jev's 32k. It scored 98.47 on the Jev Decision Index's BFCL exact-case benchmark, with Clef-flash at 98.76 and Jev at 95.75. Median latency was 209 ms for Clef, 39 ms for Clef-flash and 524 ms for Jev.
Both models use Qwen backbones, frozen at 27 billion parameters for Clef (Qwen3.8-27B) and 9 billion for Clef-flash (Qwen3.5-9B). Inference is a prefill-only pass followed by a non-autoregressive step that scores valid schema choices in parallel instead of generating text token by token.
A new RL fine-tuning service
The service is assembled from existing Cloudflare products. AI Gateway logs a customer's own requests and responses, Workers AI generates rollouts against the base model, Containers supply an RL sandbox for scoring agent actions, and a new Trainer updates the weights. Cloudflare says more than 15 years of network data make specialized classifiers practical for jobs such as support triage and Trust and Safety reviews.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.