Back
Cloudflare releases Clef: open-source decision models and an RL fine-tuning platform
SiTech AI Team2 min read

Cloudflare releases Clef: open-source decision models and an RL fine-tuning platform

Cloudflare has published Clef and Clef-flash, decision models that return typed probabilities instead of free text. The weights are on Hugging Face under Apache 2.0, and a new RL service fine-tunes them for customers.

Cloudflare has released Clef and Clef-flash, its first in-house trained decision models, hosted on Workers AI and published on Hugging Face under Apache 2.0. A new RL service lets customers fine-tune Clef for their own workloads, at first with the company's engineers and later self-serve.

Cloudflare says Clef currently leads the Jev Decision Index and that both models are fully API-compatible with Jev, the decision model from Typesafe AI that popularized the concept.

What a decision model does

A decision model turns inputs into typed, bounded outputs with probabilities attached. A support message, for instance, can be scored as urgent or not and routed to the right team, and the code around it can escalate or defer to a human. Unlike large language models, decision models aim to be cheap, fast and consistent, so agents can act without a human in the loop.

Cloudflare tested Clef inside its Threat Intelligence team, classifying website domains with Browser Run. The model fetched, rendered and classified a domain in 2.2 seconds, assigning probabilities such as 95% fashion site, 85% ecommerce and under 1% phishing. Its fastest general-purpose model, gpt-oss-120b, took 4.7 seconds on the same workflow and produced only two classifications.

Example: a decision model routes a support ticket and flags urgency

How Clef compares

Clef has two features Jev lacks: a vision encoder for images as well as text, and a 64k-token context window against Jev's 32k. It scored 98.47 on the Jev Decision Index's BFCL exact-case benchmark, with Clef-flash at 98.76 and Jev at 95.75. Median latency was 209 ms for Clef, 39 ms for Clef-flash and 524 ms for Jev.

Both models use Qwen backbones, frozen at 27 billion parameters for Clef (Qwen3.8-27B) and 9 billion for Clef-flash (Qwen3.5-9B). Inference is a prefill-only pass followed by a non-autoregressive step that scores valid schema choices in parallel instead of generating text token by token.

The fine-tuning pipeline: capturing traffic, preparing tasks, and training

A new RL fine-tuning service

The service is assembled from existing Cloudflare products. AI Gateway logs a customer's own requests and responses, Workers AI generates rollouts against the base model, Containers supply an RL sandbox for scoring agent actions, and a new Trainer updates the weights. Cloudflare says more than 15 years of network data make specialized classifiers practical for jobs such as support triage and Trust and Safety reviews.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.