
Cloudflare adds Clef-omni, lowers Clef-flash price and speeds up Clef
Cloudflare released Clef-omni, an open-weight decision model that accepts audio, video, images and text in one call. Clef-flash is now cheaper, with a shorter hosted context window, and hosted Clef is faster after serving optimizations.
Clef-omni adds audio and video
Cloudflare is introducing Clef-omni, an open-weight decision model that accepts wav or mp3 audio, mp4 or webm video, images and text in a single request. The company says this moves decision models beyond the text-only pattern that followed the debut of Jev from TypeSafe, and removes the need to chain speech-to-text or media splitting pipelines before a decision is made. The weights are available on Hugging Face.
Cloudflare says Clef-omni is built on a Qwen3-Omni-30B-A3B-Instruct mixture of experts foundation. It keeps the primary comprehension backbone and discards the text-to-speech output components. In production, the model runs a quick prefill pass across the complete payload, scores all modalities and valid parameter options at once, and skips output token generation because Clef models are not large language models. Audio and video are mapped into a unified sequence with visual frames for joint processing, then candidate values are pulled from internal embeddings through two stage attention routing with a lexical grammar for schema constrained scoring.
Pricing and context window change
Clef-flash is now listed at $0.038 per million input tokens, down from $0.09. Cloudflare says optimizations make Clef-flash cheaper than Jev while remaining performant. Clef stays at $0.24 per million input tokens, and Clef-omni launches at $0.15 per million input tokens. The hosted Clef-flash context window is now 24k instead of the previously advertised 64k. Cloudflare says Hugging Face weights are untouched and were trained to support 256k for self hosting, while usage data showed only 0.24% of requests exceed 24k input tokens. Clef remains at 64k.
Faster serving and reported results
Cloudflare reports faster hosted Clef inference on Workers AI without new model weights, crediting serving infrastructure work such as the move to SGLang. The company worked with the SGLang team on pull requests to support Clef, including PR #42721, slated for SGLang 0.5.22. Reported median and p95 latency moved from 262 and 438 ms to 152 and 351 ms at about 800 tokens, from 616 and 777 ms to 305 and 531 ms at about 3,400 tokens, and from 2,721 and 3,250 ms to 1,635 and 1,805 ms at about 16,000 tokens.
Reported Clef-omni examples include 98.2 on BFCL case exact and 73.2 on PhishNChips accuracy, while Cloudflare notes results vary across benchmarks. Internal use cases listed by Cloudflare include spam issue detection in its public GitHub docs repository, phishing moderation for plugin libraries in EmDash, data loss prevention scans for government IDs and other PII, and malicious domain detection in threat intelligence. Clef is described as fully Jev API compatible and available through AI Gateway by changing the model ID.
Sources: The Cloudflare Blog
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.