Dashboard

Cloudflare Clef: Open Decision Models Explained

Cloudflare Clef and Clef-flash are open-weight decision models that return typed probabilities, not text. The specs, pricing and where they fit.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
2 October 20261 min read

Cloudflare Clef is a model that does not write anything. Released on 1 October 2026 alongside a smaller sibling, Clef-flash, the pair read an input and a schema of typed questions, then return a probability for every allowed answer. No prose, no JSON to parse, no apology when they are unsure. Cloudflare calls them decision models, and the weights are on Hugging Face under Apache 2.0.

That is a narrower product than it first sounds, and more useful than it first sounds, because routing and classification are what most production AI systems actually spend their tokens on.

What was announced

Two models, both trained in-house at Cloudflare on top of Alibaba's Qwen family:

Clef

Clef-flash

Parameters

27 billion

9 billion

Base model

Qwen3.8-27B

Qwen3.5-9B

Median latency

209 ms

39 ms

Workers AI input price

$0.24 per million tokens

$0.09 per million tokens

Context window

64k tokens

64k tokens

Images per request

up to 4

up to 4

Output tokens are not charged on Workers AI, which makes sense given that the output is a probability distribution over a fixed answer set rather than generated text. Both models keep Qwen's vision encoder, so they can classify an image as well as text.

On accuracy, Cloudflare reports a macro-F1 of 94.20 on BANKING77 for Clef against 79.74 for the hosted competitor it benchmarks against, and cites scores on BFCL and API-Bank in the same range. Treat vendor benchmarks as vendor benchmarks: they tell you the shape of the claim, not what your data will do.

Why a model that cannot talk is interesting

Most teams wire a general chat model into the places where a decision needs making. Which department does this ticket belong to. Is this refund request inside policy. Should this task go to the cheap path or the expensive one. The chat model does it, and it mostly works, and three things go wrong quietly.

It is slow, because you are paying generation latency for a one-word answer. It is expensive, because frontier pricing applies to a problem a much smaller model could solve. And it is loose, because the answer comes back as text that you then have to parse, and occasionally the text is "I'd say this is probably Billing, though it could also relate to Technical Support."

A decision model closes all three. The answer set is fixed by the schema, so there is nothing to parse and nothing out of range. The probability comes back as a number, so you can set a threshold instead of reading hedging language. And at 39 ms, Clef-flash is fast enough to sit in a request path rather than a background job.

That last point is the one worth sitting with. A 39 ms median means routing can happen inline, before the user notices. At 209 ms, Clef is still inside the budget most teams allow for a page load. Neither is a batch job.

Where Cloudflare Clef fits against what else exists

OpenAI previewed a hosted Decisions API at DevDay on 29 September, aimed at the same problem from the opposite direction: a managed endpoint rather than weights you run. The trade is the familiar one between open-weight and closed models. Clef's weights are downloadable, so you can run them on your own hardware, air-gapped if you need to, with no per-call dependency on anyone's uptime. Whether that is cheaper depends entirely on your volume, and the arithmetic is rarely flattering below a certain scale, as the self-hosting cost comparison works through.

The useful mental model is that this is a third tier. Not the frontier model and the cheap model that most cost discussions stop at, but a category below both, for the large fraction of calls that are classifications wearing a chat interface.

What to do with this if you are building something

Look at your logs for calls where the output is always one of a small set of values. Those are your candidates. For each one, the question is whether you need language at all, or just a label and a confidence number.

If it is a label, a decision model is the right shape, and so is a plain classifier you train yourself, and so is a rule, depending on how much structure the problem has. Having a probability back rather than a sentence makes the next part easier: you can set a confidence threshold and send everything below it to a human, which is the pattern that keeps automated routing from quietly making bad calls at scale.

What this release does not do is change anything about generation. If your task is writing, summarising, or reasoning through something open-ended, none of this applies. The news here is a narrowing, not a replacement.

FAQ

Is Clef free to use?

The weights are free under Apache 2.0 and you can download and run them yourself. Running them on Cloudflare's Workers AI is paid, at $0.24 per million input tokens for Clef and $0.09 for Clef-flash, with output tokens not charged.

Can Clef replace a chat model in my app?

Only for the parts that are decisions. It returns probabilities over a fixed answer set and cannot produce free-form text, so anything that needs a written answer still needs a generative model.

What does Clef being built on Qwen mean in practice?

Cloudflare started from Qwen3.8-27B and Qwen3.5-9B and trained from there, keeping the vision encoder. That inheritance is why Clef can classify images, and it means the base model's licence and lineage are worth checking if you have procurement rules about model provenance.

How does a decision model handle something outside its schema?

It cannot return an answer outside the allowed set, so an unfamiliar input spreads probability across the options it has rather than inventing a new one. A flat or low-confidence distribution is the signal that the input does not fit, which is why thresholding matters more here than with a chat model.

Should I switch my routing to this now?

Benchmark it on your own traffic first. Vendor numbers on BANKING77 say little about your ticket categories, and the honest test is replaying a few thousand real decisions and comparing against whatever you run today, on both accuracy and cost.

Sources: Cloudflare blog, Crypto Briefing

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.