Cloudflare Clef: when decision models beat LLM classifiers
Use it for typed, repeatable routing decisions—not open-ended reasoning.
Short answerUse Clef for fast typed decisions; keep general LLMs for open-ended reasoning or generation.
By JasonPublished Oct 3, 2026Last verified Oct 3, 20265 min read

Small teams keep running into the same architecture question: should a workflow ask a general LLM to classify something, build a custom classifier, or use a newer “decision model” layer? Cloudflare’s Clef announcement is aimed directly at that gap. According to Cloudflare Blog, Clef and Clef-flash are hosted decision models on Workers AI that return typed answers with probabilities, rather than free-form prose. That matters when the output feeds code: routing a support ticket, assigning severity, deciding whether to escalate, or classifying a domain. The practical question is not whether Clef is broadly “better” than an LLM. It is narrower: when is the decision bounded enough that a structured model is safer and cheaper to operate than prompting a general model, but fluid enough that maintaining a hand-built classifier or retraining for every new label becomes annoying? The supplied source gives useful signals on latency, context length, API compatibility with Jev, open-source licensing, and Cloudflare’s planned fine-tuning path. It does not give pricing details in the excerpt, and all performance claims come from Cloudflare’s own announcement, so the decision should stay conditional.
What Cloudflare announced
Cloudflare Blog announced two decision models, Clef and Clef-flash, hosted on Workers AI. The company says both are also being open-sourced on Hugging Face under an Apache 2.0 license. Clef is positioned as the larger precision-oriented model, while Clef-flash is framed as the lower-latency option.
The key product idea is not another chat model. Cloudflare describes a decision model as a model that accepts inputs plus a set of questions, then returns typed answers with probabilities. In the example from the announcement, a support message can be classified for urgency, team ownership, and severity. That makes the output easier to wire into application logic than free-form model text.

Where Clef fits
For a small team, Clef makes the most sense when the decision space is bounded but still changes often enough that maintaining a traditional classifier is a drag. Examples from Cloudflare’s framing include support routing, escalation decisions, and domain classification.
The supplied announcement says Clef can handle images through a vision encoder and has a 64k context window. Cloudflare contrasts that with Jev, which it says currently handles text classification and has a 32k context window. The announcement also says Clef is Jev API-compatible, which matters if a team is already evaluating Jev-style decision-model workflows.
The latency argument
Cloudflare’s strongest claim is speed. In its benchmark table, Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, compared with 524.1 ms for Jev, 84.4 ms for DiffusionGemma, 51.4 ms for Kev-9B, and 5.8 ms for Laya. Its p95 latency table reports 238.6 ms for Clef and 122.4 ms for Clef-flash.
Cloudflare also gives a concrete internal workflow example: using Clef with Browser Run to classify website domains. In that case, Cloudflare says Clef took 2.2 seconds to fetch, render, and classify a site, while gpt-oss-120b took 4.7 seconds in the same workflow and returned fewer classifications.
Those numbers are useful, but they are still Cloudflare’s own figures. Treat them as a reason to evaluate Clef, not as proof that it will be faster in your workload.
Clef versus general LLM classification
Use a general LLM when the system needs open-ended reasoning, long-form explanation, generation, or tool-use planning. Use Clef when the application already knows the question and needs a structured answer that code can consume.
That is the cleanest distinction in Cloudflare’s announcement. Clef is not presented as a replacement for LLMs. Cloudflare’s own suggested architecture is to put Clef in the decision path and combine it with one of its LLMs on Workers AI when an action or generated response is needed.
Clef versus custom fine-tuning
The announcement also introduces an RL fine-tuning product for Clef. Cloudflare says it is first offering fine-tuning through its forward-deployed engineer team, with a later self-serve platform for customers to train and redeploy models on Cloudflare.
For small teams, that means the default path should probably be: start with hosted Clef or Clef-flash if the schema fits, consider fine-tuning only when the off-the-shelf model misses domain-specific distinctions, and avoid custom work until there is enough real traffic to know what is failing.
What is missing
The supplied text does not include pricing. It also does not include independent benchmark confirmation. Cloudflare says the models are enterprise-ready and that it does not read, store, or train on requests or responses unless customers use fine-tuning, but teams with compliance constraints should verify those terms directly before routing sensitive data through the service.
The short version: Clef is worth evaluating for hot-path, typed decisions. It is not yet a reason to replace every LLM classifier or build process around vendor benchmarks alone.
Clef is interesting because it narrows the job description. A general LLM is useful when the system needs to reason, write, call tools, or handle open-ended context. A decision model is more compelling when the application already knows the question and needs a typed answer that software can act on. Cloudflare’s own framing points to a practical split: use Clef for bounded decisions in an agent workflow, then use an LLM elsewhere for the messy parts. The strongest numbers in the announcement are latency-oriented: Cloudflare says Clef had 209.3 ms median latency across its evals, while Clef-flash had 38.8 ms, compared with 524.1 ms for Jev in the same table. It also says a domain-classification workflow took 2.2 seconds with Clef versus 4.7 seconds with gpt-oss-120b. But these are vendor-provided figures. Without independent tests or pricing in the supplied text, this should not be treated as a blanket migration recommendation.
Use Clef for fast typed decisions; keep general LLMs for open-ended reasoning or generation.
Cloudflare’s announcement makes Clef look most useful as a decision layer inside workflows: classify, score, route, or escalate with typed outputs. The case is weaker for teams that need generation, broad reasoning, or proven economics, because the supplied source is a single vendor announcement and does not include pricing.
Skip it for now if your classification workload is low-volume, pricing is the deciding factor, or you need independent benchmarks before adopting a new model path.
Read next
Follow new articles
Email updates are not live yet, and we are not collecting addresses. To follow new articles, use the RSS feed.