Skip to content

Cloudflare Clef API Tutorial: Build Fast AI Decision Gates

Learn how to use Cloudflare Clef with Workers AI to build fast decision gates for routing, escalation, classification, and AI-agent guardrails.

Cloudflare Clef API Tutorial: Build Fast AI Decision Gates

On this page

Cloudflare released Clef and Clef-flash on October 1, 2026, but these models are not designed to write chat responses. They take application state plus typed questions and return probabilities for the answers your code has already defined. That makes them useful for the small decisions sitting inside larger AI workflows: route a support ticket, flag a submission, ask for human review, or decide whether an agent should call a tool.

This Cloudflare Clef API tutorial shows how to put that decision layer inside a Worker using the Workers AI binding, define noul, choice, and score questions, interpret the returned probabilities, and add a simple application threshold without asking a general-purpose language model to generate text first.

Why Clef is different from a normal AI model

A conventional large language model is usually asked to produce something open-ended: an answer, a paragraph, a piece of code, or a tool call. Clef works in the opposite direction. You give it a bounded set of decisions and the information needed to make them, and it evaluates those decisions without generating a free-form response that your application has to parse. Cloudflare describes Clef as a 27-billion-parameter multimodal decision model, while Clef-flash is the smaller 9-billion-parameter version intended for lower-latency decisions.

That distinction matters when the output is going straight into application logic. Suppose an agent receives a customer message. The useful result might simply be billing, technical, or sales, plus the probability assigned to each option. Sending the same job through a general chat model can produce a much larger response than the application actually needs. Clef instead gives your program structured answers that can become the next branch in the workflow.

Cloudflare currently exposes Clef through Workers AI as @cf/cloudflare/clef and Clef-flash as @cf/cloudflare/clef-flash. The model documentation supports text and structured state, while Clef can also process images and video-related state; requests can contain up to 64 questions and up to four embedded images.

Choose Clef or Clef-flash before you build

The two models are aimed at different points in the same workflow. Cloudflare lists Clef at $0.24 per million input tokens and Clef-flash at $0.09 per million input tokens. The published model information gives both a 65,536-token context window, so the practical difference is primarily the model size, decision quality characteristics, latency, and price rather than a smaller context limit on the Flash model.

ModelParametersInput pricePublished median latencyBest fit
Clef27B$0.24 / million input tokens209.3 msMore demanding decisions
Clef-flash9B$0.09 / million input tokens38.8 msLatency-sensitive decisions

The latency figures above come from Cloudflare's own 43 benchmark runs, so they should be treated as vendor measurements rather than a promise for every deployment. Cloudflare reports median latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, compared with 524.1 milliseconds for Jev in its published comparison. The benchmark also shows that the smaller model does not lose every task: Clef-flash leads Clef on some individual benchmark rows.

Create the Worker and connect Workers AI

The simplest implementation keeps the decision model inside a Cloudflare Worker. You need a Cloudflare account, Node.js, and Wrangler, Cloudflare's command-line development and deployment tool. Cloudflare's current Worker setup uses the Create Cloudflare CLI to generate a Worker project and then adds an AI binding to the Wrangler configuration.

npm create cloudflare@latest -- clef-decision-gate cd clef-decision-gate

During setup, choose a Worker-only project and TypeScript. After the project exists, add the Workers AI binding to wrangler.jsonc. The binding name becomes the env.AI object available to your Worker code.

{ "$schema": "./node_modules/wrangler/config-schema.json", "name": "clef-decision-gate", "main": "src/index.ts", "compatibility_date": "2026-03-27", "ai": { "binding": "AI" } }

If your project already contains a Wrangler configuration, do not replace its existing settings just to add the AI binding. Add the ai section alongside the configuration you already use. Workers AI permits one AI binding per Worker project, and the binding is accessed as env.AI at runtime.

Define the decision questions Clef should answer

Clef has three useful question types. noul is a yes-or-no decision, choice selects from options you define, and score evaluates an ordered scale. The important design decision happens here: your application defines the allowed outputs instead of asking the model to invent them.

For a support workflow, you might ask whether a ticket is urgent, which team should receive it, and how severe the customer impact is. A choice question might contain billing, technical, and sales; a score question could contain four ordered levels from no impact to critical.

const response = await env.AI.run("@cf/cloudflare/clef", { model: "clef", state: { ticket: "The export button crashes the settings page in Safari. " + "It works in Chrome, but several customers use Safari." }, questions: { urgent: { type: "noul", instructions: "Is this support request urgent?" }, team: { type: "choice", instructions: "Which team should handle this request?", criteria: { billing: "Payments, invoices, and refunds", technical: "Bugs, outages, and configuration", sales: "Plans and upgrades" } }, severity: { type: "score", instructions: "How severe is the customer impact?", criteria: [ "No impact", "Minor", "Major", "Critical" ] } } }); return Response.json(response);

The model receives the ticket as state and evaluates all three questions in the same request. The result contains answers under the same question names, which means your application can address response.answers.urgent, response.answers.team, and response.answers.severity without trying to extract information from generated prose. Cloudflare's model schema allows between one and 64 questions in a request.

Turn the model response into an application decision

The probability returned by Clef is evidence for your application logic, not a command that your application must blindly obey. For example, a support system might only auto-route a ticket when the selected team has a probability above a threshold; otherwise it can send the case to a human queue. That threshold is your product rule, not something Clef automatically decides for you.

This distinction becomes especially useful when building reliable AI agents. The model can answer a narrow question such as whether a tool call should proceed, while the surrounding application retains responsibility for authentication, permissions, rate limits, and other hard rules.

const teamAnswer = response.answers.team; const urgentAnswer = response.answers.urgent; const team = teamAnswer.choice; const teamConfidence = teamAnswer.confidence; const urgentProbability = urgentAnswer.noul; let action = "human_review"; if (teamConfidence >= 0.85 && urgentProbability >= 0.80) { action = "priority_route"; } else if (teamConfidence >= 0.85) { action = "normal_route"; } return Response.json({ action, team, confidence: teamConfidence, urgentProbability });

The exact threshold in this example is deliberately an application choice rather than a claimed Clef requirement. In production, test your own traffic and decide what error rate is acceptable before allowing a probability threshold to trigger an irreversible action. A low-confidence classification should have a safe fallback instead of becoming an automatic approval, deletion, block, or payment.

Make the Worker accept real application state

Once the fixed example works, the next step is to stop hard-coding the ticket. Your Worker can accept structured JSON from the rest of your application, validate the incoming fields, pass the state to Clef, and return only the decision data your frontend or agent needs.

export interface Env { AI: Ai; } export default { async fetch(request, env): Promise { if (request.method!== "POST") { return Response.json( { error: "POST is required" }, { status: 405 } ); } const input = await request.json(); if (!input.state) { return Response.json( { error: "state is required" }, { status: 400 } ); } const response = await env.AI.run("@cf/cloudflare/clef-flash", { model: "clef-flash", state: input.state, questions: { urgent: { type: "noul", instructions: "Is this request urgent?" }, route: { type: "choice", instructions: "Which department should handle this request?", criteria: { billing: "Payments and invoices", technical: "Technical problems and outages", sales: "Purchases, upgrades, and plans" } } } }); return Response.json(response); } } satisfies ExportedHandler;

Clef-flash is a useful choice for this kind of hot-path Worker because Cloudflare positions it specifically as the lower-latency member of the pair. If your tests show that its decisions are not accurate enough for a particular workflow, changing the model identifier back to Clef does not require redesigning the question schema.

Test the decision gate locally before deployment

Run the Worker with Wrangler after adding the AI binding. Cloudflare notes that local Workers AI development still accesses your Cloudflare account and therefore incurs Workers AI usage charges; local development is not an offline simulation of the model.

npx wrangler dev

Wrangler will provide a local address for the Worker. Send a test request containing a realistic support ticket and inspect the returned answers. Do not test only obvious examples such as clearly urgent outages. Include ambiguous cases, contradictory information, short messages, long messages, and cases where none of your predefined choices is a comfortable fit.

For an agent workflow, it is also worth keeping a small evaluation set outside production. Record the expected route, severity, or escalation decision for each example and compare Clef's output against that set whenever you change the question wording, model version, or threshold. This gives you a way to detect a regression instead of discovering it through a customer ticket.

Know where Clef should not replace a language model

Clef is useful precisely because it does less than a general language model. It is not the component you want when the application needs a long explanation, creative generation, code generation, open-ended reasoning, or a response whose possible structure cannot be defined ahead of time. Those jobs still belong to a generative model.

A useful architecture is therefore a split rather than a replacement. A generative model can gather context or produce the action, while Clef sits at a decision point and answers a bounded question before the action is allowed. That makes the decision layer easier to inspect because the permitted outputs are visible in the application code rather than hidden inside a prompt.

Cloudflare's own examples include support triage, threat-intelligence classification, trust-and-safety scoring, agent guardrails, and visual classification. The company also reports a test in which Clef classified a website after fetching and rendering it with Browser Run in 2.2 seconds, compared with 4.7 seconds for its tested gpt-oss-120b workflow. That is a Cloudflare measurement for a particular workflow, not a general claim that Clef is faster for every AI task.

For a larger cloud agent, that separation can be particularly useful. The agent can remain responsible for understanding the task and choosing a broad strategy, while a small decision model handles bounded checks such as routing, escalation, or whether a particular action belongs to an allowed category.

What to check before putting Clef on the hot path

The most important production test is not whether Clef can answer your sample questions. It is whether the probabilities and resulting thresholds behave acceptably on the messy data your users actually produce. Keep a human-review path for uncertain cases, especially when an incorrect decision can delete data, block a legitimate user, expose information, or trigger a financial action.

Also remember that the published benchmark results are vendor measurements. Cloudflare reports that Clef leads on seven of ten decision benchmarks, but the same published comparison shows other models leading on individual tasks, including When2Call, BRIGHT, and PhishNChips. The benchmark table therefore supports using Clef as a serious candidate for decision workloads, not treating one aggregate result as proof that it is the right model for every application.

Finally, keep the decision schema small and explicit. If your question requires the model to invent categories, write an explanation, and then decide what to do with that explanation, you have probably pushed too much responsibility into the decision model. If the application can describe the possible outcomes in a few clear options and safely handle uncertainty, that is where Clef's design starts to make practical sense.

Deploy the Worker when the evaluation is stable

Once the test set produces acceptable results, deploy the Worker with Wrangler. The same AI binding used during development is available to the deployed Worker, so the application code does not need a separate inference implementation for production.

npx wrangler login npx wrangler deploy

At that point, the useful part of Clef is no longer the model call itself. It is the boundary you have created around a decision: defined inputs, a finite set of answers, measurable probabilities, explicit thresholds, and a fallback when the model is uncertain. That pattern can sit beside a larger generative AI system without forcing every small routing decision through a model designed to write paragraphs.

Muhammad Saleem profile photo

Written by

Muhammad Saleem

I’m Muhammad Saleem, a web developer and the owner of TechWare House, a software house focused on practical web and software solutions. With over 14 years of experience, I’ve built and managed hundreds of websites and custom , PHP/MySQL, Python, Django applications. I share hands-on insights about web development, software, technology, and digital solutions on WizTechnoz.com

78 posts published

All posts by this author

0 Comments

No comments yet. Be the first to share your thoughts.

Join the conversation

Log in or create a free account to leave a comment. You can edit or delete your own comments any time.