# Clef, Clef-flash or Clef-omni: which decision model to use

> Cloudflare's three Clef decision models compared on size, context, latency and price, with a manufacturing warranty-claim walkthrough.

Published: 2026-10-09
Author: Preetham Reddy
Tags: AI/ML, Applied AI, News
Canonical: https://www.techfabric.com/blog/clef-vs-clef-flash-vs-clef-omni

---

Cloudflare shipped Clef-omni the week after Clef and Clef-flash, and the family now covers three different jobs. A decision model doesn't generate text. It reads a state and a schema of typed questions, then returns a probability for every allowed option of every question, which your code acts on directly. No parsing, no reasoning tokens, and no retry when the JSON comes back malformed.

So choosing between the three is an engineering decision more than a matter of taste. They take the same request shape. What separates them is how much context they hold, how fast they answer, what media they can read, and what they cost.

## The three models, side by side

Clef is a 27B multimodal model with a 65,536 token context window, priced at $0.24 per M input tokens, with a vision encoder that accepts up to four images per request ([docs](https://developers.cloudflare.com/workers-ai/models/clef/)). It leads Jev on those benchmarks: 97.43 macro-F1 on CLINC150+OOS, 94.20 on BANKING77, 69.19 nDCG@10 on ToolRet ([launch post](https://blog.cloudflare.com/clef-decision-models/)).

Clef-flash is 9B, a 24,576 token context window, and $0.038 per M input tokens after a price cut from $0.09, which Cloudflare says makes it cheaper than Jev ([docs](https://developers.cloudflare.com/workers-ai/models/clef-flash/)). At the 1 October launch its median latency was 38.8 ms against Clef's 209.3 ms and Jev's 524.1 ms, measured across 43 benchmark runs, before Cloudflare made Clef faster ([changelog](https://developers.cloudflare.com/changelog/post/2026-10-01-clef-workers-ai/)). It isn't uniformly weaker, either. It scores 97.73 case exact on the home appliances benchmark where Clef gets 82.95, and 93.11 on API-Bank. It falls over on CLINC150+OOS at 66.77, which is the out-of-scope detection benchmark, so a flash model that has to recognise "none of these categories apply" needs checking before you trust it.

Clef-omni is built on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts backbone with 3B active parameters, holds 64,000 tokens, and costs $0.15 per M input tokens ([docs](https://developers.cloudflare.com/workers-ai/models/clef-omni/)). It takes audio and video alongside text and images. Cloudflare kept the comprehension backbone and threw away the text-to-speech output components, froze the Qwen3 weights, trained LoRA adapters, and applied label-smoothed cross-entropy with Brier score calibration ([announcement](https://blog.cloudflare.com/clef-faster-cheaper-multimodal/)).

```mermaid
%% caption: Choosing between the three Clef models by what the decision needs.
flowchart TD
  A["Decision to make"] --> B{"Audio or video input"}
  B -->|"Yes"| C["Clef-omni"]
  B -->|"No"| D{"Hot path under 50 ms"}
  D -->|"Yes"| E["Clef-flash"]
  D -->|"No"| F["Clef"]
  class C accent
```

## What multimodality actually buys you

Cloudflare's main claim for Clef-omni is architectural. Normally, scoring a decision against a recording means a cascade. You transcribe the speech, split the audio channel from the video, caption the frames, then feed the text to a classifier. Every stage adds latency and a failure mode, and every stage throws away information the next stage might have wanted.

Clef-omni runs one prefill pass across the whole payload and scores all modalities and all valid parameter options at once. Because it generates no output tokens, there is no transcription or captioning step to pay for. Media maps straight into a unified sequence, with video and audio synced to the visual frames, and two-stage attention routing pulls candidate values out of the internal embeddings: every option gathers evidence from the input, then field vectors cross-attend across the full context to produce confidence scores.

The latency numbers follow from that. Text-only decisions come back in about 130 ms at the median, images in about 150 ms, audio clips in a few hundred milliseconds, and a full 21-second video clip with sound scores in about 1.5 seconds, all in a single API call.

The limits are worth reading before you design around them. Each request takes up to 4 audio clips (8 MiB and 300 seconds each) and up to 2 videos (16 MiB and 60 seconds each), sampled at 2 frames per second, with audio and video together capped at 16 MiB decoded. A video's soundtrack is heard with its frames when every video in the request has one. Remote URLs are not accepted, so media goes in base64 encoded in the body.

Media costs come down to tokens. Audio runs about 780 tokens per minute. Video runs up to about 15,400 tokens per minute at maximum resolution, and a 480p clip uses about 8,600 frame tokens per minute, which is 144 tokens per second, plus about 780 tokens per minute if there is sound. Images are tokenized at one token per 32x32 pixel block plus 3 marker tokens, capped at 1,024 tokens each. Media tokens count against the context window alongside the questions, and if they exceed it the request fails outright. Under the limit, the text state is truncated to fit whatever space is left, which means a long state plus a long video silently loses the tail of the state. Put the fields the decision depends on at the front.

Lower the resolution before anything else. A 480p clip is substantially cheaper per minute than a maximum-resolution one, and for "is the fan spinning" it is plenty.

## Picking one

Use Clef-flash when the decision sits in the hot path and the schema is narrow. Agent guardrails are the canonical case: an agent asks "should I take this action?" before calling a tool, and at a median of 38.8 ms the check is invisible against the tool call it protects. Same for first-pass spam and abuse filtering, routing by intent, and anything you run on every request instead of every transaction. The real constraint is the 24,576 token context, well before the parameter count. Check out-of-scope behaviour before you ship it.

Use Clef when the decision is expensive to get wrong and happens thousands of times a day instead of millions. Underwriting-style checks, claims adjudication, trust and safety scoring against a policy rubric, anything where you want the 65,536 token window to hold the full record instead of a summary of it. The $0.24 per M input tokens is more than six times Clef-flash, which matters at volume and doesn't matter at all when the alternative is a human reading the file.

Use Clef-omni when turning the evidence into text would lose what you care about, like the rattle in a recording, the sequence of a motion, or whether the sound and the picture agree. At $0.15 per M input tokens it sits between the other two, so if your inputs are text and images, Clef-omni is not the default. Clef scores higher on several of the text benchmarks, 98.47 against 98.2 on BFCL case exact, 69.19 against 66.6 on ToolRet, 79.60 against 73.2 on PhishNChips, while Clef-omni edges ahead on BANKING77, CLINC150+OOS, API-Bank and Amazon ESCI. Clef-omni's advantage is modality coverage, and you pay for it in a point or two of text accuracy.

A useful pattern is Clef-flash as a cheap gate on everything, with Clef or Clef-omni on the fraction that the gate flags. Clef follows the System One API and is Jev-compatible, so switching an existing Jev integration is a change of endpoint and model string.

## A warranty claim in manufacturing

Imagine an appliance manufacturer running warranty claims through a dealer network, where each claim packet holds a serial number, a fault code the technician typed in, a photo of the rating plate, a phone recording of the unit running, and sometimes a short video. A claims clerk opens each one, checks the label is legible and matches the serial on file, listens for the noise the technician described, decides whether the fault is covered or wear and tear, and approves or kicks it back. Manual review of packets like this is slow and two clerks can reach different conclusions about the same recording.

Every part of that is a typed question. The Cloudflare announcement's own example is almost exactly this shape:

```bash
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-omni \
  -X POST \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "clef-omni",
    "state": "Review the installation: a photo of the unit, an audio recording of it running, and a video of the fan.",
    "images": ["data:image/png;base64,<base64-png>"],
    "audio": ["data:audio/mpeg;base64,<base64-mp3>"],
    "videos": ["data:video/mp4;base64,<base64-mp4>"],
    "questions": {
      "label_visible": {"type": "noul", "instructions": "Is the model and serial number label visible in the photo?"},
      "sounds_normal": {"type": "noul", "instructions": "Does the unit sound like it is running smoothly, without rattling or grinding?"},
      "fan_running": {"type": "noul", "instructions": "Is the fan running in the video?"}
    }
  }'
```

Extend that schema and you have the claim. A `choice` question for fault category against criteria you write, a `score` question for severity against an ordered rubric, a `noul` for whether the packet is complete enough to decide at all. Up to 64 questions per request, so the whole claim form fits in one call. The answers come back keyed by your question ids with a probability per option, and your code sets the thresholds: auto-approve above one band, auto-reject below another, route the middle to a clerk. A decision model gives you calibrated probabilities; deciding where to cut them is a business policy you own and should be able to change without redeploying anything.

The savings come from the claims that never reach a human at all, and from consistency in the ones that do. The metric to watch is the share of claims handled end to end without review, the disagreement rate between the model and the clerks on the sampled overlap, and the reversal rate on auto-approvals thirty days later, well ahead of accuracy in the abstract. Instrument those from day one, because the threshold you pick in week one will be wrong and you need the data to move it.

The same shape works in retail. A returns desk takes a photo of the item, a customer's description, and the order record, then answers whether the condition matches the claimed reason, whether this is resaleable, and whether to refund or inspect. Returns are high volume and low value per decision, which is Clef-flash territory for the first pass, with Clef-omni on the ones that come with a video.

## Write the questions before you pick the model

Start with the schema before the model. Write the questions a human currently answers on the form, in their words, with the criteria spelled out, and you will find half of them are ambiguous. Fixing that is most of the work and it pays off whichever model you run.

Then benchmark on your own data. The published benchmarks disagree with each other in useful ways, Clef-flash beating Clef on home appliances by fifteen points and losing by thirty on CLINC150+OOS, and your workload resembles one of those more than the average. Run all three against a few hundred decisions a human already made and look at the disagreements more closely than the headline score. Clef and Clef-flash are open-weight under Apache 2.0 on Hugging Face, and Clef-omni's weights are on Hugging Face too, if you want to evaluate locally before committing to the hosted endpoints.

Budget the media before you build the pipeline. Work out the tokens per claim from the published rates, multiply by your volume, and compare against the $0.038, $0.15 and $0.24 per M input tokens. Then drop your video resolution and recalculate, because that one change moves the number more than anything else you will do.

And keep the thresholds in configuration. The model returns probabilities; the business decides what to do with them. If you have been thinking about where decision models fit in a wider agent stack, our [AI practice page](/ai) covers how this connects to the rest of the platform.

## Frequently asked questions

### What is the difference between Clef, Clef-flash and Clef-omni?

Clef is a 27B model with a 65,536 token context at $0.24 per M input tokens for the highest-precision text and image decisions, Clef-flash is a 9B model with a 24,576 token context at $0.038 per M input tokens and a 38.8 ms median latency for hot-path decisions, and Clef-omni is a 30B mixture-of-experts model with 3B active parameters and a 64,000 token context at $0.15 per M input tokens that also reads audio and video. All three take the same request shape, so switching between them is a model string change.

### Can Clef-omni replace a speech-to-text pipeline?

For decisions, yes. Clef-omni takes audio and video directly and scores your questions in a single prefill pass, so there is no transcription or captioning step in front of it. It returns probabilities against your schema and never a transcript, so if you need the words themselves for a record or for search, you still need a speech recognition model alongside it.

### How much does audio and video cost with Clef-omni?

Media is converted to input tokens and billed at the model's input rate of $0.15 per M input tokens, with audio at about 780 tokens per minute and video at up to about 15,400 tokens per minute at maximum resolution, or about 8,600 frame tokens per minute at 480p plus about 780 per minute for sound. Clef models don't charge for output tokens at all.

### Is Clef compatible with an existing Jev integration?

Yes, Clef follows the System One API, so an existing Jev integration switches over by changing the endpoint and the model string. The images, audio and videos arrays are Clef extensions on top of that API, so code that only sends text and questions needs no other change.
