# AI decision models compared: Jev, Clef, Kev, Microsoft

> Jev, Clef, Clef-flash, Kev and Microsoft-Decision-1 compared on price, latency, context and hosting, with the job each one actually fits.

Published: 2026-10-10
Author: Preetham Reddy
Tags: AI/ML, Applied AI, Knowledge Share
Canonical: https://www.techfabric.com/blog/ai-decision-models-compared-jev-clef-kev-microsoft

---

A decision model takes a block of state and a set of typed questions, and returns a calibrated probability for every allowed option, with no text to parse. Four of them now share almost the same request shape, which means the choice between them is a question of price, latency, hosting and input type, and not one of rewriting your application.

Microsoft-Decision-1 became available in Microsoft Foundry and on OpenRouter on 9 October 2026, priced at $0.042 per million input tokens with output free and a 32,768 token context, [post-trained from Qwen3.5-9B](https://openrouter.ai/microsoft/microsoft-decision-1). Cloudflare released [Clef and Clef-flash on Workers AI](https://blog.cloudflare.com/clef-decision-models/), fully Jev-API compatible, with the weights open-sourced on Hugging Face under an Apache 2.0 licence. Jev, from TypeSafe AI, wrote the contract everyone else copied. And Kev is a family of open-weight models, from 0.8B to 27B, that [serves the same System One API from hardware you run yourself](https://github.com/jaredpalmer/kev).

## What a decision model returns

Every one of these models takes a `state` and a map of named `questions`. There are three question types, and you can mix them in one call. A `noul` question returns a calibrated 0 to 1 probability that a statement is true. A `choice` question picks one option from a labelled map and returns a probability for each. A `score` question places the state on an ordered scale and returns a probability-weighted value. TypeSafe's docs describe the three as [primitives](https://docs.typesafe.ai/), and the framing is right, because each question should be the kind of judgment a knowledgeable person makes in a couple of seconds, with anything broader decomposed into several questions and recombined in your own code.

This is the Workers AI call for Clef, taken from Cloudflare's model page as an illustration of the shape:

```bash
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
  -X POST \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -d '{
    "model": "clef",
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {
      "urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
          "billing": "Payments, invoices, and refunds",
          "technical": "Outages, errors, and configuration",
          "sales": "Plans and upgrades"
        }
      },
      "severity": {
        "type": "score",
        "instructions": "How severe is the customer impact?",
        "criteria": ["No impact", "Minor", "Major", "Critical"]
      }
    }
  }'
```

Swap the host and the auth header and that same body runs against Jev or against a local Kev server, and Microsoft's docs describe a similar state-plus-typed-questions pattern for Microsoft-Decision-1. The portability is the most useful property in the whole category, and it should shape how you build. Put the question definitions in one module, keep the transport behind an interface, and pin a model version so your thresholds do not shift under you when an alias moves.

```mermaid
%% caption: Where a decision model sits in a request path, and what happens when it is unsure.
flowchart TD
  A["Request arrives in your service"] --> B["Decision model scores typed questions"]
  B --> C["High confidence answer"]
  B --> D["Low confidence answer"]
  C --> E["Code acts without a human"]
  D --> F["Escalate to agent or person"]
  class B accent
```

## Price, context and where the weights live

Microsoft-Decision-1 and Jev land on the same headline number. TypeSafe lists [Jev 1.13 at $42 per billion input tokens, which is $0.042 per million](https://docs.typesafe.ai/models), with output free, a 64k budget for the state plus all questions and 32k for the state plus the single longest question, and rate limits of 100K tokens per second and 80 requests per second that the docs say can change without notice. Microsoft matches the input price and the free output, and its context is 32,768 tokens.

Clef costs more and does more. Workers AI lists it at [$0.24 per million input tokens with a 65,536 token window and vision input](https://developers.cloudflare.com/workers-ai/models/clef/), up to four embedded PNG, JPEG or WebP images per request, and between 1 and 64 questions in a call. Images are billed as tokens, one per 32 by 32 pixel block, capped at 1,024 tokens each. That is roughly six times Jev's text price, and it buys you the ability to send a scanned document or a rendered page into the same typed-question interface instead of running OCR first.

Kev is the option nobody with a compliance constraint should skip. Four sizes under Apache 2.0, running on anything from a 4 GB GPU to an 80 GB one, with Apple Silicon support through MLX. On held-out datasets the project reports Kev-27B at 52.3 on the community Decision Index against Jev's 54.0, and accuracy on new sources of 0.851 for Kev-27B against 0.857 for Jev. Those are the Kev project's own measurements, and the README says plainly that Jev's training data is unknown so it is not a controlled comparison. Still, open weights within a couple of points of the hosted leader, that you can fine-tune on your own labelled history, is a real option for anyone who cannot send the state outside their network.

## Latency decides where the model sits

Cloudflare's published run puts Clef's median latency at 209.3 ms with a p95 of 238.6 ms, Clef-flash at 38.8 ms median and 122.4 ms p95, and Jev at 524.1 ms median and 536.0 ms p95. Microsoft says Microsoft-Decision-1 had the highest accuracy in its 36-benchmark comparison covering nearly 150,000 questions and was the fastest it measured, 2.5 times quicker than H2O-Lightning-4B v1.1, the accuracy runner-up, and 35 times quicker than GPT-6 Sol, and OpenRouter lists its median on Azure at 0.21 s.

Every one of those numbers was produced by the vendor that wins it. Read them as a rough ordering of magnitude instead of a ranking. Tens of milliseconds for the small models, hundreds for the large ones, and seconds for an LLM doing the same job. Cloudflare's domain classification example is the useful one, because it is end to end. Fetching, rendering and classifying a website took Clef 2.2 s against 4.7 s for gpt-oss-120b, and the LLM returned only two classifications.

A model in the tens of milliseconds can sit in the hot path of a request. A model at half a second cannot sit in a per-keystroke path, though it is fine per ticket, per document or per agent step. Decide which of those you are building before you compare accuracy tables.

## Which one fits which job

**Support and email triage.** Route a ticket to a team, score its urgency, and gate escalation with a noul question, all in one round trip. Cloudflare reports Clef at 94.20 macro-F1 on BANKING77 against Jev's 79.74, and 97.43 on CLINC150+OOS against 89.27, which are the two benchmarks closest to intent classification on a fixed taxonomy. On Typesafe's own customer service workflow eval Clef, Clef-flash and Jev are within a point of each other, at 76.3, 77.0 and 76.0. If you are already on Azure, Microsoft-Decision-1 at the lower price with Entra ID auth and, in selected regions, a DataZoneStandard deployment that keeps inference processing inside a data zone is the path of least resistance; Microsoft's [how-to guide](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/use-foundry-models-microsoft-decision) walks the ticket classifier end to end.

**Agent guardrails.** Gate a proposed tool call before it runs. A noul question on "does this touch production data" or "does this reply promise a refund" is cheap enough to run on every step, and the probability gives you a threshold instead of a coin flip. This is the case where Clef-flash earns its place, since at a 38.8 ms median you can put a check in front of each action without anyone noticing. Jev leads Clef on When2Call accuracy, 80.97 against 72.37, so test the specific judgment you are gating on instead of assuming the general ranking holds.

**Documents and invoices.** Clef is the only one of the four that reads images natively. On Typesafe's invoice processing workflow Cloudflare reports Clef at 64.7 exact actions against Jev's 61.8 and Clef-flash's 57.1. For an accounts payable flow where the input is a scan, paying $0.24 per million tokens to skip a separate extraction step is usually the cheaper architecture overall.

**Model routing.** Use a small decision model to decide whether a prompt needs a frontier model at all, then send the simple ones to a cheap tier. The decision costs a fraction of a cent and the saving is the difference between model tiers. We wrote about the same arithmetic from the other direction in [GPT-6: choosing Astra, Sol or Luna](/blog/gpt-6-model-selection-reasoning-effort).

**Underwriting and scored queues.** Score questions are built for ordered judgments, and Microsoft's docs are honest about their limits. Treat a score as relative ordering for ranking and thresholds, and validate it on your own labelled examples before you trust it as an absolute rating.

## Start by writing the questions, not by picking the model

Take fifty examples you already have labels for, write the questions as atomically as you can, and run the same payload against two of these models before you commit to either. The API compatibility makes that a morning's work, and it will tell you more than any vendor's benchmark table. Then pick on the constraint that will not move. Azure governance and price point to Microsoft-Decision-1, images or edge latency point to Clef, and weights that cannot leave your network point to Kev.

If you want the same comparison applied to a specific workflow, [decision models in auto refinance](/blog/decision-models-auto-refinance-underwriting) walks through where the questions sit in an underwriting flow, and our [AI practice page](/ai) covers how we approach this work.

## Frequently asked questions

### What is an AI decision model?

An AI decision model takes a state and a set of typed questions and returns a calibrated probability for each allowed answer, without generating text. Jev, Clef, Kev and Microsoft-Decision-1 all expose yes/no, multiple choice and ordered score questions through one endpoint, so application code branches on a number instead of parsing prose.

### Which decision model is cheapest?

Microsoft-Decision-1 and Jev both charge $0.042 per million input tokens with output free, and Clef on Workers AI charges $0.24 per million input tokens. Kev is Apache 2.0 open weights, so its cost is whatever your own GPU time costs, which can be lower at volume and higher at low volume.

### Can decision models read images?

Clef reads images and video natively and accepts up to four embedded PNG, JPEG or WebP images per request, billed as input tokens. Jev is text only and its docs say to pre-process images, audio and video into text or structured fields before sending them as state, and Microsoft-Decision-1 accepts text or JSON.

### Do I have to rewrite my code to switch decision models?

No, because Clef and Kev both implement TypeSafe's System One request and response shape, so the same state and questions payload moves between them with a change of endpoint and key. Pin a specific model version in production, since an alias such as jev-latest moves when a new release ships and your confidence thresholds were tuned against the old weights.
