Picture a hypothetical auto refinance lender. Call it Meridian Auto Refi, a company I made up to make this concrete; nothing below is a client story or a result anyone achieved. It buys leads from three aggregators, collects the applicant's vehicle details, pulls a credit score, asks for an odometer photo, checks a driver's license and an SSN, and then packages the whole thing and submits it to one of a dozen lender partners.
The failure mode in that business usually isn't the credit decision. It's the package. A lender kicks back an application because the odometer reading in the form doesn't match the photo, because the name on the license is "Robert" and the application says "Bob", because the VIN has seventeen characters but one of them is a letter O, because the stated payoff amount is stale by three weeks. Each kickback costs a day and some fraction of the applicant's patience. Enough of them and the lender starts deprioritizing your submissions.
So the interesting question isn't "can AI underwrite the loan". It's "can software decide, on every application, whether this package is ready to submit and to which lender". That's a decision and not an essay. And decisions are exactly what this new class of model is built for.
What a decision model actually is
TypeSafe released Jev in September 2026 as what they call a System One model: unstructured state in, typed probabilistic decisions out, with no string generation at all. Because the possible outputs are defined in advance, the model can't make a type error and can't hallucinate a field, and every answer arrives with a calibrated confidence score. TypeSafe quotes end-to-end response times of 70ms to 500ms and input pricing of $0.042 per million tokens with output tokens free.
Cloudflare followed with Clef and Clef-flash, open-sourced on Hugging Face under Apache 2.0 and hosted on Workers AI, API-compatible with Jev. Two things in that announcement matter for a document-heavy workflow: Clef has a vision encoder, so it can classify images, and it has a 64k context window against Jev's 32k. Cloudflare's published median latency for Clef is 209.3ms and 38.8ms for Clef-flash.
The mental model I keep coming back to is the one Cloudflare uses: a smart if-statement. You hand it the state, you hand it a fixed set of questions with typed answers, and your ordinary code branches on the result. The model's freedom is constrained by the shape of the question, which is why you can put it in the hot path of a production workflow and sleep.
The five places I would put one
Walking the Meridian pipeline end to end, here is where I'd reach for a decision model instead of writing another regex or another rules table.
Lead triage at purchase. State is the raw lead payload plus whatever enrichment you have. The questions are whether this lead is likely to produce a fundable application, which product bucket it belongs in, and how confident you are. You're spending money on every lead; a 40ms decision that defers the ambiguous ones to a human queue is cheap.
Document classification and completeness. The applicant uploads four photos. Which one is the license, which is the odometer, which is the insurance declaration page, which is a blurry picture of a parking garage. This is where Clef's vision encoder earns its place, because the alternative is a general LLM taking seconds per image.
Cross-field consistency. This is the big one. State is the whole application as a dense paragraph covering stated mileage, odometer reading extracted from the photo, VIN, year and model, payoff quote and its date, employer, stated income. Questions are the ones an underwriter asks in their head. Does the odometer reading support the stated mileage. Is the name on the license the same person as the applicant. Is the payoff quote current enough for this lender's window.
Lender routing. Each lender partner has preferences they will never write down, including LTV tolerance, vehicle age limits, and how they feel about a thin file. A choice-typed question over the lender list, with criteria you maintain in text instead of in a stored procedure, is easier to change on a Tuesday afternoon than the rules engine was.
Pre-submission gate. The last call before you hand the package over. Is this ready, what's the probability of a stipulation request, which one. Low confidence means a human looks at it. That's the whole value: the model tells you when it doesn't know.
An illustrative query
This is illustration and not code from a running system. It's the Clef shape from Cloudflare's own example, with auto refinance questions dropped in, so you can see what the typed interface feels like.
{
"model": "clef",
"state": "Applicant stated 48,500 miles on a 2021 sedan. Odometer photo OCR read 84,500. Payoff quote dated 22 days ago. License name Robert; application name Bob.",
"questions": {
"mileage_consistent": {
"type": "noul",
"instructions": "Does the odometer evidence support the stated mileage?"
},
"blocking_issue": {
"type": "choice",
"instructions": "What is the most likely reason an underwriter rejects this package?",
"criteria": {
"mileage": "Stated and observed mileage disagree",
"identity": "Name or identity fields do not reconcile",
"stale_payoff": "Payoff quote outside the lender window",
"none": "No blocking issue found"
}
},
"readiness": {
"type": "score",
"instructions": "How ready is this package for submission?",
"criteria": ["Not ready", "Needs one fix", "Minor cleanup", "Submit now"]
}
}
}
Note the noul type and the score with named bands. Cloudflare's published example uses exactly those; the point of the typed schema is that your downstream code never parses prose and never handles a field the model invented.
Where Databricks does the unglamorous work
A decision model is a function. It has no memory, no history and no opinion about your business. Everything that makes the call good comes from the state you assemble and the evidence you keep, and that's platform work.
The state is a join. Lead payload, credit pull, OCR output, vehicle data, prior applications from the same applicant, the lender's recent behavior. On a lakehouse that's a feature table with a primary key on application ID, served online so the gate call doesn't wait on a batch job. Garbage state produces a confidently wrong decision in 40 milliseconds, which is worse than a slow wrong one because nobody sees it happen.
Every call gets logged, both sides. Request state, every typed answer, every confidence score, model name and version, timestamp. One Delta table. This is the thing teams skip and then regret, because without it you cannot answer the only question that matters six weeks in: when the model said 0.91, how often was it right.
The label comes back from the lender. Approved, declined, stipulation requested, and which stipulation. Join that to the decision log on application ID and you have a calibration dataset you didn't have to build. Now you can plot predicted readiness against actual kickback rate and see whether the confidence scores mean anything on your data. Calibration is a claim the vendor makes about their training; whether it holds on your population is a measurement you run.
Thresholds are a parameter you can turn, not a belief. Once you can measure, you can set the auto-submit cutoff deliberately. Submit above 0.9, human review between 0.6 and 0.9, hard stop below. Move the numbers as the evidence accumulates. Unity Catalog gives you the governance story for the PII sitting in that state, which in a workflow handling SSNs and license images is mandatory. We wrote about the shape of that in an AI governance framework is four answers you can query.
Pick the model by the measurement. Clef and Clef-flash are Jev-API compatible, which means you can route a slice of traffic to each and compare on your own labels. Cloudflare's own numbers show Clef-flash beating Clef on some evals and losing badly on others, CLINC150 macro-F1 at 66.77 against 97.43. Published benchmarks tell you which models are worth testing. They don't tell you which one to ship.
What I'd measure, and what I wouldn't promise
If I were standing this up, these are the four numbers on the dashboard from week one. Kickback rate per lender per submission reason. Share of applications auto-submitted versus deferred to a human. Time from lead purchase to submission. Calibration: predicted versus observed, bucketed by confidence decile.
Notice those are things you measure, and nothing I'm claiming anyone achieved. Meridian doesn't exist. I have no result to report, and if I invented one it would be the least useful paragraph in this post.
The honest risk is that a fast decision becomes an unexamined one. A rules engine fails loudly and in a place you can read. A decision model returns a clean typed answer with a confidence score whether or not the state it was handed made any sense, and it will do that four hundred times a minute. The defense is the log and the label loop, which is why the Databricks half of this matters as much as the model call.
What I'd tell someone starting this
Start with the gate before you touch the whole pipeline. One decision, the pre-submission readiness check, running in shadow mode next to whatever humans do today. Log both answers for two weeks and look at the disagreements; they will teach you more about your lenders than the lenders will.
Write the questions before you write the code. If you can't express the decision as a bounded set of typed answers, it isn't a System One task and you should leave it with a person or an LLM.
Don't start by replacing your rules. The deterministic checks that work, VIN checksums, date arithmetic, field presence, should stay deterministic. Decision models are for the judgment calls the rules engine was always bad at, the ones full of "usually" and "depends".
And budget for the boring half. Feature table, decision log, label join, calibration report. That's where the accuracy actually comes from. If you want to talk through how it maps onto a lakehouse, that's the kind of thing we do on Databricks.