·23 min read

Decision Models, Explained From the Jev Hype Down

llmdecision-modelsai-engineeringclassificationagents

If you've been anywhere near AI Twitter or Hacker News in the last three weeks, you've heard of Jev. Maybe you saw the latency numbers. Maybe you saw the price, $0.042 per million input tokens, output free, and assumed it was a typo. Maybe someone told you it was "a new kind of model" and you nodded along.

This post is for that reader. You know the name and the hype. You might not know what the thing actually does differently, or whether "different" means "better." I didn't either, so I went and read everything I could find, including the benchmarks that aren't flattering.

It's long. Get a coffee.

Chapter 1: A question with a short answer

Start with a support ticket.

"Hi, I was charged twice for my March invoice and I need one of them refunded before my card statement closes on Friday."

Your software needs to know three things about it. Which team owns it (billing, technical, sales)? Is the customer asking for a refund, yes or no? How urgent is it, on a scale of 0 to 3?

None of those answers is a paragraph. Each one is a pick from a list you already wrote. And yet for the last three years, the standard way to get those picks has been to send the ticket to a large language model, ask it to "respond in JSON," and hope it spells billing the same way every time.

A decision model is a model built for exactly this shape of question and nothing else. You give it three things:

  1. State: the text, JSON, or image you want judged. The ticket.
  2. Questions: what you want to know about it.
  3. Options: the allowed answers for each question.

It gives you back, for every question, a probability for every option. Not a sentence. Not JSON that might be malformed. Numbers that add up to one.

Jev's API describes three kinds of questions, and the other decision models have mostly copied the vocabulary:

  • Choice picks one option from a list ("which team?"), with a probability on each.
  • Noul asks whether a statement is true and returns P(true) between 0 and 1 ("is this a refund request?"). The name is TypeSafe's, and it stuck.
  • Score places the state on an ordered scale ("urgency 0 to 3"), returning probabilities over the levels.

A request looks roughly like this (simplified):

{
  "state": "Hi, I was charged twice for my March invoice and I need one of them refunded before Friday.",
  "questions": {
    "team":    { "type": "choice", "options": ["billing", "technical", "sales"] },
    "refund":  { "type": "noul",   "statement": "The customer is asking for a refund." },
    "urgency": { "type": "score",  "levels": [0, 1, 2, 3] }
  }
}

And the answer comes back like this:

{
  "team":    { "value": "billing", "probability": 0.97, "confidence": 0.91 },
  "refund":  { "value": true,      "probability": 0.98 },
  "urgency": { "value": 2,         "probability": 0.64 }
}

That's the whole product. The rest of this post is about why such a small idea got a three-week news cycle.

Chapter 2: How an LLM answers the same question

To see what's new, you have to look at what a normal LLM does with that ticket.

A chat model is autoregressive. It writes one token at a time, and each token depends on all the ones before it. Answering the ticket happens in two phases:

  1. Prefill. The model reads your whole prompt in one go. Every input token is processed in parallel on the GPU. This part is fast, because GPUs are very good at doing the same math on many tokens at once.
  2. Decode. Now it writes. One token, then feed that token back in, then the next token, then feed it back in. Each step has to stream the model's weights through memory again to produce a single token. For a big model this is slow and it can't be parallelized within one answer, because token 40 depends on token 39.

If the model is a reasoning model, decode includes all the thinking: hundreds or thousands of tokens of "let me consider whether this is a billing issue" before it writes {"team": "billing". Then your code parses the string, hopes it's valid JSON, and checks the label is one you allowed.

A generative model loops through decode until it decides it's finished, then your code parses the text. A decision model reads the input once and the answer is already sitting in the output layer.

Here's the uncomfortable part for anyone who has built a classifier on top of a chat model: most of that loop is wasted. You knew the possible answers before you sent the request. The model spent its decode budget spelling out a word you had already written down.

Chapter 3: So how is it so fast?

The short answer: a decision model is all prefill and no decode.

When the model reads your state and your questions, it already computes, internally, a score for every token in its vocabulary at every position. A chat model throws almost all of that away and samples one token. A decision model keeps it. It looks at the specific positions where the answers go, looks only at the scores for the options you allowed, and turns those into probabilities with a softmax. Done. One forward pass.

The small open models make this concrete. The decider-4b model card on Hugging Face lays it out plainly: the prompt ends each question with an answer slot like Answer: (, the options are lettered A, B, C, and the hidden state at each slot is projected onto just the rows of the output layer for those letters, then softmaxed. No sampling, no parsing. On a GPU with CUDA graphs, one support ticket (228 tokens, 3 questions) takes 5.2 ms, and batched it does 1,178 decisions per second.

That one trick explains almost everything in the hype cycle:

Why output is free. The expensive part of serving an LLM is decode, and there isn't any. You only pay for what the model reads. That's why Jev, Perplexity's Decider, and Cloudflare's Clef all price by input tokens and charge nothing for output.

Why many questions cost about the same as one. Every question reads the same state. Ask ten questions about one ticket and the ticket is processed once. Jev's docs say each question is evaluated independently, so adding one doesn't change the others' answers.

Why it batches so well. Prefill is the part of LLM inference that GPUs handle best. A server can pack many short requests together and push them through at once.

Why the model can be small. If the job is "pick from a list," you don't need 400B parameters. You can use an encoder in the BERT family and get very low latency.

Some numbers, collected from the published material:

ModelSizeReported latency
decider-4b (open)4B5.2 ms per request, local GPU
Laya (open)421M38.4 ms p50, one question
DiffusionGemma on Cloud Run26B total, 4B active35 to 60 ms per step
Jev (hosted)not published70 to 500 ms, US West Coast
Perplexity Decider (hosted)27Bunder 2 s for a few hundred tokens

Hosted numbers include the network round trip, so they aren't comparable with local ones. The pattern still holds: the big hosted models land in the hundreds of milliseconds, and the small open ones land in the tens.

Google's take on this is the strangest one, so it gets its own paragraph. DiffusionGemma is a diffusion language model: instead of writing left to right, it fills in a whole "canvas" of tokens at once and refines them over several steps. To use it as a decision model, you pre-fill the canvas with an answer template (urgent: @, category: @), mark everything except the answer slots read-only, and run exactly one denoising step. The logits at each slot are your probabilities. The constraint is that every answer must be a single token, so you map moderation_spam to B on the client side.

Chapter 4: Is it smarter than an LLM?

No. And yes, that's a slightly unfair answer, so let me unpack it.

Look at what these models are made of. Perplexity's Decider is a fine-tune of Qwen3.8-27B. Cloudflare's Clef is a fine-tune of the same base, and Clef-flash is a fine-tune of Qwen3.5-9B. Google's version is DiffusionGemma. OpenAI's Decisions API runs on GPT-6 Luna. These are LLMs with a different output head and different training. They don't know more facts or read text more deeply than the models they came from. Same brain, different mouth.

What a decision model gains:

  • Calibration. This is the real headline, and it got buried under the latency numbers. TypeSafe trains Jev with a method they call RLCD (Reinforcement Learning for Calibrated Decisions), which rewards probabilities that match real outcomes. Put plainly, if Jev says 0.9, it should be right about 90% of the time. Chat models will happily tell you "confidence: 95%" in JSON, and that number is generated text. It means nothing.
  • No format errors. The answer is an index into your list. It can't be misspelled, can't be a label you didn't define, can't be wrapped in a markdown fence.
  • Volume. At these prices you can ask questions you would never have paid an LLM to answer, like checking every message in an agent's trace instead of a sample.

What it gives up is room to think. A reasoning model gets better at hard questions by writing out intermediate steps. A decision model gets one pass. Kahneman's System 1 and System 2 are the obvious comparison, and TypeSafe uses them on purpose: Jev is marketed as a "System One" model. Fast, intuitive, and wrong in predictable ways.

The independent testing shows where it breaks:

  • The Towards Data Science review found Jev "struggles with counting, arithmetic, and date comparisons." Anything you'd solve with a scratchpad, it can't.
  • When the information needed to answer was missing from the state, accuracy fell to 44.7% while average confidence stayed at 0.74. That's the worst way to be wrong: confidently.
  • TypeSafe's own workflow benchmark (711 cases) reports only 67.8% average agreement with GPT-6 Astra and Fable 5.1. On messy multi-step judgments, the frontier reasoning models and the fast model disagree a third of the time.

And where it wins:

  • On Banking77 (3,080 customer messages, 77 intents), the same reviewer measured Jev at 81.1% against 76.4% for Qwen3-Coder-Next-80B, a 4.7-point lead for a much cheaper call.
  • On PriorBench, a 400-item priority-classification set, Jev scored 95.9% against 77.2% for keyword rules.

So the honest version is this. On a question that a person could answer at a glance, a decision model is about as good as an LLM, sometimes better, and its probabilities are worth trusting. On a question a person would need to work out, it's worse, and its probabilities can be wrong too.

Even calibration has an asterisk. In the Banking77 test, answers where Jev reported exactly 1.00 confidence were right 97.1% of the time, which is great. But the band where it reported an average of 0.81 was right only 53% of the time. The model is well calibrated at the extremes and overconfident in the middle. Thresholds still need to be tuned on your own data.

Chapter 5: Haven't we seen this before? (Yes. It was called BERT.)

If you trained models before 2022, the last two chapters probably sounded familiar. "Read the text once, output a probability over fixed labels" is how text classification worked for years.

BERT (2018) is an encoder. Unlike GPT-style decoders, it reads the whole input bidirectionally, with every token attending to every other token, and produces a vector per token. To classify, you bolt a small head on top and fine-tune on labeled data. It's fast, cheap, and usually very accurate on the task you trained it for.

The catch is in that last sentence. The labels are baked into the head. If you want a new category, you collect examples and retrain. Want to ask "is this ticket mentioning a competitor?" next Tuesday? That's a labeling project.

People did try to make BERT-style models zero-shot. The famous trick was natural language inference: bart-large-mnli takes the text plus a hypothesis ("This text is about billing") and scores whether the text entails it. Run it once per label and pick the winner. It worked, sort of, but it was slow for large label sets and the scores weren't calibrated.

A decision model is that idea finished properly. The labels move out of the trained head and into the input. The model is trained on enough varied tasks that it handles questions it has never seen, and trained for calibration so the probabilities mean something.

The fun part is that the BERT lineage is still in the race. Laya is a 421M-parameter decision model built on ModernBERT-large (a 2024 rewrite of BERT), with a small decision transformer on top, and the author reports 38 ms latency. Fastino's GLiNER2.5-Decide is a 340M encoder that runs on CPU.

BERT classifierDecision modelChat LLM
OutputOne of the trained labelsProbabilities over labels you pass inFree text
New labelRetrainChange the requestChange the prompt
Calibrated?Usually overconfidentTrained to beNot really
Can reason step by step?NoNoYes
Typical latencyTens of msTens to hundreds of msHundreds of ms to many seconds
Cost driverYour own hardwareInput tokensInput plus output tokens

And now the part the launch posts left out. Red Hat published a guardrail benchmark on October 2 that pitted decision models against the old guard on prompt-injection and content-safety detection:

Prompt injectionAccuracyMedian latency
Qwen3.6-35B (LLM as judge)89.31%312.5 ms
deberta-v3-base (fine-tuned)89.01%54.1 ms
DiffusionGemma87.72%561.7 ms
Jev86.35%348.1 ms
Content safetyAccuracyMedian latency
Jev86.20%360.4 ms
DiffusionGemma85.53%499.3 ms
Qwen3.6-35B (LLM as judge)85.47%307.6 ms
granite-guardian-hap-125m (fine-tuned)80.27%33.2 ms

Their conclusion was blunt: decision models "do not reliably outperform LLM-as-a-judge, pre-trained predictive models, or open source decision models in speed or accuracy." A fine-tuned DeBERTa from 2021 matched the best result on prompt injection, at a sixth of the latency.

I don't read that as "decision models are hype." I read it as: if you have a fixed task, labeled data, and a stable set of labels, a fine-tuned encoder is still very hard to beat. Decision models win when you don't have those things, which is most of the time when you're building something new. They're the zero-shot option, not the best-possible option.

Chapter 6: Context is the whole bill

With a chat model, "context" covers a lot: the system prompt, the conversation, the retrieved documents, the tool results, and the model's own output piling up. With a decision model, context is simpler and matters more. It's the state, plus your questions, and that's the entire input. There's no output to speak of. So everything you pay for and everything the model knows is in the context you send.

That has some practical consequences.

Window sizes vary a lot. From the published specs:

ModelContext window
Jev 1.1332K tokens
GLiDE (Fastino)40K per question
Clef and Clef-flash (Cloudflare)64K tokens
pplx-decider-v1-27b (Perplexity)262K tokens

Latency scales with the state, not the answer. Since everything is prefill, a longer input means a longer wait, and it adds up. Perplexity's Decider answers in under two seconds for a few hundred tokens and takes about 23 seconds near its 262K ceiling. A 262K window is useful for "does this contract contain a non-standard indemnity clause," but it isn't the 50 ms experience from the launch demos.

The model can't ignore what you send. A reasoning model can write "the second paragraph is irrelevant" and move on. A single forward pass can't. The Towards Data Science review notes Jev is sensitive to irrelevant context and to prompt injection. Trimming the state down to what the question actually needs is the single biggest quality lever you have. If you're judging an agent's next tool call, send the call and the bit of the plan it serves, not the whole 30-turn transcript.

Missing context produces confident nonsense. That 44.7%-accuracy-at-0.74-confidence result is a context problem. If the answer isn't in the state, the model still has to put probability somewhere, and it doesn't have an "I don't know" unless you give it one. Always include an other or not enough information option. It's a cheap fix and it's the one people forget.

Context gets reused across questions. Because every question reads the same state, the cost of a long document is paid once, no matter how many questions you ask about it. Asking twelve questions about one contract is close to the price of asking one. Design requests that way: one state, many questions.

Chapter 7: The three weeks the frontier labs noticed

Here's how it played out, in order.

June 10. Google releases DiffusionGemma. At launch it wasn't pitched as a decision model. It's an experimental open-weights diffusion LLM (26B total, 4B active, Apache 2.0). After Jev showed up, Google's Gemma account promoted "DiffusionGemma as Jev," and Google Cloud published a vLLM recipe for running it as a one-step decision engine. It's the only multimodal option here that you can run yourself with Google's backing.

September 15. TypeSafe launches Jev. TypeSafe was founded by Diogo Almeida, formerly of OpenAI. Jev is closed, the size is unpublished, and it costs $0.042 per million input tokens with free output. TypeSafe framed it as "System One" and introduced the choice / noul / score vocabulary. The name refers to William Stanley Jevons, whose paradox says efficiency increases total consumption. That's a bet that cheap decisions will lead to a lot more decisions. According to Clouded Judgement, about 13% of paid teams on Vercel's AI Gateway used it within 24 hours.

September 24 and 30. Fastino ships GLiNER2.5-Decide and GLiDE. The first is an open 340M encoder you can self-host on CPU. The second is a hosted, closed model.

September 29. OpenAI announces the Decisions API at DevDay. It runs on a specialized version of GPT-6 Luna, takes text or images plus questions with fixed answer lists, and returns one answer per question. Sam Altman's framing: "By focusing the model on that choice, we can make it extremely fast while keeping capabilities like image understanding, broad language support, and safety protections." It's a limited preview. As of October 2, pricing, schema, whether probabilities come back, context limits, and latency were all unpublished, and a regular API key got a 403. TechCrunch's headline called it a Jev clone, which is fair.

October 1. Cloudflare and Perplexity release open weights the same day. Cloudflare's Clef (27B, from Qwen3.8-27B) and Clef-flash (9B, from Qwen3.5-9B) are Apache 2.0, priced at $0.24 and $0.09 per million input tokens on Workers AI. Perplexity's pplx-decider-v1-27b is also Apache 2.0 from Qwen3.8-27B, multimodal, with a 262K context, priced at $0.04 per million input tokens. On Perplexity's 11-test panel (7,210 samples), Decider scored 85.71% overall versus Jev's 84.51%, though Jev won 6 of the 11 individual tests.

Around the same time, Databricks added decision models to its SQL AI functions, which means you can run a decision over every row of a table. Classifying a million support tickets becomes one SQL query.

Anthropic hasn't announced a decision model or a decisions endpoint that I could find. The closest thing is the server-side permission classifier Claude Code switched to for auto mode in September, which is an internal classifier, not a product.

ModelMakerOpen?BasePrice per 1M input
Jev 1.13TypeSafeClosedUnpublished$0.042
Decisions APIOpenAIClosed, previewGPT-6 LunaUnpublished
pplx-decider-v1-27bPerplexityApache 2.0Qwen3.8-27B$0.04
ClefCloudflareApache 2.0Qwen3.8-27B$0.24
Clef-flashCloudflareApache 2.0Qwen3.5-9B$0.09
DiffusionGemmaGoogleApache 2.0Gemma 4 MoESelf-host
GLiNER2.5-DecideFastinoOpen340M encoderSelf-host (CPU)
LayaIndependentApache 2.0ModernBERT-largeSelf-host

Two things stand out to me in that table. First, three of the open models are Qwen fine-tunes, and two share the exact same 27B base, so the moat isn't architecture. TypeSafe's founder told TechCrunch their moat is the synthetic data used to train calibration. Second, open weights caught up with the closed launch in sixteen days. That's faster than usual, even by 2026 standards.

Chapter 8: Where you'd actually use one

The rule I've landed on: if you could write the possible answers down before you see the input, and a person could answer at a glance, use a decision model. Otherwise use an LLM.

Good fits:

  • Routing. Which team gets this ticket, which model handles this prompt, which tool does the agent need next. Model routing alone could be a big market: one investor newsletter estimates decision models could take 10 to 15% of tokens currently generated.
  • Agent guardrails. Before an agent runs rm, sends an email, or makes a purchase, ask "is this action consistent with the user's request?" and block below a threshold. TechCrunch cited a Jev comparison where monitoring an agent swarm cost $2.94 versus $372 with frontier LLMs. That price difference is what makes checking every action possible instead of sampling.
  • Moderation and policy checks. Brand compliance, toxicity, PII, "does this contract contain non-standard terms." Questions with clear answers that you want asked about everything.
  • Bulk labeling and analytics. Tag every row in a warehouse table. This is the Databricks use case, and it's where the free-output pricing matters most.
  • CI and scripts. Jev's CLI turns a yes/no into an exit code, so "does this diff touch auth code without a test?" can gate a pull request.

Bad fits:

  • Anything that needs arithmetic, counting, or comparing dates.
  • Anything where the answer isn't in the input.
  • Anything where you need to know why. Decision models give no explanation, so if a human has to audit the reasoning, you need text.
  • Fixed, high-volume tasks with lots of labeled data. A fine-tuned encoder will be faster and probably as accurate (see Red Hat's numbers).

Where I've landed

Strip off the branding and a decision model is three ideas we already had, combined properly for the first time: read the input once (BERT), put the labels in the prompt (zero-shot NLI), and use a big pretrained language model as the reader (LLMs). The new ingredient is training for calibration, so the probability you get back is a number you can set a threshold on.

It isn't a smarter model. It's a cheaper, more honest answer to a smaller question. I think that's the more interesting story anyway. Most of the "AI" in a production system isn't writing essays. It's a thousand small yes/no checks running between the essays, and until now we've been paying essay prices for them.

What I'd watch next: whether OpenAI publishes calibration numbers for Luna's decisions, whether the open 27B models get distilled into something in the 1B range that runs on a laptop, and whether anyone builds the obvious hybrid that lets a decision model ask for a few reasoning tokens when it isn't sure.

Sources