Dev

Decision models vs LLM-as-a-judge: Jev, OpenAI's Decisions API, and who calibrates the probability

Seven places vendors want a typed answer to replace a free-text judge, what Red Hat's benchmark found, and why every vendor in the category, including the one that sells calibration, ends its documentation by handing the threshold back to you.

A decision model is a language model that has been forbidden to write. You give it a document and a typed question, and it returns a typed answer: a probability that a statement is true, a distribution over the options you listed, or a position on a scale you defined. There is no free-form prose to parse, though the answer can still be wrong. On October 6 OpenAI shipped one as a dedicated endpoint, the Decisions API, running a single model called gpt-6-luna. It joins TypeSafe AI's Jev, which opened in September, and two open-weight models, Convai Innovations' Laya and Strands Decider 2B from AWS Strands Labs.

The vendors sell speed and price, and both are real. The part I want to examine is the number attached to every answer. A probability looks like something software can act on directly: route if above 0.8, block if above 0.95. Before I let code branch on 0.8, I want to know who checked that 0.8 means eight in ten. Read the four vendors' own documentation to the end and each gives the same answer, including the one whose product is calibration: you do.

What is a decision model, and how do Jev, Laya, Strands Decider and OpenAI's Decisions API differ?

A decision model takes unstructured input plus one or more typed questions and returns typed decisions with numeric scores. The four compared here share three primitives: a yes/no predicate, a pick-one choice, and an ordinal score. They differ on weights, input types and price.

OpenAI Decisions APIJev (TypeSafe AI)Laya (Convai Innovations)Strands Decider 2B (AWS)
Weightsclosed, API onlyclosed, API onlyopen, Apache 2.0open, Apache 2.0
Inputtext and imagestexttexttext
Price per 1M input tokens$0.10, no output charge$0.042, output freeself-hostedself-hosted
Sizenot publishednot published421M (ModernBERT-large base)about 2B (Qwen3.5-2B base)
Statuspublic betaearly access since Sept 15releasedreleased Oct 1

OpenAI's guide says the API returns typed answers "about 10x faster than the Responses API," a vendor figure with no published method. It also tells you when not to use it: Structured Outputs for generating your own JSON schema, function calling for tool calls. TypeSafe's launch post states the trade plainly: Jev "gives up string generation." Strands Decider is a Qwen3.5-2B whose text-generating head was removed and replaced with a pointer head of about a million parameters that scores the options you supply.

Where do decision models fit? Seven patterns the vendors propose, checked against Red Hat's benchmark

These are the uses the vendors name. Red Hat's comparative benchmark bears on three of them; the rest are reported patterns. Where a claim is the vendor's, I say so.

Can a decision model replace an if-statement in a workflow?

Yes, and that is the use TypeSafe leads with: "smart if-statements" and map-reducing a question over a large dataset. At $0.042 per million input tokens with free output, asking one yes/no question of a million tokens of support tickets costs about four cents, and the answer arrives as a number your code can branch on. What keeps the pitch grounded is Jev's own weakness page, which lists arithmetic, counting, date comparison and hex or RGB values as unreliable and tells you to do those in code. The decision model answers the judgment; your code still does the math.

Can a decision model route requests and pick tools in an agent?

The vendors say yes. Strands names "model routing, tool selection, evaluations, guardrails, memory, context management, and policy classification" as the uses it has seen work. It extends the point I made about worker models in August: match the model to the seat. A seat that only has to choose does not need a model that can write. The caveat is in the Strands model card: "With the state and options fixed, a changed question often gets the same answer." A router that ignores the question you asked looks fine on a test set built from one question. Test it by changing the question, not only the input.

Does a decision model work for content moderation against a plain-English policy?

This is where Red Hat's numbers are most favorable. In its comparison of nine models on class-balanced English datasets, Jev scored highest on content safety at 86.2%, ahead of Qwen3.6-35B as a judge at 85.5% and IBM's 125M Granite Guardian HAP classifier at 80.3%. The Granite model answered in a median 33 ms against Jev's 360 ms, and Jev's figure includes a network round trip the local models did not pay (the authors put it at about 56 ms minimum). TechCrunch reports a moderation-specific entrant, Musubi's open-weight PolicyLM-1.7B, pitched on policies written in plain English that change without retraining. Hold that pitch against the policy section below.

Is a decision model a good prompt-injection filter?

As a filter it is competitive and not the best. On Red Hat's prompt-injection set, Qwen3.6-35B as a judge led at 89.3%, Protect AI's 200M DeBERTa-v3 classifier trained for prompt injection scored 89.0% (54 ms median in the summary table; an appendix table gives 80 ms), and Jev scored 86.4% at a 348 ms median. TypeSafe documents the deeper limit itself:

text
State is data, and jev-1.13 does not treat it as hostile by default.
Content written to adversarially steer the model, whether that is an
injected instruction, a deliberately misleading framing, or text that
argues for its own classification, can move the answer.

That is a vendor being honest, and it is the vendor's own version of the argument in Your approval gate is a guess now: a model that judges by resemblance can rank and filter, and cannot be the boundary. A decision model is the same guess with a cleaner API. The float is easier to threshold than a paragraph, but it is still a judgment about meaning, and meaning is the side of the line where no detector closes the gap. Use it to sort what reaches a person. Gate the actions on what the harness grants.

Can a decision model replace LLM-as-a-judge for evals and grading?

Partly. The score primitive is built for rubric grading, and Strands lists evaluations among its uses. But the failure modes are the familiar ones in a quieter form. Jev's documentation says it "leans toward the option that comes first" in a choice and recommends reordering options to check the answer holds. Laya's card says its yes/no primitive "can follow its option labels instead of the state," and that ordinal scores are its weakest primitive. A judge that writes a paragraph at least gives you a rationale to inspect, though that rationale is not guaranteed to be the reason it decided. A decision model gives you nothing to inspect. As Simon Willison put it about Jev, "the only thing you’re going to get back is a floating point number." For an eval you will audit, that is a cost, not a feature.

Can a decision model triage images?

Among these four, only OpenAI's can. gpt-6-luna accepts images as inline base64 (hosted URLs and file IDs are not supported), and the guide's examples include damage detection. Simon Willison's plugin post shows the shape of an answer, a pelican photo asked whether it contains any mammals:

json
{"type": "predicate", "name": "evaluation", "probability": 0.0}

A clean zero on an easy question is the demo. Whether 0.3 on a hard photo means a 30% chance is the question the rest of this piece is about.

Should you fine-tune an open decision model instead of calling an API?

If you have labels, the open models are built for it. Laya's model card is unusually direct: "Laya is a fast base to specialise, not a zero-shot decision engine." Its base checkpoint scores 0.362 on its own typed-decisions benchmark, below the 0.461 majority-class baseline; the version fine-tuned on that benchmark's training split scores 0.766. Red Hat ran Laya with a policy prompt rather than fine-tuning it, and it came last on content safety at 57.9%. Those are different tasks and different setups, so read them side by side, not as a verdict. Strands published its training scripts, and its model card puts the complete retraining recipe at about 70 minutes on eight H100s.

Can you set a threshold on a decision model's probability?

Not on the vendor's word. Every vendor in the category tells you to validate on your own data before you act on the number, and two go further and tell you to refit the probabilities themselves. Read their statements in order, from the most promised to the least.

TypeSafe sells calibration as the point of the product, trained with a method it calls Reinforcement Learning for Calibrated Decisions:

text
Always communicates confidence and uncertainty with every output.
Calibrated: higher confidence means higher accuracy.

Its launch post and weakness page publish no calibration error figure behind that, and the weakness page's advice is to test your integration thoroughly before deploying it. Strands fitted its calibration on one kind of data and says so:

text
Calibration is one temperature per primitive, fitted on held-out short
classification. The confidence bands are established there only:
measure on your own traffic before you trust a threshold.

Laya publishes the numbers, which is why its card is the most useful of the four:

text
Ships over-confident: Refitting one temperature per (question type,
option count) moves mean ECE 0.466 → 0.081 (laya) and 0.314 → 0.106
(laya-multilingual). Do this on your own data before trusting the
probabilities.

The same card says Laya's act/escalate probability "carries no usable signal yet": it reads 1.0 for almost every input, and its raw signal runs against correctness (AUROC 0.30 on 396 labelled decisions). A field named for exactly the decision you want to automate is the one to ignore.

The card also relays a third-party report, not its authors' own measurement, that Jev assigned zero probability to the true label on 16% of DAIR Emotion examples. A calibrated model can be wrong; a zero says it could not be.

OpenAI says the least. Its guide never defines how the separate confidence field is computed, and its advice on thresholds is brief:

text
Use labeled examples from your application to set thresholds for
routing, filtering, or review.

Three different instructions sit in those quotes, and it helps to keep them apart. Validation tells you how often the model is right on your traffic. Calibration changes what a score means, so that 0.8 really is eight in ten. A threshold decides what you do at a score, and a useful one does not need perfect calibration, only enough labeled examples to see what each cutoff costs you. All three need the same raw material. A decision model still needs labeled data; it needs it at threshold time instead of training time. Refitting a temperature is a far smaller estimation problem than training a classifier, so it needs fewer labels, but it needs them on day one and they have to look like your traffic. What it buys you is a usable first answer before you have trained anything, and a policy you can change by editing text.

Does the policy text matter more than the decision model?

In Red Hat's setup, the policy moved accuracy by more than the gap between most of the models, so treat a policy edit as a model change. Restructuring Laya's content-safety policy, splitting it into separate yes/no questions per harm category and blocking if any answer exceeded 0.5, lifted it from 57.87% to 75.20%, 17.33 points, and more than doubled its median latency, from 118 ms to 289 ms. The same tuned policy, given to Jev, cost Jev 3.67 points.

So a policy is not portable across decision models, and it is not neutral within one. That is the catch in the "no retraining" pitch: the weights do not change, but the accuracy does, in a direction you only learn by measuring. If a policy can be edited in plain English by someone who is not looking at the eval numbers, the eval set has to run on every edit, the way a test suite runs on every commit. One more limit from the same benchmark: it is English only, and it was published four days before the Decisions API went into public beta and did not test OpenAI's model.

Decision model, fine-tuned classifier or LLM judge: which one should you pick this week?

Pick by what the task does to you when it is wrong, and build the labeled set before you pick anything. My read of the evidence so far follows.

A fixed task under attack, prompt injection or a known abuse class, wants a small fine-tuned classifier. In Red Hat's test it came within a third of a point of the best judge at a fraction of the latency, and its behavior does not move when someone edits a prompt. Keep it as a filter, never as the boundary.

A policy that changes often, where an error costs a human review, is where a decision model earns its place. You edit the policy, rerun the eval set, and ship the change the same afternoon, with no training run in between.

A decision you will have to explain, an audit or a grade someone can appeal, wants an LLM judge, or a decision model with a judge behind it for the disputed cases. The judge's rationale is not proof of why it decided, but a bare float gives the person appealing nothing to argue with.

The labeled set is the part that outlives the choice. It sets the decision model's threshold and evaluates any classifier or judge you try against it, as long as you hold part of it back from any fitting, and revisit the labels when the policy changes. These products are weeks old and their models will be replaced; the examples you labeled from your own traffic are the one asset that transfers to whichever model wins. Every vendor's documentation ends by asking you for them, so build them first.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles

Agent memory as Markdown documents: a readable memory is still a memory

Moving agent memory out of a vector store and into Markdown files makes it easy to inspect, and that is worth having. Staleness, conflicting writers and the page nobody opened were never problems of storage format, so they move into the folder intact, where a clean file can pass for a checked one.