Something odd happened at the end of September. Three companies shipped AI models that can't write a sentence. TypeSafe AI released Jev on September 15. Cloudflare followed with Clef on October 1. And on October 6, OpenAI moved its Decisions API into public beta.
You give one of these models a situation and a handful of typed questions. It gives you back a choice, a score, or a probability. No prose, no explanation.
That sounds like a step backwards until you look at what enterprise AI systems actually spend their calls on. Most of it isn't writing. It's sorting: which team gets this ticket, is this invoice a duplicate, does this message need a human. Most teams handle that by asking a large language model to answer in JSON and paying for every token of the answer. It works. It's also slow, the confidence numbers don't mean much, and every so often the JSON comes back broken.
Decision models skip the writing step. They answer several questions in one pass, they return probabilities you can put a real threshold on, and none of the three charge for output tokens. Below is what each one looks like on the wire, using the same support-ticket example throughout.
Jev: the first mover
TypeSafe was founded by Diogo Almeida, a former OpenAI researcher, and Jev is a hosted, text-only API. A request has three parts: a state (the thing you're judging), a model, and a map of named questions. Each question is one of three types. A choice picks one of up to 255 options, a score rates something on an ordered scale of 2 to 10 levels, and a noul (TypeSafe's name for a yes/no question) returns the probability that the answer is yes.
curl https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": "Our payouts have failed for three days and finance is asking why.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices, refunds, payouts",
"technical": "Bugs, outages, integrations",
"other": "Anything else"
}
},
"is_urgent": {
"type": "noul",
"instructions": "Does this need a response today?"
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Annoyed", "Very angry"]
}
}
}'Answers come back keyed by the same names. A choice returns the pick plus a probability for every option and a confidence value. A score returns a probability-weighted value, so 1.4 means "somewhere between Annoyed and Very angry", along with the distribution behind it. That distribution is the part you'll actually use, because it's what you set thresholds on.
"frustration": {
"type": "score",
"score": 1.4,
"legend": { "0": "Calm", "1": "Annoyed", "2": "Very angry" },
"probabilities": { "0": 0.05, "1": 0.5, "2": 0.45 },
"confidence": 0.62
}Jev is the cheapest of the three at $0.042 per million input tokens, and TypeSafe says answers come back in 70 to 500 milliseconds. We'd read TypeSafe's own benchmark carefully, though. It grades Jev on how often it agrees with answers from frontier models, which isn't the same as being right. Independent results so far are mixed. In one test, letting Jev decide first and passing only the low-confidence cases to a frontier model matched that model's accuracy for about a quarter of the cost. In a phishing test, Jev finished well behind a small general-purpose model. Our read: a good first filter, a bad last word.
Clef: the one you can download
Cloudflare released two sizes, a 27-billion-parameter Clef built on Qwen and a 9-billion-parameter Clef-flash, under Apache 2.0 on Hugging Face. Both also run on Workers AI, take up to 64K tokens of context, and accept up to four images per request. Cloudflare matched Jev's request format on purpose. The body above works almost unchanged. You swap the URL, the key, and the model name.
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-flash \
-H "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
-H "Content-Type: application/json" \
-d @request.json # same state + questions, with "model": "clef-flash"Leave the model field as jev-latest and Workers AI rejects the request, so that's the one line people forget. The REST response wraps the answers in Cloudflare's usual result envelope. Inside a Worker, the binding hands you the answers directly:
const { answers } = await env.AI.run("@cf/cloudflare/clef-flash", {
model: "clef-flash",
state: ticket.body,
questions,
});
if (answers.is_urgent.noul > 0.8) await pageOnCall(ticket);Clef costs more: $0.24 per million input tokens, or $0.09 for Clef-flash. Cloudflare puts median latency at 209 ms and 39 ms, and says Clef beats Jev on classification, routing, and tool selection, while Jev still wins the reasoning-heavy tests. Two caveats. Cloudflare ran those numbers itself and nobody has reproduced them yet, and developers testing from outside the US have seen noticeably slower responses once network time is counted. Cloudflare also hasn't released the training data. If you want to host Clef yourself, budget for real GPUs: about 85 GB of memory for Clef and 41 GB for Clef-flash, by Cloudflare's own figures.
OpenAI's Decisions API: an endpoint, not a new model
OpenAI took a different route. The Decisions API is an endpoint, POST /v1/decisions, that runs on gpt-6-luna. It was announced at DevDay on September 29 and went into public beta on October 6. The ideas are the same, but the request format isn't. Questions are an array rather than a map, the yes/no type is called predicate, and choices and levels are lists of objects instead of plain maps.
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "Our payouts have failed for three days and finance is asking why.",
"questions": [
{
"type": "choice",
"name": "team",
"instructions": "Which team should handle this?",
"choices": [
{ "value": "billing", "description": "Payments, invoices, refunds, payouts" },
{ "value": "technical", "description": "Bugs, outages, integrations" },
{ "value": "other", "description": "Anything else" }
]
},
{
"type": "predicate",
"name": "is_urgent",
"instructions": "Does this need a response today?"
}
]
}'The response is an answers array, and each answer carries the question's name. One thing to handle from day one: an answer can come back with type set to refusal, which means the model declined that question. Check for it before you read the probability, or your router will quietly treat a refusal as a zero.
It takes text and images, though images have to be sent inline as base64 data URLs. OpenAI says it's about 10x faster than doing the same job through the Responses API, but it hasn't published latency numbers. Pricing, as reported from its documentation, is $0.10 per million input tokens, with output and cache tokens free and the usual regional and long-context multipliers on top.
For an enterprise, the price isn't the interesting part. The controls are. Eligible customers get Zero Data Retention and HIPAA eligibility, and data can stay in the US or in the EEA and Switzerland. If your legal team has already approved OpenAI, that's a much shorter conversation than onboarding a new vendor. Just remember it's a beta, and the request format can still change.
Picking one
We'd ask three questions, in this order. Where does the data have to live? If the answer is inside your own infrastructure, Clef is the only option of the three. What are you already signed up for? A company running on OpenAI under an enterprise agreement gets a decision layer without a new procurement cycle. Only then: what does it cost? At about 400 input tokens per decision, 100,000 decisions come to roughly $1.68 on Jev, $3.60 on Clef-flash, $4.00 on OpenAI, and $9.60 on Clef at list prices. That's small enough that cost almost never settles it. How well the model does on your own data does.
Don't marry a provider
Two of the three share a format and the third doesn't, and one of them is in beta. That's a good reason to describe your questions once in your own code and translate at the edge. The adapter is short:
type Question =
| { kind: "yesno"; name: string; instructions: string }
| { kind: "choice"; name: string; instructions: string; options: Record<string, string> }
| { kind: "score"; name: string; instructions: string; levels: string[] };
// Jev and Clef: a map of named questions
export function toSystemOne(questions: Question[]) {
return Object.fromEntries(
questions.map((q) => {
if (q.kind === "yesno") return [q.name, { type: "noul", instructions: q.instructions }];
if (q.kind === "choice") return [q.name, { type: "choice", instructions: q.instructions, criteria: q.options }];
return [q.name, { type: "score", instructions: q.instructions, criteria: q.levels }];
}),
);
}
// OpenAI: an array, with lists of objects for options and levels
export function toOpenAI(questions: Question[]) {
return questions.map((q) => {
if (q.kind === "yesno") return { type: "predicate", name: q.name, instructions: q.instructions };
if (q.kind === "choice") {
const choices = Object.entries(q.options).map(([value, description]) => ({ value, description }));
return { type: "choice", name: q.name, instructions: q.instructions, choices };
}
const levels = q.levels.map((label) => ({ label, description: label }));
return { type: "score", name: q.name, instructions: q.instructions, levels };
});
}Normalize the answers on the way back too, so the rest of your code only ever sees a value and a confidence. That one decision is what makes switching providers, or running two side by side, an afternoon of work instead of a migration.
Cascade, don't replace
Where should a decision model sit in your system? In the middle. Plain code should still own the rules: permissions, policy, arithmetic, anything that actually executes. A language model should still own the writing. The decision model takes the judgment calls in between, and it should be allowed to say "I'm not sure".
const THRESHOLD = 0.85; // calibrate per question on labeled data
const decision = await decide(provider, ticket.body, questions);
if (decision.team.confidence >= THRESHOLD) {
await routeTo(decision.team.value);
} else {
// a larger model or a person makes the call; log both for evals
await escalate(ticket, { suggested: decision.team });
}That threshold isn't a magic number. Set it per question, from labeled production examples, based on what a wrong answer costs in each direction. A missed fraud flag and an annoyed customer aren't the same mistake. And keep the hard limits in code: these models are weak at math, dates, and negation, and text hidden in the input can push them around. We wouldn't let one be the only thing standing between a request and a payment, an access change, or anything safety-critical.
How we'd roll one out
Start with the boring part. Find every model call in your system whose answer ends up as a yes/no or a pick from a list. Pull a few hundred real examples from production and label them. Run the decision model in shadow mode next to whatever you use today, log both answers, and don't switch anything until you've compared them. Then turn on the cascade one question at a time.
It's worth remembering how new all of this is. Most of the benchmarks come from the vendors, one of the three products is still in beta, and prices haven't had time to settle. The teams that get the most out of decision models probably won't be the ones that redesign everything around them. They'll be the ones that quietly replace their slowest, most repetitive judgment calls, check the numbers, and go from there.
Where we stand: ShipLabs is an OpenAI Select Partner. Everything here comes from public documentation and third-party reporting as of October 10, 2026. Latency and benchmark figures are the vendors' own unless we say otherwise. Request examples follow each vendor's published format at the time of writing; check the current docs before shipping, especially for the Decisions API while it's in beta.
