BackShipLabs
Strategic Briefs
AI Architecture
October 10, 202610 min read

Decision Models for the Enterprise: Jev, Clef, and OpenAI's Decisions API

Focus Area
Decision Layers
System Architecture
Jev · Clef · Decisions API
Publisher
ShipLabs Engineering

Something odd happened at the end of September. Three companies shipped AI models that can't write a sentence. TypeSafe AI released Jev on September 15. Cloudflare followed with Clef on October 1. And on October 6, OpenAI moved its Decisions API into public beta.

You give one of these models a situation and a handful of typed questions. It gives you back a choice, a score, or a probability. No prose, no explanation.

That sounds like a step backwards until you look at what enterprise AI systems actually spend their calls on. Most of it isn't writing. It's sorting: which team gets this ticket, is this invoice a duplicate, does this message need a human. Most teams handle that by asking a large language model to answer in JSON and paying for every token of the answer. It works. It's also slow, the confidence numbers don't mean much, and every so often the JSON comes back broken.

Decision models skip the writing step. They answer several questions in one pass, they return probabilities you can put a real threshold on, and none of the three charge for output tokens. Below is what each one looks like on the wire, using the same support-ticket example throughout.

Jev: the first mover

TypeSafe was founded by Diogo Almeida, a former OpenAI researcher, and Jev is a hosted, text-only API. A request has three parts: a state (the thing you're judging), a model, and a map of named questions. Each question is one of three types. A choice picks one of up to 255 options, a score rates something on an ordered scale of 2 to 10 levels, and a noul (TypeSafe's name for a yes/no question) returns the probability that the answer is yes.

Jev · three questions, one callbash
curl https://api.typesafe.ai/v1/systemone \
  -H "Authorization: Bearer $TYPESAFE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-latest",
    "state": "Our payouts have failed for three days and finance is asking why.",
    "questions": {
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "Payments, invoices, refunds, payouts",
          "technical": "Bugs, outages, integrations",
          "other": "Anything else"
        }
      },
      "is_urgent": {
        "type": "noul",
        "instructions": "Does this need a response today?"
      },
      "frustration": {
        "type": "score",
        "instructions": "How frustrated is the customer?",
        "criteria": ["Calm", "Annoyed", "Very angry"]
      }
    }
  }'

Answers come back keyed by the same names. A choice returns the pick plus a probability for every option and a confidence value. A score returns a probability-weighted value, so 1.4 means "somewhere between Annoyed and Very angry", along with the distribution behind it. That distribution is the part you'll actually use, because it's what you set thresholds on.

Jev · shape of a score answer (values illustrative)json
"frustration": {
  "type": "score",
  "score": 1.4,
  "legend": { "0": "Calm", "1": "Annoyed", "2": "Very angry" },
  "probabilities": { "0": 0.05, "1": 0.5, "2": 0.45 },
  "confidence": 0.62
}

Jev is the cheapest of the three at $0.042 per million input tokens, and TypeSafe says answers come back in 70 to 500 milliseconds. We'd read TypeSafe's own benchmark carefully, though. It grades Jev on how often it agrees with answers from frontier models, which isn't the same as being right. Independent results so far are mixed. In one test, letting Jev decide first and passing only the low-confidence cases to a frontier model matched that model's accuracy for about a quarter of the cost. In a phishing test, Jev finished well behind a small general-purpose model. Our read: a good first filter, a bad last word.

Clef: the one you can download

Cloudflare released two sizes, a 27-billion-parameter Clef built on Qwen and a 9-billion-parameter Clef-flash, under Apache 2.0 on Hugging Face. Both also run on Workers AI, take up to 64K tokens of context, and accept up to four images per request. Cloudflare matched Jev's request format on purpose. The body above works almost unchanged. You swap the URL, the key, and the model name.

Clef · same body, different hostbash
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-flash \
  -H "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d @request.json   # same state + questions, with "model": "clef-flash"

Leave the model field as jev-latest and Workers AI rejects the request, so that's the one line people forget. The REST response wraps the answers in Cloudflare's usual result envelope. Inside a Worker, the binding hands you the answers directly:

Clef · from a Workerts
const { answers } = await env.AI.run("@cf/cloudflare/clef-flash", {
  model: "clef-flash",
  state: ticket.body,
  questions,
});

if (answers.is_urgent.noul > 0.8) await pageOnCall(ticket);

Clef costs more: $0.24 per million input tokens, or $0.09 for Clef-flash. Cloudflare puts median latency at 209 ms and 39 ms, and says Clef beats Jev on classification, routing, and tool selection, while Jev still wins the reasoning-heavy tests. Two caveats. Cloudflare ran those numbers itself and nobody has reproduced them yet, and developers testing from outside the US have seen noticeably slower responses once network time is counted. Cloudflare also hasn't released the training data. If you want to host Clef yourself, budget for real GPUs: about 85 GB of memory for Clef and 41 GB for Clef-flash, by Cloudflare's own figures.

OpenAI's Decisions API: an endpoint, not a new model

OpenAI took a different route. The Decisions API is an endpoint, POST /v1/decisions, that runs on gpt-6-luna. It was announced at DevDay on September 29 and went into public beta on October 6. The ideas are the same, but the request format isn't. Questions are an array rather than a map, the yes/no type is called predicate, and choices and levels are lists of objects instead of plain maps.

OpenAI · the same ticketbash
curl https://api.openai.com/v1/decisions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    "input": "Our payouts have failed for three days and finance is asking why.",
    "questions": [
      {
        "type": "choice",
        "name": "team",
        "instructions": "Which team should handle this?",
        "choices": [
          { "value": "billing", "description": "Payments, invoices, refunds, payouts" },
          { "value": "technical", "description": "Bugs, outages, integrations" },
          { "value": "other", "description": "Anything else" }
        ]
      },
      {
        "type": "predicate",
        "name": "is_urgent",
        "instructions": "Does this need a response today?"
      }
    ]
  }'

The response is an answers array, and each answer carries the question's name. One thing to handle from day one: an answer can come back with type set to refusal, which means the model declined that question. Check for it before you read the probability, or your router will quietly treat a refusal as a zero.

It takes text and images, though images have to be sent inline as base64 data URLs. OpenAI says it's about 10x faster than doing the same job through the Responses API, but it hasn't published latency numbers. Pricing, as reported from its documentation, is $0.10 per million input tokens, with output and cache tokens free and the usual regional and long-context multipliers on top.

For an enterprise, the price isn't the interesting part. The controls are. Eligible customers get Zero Data Retention and HIPAA eligibility, and data can stay in the US or in the EEA and Switzerland. If your legal team has already approved OpenAI, that's a much shorter conversation than onboarding a new vendor. Just remember it's a beta, and the request format can still change.

Picking one

We'd ask three questions, in this order. Where does the data have to live? If the answer is inside your own infrastructure, Clef is the only option of the three. What are you already signed up for? A company running on OpenAI under an enterprise agreement gets a decision layer without a new procurement cycle. Only then: what does it cost? At about 400 input tokens per decision, 100,000 decisions come to roughly $1.68 on Jev, $3.60 on Clef-flash, $4.00 on OpenAI, and $9.60 on Clef at list prices. That's small enough that cost almost never settles it. How well the model does on your own data does.

Don't marry a provider

Two of the three share a format and the third doesn't, and one of them is in beta. That's a good reason to describe your questions once in your own code and translate at the edge. The adapter is short:

One question definition, two wire formatsts
type Question =
  | { kind: "yesno"; name: string; instructions: string }
  | { kind: "choice"; name: string; instructions: string; options: Record<string, string> }
  | { kind: "score"; name: string; instructions: string; levels: string[] };

// Jev and Clef: a map of named questions
export function toSystemOne(questions: Question[]) {
  return Object.fromEntries(
    questions.map((q) => {
      if (q.kind === "yesno") return [q.name, { type: "noul", instructions: q.instructions }];
      if (q.kind === "choice") return [q.name, { type: "choice", instructions: q.instructions, criteria: q.options }];
      return [q.name, { type: "score", instructions: q.instructions, criteria: q.levels }];
    }),
  );
}

// OpenAI: an array, with lists of objects for options and levels
export function toOpenAI(questions: Question[]) {
  return questions.map((q) => {
    if (q.kind === "yesno") return { type: "predicate", name: q.name, instructions: q.instructions };
    if (q.kind === "choice") {
      const choices = Object.entries(q.options).map(([value, description]) => ({ value, description }));
      return { type: "choice", name: q.name, instructions: q.instructions, choices };
    }
    const levels = q.levels.map((label) => ({ label, description: label }));
    return { type: "score", name: q.name, instructions: q.instructions, levels };
  });
}

Normalize the answers on the way back too, so the rest of your code only ever sees a value and a confidence. That one decision is what makes switching providers, or running two side by side, an afternoon of work instead of a migration.

Cascade, don't replace

Where should a decision model sit in your system? In the middle. Plain code should still own the rules: permissions, policy, arithmetic, anything that actually executes. A language model should still own the writing. The decision model takes the judgment calls in between, and it should be allowed to say "I'm not sure".

Confident cases go straight through, the rest escalatets
const THRESHOLD = 0.85; // calibrate per question on labeled data

const decision = await decide(provider, ticket.body, questions);

if (decision.team.confidence >= THRESHOLD) {
  await routeTo(decision.team.value);
} else {
  // a larger model or a person makes the call; log both for evals
  await escalate(ticket, { suggested: decision.team });
}

That threshold isn't a magic number. Set it per question, from labeled production examples, based on what a wrong answer costs in each direction. A missed fraud flag and an annoyed customer aren't the same mistake. And keep the hard limits in code: these models are weak at math, dates, and negation, and text hidden in the input can push them around. We wouldn't let one be the only thing standing between a request and a payment, an access change, or anything safety-critical.

How we'd roll one out

Start with the boring part. Find every model call in your system whose answer ends up as a yes/no or a pick from a list. Pull a few hundred real examples from production and label them. Run the decision model in shadow mode next to whatever you use today, log both answers, and don't switch anything until you've compared them. Then turn on the cascade one question at a time.

It's worth remembering how new all of this is. Most of the benchmarks come from the vendors, one of the three products is still in beta, and prices haven't had time to settle. The teams that get the most out of decision models probably won't be the ones that redesign everything around them. They'll be the ones that quietly replace their slowest, most repetitive judgment calls, check the numbers, and go from there.

Where we stand: ShipLabs is an OpenAI Select Partner. Everything here comes from public documentation and third-party reporting as of October 10, 2026. Latency and benchmark figures are the vendors' own unless we say otherwise. Request examples follow each vendor's published format at the time of writing; check the current docs before shipping, especially for the Decisions API while it's in beta.

Verified Performance Outcomes

"List price per 1M input tokens: Jev $0.042 · Clef-flash $0.09 · OpenAI Decisions API $0.10 · Clef $0.24 · Output tokens free on all three"

Underlying Core Architecture

Deterministic Agentic Workflows

Explore how we separate deterministic policy code, bounded decision steps, and generative models inside production agent workflows.

Explore Technical Architecture
Related Strategic Briefs
Finance Operations
Invoice Reconciliation & Payment Controls at Scale

Enterprise finance teams spend thousands of manual hours reconciling supplier invoices against purchase orders, goods receipts, and payment ledgers. Here is how we automated it.

Read Brief
Commerce Operations
Multilingual Voice Orders & Demand Forecasting

High-volume restaurants and service networks lose order potential during peak traffic because their teams cannot answer every call. We built low-latency multilingual voice systems linked directly to demand and procurement forecasts.

Read Brief
Risk & Governance
Regulatory Change Management with Deterministic RAG

Regulatory requirements can change overnight. We built deterministic Graph RAG architectures to monitor authoritative updates and automatically suggest redlines to existing policies and agreements.

Read Brief
Direct Consult

Have similar structural bottlenecks?

We deploy highly deterministic agent systems directly into enterprise database and pipeline topologies. Drop us your specifications or book a technical strategy sync.

Technical Strategy Sync

Book a direct 30-minute sync to dissect your pipeline schemas:

Book 30-Min Call
Drop a Brief

Send details of your architecture directly to our developer team:

hello@shiplabs.app