Pick One
One company shipped a decision model on 15 September. Seven others had shipped their own by 1 October. What decision models are, and which parts of the launch coverage do not survive checking.

One company shipped a new API surface on 15 September. Seven others had shipped their own by 1 October. What decision models are, and which parts of the launch coverage do not survive checking.
On 15 September TypeSafe came out of stealth with a model called Jev and $40m from DCVC. Within sixteen days Convai, Together AI, Fastino, InternLM, Cloudflare, AWS and Perplexity had all shipped one, OpenAI had previewed a Decisions API at DevDay, and the community leaderboard listed around seventy open entrants.
The coverage called it convergence. The vendors’ own posts read more like copying. Cloudflare names Jev in its opening paragraph, calls Clef “fully Jev-API compatible” and scores itself on Jev’s benchmark. AWS describes the category as one that “has been gaining a lot of attention since TypeSafe AI’s launch of Jev earlier this month”. That tells you more than parallel invention would, because it says the interface was the only part anyone had to work out.
What a call looks like
You hand the model a piece of state, a support ticket, a log line, a signup form, or in Cloudflare’s case an image. You hand it a question and the complete list of answers you will accept. It returns a probability for each one. No prose, and no model-written JSON that can come back broken. Three question shapes have settled across every vendor that publishes a schema. Pick one option from a set you define, score against an ordered rubric, or answer yes or no with a probability.
Here is the whole thing, routing a support ticket with TypeSafe’s Python SDK. State goes in, a typed question for each decision you want made, and a probability distribution comes back.
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient()
response = client.system_one(
state="Hi, I've been trying to connect my Stripe account for 3 days "
"and the integration keeps failing. I'm losing sales. Please help ASAP.",
questions={
"department": Choice(
instructions="Which team should handle this",
criteria={
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions",
},
),
"frustration": Score(
instructions="How frustrated the customer appears",
criteria=[
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language",
],
),
"is_urgent": Noul(
instructions="The message conveys urgency or time-sensitivity",
),
},
)
answer = response.answers["department"]
print(answer.choice) # "technical"
print(answer.probabilities) # the distribution across billing / technical / sales
print(answer.confidence) # how peaked that distribution is
print(response.answers["frustration"].score) # 1.0
print(response.answers["is_urgent"].noul) # 1.0
The pick is constrained to your three departments, so it cannot come back as a fourth thing, and the number behind it is what you threshold on. Nothing to parse.
Three tickets through that same call, to show how the distribution moves. The states are realistic, the probabilities are illustrative, since TypeSafe publishes no example numbers.
1) "Your API has returned 500s on every call since 9am and checkout is down."
department -> "technical" {technical: 0.95, billing: 0.03, sales: 0.02} confidence 0.95
frustration -> 1.8 {0: 0.04, 1: 0.12, 2: 0.84} confidence 0.84
is_urgent -> 0.98
verdict: confident. Route to technical, flag urgent, no human needed.
2) "Do you offer annual billing, and is there a nonprofit discount? No rush."
department -> "sales" {sales: 0.90, billing: 0.08, technical: 0.02} confidence 0.90
frustration -> 0.1 {0: 0.90, 1: 0.09, 2: 0.01} confidence 0.90
is_urgent -> 0.04
verdict: confident. Route to sales.
3) "I got charged twice this month and now the usage webhook isn't firing."
department -> "billing" {billing: 0.52, technical: 0.46, sales: 0.02} confidence 0.52
frustration -> 1.1 {0: 0.30, 1: 0.50, 2: 0.20} confidence 0.50
is_urgent -> 0.60
verdict: 0.52 is a coin toss between billing and technical. Send it to a person.
The third is the case the whole post is about. The pick is still a valid option, but the number behind it is barely a majority, and that is the signal to escalate instead of act.
Why it is fast and cheap
Generating text means a pass over your input and then one decode step per output token, each step dragging the weights and a growing cache through memory. A decision model does the first pass and reads the answer off the hidden states. One pass, no loop. Prices run $0.042 per million input tokens for Jev and $0.24 for Clef, with nothing charged on output because there is no output. Strands Decider is free if you can host 2B parameters, which you can on a laptop.
The claim worth testing
Speed and price got the headlines. Calibration is the actual pitch. Cloudflare’s post-training “utilizes label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration”. TypeSafe trains for the same property with a method it calls Reinforcement Learning for Calibrated Decisions. The probability comes out of training rather than out of a sentence the model wrote about itself.
That argument proves less than it sounds like. A proper scoring rule reports true probabilities at the population optimum, and that does not survive the trip through a finite model and gradient descent. Błasiok and colleagues called exactly this inference folklore at NeurIPS 2023. Cross-entropy is proper too, and networks trained on it are famously overconfident. So treat calibration as something to measure rather than something you bought.
Measure it against what, though. Here the category has a real case. Ryan Porter at Anthus ran 3,600 reasoning problems across ten models on 1 October. On the answers where GPT-6 Luna’s confidence was 99% or higher, it was right 68% of the time. Jev was right 98.9%. Luna’s expected calibration error came out at 0.32 against Jev’s 0.03 to 0.04, and by five reasoning steps Luna’s AUROC had fallen to 0.51, the coin-flip line.
Porter did not ask Luna how sure it was. He pulled token log-probabilities off the standard chat API, which is the fairer measure and usually the better-calibrated one, so the result is worse than it looks. One author, one run, not replicated, and Anthus sells implementation work in this space. But if you route work on a confidence score you read off a chat model, that is the finding to act on this week, and acting on it costs nothing.
The part nobody tells you
Everyone’s advice, mine included, has been to put a “none of these” option on the menu so the model has somewhere to go when nothing on your list fits.
A paper published on 30 September tested that. The setup is pairs of problems. In one the right answer is among the options. In the other it has been removed, so the rejection option is the correct response. With that option sitting right there, Jev picked the right answer 99% of the time when it was present, and picked the rejection option 7% of the time when it was not. The other 93% it chose a wrong option from the list, and the gap held across number sizes, problem depths and different wordings for the label.
The reason is mechanical. These models return a probability for every option and the caller takes the highest one. A highest one always exists, even when every option is wrong, so nothing in that step can refuse.
What catches it is the score itself. With the right answer missing, no option scores highly, and that spread is visible in the output even though it never reaches the choice. Look at the number behind the pick and treat anything below a cutoff as a rejection. Correct rejection goes from 7% to 79%.
So the escape hatch gives the model somewhere to go and the model does not go there. The threshold is what does the work, and you can only set a threshold from your own labelled decisions, which is the one input no vendor can ship you.
What I would do
Four ways to respond. Keep your frontier model and Structured Outputs, and keep paying decode latency on every trivial decision with a confidence number that may mean nothing. Buy hosted, which is the lowest effort, with your state leaving your network on every call and no published architecture to audit. Self-host Clef-flash or Strands Decider, Apache 2.0, a version you pin, with GPU capacity and the evaluation on you. Or treat it as an interface and run a stock small model or a zero-shot encoder, spending nothing on a vendor.
What I would actually do is keep testing. Run these against my own problems, on my own data, and look at what comes back. Make it a standing habit rather than a one-off procurement exercise, because the field is weeks old and today’s answer will not be next quarter’s.
So pick one decision you already make thousands of times a day. Instrument it for a month until you have a labelled log. Run Clef-flash or Strands Decider against it, both free, and measure the calibration error on your own data rather than trusting a leaderboard a competitor ran. Set the threshold from what you find. Then do it again when the next model lands.
The use case I believe in is narrow, and it is everywhere. A problem-specific solution with low latency and high accuracy on a standard classification problem has a very large number of homes in any real system. Routing, triage, admission, scoring, flagging. That is where this pays, and it pays whichever vendor you end up with.
Type safety is a nice to have. It is pleasant to work with, and it will find its way into the larger models as a convenience feature soon enough. Do not build a strategy on it.
On accuracy, be realistic in both directions. None of these will ever be a 100% solution. Neither is a frontier LLM. What you get depends on training and optimisation, on your data and your thresholds, and that holds for every option on the list above. Judge them on your own numbers rather than on the category they belong to.
And keep the age of all this in mind. Sixteen days from one launch to seven clones, and the category is barely a month old. Whatever is true today about who is fastest, cheapest or best calibrated has a short shelf life. There is more coming.
Sources
- Cloudflare, “Introducing Clef: our open-source decision models, and new RL fine-tuning platform”, 1 October 2026. “Fully Jev-API compatible”, label-smoothed cross-entropy plus a Brier loss, the non-autoregressive decision step, Apache 2.0.
- Cloudflare Workers AI pricing. Clef $0.240 and Clef-flash $0.090 per M input tokens, output free.
- Strands Agents, “Introducing Strands Decider”, 1 October 2026. The “gaining a lot of attention since TypeSafe AI’s launch of Jev” framing; a 2B self-hosted decision model.
- Hugging Face,
StrandsAgents/strands-decider-2B-hobson-v19. Apache 2.0; small enough to run on a laptop. - TypeSafe, “Introducing System One Models and Jev”, and docs including the quickstart and primitives pages. Out of stealth on 15 September with a $40m seed led by DCVC; the
system_oneSDK example; $0.042 per M input tokens, output free; the Reinforcement Learning for Calibrated Decisions method; the definitions ofprobabilitiesandconfidence. - OpenAI, DevDay 2026 recap. The Decisions API preview on GPT-6 Luna; no published schema, price or limits.
- OpenAI, “Introducing Structured Outputs in the API”, 6 August 2024. The schema-valid-output baseline decision models are weighed against.
- Błasiok, Gopalan, Hu & Nakkiran, “When Does Optimizing a Proper Loss Yield Calibration?”, NeurIPS 2023 (arXiv 2305.18764). Proper-loss-implies-calibration named as folklore, with a counterexample.
- Anthus (Ryan Porter), “The OpenAI Decisions API Needs a Confidence You Can Trust”, 1 October 2026. 3,600 problems across ten models: GPT-6 Luna ECE 0.32 vs Jev 0.03 to 0.04, 68% correct at 99%+ confidence, AUROC falling to 0.51 by five reasoning steps, confidence read from token log-probabilities. Single author, unreplicated; Anthus sells implementation work in the space.
- Zhong, Li & Lai, “When the Right Answer Is Missing” (arXiv 2609.39496), 30 September 2026. With a rejection option supplied, 99% answer-present accuracy and 7% correct rejection; a decision threshold moved rejection to 79%.
- Launch-window roster and Nadir vendor tally: Convai Laya, Together Tev1, Fastino GLiNER2.5-Decide and GLiDE, InternLM Intern-Decision, Perplexity pplx-decider-v1-27b; roughly 70 open entrants on the community index by 2 October.


