Monday, 8:40. There are 28 new emails in a customer service inbox. Three of them:
"Hi, could you tell me when order 4471 will ship? Thanks, Marc."
"Hello, order 5203 was due last week. Can you give me a new date, please?"
"Great, the third delay this month. Thanks a lot."
All three ask the same question, and all three are polite, word by word. Only one comes from a customer who is about to try another supplier. Anyone who has worked in customer service sees which one in a second, once they get to it.
That second is a decision: one input, a short list of possible answers (neutral, frustrated, angry), made hundreds of times a day. People make it well and fast, but only once they open the email. If the third one is number 26 in the pile, it waits until the afternoon.
Until recently, handing this call to an AI meant a chatbot: slow, paid per word, and now and then an answer that is not on your list. On 15 September TypeSafe released Jev, a model that cannot write a sentence. It reads a record, picks an answer from a list you define, and says how sure it is. By 2 October an independent registry counted 30 of these decision models from 30 makers, and most of them accept Jev's request format.
I have written about Jev twice. This post steps back: what decision models are, how they differ from the chat models we know, which ones are on the market, and ten places they fit in supply chain work. At the end we go back to that Monday inbox.
What decision models are
A decision model takes two things: a block of information (an email, a record, a JSON object, sometimes a photo) and a set of questions whose answers you fix in advance. It answers every question in one pass and gives a probability for each allowed answer. It writes nothing else.
There are three kinds of question, and nearly every model in this post supports all three:
- Choice: pick one option from a list, such as a reason code or a queue.
- Score: place the input on an ordered scale, such as low, medium, high.
- Yes/no: the probability that a statement is true.
The difference with a chat model is the shape of the answer, and everything follows from that.
| Chat model | Decision model | |
|---|---|---|
| Output | Text you then have to parse | One of your answers, with a probability |
| Speed | Seconds | Tens to hundreds of milliseconds |
| Cost | You pay for every word it writes | Input only; output is free |
| Typical failure | An answer that is not on your list | A confident answer that is wrong |
| Best at | Explaining, drafting, reasoning | Sorting, routing, flagging at volume |
That last line matters. A decision model cannot invent a category or misspell a status code. It can still be sure and wrong, which is why the probability is the most important part of the answer.
In practice, that gives them five jobs they do well:
- Classify: put an email, a complaint or an item record in the right category.
- Score: rate urgency, severity or risk on a scale you define.
- Flag: say how likely it is that a statement is true ("this invoice is a duplicate").
- Route: choose the queue, team or tool that handles the next step.
- Look: some models also read photos, such as a damaged pallet or a scanned delivery note.
They do all of this in a single pass, so you can ask five questions about one record for the price of reading it once.
The models so far
This is the field as of 4 October 2026, when the System One Models registry listed 30 decision models from 30 makers. The table shows the ones a supply chain team is most likely to meet. Dates, prices and context sizes come from the makers or from the System One Models registry, and all benchmark figures are self-reported.
| Model | Maker | Released | How you get it | Reads up to | Price per million input tokens |
|---|---|---|---|---|---|
| Jev 1.13 | TypeSafe AI | 15 Sep | Hosted API | 32k tokens, text | $0.042 |
| Decider 1 | meraGPT | 22 Sep | Hosted API | 4k tokens | $0.03 |
| Solar Decide | Upstage | 22 Sep, beta | Hosted API | 512k tokens | Beta pricing |
| Tev1 (4B and 0.8B) | Together AI | 23 Sep | Hosted, and locally in Ollama | Not published | $0.042 hosted |
| GLiNER2.5-Decide | Fastino Labs | 24 Sep | Open weights, 340M parameters | Not published | Free to run |
| Nimble (9B) | Bespoke Labs | Late Sep | Open weights, locally in Ollama | 8k tokens | Free to run |
| d1 | Liquid AI | 29 Sep | Hosted API, no weights | Not published | Free tier, paid price not published |
| Jeff 1.1 (0.8B to 2B) | Independent | 29 Sep | Open weights | Small models | Free to run |
| Decisions API | OpenAI | 29 Sep, preview | Limited preview, on GPT-6 Luna | Not published | Not published |
| Mercury Decide | Inception | 30 Sep | Hosted, via OpenRouter | 32k tokens | Free in early access |
| Clef / Clef-flash (27B / 9B) | Cloudflare | 1 Oct | Open weights, and hosted on Workers AI | 64k tokens, text and images | $0.24 / $0.09 |
| Strands Decider 2B | AWS Strands Labs | 1 Oct | Open weights, Apache 2.0, runs on a laptop | Not published | Free to run |
| Kev 1.0 (0.8B to 27B) | Independent | 1 Oct | Open weights | 8k tokens (27B: 64k) | Free to run |
Two tools sit around the models. Ollama 0.35 (28 September) runs decision models on your own laptop. OpenRouter serves several hosted ones behind a single account. Both use the same request format as Jev, so in principle swapping models is a configuration change, not a rebuild. That shared format is what turns a single product into a category.
How they differ from each other
They all answer the same three kinds of question. They differ in five ways that matter for a supply chain team.
- Where it runs. Jev, d1, Decider 1 and Solar Decide only run on the maker's servers. Clef, Kev, Nimble and Jeff can run on your own machine, so supplier prices or customer data never leave the building.
- How much it reads. Decider 1 reads about 4,000 tokens, a long email. Solar Decide reads 512,000, a full contract with its annexes. Pick the model for your longest realistic document, not your average one.
- Whether it sees pictures. Clef accepts up to four images per request on Workers AI. Cloudflare notes that Jev handles text only today. A damaged pallet or a scanned delivery note needs the first kind.
- Whether you can teach it. Nimble was trained on 2,676 labelled examples. Jeff's maker reports one fine-tune on about 11,000 examples took half an hour on one GPU, lifting accuracy from 31.7% to 95.8%. The hosted models offer no fine-tuning today.
- Speed and price. In Cloudflare's own test across 43 benchmarks, Clef-flash took a median 39 milliseconds, Clef 209 and Jev 524. AWS reports 115 milliseconds for Strands Decider on a gaming graphics card. Prices run from free to $0.24 per million input tokens. At those prices the model is never the expensive part of the project.
- Sorting versus judging. An independent comparison by BERI found Clef far ahead on sorting work, with 94.2 against Jev's 79.7 on recognising banking intents. Jev stayed ahead on judgement calls, such as whether an agent should act at all (81.0 against 72.4). Strands Decider was the weakest on hard reasoning. A small model can sort emails well and still be the wrong choice for "should we release this order?".
What it means for supply chain
Here is the frame I use to decide where a decision model belongs. Ask two questions about the decision: how often do we make it, and what does a wrong answer cost?
| Wrong answer is cheap to fix | Wrong answer is expensive | |
|---|---|---|
| Hundreds a day | The model decides, a person checks a sample | The model sorts, a person signs |
| A few a week | Not worth building | Keep it human |
The top row is where these models earn their place.
Ten use cases in supply chain
Each of these is a decision made hundreds of times a day, with a fixed list of answers.
- Customer service mailbox. Intent (order status, change, complaint, quote), urgency, and "is an order number present?", answered for every message before a person opens it.
- Supplier order confirmations. Does the confirmation change the date, the quantity or the price compared with the PO? Only the changed ones reach the buyer.
- Invoice exceptions. Why did the three-way match fail: price, quantity, missing goods receipt or likely duplicate? Each reason goes to the person who can fix it.
- Carrier status messages. Turn free-text updates from carriers ("truck broke down near Lyon, new ETA tomorrow") into your own event codes, with an urgency score.
- Supplier risk news. Does this news item concern one of our supplier sites, and how severe is it on a three-point scale?
- Master data checks. Does this item description fit its commodity group and unit of measure? Thousands of records checked overnight.
- Purchase requisitions. Catalogue, existing contract or new sourcing event, and which category buyer takes it.
- Quality complaints with photos. Defect type and severity from the customer's picture, with an image-capable model like Clef.
- Licence and sanctions pre-screening. Does this product description or end use need a compliance check? Anything above a low probability goes to the compliance desk.
- Inside an agent. Choosing which tool, queue or team handles the next step, so the expensive chat model only runs when it must.
The worked example below goes back to Monday's inbox and builds the first of these, with one extra question that people usually answer by feel: how is this customer feeling?
Worked example: reading the tone of customer emails at Upshift
Upshift is a fictional bicycle maker. Its customer service team receives about 400 emails a day from dealers and consumers, in English, Dutch and French (all numbers here are illustrative). Most are routine. A few are from customers who are one bad reply away from cancelling an order or posting a review. Today those surface when someone happens to open them, which can be the next afternoon.
The team asks four questions of every email, in one call:
| Question | Type | Allowed answers |
|---|---|---|
| What does the customer want? | Choice | Order status, change or cancel, complaint, quote, invoice, other |
| What is the tone? | Score | Neutral, frustrated, angry |
| Is there an escalation signal: a threat to cancel, a competitor, a lawyer, social media? | Yes/no | Probability |
| Is an order number present? | Yes/no | Probability |
Notice what is not on the list: "How late is the order?" That is a lookup, not a decision. The system reads the order number, fetches the promised and current delivery dates from the ERP, and adds "order 6118, 6 days late, second delay" to the input as a fact. A polite email about an order that has slipped twice deserves different handling from a polite email about an order on time, and the model can only weigh that if it is told.
Tone is where a decision model needs testing most. The sarcastic email from Monday is angry, though every word in it is polite. Before going live, the team labels 300 past emails by hand: 100 per language, deliberately including the ones that later turned into a complaint or a cancellation. They check two things. Are answers given at 80% confidence right about eight times in ten? And does that hold in each language, not just on average? If the model is reliable in English and weak in French, French emails get a lower bar for review.
The test sets the routing:
| Route | Rule | Emails a day |
|---|---|---|
| Call back today | Escalation signal above 0.30, or tone angry above 0.60 | 20 |
| Priority queue | Tone frustrated above 0.60, or the order is late for the second time | 60 |
| Standard queue | Everything else, sorted by intent | 300 |
| Automatic reply | Order status, neutral tone above 0.90, order number present and order on time | 20 |
Look at the escalation threshold. It sits at 0.30, not 0.50, because a missed angry customer costs far more than a team lead reading an email that turns out to be fine. Look also at what is automated: only the calmest, simplest case. An angry customer never gets an automatic reply, however sure the model is about the intent.
At about 800 tokens per email, 400 emails is 0.32 million tokens a day. That is under 8 cents on Clef or under 2 cents on Jev. The difference is not the cost. It is that the 20 customers most likely to walk away get a phone call the same day instead of a reply the next afternoon. And because the request format is shared, the team can run the same 300 labelled emails through a second model next month and switch if it reads tone better.
What it can't do, and the traps
- It doesn't calculate or explain. No dates, no totals, no emails. Keep the maths in code and pair it with a chat model for anything that needs words.
- Published benchmarks don't compare. The registry itself warns that makers report accuracy on different tests. Jeff's 2B model scores 82.0 against Jev's 83.0 on its maker's panel, yet one practitioner reported 70% against Jev's 94% on their own task.
- Good averages hide blind spots. A September study of four decision models found Nimble ranked first at spotting prompt injection overall, while missing 92.2% of attacks without direct instructions. Test on your own awkward cases, by category.
- A second model rarely catches the first one's confident mistakes. In the same study, a second judge corrected only 8 to 32 of 166 unsafe misses. Two models are not a control; a sample checked by a person is.
- Thirty models in three weeks means many will disappear. Most are early access, and prices are launch prices. Build on the shared format, keep your labelled test set, and be ready to swap.
Pick the question before the model
This week, choose one decision your team makes more than a hundred times a day. Write down the exact answers people choose between, and label 200 past cases with the right one. That list and that test set are worth more than any model in the table above, because they let you test all of them.
The sarcastic email on Monday morning did not need a cleverer reader. It needed to be read first. That is the job a decision model does well: a fixed list of answers, a hit rate you have measured on your own cases, and a person for the doubtful ones. It will not replace the people who make these calls, but it can take the clear-cut ones and put the urgent ones on top, so people spend their attention where it is needed.
Which decision in your operation would you hand to one of these models first?
Sources
- TypeSafe AI, Introducing System One models and Jev
- eesel AI, TypeSafe Jev review
- System One Models registry
- System One Models, every model, specs and prices
- Awesome decision models list on GitHub
- Cloudflare, Clef decision models, 1 Oct 2026
- MarkTechPost, Cloudflare releases Clef and Clef-flash, 1 Oct 2026
- MarkTechPost, Liquid AI releases d1, 29 Sep 2026
- AlphaSignal, Inception releases Mercury Decide, 30 Sep 2026
- Ollama v0.35.0 release notes
- Bespoke Labs, Nimble on GitHub
- Developers Digest, Jeff: Jev-style decision models trained on one GPU, 30 Sep 2026
- Kev 1.0 release on GitHub, 1 Oct 2026
- AI TL;DR, AWS Strands Decider 2B, 1 Oct 2026
- BERI, Cloudflare Clef and Amazon Strands Decider against Jev
- The New Stack, OpenAI answers Jev with a Decision API built on Luna
- Liu, Evaluating System One models for agent security decisions, arXiv, 27 Sep 2026




