Everyone's testing Jev with flashy experiments. We put it to work in logistics.
10 minutes read
Jev came out on September 15. Within a week, every feed we follow was full of it. Sorting GitHub issues. Classifying tweets. Benchmarks of other people's benchmarks. (Fun to watch.)
We wanted a less viral answer: what does it do on a desk where the decisions cost money? So we put it on the work we know best. A freight exception desk.
Here is what lands on that desk, all day long:
"Truck on MF-48213 running about 3h behind, new ETA 07:40."
One email. Eight questions, before anyone writes a reply.
Which open load is it about. Is it a delay, damage, or a document hold. How urgent is it. Do the TMS, carrier and portal ETAs still agree. Does the invoice match the rate confirmation. Did the delivery note report damage. Does this need a person. Is the agent's next step safe to run.
Not one of these needs a model that writes. They need a fast, right answer from a fixed list. That is exactly what Jev was built for.
What Jev actually is
A model that answers questions and never writes a word. (Takes a minute to get used to, after two years of chatbots.)
It comes from TypeSafe AI, released on September 15, 2026. TypeSafe calls it a "System One model". You send it a state (text, a JSON object, or a list of messages) and a set of typed questions. It sends back an answer and a probability for each one.
Three question types, that is all:
- Choice: pick one option from a list you define, up to 255 options.
- Score: place the input on ordered levels, such as low, medium, high.
- Noul: a yes or no question, answered with the probability that it is true.
Everything in one request runs in parallel, so eight questions take about as long as one. At the time of our test it cost $0.042 per million input tokens, with output free. No replies. No plans. No tool calls.
The test
185 exception cases, written by us and modeled on the exceptions our logistics clients handle every day: carrier emails, EDI 214 messages, driver texts, TMS alerts, invoices, delivery notes. 67 of them hard on purpose. Typos. Two load numbers in one message. PO numbers one digit apart. ETAs 25 minutes apart across time zones.
Jev against Claude Haiku 4.5. Same cases, same eight questions, three repeats each. 555 calls per model.
| Decision | Jev question type | What the answer drives |
|---|---|---|
| Exception type (12 types) | Choice | Which playbook runs |
| Urgency | Score | Queue order and alerts |
| Needs a person (the stop rule) | Noul | Human review before any action |
| Which open load the message is about | Choice | The record the agent works on |
| Do the TMS, carrier and portal ETAs agree within 30 minutes | Noul | Rebook now or wait |
| Does the carrier invoice match the rate confirmation (within $1) | Noul | Pay, short-pay, or dispute |
| Does the delivery note report damage, shortage or overage | Noul | Open an OS&D claim |
| Is the agent's proposed action safe to run without review | Noul | Execute or hold |
One request with three of those questions looks like this:
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient() # reads TYPESAFE_API_KEY
case = {
"channel": "carrier-email",
"body": "Truck on MF-48213 running about 3h behind, new ETA 07:40.",
"etas": {
"tms_eta": "2026-09-18T04:30",
"carrier_eta": "2026-09-18T07:40",
"portal_eta": "2026-09-18T04:30",
},
}
r = client.system_one(
state=case,
questions={
"exception_type": Choice(
instructions="What kind of exception is this?",
criteria={
"late-eta": "The carrier reports an ETA later than the window.",
"carrier-no-show": "The carrier will not make the pickup or delivery.",
"other": "None of the listed types.",
},
),
"urgency": Score(
instructions="How urgent is this case?",
criteria=["Low", "Medium", "High", "Critical"],
),
"etas_agree": Noul(
instructions="The TMS, carrier and portal ETAs are within 30 minutes of each other.",
),
},
)
print(r.answers["exception_type"].choice)
print(r.answers["urgency"].score)
print(r.answers["etas_agree"].noul) # probability the statement is true
4x faster. 28x cheaper.
Time and cost per case
Each row is scaled to Haiku's value. Cost at list prices on September 21, 2026.
Scale that to a busy desk and the gap stops being a rounding error. Jev also gave the same answer across its three repeats more often: 97.9% of cases, against 93.9% for Haiku.
Fast and cheap is worth nothing if it is wrong, though. So, accuracy.
Where Jev wins: anything you can check against the data
Accuracy per decision
Share of correct answers per decision, across all repeats.
The two biggest gaps were the checks that compare fields. ETA agreement: Jev 90.0%, Haiku 56.7%. Invoice against rate confirmation: Jev 100%, Haiku 66.7%.
The Haiku number surprised us, so we looked closer. It said "no" almost every time. 28 of 30 ETA cases. 25 of 30 invoices. The right split was 15 and 15.
It had to answer at once, with no step to work out the time gap or add up the totals. (With a reasoning step it would likely do better. It would also be slower and cost more. We did not test that.)
Jev also led on the two safety questions. Is the agent's proposed action safe to run: 83.2% against 65.2%. Does this case need a person: 77.7% against 59.6%. On which load and on damage, both models were perfect, hard cases included. Jev just got there in about a quarter of the time.
Where it loses: reading between the lines
Urgency: Haiku 88.1%, Jev 76.0%. Exception type: Haiku 94.2%, Jev 88.1%. On the hard cases Haiku got the type right 99.5% of the time.
Tone. A deadline buried in the third paragraph. A polite email that hides a real problem. That is where a language model still earns its money.
And anything that writes or acts stays with the LLM anyway: replies to carriers and customers, TMS notes, claim descriptions, rebooking a load. Jev answers. It does not act.
The part nobody posts: "does this need a person?" is not a yes or no
At the default cut of 0.5, neither model is safe as the only gate. Jev let through 34% of the cases that needed a person. Haiku let through 42%.
Jev returns a probability, so the cut can move. (Haiku, the way we ran it, returned true or false. Nothing to move.) Drag it and watch what happens:
Guardrail (Safe to run) at a 0.7 cut: 1 of 129 risky cases let through (0.8%), 54.4% of cases sent to a person.
The interactive slider needs JavaScript. See the table below for every cut.
| Check | Cut | Risky cases let through | Cases sent to a person |
|---|---|---|---|
| Safe to run (guardrail) | 0.5 | 18 of 129 (14.0%) | 33.3% |
| Safe to run (guardrail) | 0.6 | 5 of 129 (3.9%) | 47.0% |
| Safe to run (guardrail) | 0.7 | 1 of 129 (0.8%) | 54.4% |
| Needs a person | 0.5 | 42 of 132 (31.8%) | 31.7% |
| Needs a person | 0.3 | 26 of 132 (19.7%) | 45.9% |
| Needs a person | 0.2 | 12 of 132 (9.1%) | 62.7% |
On a concrete action, a stricter cut catches almost every unsafe one, and about half the actions still run without review. On the open question "does this need a person", Jev needs a cut so strict that most cases go to a person anyway. These numbers come from a second Jev run on the final test set (555 calls).
So our reading is simple. Use Jev as a guardrail on specific actions. Write the stop rule as plain rules on money, reversibility and holds. Not as one yes or no question to a model.
The desk we would build
- JevEvery inbound message gets its exception type, its load and the data checks, in one call, in under half a second.
- RulesExplicit rules on money, reversibility and holds decide what needs a person.
- LLM agentReads the case, writes the replies and proposes the next action.
- JevChecks each proposed action before it runs.
- PersonAnything below the cut waits for review.
How to start
- Pick one structured, frequent decision. An ETA check or an invoice match is a good first one. You can count how often it happens and check every answer against your data.
- Write the answer set first. Jev needs the options up front. Writing them down also shows where your own team disagrees.
- Run it in shadow. Send last month's cases through Jev and compare its answers with what your team did. Do not act on its answers yet.
- Set the cut from that log. Choose the probability below which a case goes to a person or to an LLM, based on the misses you can accept.
- Keep the LLM for the rest. Let Jev do the fast checks and the guardrail, and let the LLM read, reason and write.
How we tested
Show the full method
The cases
185 exception cases, 67 of them hard, written by us and modeled on the exceptions our logistics clients handle every day. They cover carrier emails, EDI 214 status messages, driver texts, TMS alerts, customer emails, delivery notes, and invoice and rate confirmation pairs.
Hard cases include typos and broken English, two load numbers in one message, PO numbers one digit apart, ETAs 25 to 35 minutes apart, time zones, invoices off by rounding against invoices missing an accessorial, and damage reported and then taken back. All names and companies are fictional.
The correct answers
For load match, ETA agreement and invoice match, code computes the answer from the data. For the other five decisions, we wrote the labeling rules before any model ran, labeled every case by those rules, and ran two blind label reviews on 60 cases.
What the models see
Only what a real TMS, email inbox, portal or invoice system would hold: the message text, load IDs, ETAs, amounts and line items, temperature readings, deadlines, and the action the agent proposes.
Our first run leaked answers. A few fields in the test data were labels, such as a field that said whether the invoice matched. We found them by reading the cases where both models failed and then checking every field the models see. We moved every label out of that data and added tests that fail if one comes back. If you build your own test, check this first.
The runs
Both models got the same case data and the same question text. One call per case with all of that case's questions. Three repeats per model, 555 calls each, sent one at a time from one machine. The load question counts 525 answers and the ETA, invoice and damage checks 90 each, because not every case carries every question.
Haiku was claude-haiku-4-5 with structured output and no extended thinking. Cost comes from the token usage each API reported, at list prices on September 21, 2026 (Haiku 4.5: $1 per million input tokens, $5 per million output tokens).
Limits of this test
- The cases are ours. 185 cases written by Cryenx and modeled on real exception work. They are not client data. Your cases will differ, so run your own before you rely on these numbers.
- Haiku answered without a reasoning step. With extended thinking, Haiku would likely do better on the ETA and invoice arithmetic, and it would be slower and cost more. We did not test that.
- One LLM baseline. We compared against Claude Haiku 4.5 only.
- Timing is from one machine. Calls were sent one at a time and include network time. Your numbers will vary with location and load.
- Jev is in early access. TypeSafe says rate limits can change without notice. Prices are as of September 21, 2026.
- Four ETA cases had a hint in the text. In the run above, 4 of the 30 ETA cases still had a phrase such as "count the minutes" in the message. Both models saw it. Jev's second run without those phrases scored 91.1% on ETA agreement, against 90.0% with them.
If you run a freight desk and want to test this on your own cases, talk to us.


