The question everyone asks about Jev, the decision model TypeSafe AI released on 15 September 2026, is whether it is as good as a real LLM. Most answers so far come from public classification datasets or from TypeSafe's own workflows. We had something better to hand: 317 cases from real web work, labelled by hand earlier the same day, where a wrong answer costs money. On 28 September 2026 we sent every case, with the identical question, to Jev and to five LLMs. Accuracy was a tie. The bill was not. And the most useful difference was one nobody puts in a leaderboard: where each model makes its mistakes.
The test
The 317 cases come from two benchmarks we published today:
- 134 scraper responses from 70 heavily protected sites, labelled usable or not (a block page, a region notice, the wrong page). Details in our block page test.
- 183 businesses from a real scraped lead list, labelled as the buyer's target or not. Details in our lead scoring test.
Every model saw the same state (page text, or the lead's name, categories, description and domain) and the same one-sentence statement. Jev answered through OpenRouter's Decisions API. The five LLMs answered through OpenRouter's chat API with a strict JSON schema asking for a true or false answer and the probability that the statement is true, temperature 0 where supported, and each provider's default reasoning setting. Costs are what OpenRouter billed per call; latency is the time for each call from Europe.
| Model | Scraper pages (134) | Leads (183) | Total correct | Cost per 1,000 | Median latency | 90th percentile |
|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 134 | 175 | 309 | $0.339 | 1.54 s | 4.28 s |
| GPT-6 Luna | 133 of 133 | 175 | 308 of 316 | $0.112 | 2.47 s | 3.45 s |
| GPT-6 Sol | 133 | 175 | 308 | $2.081 | 2.79 s | 3.84 s |
| Claude Haiku 4.5 | 132 | 175 | 307 | $1.071 | 1.21 s | 1.57 s |
| Gemini 3.5 Flash-Lite | 131 | 176 | 307 | $0.273 | 0.86 s | 1.07 s |
| Jev 1.13 | 131 | 175 | 306 | $0.049 | 0.33 s | 0.41 s |
One GPT-6 Luna call failed after retries and is left out of its total. The spread between the best and the worst model is three cases out of 317, less than 1%, and several of the shared mistakes are cases where our own label is arguable (a body shop whose name says it is also an auto electrician was "wrong" for every model). On this kind of work, accuracy does not separate these models.
Cost and speed
What does separate them is the bill. For the same 317 decisions Jev cost $0.0154. GPT-6 Luna cost 2.3 times as much, Gemini 3.5 Flash-Lite 5.6 times, DeepSeek V4.1 Flash 7 times, Claude Haiku 4.5 22 times and GPT-6 Sol 42 times. Jev bills only input tokens; the LLMs also bill their output and, where they reason before answering, the reasoning tokens, which were a median of 100 per call for DeepSeek and 31 for GPT-6 Luna.
Speed follows the same shape. Jev's median answer took 328 ms and its slowest tenth stayed under 405 ms. The fastest LLM, Gemini Flash-Lite, took 862 ms at the median; DeepSeek's slowest tenth took over four seconds. For a batch job that difference is a coffee break. For a decision inside a request, before a page is stored or a lead is shown to a salesperson, it is the difference between usable and not.
Where the mistakes were
This is the part a leaderboard hides. A classifier that is right 97% of the time is only useful if you can tell which 3% to check. So for every wrong answer we looked at the probability the model attached to it.
| Model | Wrong answers | Of those, scored between 0.3 and 0.8 | Scored 0.9 or higher (or 0.1 or lower) |
|---|---|---|---|
| Jev 1.13 | 11 | 11 | 0 |
| GPT-6 Sol | 9 | 2 | 1 |
| DeepSeek V4.1 Flash | 8 | 0 | 2 |
| GPT-6 Luna | 8 | 0 | 6 |
| Claude Haiku 4.5 | 10 | 0 | 5 |
| Gemini 3.5 Flash-Lite | 10 | 0 | 7 |
Every one of Jev's mistakes came with a hesitant score. None of the small LLMs' mistakes did. Apply the policy from our earlier tests (accept above 0.8, reject below 0.3, send the rest to a human or a second model) and Jev decides 289 of 317 cases with no errors, leaving 28 for review. The LLMs decide almost everything, 313 to 317 cases, and carry 7 to 10 errors that nothing in their output flags. You get the same accuracy either way; with Jev you know where the errors are.
The reason is visible in the numbers themselves. Jev returned 63 distinct probability values across the 317 cases. The LLMs returned 13 to 39, clustered on round numbers such as 0.95, 0.9 and 0.05, and put almost nothing in the middle: 0 to 3 cases between 0.3 and 0.8, against Jev's 28. An LLM asked for a probability tends to report how sure it sounds, not how likely it is to be right.
An LLM's "probability" does not even mean the same thing
We asked every LLM for "the probability that the statement is true". Three of them often answered something else. Their answer said false and their probability said 0.95: they were reporting confidence in their own answer, not the probability of the statement. That happened on 149 of 316 cases for GPT-6 Luna, 130 of 317 for Claude Haiku 4.5 and 113 of 317 for DeepSeek V4.1 Flash; Gemini did it 5 times and GPT-6 Sol never. We corrected for it before scoring, but a pipeline that trusted the field as documented would have read those as the opposite of what the model meant. Jev has no such ambiguity: a Noul is defined as the probability the statement is true, and a Choice returns a distribution over your options.
On overall calibration, measured by the Brier score (lower is better), GPT-6 Sol was best at 0.018 and Jev scored 0.029, in the middle of the pack. Jev loses points for being conservative: it never went above 0.97 or below 0.01, and when it said 0.9 or more it was right on all 131 such cases in this set. That is the opposite of the complaint usually made about AI models, and for routing decisions it is the useful failure mode.
What to use when
| If you need | Use | Why, from this test |
|---|---|---|
| Thousands of yes-or-no or pick-one decisions on scraped data | Jev | Same accuracy, a fraction of the cost, and uncertain cases flagged for you |
| A decision inside a user-facing request | Jev | 0.33 s median and 0.41 s at the 90th percentile |
| A second opinion on Jev's uncertain 9% | A frontier LLM | GPT-6 Sol had the best calibration here, and on 9% of the traffic it is affordable |
| An explanation, an extracted value, or a rewrite | An LLM | Jev cannot produce text |
We ran the cascade in the third row on the same data: Jev on everything, GPT-6 Sol only on the 28 cases Jev scored between 0.3 and 0.8. It got 310 of 317 right, better than any single model, and every one of its seven errors was a case a human would have been shown first. The total cost was $0.083, about $0.26 per 1,000 decisions: more than GPT-6 Luna alone, less than Gemini Flash-Lite or DeepSeek alone, and an eighth of GPT-6 Sol alone.
All of this depends on feeding models text they can judge. The states in this test came from our web scraping API and from lead records with two sources behind them; if you are building agents on top of web data, our web data API for AI and MCP server deliver pages as clean Markdown that fits Jev's 32,000-token state, which raw HTML usually does not (we measured that too).
Limits of this test
- 317 cases from two tasks, both labelled by one person. Differences of one to three cases between models are noise.
- Each LLM ran with its provider's default reasoning setting. More reasoning would change cost and latency, and possibly accuracy.
- Our account was on OpenRouter's new-account rate limit, so the LLM calls were spaced out; that does not change per-call latency, but it means we did not test throughput.
- The 0.3 and 0.8 cut-offs were chosen on data from the same day. Check them on a labelled sample of your own before relying on them.