Documentation Python quickstart Blog Free tools hello@quanticdata.ioLog in

Jev vs GPT, Claude, Gemini on Real Web Data

Cost per 1,000 classifications for Jev and five LLMs on the same 317 hand-labelled scraping and lead cases: Jev 0.049 dollars up to GPT-6 Sol at 2.081 dollars, with nearly identical accuracy
Cost per 1,000 classifications for Jev and five LLMs on the same 317 hand-labelled scraping and lead cases: Jev 0.049 dollars up to GPT-6 Sol at 2.081 dollars, with nearly identical accuracy

The question everyone asks about Jev, the decision model TypeSafe AI released on 15 September 2026, is whether it is as good as a real LLM. Most answers so far come from public classification datasets or from TypeSafe's own workflows. We had something better to hand: 317 cases from real web work, labelled by hand earlier the same day, where a wrong answer costs money. On 28 September 2026 we sent every case, with the identical question, to Jev and to five LLMs. Accuracy was a tie. The bill was not. And the most useful difference was one nobody puts in a leaderboard: where each model makes its mistakes.

The test

The 317 cases come from two benchmarks we published today:

  • 134 scraper responses from 70 heavily protected sites, labelled usable or not (a block page, a region notice, the wrong page). Details in our block page test.
  • 183 businesses from a real scraped lead list, labelled as the buyer's target or not. Details in our lead scoring test.

Every model saw the same state (page text, or the lead's name, categories, description and domain) and the same one-sentence statement. Jev answered through OpenRouter's Decisions API. The five LLMs answered through OpenRouter's chat API with a strict JSON schema asking for a true or false answer and the probability that the statement is true, temperature 0 where supported, and each provider's default reasoning setting. Costs are what OpenRouter billed per call; latency is the time for each call from Europe.

ModelScraper pages (134)Leads (183)Total correctCost per 1,000Median latency90th percentile
DeepSeek V4.1 Flash134175309$0.3391.54 s4.28 s
GPT-6 Luna133 of 133175308 of 316$0.1122.47 s3.45 s
GPT-6 Sol133175308$2.0812.79 s3.84 s
Claude Haiku 4.5132175307$1.0711.21 s1.57 s
Gemini 3.5 Flash-Lite131176307$0.2730.86 s1.07 s
Jev 1.13131175306$0.0490.33 s0.41 s

One GPT-6 Luna call failed after retries and is left out of its total. The spread between the best and the worst model is three cases out of 317, less than 1%, and several of the shared mistakes are cases where our own label is arguable (a body shop whose name says it is also an auto electrician was "wrong" for every model). On this kind of work, accuracy does not separate these models.

Cost and speed

What does separate them is the bill. For the same 317 decisions Jev cost $0.0154. GPT-6 Luna cost 2.3 times as much, Gemini 3.5 Flash-Lite 5.6 times, DeepSeek V4.1 Flash 7 times, Claude Haiku 4.5 22 times and GPT-6 Sol 42 times. Jev bills only input tokens; the LLMs also bill their output and, where they reason before answering, the reasoning tokens, which were a median of 100 per call for DeepSeek and 31 for GPT-6 Luna.

Speed follows the same shape. Jev's median answer took 328 ms and its slowest tenth stayed under 405 ms. The fastest LLM, Gemini Flash-Lite, took 862 ms at the median; DeepSeek's slowest tenth took over four seconds. For a batch job that difference is a coffee break. For a decision inside a request, before a page is stored or a lead is shown to a salesperson, it is the difference between usable and not.

Where the mistakes were

This is the part a leaderboard hides. A classifier that is right 97% of the time is only useful if you can tell which 3% to check. So for every wrong answer we looked at the probability the model attached to it.

ModelWrong answersOf those, scored between 0.3 and 0.8Scored 0.9 or higher (or 0.1 or lower)
Jev 1.1311110
GPT-6 Sol921
DeepSeek V4.1 Flash802
GPT-6 Luna806
Claude Haiku 4.51005
Gemini 3.5 Flash-Lite1007

Every one of Jev's mistakes came with a hesitant score. None of the small LLMs' mistakes did. Apply the policy from our earlier tests (accept above 0.8, reject below 0.3, send the rest to a human or a second model) and Jev decides 289 of 317 cases with no errors, leaving 28 for review. The LLMs decide almost everything, 313 to 317 cases, and carry 7 to 10 errors that nothing in their output flags. You get the same accuracy either way; with Jev you know where the errors are.

The reason is visible in the numbers themselves. Jev returned 63 distinct probability values across the 317 cases. The LLMs returned 13 to 39, clustered on round numbers such as 0.95, 0.9 and 0.05, and put almost nothing in the middle: 0 to 3 cases between 0.3 and 0.8, against Jev's 28. An LLM asked for a probability tends to report how sure it sounds, not how likely it is to be right.

An LLM's "probability" does not even mean the same thing

We asked every LLM for "the probability that the statement is true". Three of them often answered something else. Their answer said false and their probability said 0.95: they were reporting confidence in their own answer, not the probability of the statement. That happened on 149 of 316 cases for GPT-6 Luna, 130 of 317 for Claude Haiku 4.5 and 113 of 317 for DeepSeek V4.1 Flash; Gemini did it 5 times and GPT-6 Sol never. We corrected for it before scoring, but a pipeline that trusted the field as documented would have read those as the opposite of what the model meant. Jev has no such ambiguity: a Noul is defined as the probability the statement is true, and a Choice returns a distribution over your options.

On overall calibration, measured by the Brier score (lower is better), GPT-6 Sol was best at 0.018 and Jev scored 0.029, in the middle of the pack. Jev loses points for being conservative: it never went above 0.97 or below 0.01, and when it said 0.9 or more it was right on all 131 such cases in this set. That is the opposite of the complaint usually made about AI models, and for routing decisions it is the useful failure mode.

What to use when

If you needUseWhy, from this test
Thousands of yes-or-no or pick-one decisions on scraped dataJevSame accuracy, a fraction of the cost, and uncertain cases flagged for you
A decision inside a user-facing requestJev0.33 s median and 0.41 s at the 90th percentile
A second opinion on Jev's uncertain 9%A frontier LLMGPT-6 Sol had the best calibration here, and on 9% of the traffic it is affordable
An explanation, an extracted value, or a rewriteAn LLMJev cannot produce text

We ran the cascade in the third row on the same data: Jev on everything, GPT-6 Sol only on the 28 cases Jev scored between 0.3 and 0.8. It got 310 of 317 right, better than any single model, and every one of its seven errors was a case a human would have been shown first. The total cost was $0.083, about $0.26 per 1,000 decisions: more than GPT-6 Luna alone, less than Gemini Flash-Lite or DeepSeek alone, and an eighth of GPT-6 Sol alone.

All of this depends on feeding models text they can judge. The states in this test came from our web scraping API and from lead records with two sources behind them; if you are building agents on top of web data, our web data API for AI and MCP server deliver pages as clean Markdown that fits Jev's 32,000-token state, which raw HTML usually does not (we measured that too).

Limits of this test

  • 317 cases from two tasks, both labelled by one person. Differences of one to three cases between models are noise.
  • Each LLM ran with its provider's default reasoning setting. More reasoning would change cost and latency, and possibly accuracy.
  • Our account was on OpenRouter's new-account rate limit, so the LLM calls were spaced out; that does not change per-call latency, but it means we did not test throughput.
  • The 0.3 and 0.8 cut-offs were chosen on data from the same day. Check them on a labelled sample of your own before relying on them.

Sources & further reading

FAQ

Quick answers on jev vs llm.

Something else? Ask us →

Is Jev as accurate as GPT or Claude?

On our 317 hand-labelled scraping and lead cases, yes: Jev got 306 right, and five LLMs including GPT-6 Sol and Claude Haiku 4.5 got 307 to 309. Differences of one to three cases out of 317 are within noise.

How much cheaper is Jev than an LLM?

Per 1,000 decisions in our test Jev cost $0.049, GPT-6 Luna $0.112, Gemini 3.5 Flash-Lite $0.273, DeepSeek V4.1 Flash $0.339, Claude Haiku 4.5 $1.071 and GPT-6 Sol $2.081, as billed by OpenRouter on 28 September 2026.

Is Jev faster than an LLM?

Yes. Its median answer took 328 ms and its 90th percentile 405 ms. The fastest LLM we tested, Gemini 3.5 Flash-Lite, had a median of 862 ms; the slowest tenth of DeepSeek V4.1 Flash calls took over four seconds.

Are Jev's probabilities calibrated?

In our test they were conservative: Jev never went above 0.97, and when it said 0.9 or more it was right on all 131 cases. All 11 of its errors had scores between 0.3 and 0.8, so a review band catches them. Its Brier score was 0.029, behind GPT-6 Sol at 0.018.

Can I trust the probability an LLM gives me?

Not blindly. Asked for the probability that a statement is true, GPT-6 Luna, Claude Haiku 4.5 and DeepSeek V4.1 Flash often returned their confidence in their own answer instead, on 113 to 149 of about 317 cases. Their probabilities also clustered on round numbers and almost never landed between 0.3 and 0.8.

Should I replace my LLM with Jev?

For high-volume yes-or-no and pick-one decisions, it makes sense to run Jev first and send only its uncertain cases to an LLM: on our data that cascade got 310 of 317 right, better than any single model, for about $0.26 per 1,000 decisions. For anything that needs generated text, an explanation or an extracted value, you still need an LLM, because Jev cannot produce text.

Give your models clean text to judge

Every case in this test started as a page or a lead record from our API. Fetch pages as Markdown that fits a 32,000-token state, or collect leads from two sources, and pay only for results that come back usable. Every account gets $2 of free usage each month.

Related reading