Jev, the model TypeSafe AI released on 15 September 2026, does not write text. It reads a block of state, answers typed questions about it, and bills only the tokens it reads. That makes it an obvious fit for a scraping pipeline, where the question is rarely "summarise this page" and usually "is this the page I wanted, and should I keep it". The catch nobody has measured is the input. Jev accepts at most 32,000 tokens of state per request. On 28 September 2026 we fetched 20 real web pages and counted: raw HTML fit on 4 of them, with a median of 152,277 tokens. The same pages as Markdown fit on all 20, with a median of 4,512.
What Jev is, in one paragraph
Jev is what TypeSafe calls a System One model. You send a state (a string, a JSON object or an array of text) and a set of questions, and every question is answered in parallel in one pass. There are three question types: a Choice picks one option from a list you define and returns a probability for each, a Score places the state on a rubric you write, and a Noul returns the probability that a statement is true. It cannot return a value outside the options you supplied, and it cannot produce a string, so it cannot extract a field for you. TypeSafe lists the model at $0.042 per million input tokens with output free, 250,000 tokens per second and 1,200 requests per minute, and a context of 64,000 tokens per request, of which the state plus the single longest question may use 32,000. The input is text only.
That last line is the one that matters for anyone pointing Jev at the web. A web page is not text. It is HTML, and most of its bytes are scripts, styles, tracking and inline JSON that no classifier needs.
What we measured
We fetched 25 public URLs across the page types a scraping pipeline actually meets: news homepages and articles, product pages, a pricing page, docs, a GitHub repository, a subreddit, a job search, a directory, a real estate search, a travel listing, a video page, a streaming channel and a government homepage. Every fetch went through our web scraping API over plain HTTP with a browser TLS fingerprint and no JavaScript, which is what a classification step would normally see. Two URLs timed out and are left out; 23 answered, and 3 of those were block pages. For the 20 that returned content we counted tokens four ways:
- Raw HTML, exactly as served.
- Smart Markdown: the whole page minus navigation, footer and cookie chrome, tables kept, links absolute.
- Markdown without links: the same, with link and image URLs removed and their text kept.
- Article mode: only the main article block.
TypeSafe has not published Jev's tokenizer, so we counted with o200k_base, a widely used modern tokenizer. Jev's own counts will differ somewhat; the ratios between formats will not change much, because they are driven by how much markup there is, not by the tokenizer.
The results, page by page
| Page type | Site | Status | Raw HTML tokens | Smart Markdown | No links | Raw HTML fits 32k? |
|---|---|---|---|---|---|---|
| Travel page | tripadvisor.com | 200 | 862,087 | 29,825 | 10,728 | no |
| Product page | amazon.com | 200 | 686,504 | 10,088 | 6,328 | no |
| Video page | youtube.com | 200 | 488,768 | 225 | 30 | no |
| News homepage | theverge.com | 200 | 406,476 | 19,500 | 8,182 | no |
| News homepage | reuters.com | 200 | 371,622 | 5,602 | 2,163 | no |
| Product page | ikea.com | 200 | 333,492 | 15,747 | 6,648 | no |
| SaaS pricing | stripe.com | 200 | 326,329 | 9,420 | 6,455 | no |
| Streaming channel | kick.com | 200 | 214,066 | 367 | 102 | no |
| Recipe | allrecipes.com | 200 | 212,650 | 7,546 | 5,122 | no |
| Subreddit | reddit.com | 200 | 191,827 | 2,041 | 1,273 | no |
| GitHub repo | github.com | 200 | 112,726 | 2,802 | 1,893 | no |
| Company blog post | typesafe.ai | 200 | 79,800 | 3,421 | 3,143 | no |
| News article | techcrunch.com | 200 | 78,202 | 3,063 | 1,825 | no |
| Job listing | linkedin.com | 200 | 75,342 | 8,815 | 1,956 | no |
| Wikipedia article | en.wikipedia.org | 200 | 71,571 | 12,443 | 6,953 | no |
| Docs page | developer.mozilla.org | 200 | 56,747 | 1,289 | 485 | no |
| Docs page | docs.python.org | 200 | 31,933 | 9,043 | 6,637 | yes |
| Real estate listing | realtor.com | 200 | 22,047 | 37 | 37 | yes |
| Government page | usa.gov | 200 | 12,836 | 1,267 | 805 | yes |
| Product page (sandbox) | books.toscrape.com | 200 | 2,123 | 469 | 370 | yes |
| Local directory | yellowpages.com | blocked (403) | 1,592 | 148 | 134 | n/a |
| Movie page | imdb.com | blocked (202) | 766 | 40 | 40 | n/a |
| Search results (e-commerce) | ebay.com | blocked (403) | 603 | 77 | 68 | n/a |
| Forum thread | news.ycombinator.com | blocked (undefined) | 0 | 0 | 0 | n/a |
| Q&A thread | stackoverflow.com | blocked (undefined) | 0 | 0 | 0 | n/a |
Summarised over the 20 content pages:
| Format | Median tokens | Largest page | Fits the 32k state | Jev cost per 1,000 pages at the median |
|---|---|---|---|---|
| Raw HTML | 152,277 | 862,087 | 4 of 20 | $6.40, if it fit |
| Smart Markdown | 4,512 | 29,825 | 20 of 20 | $0.19 |
| Markdown without links | 2,060 | 10,728 | 20 of 20 | $0.09 |
| Article mode | 2,647 | 30,037 | 20 of 20 | $0.11 |
Three things follow. First, raw HTML is not an option for most of the web: the median page is almost five times the limit, and the four that fit include a sandbox built for scraping practice and a Python docs page that squeezes under by 67 tokens, before any question is added. Second, cleaning is where the savings are: the median page shrank 28 times from HTML to smart Markdown and 50 times once link URLs were removed. Across all 20 pages, 4,637,148 tokens of HTML became 71,135 tokens without links. Third, links are the cheapest thing to drop. On the LinkedIn job search they were 78% of the Markdown (8,815 tokens down to 1,956), because every card carries a long tracking URL. Unless your question is about the links themselves, they are noise that you pay for.
Pages that fit and still tell Jev nothing
Fitting the limit is necessary, not sufficient. Three pages in our sample were almost empty without JavaScript, and a classifier will answer confidently about an empty page because it has to pick one of your options.
- YouTube video page: 488,768 tokens of HTML, 225 tokens of Markdown, 30 without links. The video's substance lives in scripts.
- Realtor.com search: 22,047 tokens of HTML, 37 tokens of Markdown. A shell that fills in after load.
- Kick channel: 214,066 tokens of HTML, 367 of Markdown. We measured this one in detail in our Kick post: the data is in a separate JSON endpoint.
The fix is a check in code before the call, not a question to the model: if the cleaned page is under a few hundred tokens, render it or fetch its data endpoint, then classify. That costs a render ($0.001 per page on our API) instead of a wrong answer.
The same goes for block pages. eBay answered 403, Yellow Pages served a Cloudflare challenge with a 403, and IMDb returned a 202 with a "verify that you're not a robot" page. The status code catches two of the three. The third is exactly the kind of case where one Noul ("this page is an anti-bot challenge or an access-denied notice, not the requested content") earns its fraction of a cent, because the text is unambiguous even when the status is not. The reverse trap exists too: Kick's normal pages contain Cloudflare's challenge-platform script, and a string search calls every one of them blocked.
A pipeline that respects the limit
The pattern that holds up, and the one TypeSafe's own documentation points toward for extraction, is that code proposes and Jev decides. Fetch the page as clean Markdown, drop links unless you need them, run the cheap structural checks in code, and only then ask Jev the questions that live in the language of the page.
import requests
QD_KEY = "your-quanticdata-key"
TS_KEY = "your-typesafe-key"
def fetch_markdown(url):
r = requests.post(
"https://api.quanticdata.io/v1/scrape",
headers={"Authorization": "Bearer " + QD_KEY},
json={"url": url, "links_mode": "strip", "images_mode": "strip"},
timeout=90,
)
p = r.json()["payload"]
return p["status"], p.get("content") or ""
def judge(url):
status, md = fetch_markdown(url)
if status in (401, 403, 429):
return {"url": url, "verdict": "blocked", "by": "status code"}
if len(md) < 1200: # roughly 300 tokens: an empty shell
return {"url": url, "verdict": "render first", "by": "size check"}
body = {
"model": "jev-latest",
"state": {"url": url, "page": md[:100000]},
"questions": {
"is_block_page": {"type": "noul",
"instructions": "The page is an anti-bot challenge or an access-denied notice, not the requested content"},
"page_kind": {"type": "choice",
"instructions": "What kind of page this is",
"criteria": {"product": "A single product with a price",
"listing": "A list or search results of several items",
"article": "An editorial article or blog post",
"other": "None of the above"}},
"has_price": {"type": "noul",
"instructions": "The page states at least one price for something that can be bought"},
},
}
r = requests.post("https://api.typesafe.ai/v1/systemone",
headers={"Authorization": "Bearer " + TS_KEY}, json=body, timeout=30)
return {"url": url, "answers": r.json()["answers"]}
Three details in that code come straight from the measurements. The size check runs before Jev because empty shells fit the limit and mislead it. The character cap on the page (about 25,000 tokens of English) leaves room for the questions inside the 32,000-token budget; the largest page in our sample needed 10,728 tokens without links, so the cap only bites on outliers. And the Choice has an explicit other option, because a model that must pick from your list will pick something even when nothing fits.
What it costs end to end
Per 1,000 median pages: fetching them as Markdown on our API costs $0.20 (from $0.0002 per successful page, blocks not billed), and Jev reading them without links costs about $0.09 at $0.042 per million tokens. Call it $0.29 per thousand pages to fetch and classify, of which the classification is under a third. Had the raw HTML fit, Jev alone would have cost $6.40 per thousand at the median and $36 per thousand for pages the size of the Tripadvisor listing, which is the entire argument for cleaning first.
If you drive this from an agent rather than a script, our MCP server exposes the same scrape with links_mode and content modes, so an agent in Claude or Cursor can hand Jev a page that fits. For a one-off check of what a page looks like after cleaning, the website to Markdown converter shows the result in the browser.
Where Jev fits in a scraping pipeline, and where it does not
| Job | Use Jev? | Why |
|---|---|---|
| Extracting a price, a name or a date | No | It cannot produce a string. Extract with selectors or an LLM, then let Jev choose between candidates if there are several |
| Is this a block page, a soft 404, an empty shell? | After code checks | Status codes and size catch most cases for free; a Noul settles the rest |
| What kind of page is this? | Yes, when structure is missing | Read JSON-LD type, og:type and the URL first; ask Jev only where none exists |
| Is this lead, listing or review relevant to my criteria? | Yes | A judgment that lives in the language, repeated thousands of times, which is the shape Jev was built for |
| Summarising or rewriting a page | No | That is text generation; use an LLM |
Access is its own constraint for now. TypeSafe opened signups on 20 September and paused new ones two days later under demand; existing accounts keep working, and the model is also listed on Vercel AI Gateway and OpenRouter. The pipeline above does not depend on which route you use: whatever sends the request, the state still has to fit.