# Jev for Web Scraping: Will Your Page Fit?

> Jev reads 32k tokens of state. We measured 20 real web pages: raw HTML fit on 4, with a median of 152,277 tokens. Clean Markdown fit on all 20, median 4,512.

[Home](https://quanticdata.io/)/[Blog](https://quanticdata.io/blog/)/Jev for Web Scraping: Will Your Page Fit?

# Jev for Web Scraping: Will Your Page Fit?

AI scrapingSep 28, 2026·10 min read·By [Aldo Morese](https://quanticdata.io/about/), founder of QuanticData

Median size of 20 real web pages in tokens against the 32,000-token state limit of TypeSafe's Jev model, measured 28 September 2026: raw HTML 152,277, smart Markdown 4,512, Markdown with links stripped 2,060

On this page [What Jev is, in one paragraph](/blog/jev-web-scraping/#what-jev-is-in-one-paragraph) [What we measured](/blog/jev-web-scraping/#what-we-measured) [The results, page by page](/blog/jev-web-scraping/#the-results-page-by-page) [Pages that fit and still tell Jev nothing](/blog/jev-web-scraping/#pages-that-fit-and-still-tell-jev-nothing) [A pipeline that respects the limit](/blog/jev-web-scraping/#a-pipeline-that-respects-the-limit) [What it costs end to end](/blog/jev-web-scraping/#what-it-costs-end-to-end) [Where Jev fits in a scraping pipeline, and where it does not](/blog/jev-web-scraping/#where-jev-fits-in-a-scraping-pipeline-and-where-it-does-not)

Jev, the model TypeSafe AI released on 15 September 2026, does not write text. It reads a block of state, answers typed questions about it, and bills only the tokens it reads. That makes it an obvious fit for a scraping pipeline, where the question is rarely "summarise this page" and usually "is this the page I wanted, and should I keep it". The catch nobody has measured is the input. Jev accepts at most 32,000 tokens of state per request. On 28 September 2026 we fetched 20 real web pages and counted: raw HTML fit on 4 of them, with a median of 152,277 tokens. The same pages as Markdown fit on all 20, with a median of 4,512.

## What Jev is, in one paragraph

Jev is what TypeSafe calls a System One model. You send a *state* (a string, a JSON object or an array of text) and a set of *questions*, and every question is answered in parallel in one pass. There are three question types: a **Choice** picks one option from a list you define and returns a probability for each, a **Score** places the state on a rubric you write, and a **Noul** returns the probability that a statement is true. It cannot return a value outside the options you supplied, and it cannot produce a string, so it cannot extract a field for you. TypeSafe lists the model at $0.042 per million input tokens with output free, 250,000 tokens per second and 1,200 requests per minute, and a context of 64,000 tokens per request, of which the state plus the single longest question may use 32,000. The input is text only.

That last line is the one that matters for anyone pointing Jev at the web. A web page is not text. It is HTML, and most of its bytes are scripts, styles, tracking and inline JSON that no classifier needs.

## What we measured

We fetched 25 public URLs across the page types a scraping pipeline actually meets: news homepages and articles, product pages, a pricing page, docs, a GitHub repository, a subreddit, a job search, a directory, a real estate search, a travel listing, a video page, a streaming channel and a government homepage. Every fetch went through our [web scraping API](https://quanticdata.io/web-scraping-api/) over plain HTTP with a browser TLS fingerprint and no JavaScript, which is what a classification step would normally see. Two URLs timed out and are left out; 23 answered, and 3 of those were block pages. For the 20 that returned content we counted tokens four ways:

- **Raw HTML**, exactly as served.

- **Smart Markdown**: the whole page minus navigation, footer and cookie chrome, tables kept, links absolute.

- **Markdown without links**: the same, with link and image URLs removed and their text kept.

- **Article mode**: only the main article block.

TypeSafe has not published Jev's tokenizer, so we counted with o200k_base, a widely used modern tokenizer. Jev's own counts will differ somewhat; the ratios between formats will not change much, because they are driven by how much markup there is, not by the tokenizer.

## The results, page by page

| Page type | Site | Status | Raw HTML tokens | Smart Markdown | No links | Raw HTML fits 32k? |
| --- | --- | --- | --- | --- | --- | --- |
| Travel page | tripadvisor.com | 200 | 862,087 | 29,825 | 10,728 | no |
| Product page | amazon.com | 200 | 686,504 | 10,088 | 6,328 | no |
| Video page | youtube.com | 200 | 488,768 | 225 | 30 | no |
| News homepage | theverge.com | 200 | 406,476 | 19,500 | 8,182 | no |
| News homepage | reuters.com | 200 | 371,622 | 5,602 | 2,163 | no |
| Product page | ikea.com | 200 | 333,492 | 15,747 | 6,648 | no |
| SaaS pricing | stripe.com | 200 | 326,329 | 9,420 | 6,455 | no |
| Streaming channel | kick.com | 200 | 214,066 | 367 | 102 | no |
| Recipe | allrecipes.com | 200 | 212,650 | 7,546 | 5,122 | no |
| Subreddit | reddit.com | 200 | 191,827 | 2,041 | 1,273 | no |
| GitHub repo | github.com | 200 | 112,726 | 2,802 | 1,893 | no |
| Company blog post | typesafe.ai | 200 | 79,800 | 3,421 | 3,143 | no |
| News article | techcrunch.com | 200 | 78,202 | 3,063 | 1,825 | no |
| Job listing | linkedin.com | 200 | 75,342 | 8,815 | 1,956 | no |
| Wikipedia article | en.wikipedia.org | 200 | 71,571 | 12,443 | 6,953 | no |
| Docs page | developer.mozilla.org | 200 | 56,747 | 1,289 | 485 | no |
| Docs page | docs.python.org | 200 | 31,933 | 9,043 | 6,637 | yes |
| Real estate listing | realtor.com | 200 | 22,047 | 37 | 37 | yes |
| Government page | usa.gov | 200 | 12,836 | 1,267 | 805 | yes |
| Product page (sandbox) | books.toscrape.com | 200 | 2,123 | 469 | 370 | yes |
| Local directory | yellowpages.com | blocked (403) | 1,592 | 148 | 134 | n/a |
| Movie page | imdb.com | blocked (202) | 766 | 40 | 40 | n/a |
| Search results (e-commerce) | ebay.com | blocked (403) | 603 | 77 | 68 | n/a |
| Forum thread | news.ycombinator.com | blocked (undefined) | 0 | 0 | 0 | n/a |
| Q&A thread | stackoverflow.com | blocked (undefined) | 0 | 0 | 0 | n/a |

Summarised over the 20 content pages:

| Format | Median tokens | Largest page | Fits the 32k state | Jev cost per 1,000 pages at the median |
| --- | --- | --- | --- | --- |
| Raw HTML | 152,277 | 862,087 | 4 of 20 | $6.40, if it fit |
| Smart Markdown | 4,512 | 29,825 | 20 of 20 | $0.19 |
| Markdown without links | 2,060 | 10,728 | 20 of 20 | $0.09 |
| Article mode | 2,647 | 30,037 | 20 of 20 | $0.11 |

Three things follow. First, raw HTML is not an option for most of the web: the median page is almost five times the limit, and the four that fit include a sandbox built for scraping practice and a Python docs page that squeezes under by 67 tokens, before any question is added. Second, cleaning is where the savings are: the median page shrank 28 times from HTML to smart Markdown and 50 times once link URLs were removed. Across all 20 pages, 4,637,148 tokens of HTML became 71,135 tokens without links. Third, links are the cheapest thing to drop. On the LinkedIn job search they were 78% of the Markdown (8,815 tokens down to 1,956), because every card carries a long tracking URL. Unless your question is about the links themselves, they are noise that you pay for.

## Pages that fit and still tell Jev nothing

Fitting the limit is necessary, not sufficient. Three pages in our sample were almost empty without JavaScript, and a classifier will answer confidently about an empty page because it has to pick one of your options.

- **YouTube video page:** 488,768 tokens of HTML, 225 tokens of Markdown, 30 without links. The video's substance lives in scripts.

- **Realtor.com search:** 22,047 tokens of HTML, 37 tokens of Markdown. A shell that fills in after load.

- **Kick channel:** 214,066 tokens of HTML, 367 of Markdown. We measured this one in detail in [our Kick post](https://quanticdata.io/blog/kick-proxies/): the data is in a separate JSON endpoint.

The fix is a check in code before the call, not a question to the model: if the cleaned page is under a few hundred tokens, render it or fetch its data endpoint, then classify. That costs a render ($0.001 per page on our API) instead of a wrong answer.

The same goes for block pages. eBay answered 403, Yellow Pages served a Cloudflare challenge with a 403, and IMDb returned a 202 with a "verify that you're not a robot" page. The status code catches two of the three. The third is exactly the kind of case where one Noul ("this page is an anti-bot challenge or an access-denied notice, not the requested content") earns its fraction of a cent, because the text is unambiguous even when the status is not. The reverse trap exists too: Kick's normal pages contain Cloudflare's challenge-platform script, and a string search calls every one of them blocked.

## A pipeline that respects the limit

The pattern that holds up, and the one TypeSafe's own documentation points toward for extraction, is that code proposes and Jev decides. Fetch the page as clean Markdown, drop links unless you need them, run the cheap structural checks in code, and only then ask Jev the questions that live in the language of the page.

```
import requests

QD_KEY = "your-quanticdata-key"
TS_KEY = "your-typesafe-key"

def fetch_markdown(url):
    r = requests.post(
        "https://api.quanticdata.io/v1/scrape",
        headers={"Authorization": "Bearer " + QD_KEY},
        json={"url": url, "links_mode": "strip", "images_mode": "strip"},
        timeout=90,
    )
    p = r.json()["payload"]
    return p["status"], p.get("content") or ""

def judge(url):
    status, md = fetch_markdown(url)
    if status in (401, 403, 429):
        return {"url": url, "verdict": "blocked", "by": "status code"}
    if len(md) < 1200:                     # roughly 300 tokens: an empty shell
        return {"url": url, "verdict": "render first", "by": "size check"}
    body = {
        "model": "jev-latest",
        "state": {"url": url, "page": md[:100000]},
        "questions": {
            "is_block_page": {"type": "noul",
                "instructions": "The page is an anti-bot challenge or an access-denied notice, not the requested content"},
            "page_kind": {"type": "choice",
                "instructions": "What kind of page this is",
                "criteria": {"product": "A single product with a price",
                             "listing": "A list or search results of several items",
                             "article": "An editorial article or blog post",
                             "other": "None of the above"}},
            "has_price": {"type": "noul",
                "instructions": "The page states at least one price for something that can be bought"},
        },
    }
    r = requests.post("https://api.typesafe.ai/v1/systemone",
                      headers={"Authorization": "Bearer " + TS_KEY}, json=body, timeout=30)
    return {"url": url, "answers": r.json()["answers"]}
```

Three details in that code come straight from the measurements. The size check runs before Jev because empty shells fit the limit and mislead it. The character cap on the page (about 25,000 tokens of English) leaves room for the questions inside the 32,000-token budget; the largest page in our sample needed 10,728 tokens without links, so the cap only bites on outliers. And the Choice has an explicit `other` option, because a model that must pick from your list will pick something even when nothing fits.

## What it costs end to end

Per 1,000 median pages: fetching them as Markdown on our API costs $0.20 (from $0.0002 per successful page, blocks not billed), and Jev reading them without links costs about $0.09 at $0.042 per million tokens. Call it $0.29 per thousand pages to fetch and classify, of which the classification is under a third. Had the raw HTML fit, Jev alone would have cost $6.40 per thousand at the median and $36 per thousand for pages the size of the Tripadvisor listing, which is the entire argument for cleaning first.

If you drive this from an agent rather than a script, our [MCP server](https://quanticdata.io/mcp-server/) exposes the same scrape with `links_mode` and content modes, so an agent in Claude or Cursor can hand Jev a page that fits. For a one-off check of what a page looks like after cleaning, the [website to Markdown converter](https://quanticdata.io/tools/website-to-markdown/) shows the result in the browser.

## Where Jev fits in a scraping pipeline, and where it does not

| Job | Use Jev? | Why |
| --- | --- | --- |
| Extracting a price, a name or a date | No | It cannot produce a string. Extract with selectors or an LLM, then let Jev choose between candidates if there are several |
| Is this a block page, a soft 404, an empty shell? | After code checks | Status codes and size catch most cases for free; a Noul settles the rest |
| What kind of page is this? | Yes, when structure is missing | Read JSON-LD type, og:type and the URL first; ask Jev only where none exists |
| Is this lead, listing or review relevant to my criteria? | Yes | A judgment that lives in the language, repeated thousands of times, which is the shape Jev was built for |
| Summarising or rewriting a page | No | That is text generation; use an LLM |

Access is its own constraint for now. TypeSafe opened signups on 20 September and paused new ones two days later under demand; existing accounts keep working, and the model is also listed on Vercel AI Gateway and OpenRouter. The pipeline above does not depend on which route you use: whatever sends the request, the state still has to fit.

### Sources & further reading

- [TypeSafe AI docs: Models (limits and pricing)](https://docs.typesafe.ai/models)

- [TypeSafe AI docs: State](https://docs.typesafe.ai/concepts/state)

- [Introducing System One Models and Jev, TypeSafe AI blog](https://typesafe.ai/blog/introducing-system-one-models-and-jev)

- [Jev (AI model), Wikipedia](https://en.wikipedia.org/wiki/Jev_(AI_model))

- [A new kind of AI model from a ChatGPT inventor is thrilling developers, TechCrunch](https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/)

- [Building a Harness with Jev, LangChain blog](https://www.langchain.com/blog/building-a-harness-with-jev)

## FAQ

Quick answers on jev web scraping.

[Something else? Ask us →](mailto:hello@quanticdata.io)

### What is Jev?

Jev is a System One model released by TypeSafe AI on 15 September 2026. Instead of generating text it answers typed questions about a state you send: a Choice from your options, a Score on your rubric, or a Noul, the probability that a statement is true. It is priced at $0.042 per million input tokens with output free.

### Can Jev scrape or extract data from a web page?

No. Jev cannot produce a string, so it cannot fill in a field. It can decide things about a page you have already fetched: whether it is a block page, what kind of page it is, whether it matches your criteria, or which of several extracted candidates is correct.

### How big a web page can Jev read?

TypeSafe lists 64,000 tokens per request, with 32,000 for the state plus the longest question. In our 28 September 2026 sample of 20 pages, raw HTML fit on 4, with a median of 152,277 tokens. The same pages as Markdown fit on all 20, median 4,512 tokens, or 2,060 with links removed.

### How much does it cost to classify web pages with Jev?

At the median page in our sample, Markdown without links is about 2,060 tokens, which is roughly $0.09 per 1,000 pages at $0.042 per million tokens. Fetching those pages as Markdown on our API adds $0.20 per 1,000, so about $0.29 per 1,000 pages to fetch and classify.

### Should I send raw HTML to Jev?

Almost never. On 16 of the 20 pages we measured it does not fit the limit, and where it does, most of the tokens are scripts and markup you pay for without improving the answer. Convert to Markdown first and strip link URLs unless the question is about the links.

### How do I get access to Jev?

TypeSafe paused new direct signups on 22 September 2026 under demand, with existing accounts unaffected. The model is also listed on Vercel AI Gateway and on OpenRouter. The request shape is the same: a state and a map of typed questions.

## Hand Jev pages that fit

Fetch any page as clean Markdown with links stripped, over plain HTTP or rendered, and pay only for pages that come back usable. Every account gets $2 of free API usage each month.

[Start free — $2/month included](https://quanticdata.io/signup/)[Explore Web Scraping API](https://quanticdata.io/web-scraping-api/)

## Related reading

[AI scraping Can AI Work Without Data? Two different questions hide in one query: AI runs offline just fine, but AI without data is a contradiction — and stale data is a slower version of none. Read →](https://quanticdata.io/blog/can-ai-work-without-data/) [AI scraping How to Build an AI Web Scraper Where the LLM actually belongs in a scraping pipeline, how prompt-based extraction survives redesigns, and how to keep token costs from eating the project. Read →](https://quanticdata.io/blog/how-to-build-an-ai-web-scraper/) [AI scraping How to Stop Web Scraping (Honestly) A defender's honest guide: the tactics that stop casual scrapers, the ones that only add friction, and why protecting logins and PII beats blocking public prices. Read →](https://quanticdata.io/blog/how-to-stop-web-scraping/)

## Also on this site

Quantic**Data**

Residential proxies & web data APIs for AI.

#### Proxies

- [Residential Basic](https://quanticdata.io/residential-proxies/#basic)

- [Residential Premium](https://quanticdata.io/residential-proxies/#plans)

- [Cheap Residential](https://quanticdata.io/cheap-residential-proxies/)

- [Mobile Proxies](https://quanticdata.io/mobile-proxies/)

- [Datacenter Proxies](https://quanticdata.io/datacenter-proxies/)

- [ISP Proxies](https://quanticdata.io/isp-proxies/)

- [Rotating Proxies](https://quanticdata.io/rotating-proxies/)

- [Sneaker Proxies](https://quanticdata.io/sneaker-proxies/)

- [SOCKS5 Proxies](https://quanticdata.io/socks5-proxies/)

- [IPv6 Proxies](https://quanticdata.io/ipv6-proxies/)

- [Proxy locations](https://quanticdata.io/proxies/)

#### Data APIs

- [MCP Server](https://quanticdata.io/mcp-server/)

- [Web Scraper API](https://quanticdata.io/web-scraping-api/)

- [SERP API](https://quanticdata.io/serp-api/)

- [Collectors](https://quanticdata.io/collectors/)

- [Web Data for AI](https://quanticdata.io/web-data-api-for-ai/)

- [Quantic AI](https://quanticdata.io/ai-web-scraping-service/)

- [Crawl & Map](https://quanticdata.io/crawl-map/)

- [SEO Audit](https://quanticdata.io/seo-audit/)

#### Use cases

- [Company data](https://quanticdata.io/scrape-company-data/)

- [Price monitoring](https://quanticdata.io/competitor-price-monitoring/)

- [Market research](https://quanticdata.io/market-research-data/)

- [Real estate data](https://quanticdata.io/real-estate-data-scraping/)

- [Scrape job postings](https://quanticdata.io/scrape-job-postings/)

#### Company

- [Documentation](https://quanticdata.io/docs/)

- [Blog](https://quanticdata.io/blog/)

- [Free tools](https://quanticdata.io/tools/)

- [Partners](https://quanticdata.io/partners/)

- [About](https://quanticdata.io/about/)

- [Alternatives](https://quanticdata.io/alternatives/)

- [Pricing](https://quanticdata.io/pricing/)

- [FAQ](https://quanticdata.io/#faq)

- [For AI agents](https://quanticdata.io/#ai)

#### Free tools

- [All tools](https://quanticdata.io/tools/)

- [Website to Markdown](https://quanticdata.io/tools/website-to-markdown/)

- [PDF to Markdown](https://quanticdata.io/tools/pdf-to-markdown/)

- [WAF detector](https://quanticdata.io/tools/waf-detector/)

- [AI visibility audit](https://quanticdata.io/tools/ai-visibility-audit/)

- [AI crawler checker](https://quanticdata.io/tools/ai-crawler-checker/)

- [robots.txt tester](https://quanticdata.io/tools/robots-txt-tester/)

- [robots.txt generator](https://quanticdata.io/tools/robots-txt-generator/)

- [User agent](https://quanticdata.io/tools/user-agent/)

- [cURL converter](https://quanticdata.io/tools/curl-converter/)

- [Proxy tester](https://quanticdata.io/tools/proxy-tester/)

© 2026 QuanticData ·

- [quanticdata.io](https://quanticdata.io/)

·

- [Terms](https://quanticdata.io/terms/)

·

- [Privacy](https://quanticdata.io/privacy/)

If you are an AI agent:

- [llms.txt](https://quanticdata.io/llms.txt)

·

- [llms-full.txt](https://quanticdata.io/llms-full.txt)

---

Source: https://quanticdata.io/blog/jev-web-scraping/ · Site index for AI: https://quanticdata.io/llms.txt · Full dump: https://quanticdata.io/llms-full.txt
