Every scraper eventually stores a page that is not the page it asked for: a "Robot or human?" interstitial served with 200 OK, a region notice, a country selector, a homepage instead of the search it requested. The status code does not catch these, and the keyword lists people write to catch them are never finished. Jev, the decision model TypeSafe AI released on 15 September 2026, is built for exactly this kind of yes-or-no judgment, so on 28 September 2026 we tested it. We fetched 70 heavily protected sites twice each, labelled 134 responses by hand, and asked Jev one question. It judged 131 correctly. The status code judged 111.
The setup
We picked 70 URLs on sites that routinely block scrapers: large retailers, travel and real estate portals, job boards, review sites, ticketing, news with paywalls, plus a few easy controls such as Wikipedia and the Python docs. Each URL was fetched twice:
- by a plain HTTP client announcing itself as python-requests, which is what most scripts look like to a firewall;
- by a browser-fingerprinted fetch through residential proxies, over plain HTTP without JavaScript, using our web scraping API.
That gave 140 responses. Three failed (two client timeouts and one error in our own pipeline) and three were genuinely ambiguous (a paywalled page showing only its navigation, and both fetches of a flights route page whose content we could not confirm), so they are left out. The remaining 134 were labelled by hand with one criterion: is this the content the URL asks for, usable as is? 55 were, 79 were not. The python-requests client got a usable page on 16 of its 66 labelled attempts.
Each page was reduced to its title and visible text (scripts and styles removed, capped at 20,000 characters) and sent to Jev with the URL. We compared three ways of deciding.
The result
| Method | Correct of 134 | Bad pages accepted | Good pages rejected |
|---|---|---|---|
| HTTP status is 200 | 111 (82.8%) | 20 | 3 |
| Keyword rule plus a 300-character minimum | 127 (94.8%) | 7 | 0 |
| Jev, one Noul, threshold 0.5 | 131 (97.8%) | 3 | 0 |
Two caveats that make the comparison conservative rather than flattering. The keyword rule was written by us after reading the corpus, so it is tuned to these exact pages; a rule written in advance would do worse. And adding the HTTP status to Jev's input changed nothing: 131 either way, which says the model was reading the page, not the code.
Jev's three mistakes are the interesting part. All three were real pages about the wrong thing: Expedia answered a request for Austin hotels with a list of hotels in Orangeburg, South Carolina, and Hotels.com and Vrbo returned their homepages instead of the Austin search. Neither a status code nor a keyword can see that. Jev let them through, but not confidently.
The probabilities are the useful part
Jev does not just say yes or no. The Noul returns a probability, and TypeSafe trains it to be calibrated. On this corpus the scores separated almost completely:
| Group | Lowest score | Median | Highest score |
|---|---|---|---|
| Usable pages (55) | 0.62 | 0.95 | 0.98 |
| Not usable (79) | 0.00 | 0.02 | 0.71 |
Which suggests a policy rather than a threshold: accept above 0.8, reject below 0.5, and send everything in between to a second opinion. On our 134 pages that policy accepted 52 pages and rejected 76 with no errors in either group, and left 6 in the middle: the three wrong-city pages, an Amazon product page, and two Rotten Tomatoes pages that were real but heavy with sign-in chrome. Six of 134 is 4.5% of the traffic going to a slower check, a render or a larger model, which is what you would want. A word of honesty: we picked 0.8 by looking at these same scores. Use it as a starting point and tune it on your own labelled sample.
The traps, case by case
| Response | Status | What it really was | Keyword rule | Jev score | Jev's label |
|---|---|---|---|---|---|
| walmart.com, python client | 200 | "Robot or human?" press-and-hold challenge | caught | 0.02 | anti-bot |
| adidas.com | 200 | 32 characters: a bot-protection stub | caught by length | 0.14 | empty |
| bestbuy.com | 200 | International country selector | missed | 0.03 | elsewhere |
| realtor.com | 200 | "Not available in your region" | missed | 0.02 | elsewhere |
| footlocker.com, python client | 200 | Redirected to the European site's cookie page | missed | 0.06 | elsewhere |
| autotrader.com | 200 | "Site currently unavailable" | caught by length | 0.02 | error |
| ebay.com | 403 | "Something went wrong on our end" (a block dressed as an error) | caught by length | 0.01 | error |
| ticketmaster.com, python client | 403 | The JSON body {"response":"block"} | caught by length | 0.09 | anti-bot |
| stackoverflow.com | 403 | The full question list, 12,219 characters | kept | 0.94 | content |
| chewy.com | 429 | The full category page, 38,644 characters | kept | 0.96 | content |
| cars.com | 403 | The full listing page, 58,730 characters | kept | 0.96 | content |
| doordash.com | 200 | Real page that also says "Checking if the site connection is secured" | kept | 0.85 | content |
| expedia.com | 200 | Hotels in the wrong city | missed | 0.56 | content |
The middle three rows are why "status is 200" is a bad rule in both directions: some sites return the full page with a 403 or 429 attached, so a status filter throws away good data as well as letting bad data in. The same thing, in reverse, is what we found on Kick, whose normal pages contain Cloudflare's challenge script and fool string matching (details in our Kick post).
Jev's second question, a Choice between content, anti-bot, error, elsewhere and empty, labelled the 79 unusable pages as 45 anti-bot, 17 empty, 10 error, 4 elsewhere and 3 content (the wrong-city trio). That breakdown is what tells you what to do next: an anti-bot page needs a different fetch, an empty shell needs a render, a region notice needs an exit in another country, and an error page may just need a retry.
One more question made it worse
To catch the wrong-city pages we added a second Noul: "the page is specifically about the place or search named in the URL, not a generic homepage". Requiring both to pass dropped accuracy to 126 of 134. It still let Hotels.com and Vrbo through, and it now rejected six good pages, including the DoorDash and Instacart homepages, because we had asked for homepages to fail and those URLs were homepages. The lesson is the one TypeSafe's own docs make about decomposing questions: every question is another classifier with its own errors, and a badly scoped one costs more than it catches. Write questions that are true or false on their own, and test each one.
Speed and cost
We called Jev through OpenRouter's Decisions API, which needs no TypeSafe account (direct signups have been paused since 22 September). Over 134 calls the median input was 700 tokens and the largest 7,092; median latency from Europe was 324 ms and the 90th percentile 393 ms. The whole run cost $0.0107, which is about $0.08 per 1,000 pages. That is small next to the cost of fetching the pages in the first place, and very small next to storing a few thousand challenge pages in a dataset someone paid for.
import re, requests
OR_KEY = "your-openrouter-key"
def visible_text(html):
html = re.sub(r"(?is)<(script|style|noscript)[^>]*>.*?</\1>", " ", html)
return re.sub(r"\s+", " ", re.sub(r"(?s)<[^>]+>", " ", html)).strip()
def is_usable(url, title, html):
body = {
"model": "typesafe/jev-1.13",
"state": {"url": url, "title": title, "page_text": visible_text(html)[:20000] or "(no text)"},
"questions": {
"usable": {"type": "noul", "instructions":
"The page is the content that was requested at this URL, not an anti-bot challenge, "
"access-denied notice, error page, region notice, redirect, a different page, or an empty shell."},
"kind": {"type": "choice", "instructions": "What did the server actually return?",
"criteria": {"content": "The requested content itself",
"antibot": "An anti-bot challenge, captcha, or access-denied notice",
"error": "An error, not-found or site-unavailable page",
"elsewhere": "A redirect, region notice, country selector or a different page",
"empty": "A title or script shell with no readable content"}},
},
}
r = requests.post("https://openrouter.ai/api/alpha/decisions",
headers={"Authorization": "Bearer " + OR_KEY}, json=body, timeout=30)
a = r.json()["answers"]
p = a["usable"]["noul"]
if p > 0.8:
return "accept", a["kind"]["choice"]
if p < 0.5:
return "reject", a["kind"]["choice"]
return "review", a["kind"]["choice"]
For long pages, convert to Markdown before sending rather than stripping tags by hand: in our measurement of 20 real pages, raw HTML fit Jev's 32,000-token state on only 4, while Markdown fit on all of them.
Where this belongs in a pipeline
- Free checks first. Network errors, empty bodies and obvious status codes such as 404 need no model.
- Jev on everything else. One call per page, about $0.00008 at our median size, 0.3 seconds.
- Act on the label. Anti-bot: refetch through a better exit, such as a residential proxy or a rendered request. Empty: render. Elsewhere: change the country. Error: retry later. Our guides to 403 and 429 responses cover the fetch side.
- Review the middle band. Scores between 0.5 and 0.8 go to a render, a larger model or a human, which on our sample was 4.5% of pages.
Our API already decides this for the fetch itself: it bills only pages that come back usable and retries blocks for free. A Jev check on top is worth it for the cases a fetcher cannot see, such as the right site answering with the wrong page.