Anthropic released Claude Sonnet 5.5 on 28 September 2026, and within hours we had a Claude Sonnet 5.5 web scraping number nobody else has: on 134 real scraper responses labelled by hand, it judged 133 correctly, for $0.56 in total or $4.17 per 1,000 pages. On the same pages GPT-6 Luna and DeepSeek V4.1 Flash judged all 134 correctly for $0.147 and $0.236 per 1,000. Claude Haiku 4.5 judged 125. If you need a model behind a scraper to say whether a page is the page you asked for, the answer today is GPT-6 Luna, and the whole classifier costs less than the fetch it protects.
Two models got all 134 right, and neither is the one that launched today
The corpus is the one we built for ../jev-block-page-detection/: 70 heavily protected sites fetched twice each, once by a plain Python client and once through our web scraping API over residential exits, reduced to 134 responses and labelled by hand. 55 are usable, meaning the page is the listing, article, product or search the URL asks for. 79 are not: an empty shell, a sign-in page, a consent page, an error page, or a real page that is the wrong page. The 134 pages weigh 79,571,184 bytes together, a median of 112,591 bytes each.
Each model saw the URL, the title and up to 12,000 characters of visible text, and one question: is this the content the URL asks for? It had to answer as JSON with a true or false and a probability, at temperature 0, with 200 output tokens allowed. Every call went through OpenRouter from Europe on 29 September 2026 with two requests in flight. Cost is what OpenRouter billed per call.
| Model | Answered | Correct | Bad pages accepted | Good pages rejected | Cost per 1,000 | Median latency | 90th percentile | Mean input tokens |
|---|---|---|---|---|---|---|---|---|
| GPT-6 Luna | 134 | 134 | 0 | 0 | $0.147 | 1.96 s | 10.06 s | 1,103 |
| DeepSeek V4.1 Flash | 134 | 134 | 0 | 0 | $0.236 | 1.38 s | 3.05 s | 1,183 |
| Claude Sonnet 5.5 | 134 | 133 | 1 | 0 | $4.170 | 1.46 s | 9.74 s | 1,979 |
| Claude Haiku 4.5 | 134 | 125 | 8 | 1 | $1.421 | 1.47 s | 9.61 s | 1,311 |
| Jev 1.13, two questions | 134 | 126 | 2 | 6 | $0.075 | not in this run | not in this run | not billed the same way |
Every model answered every page; no call failed after retries. The two 134s cost $0.0197 and $0.0317 for the whole corpus. Sonnet 5.5 cost $0.5588, which is 28 times the GPT-6 Luna bill for one fewer correct answer. Jev's row comes from the two-question configuration we published in the Jev post, where the page has to pass both "usable" and "matches the URL" above 0.5; its latency was not recorded in that file, and the earlier run in ../jev-vs-llm-benchmark/ measured a median of about 320 ms.
Claude Sonnet 5.5 web scraping check: one miss, and it was a real page
Sonnet 5.5's single error is the most instructive page in the corpus. The URL asked a travel site for hotels in Austin, Texas. The site answered with a complete, real, 1,621,876-byte hotel listing for a different city in South Carolina: prices, a map, a travel guide, everything a usable page has, about the wrong place. Sonnet 5.5 accepted it with a probability of 0.85. Haiku 4.5, GPT-6 Luna and DeepSeek V4.1 Flash all rejected it, and Jev's second question, the one that asks whether the page matches the place named in the URL, scored it 0.38 and rejected it too.
That page was also Sonnet 5.5's most expensive call: 4,581 input tokens, 171 output tokens, $0.0109 and 4.2 seconds, against a median call of 448 input tokens and $0.0011. It read more of that page than of any other and still let it through. A wrong-city page is the hardest case in the set because nothing about it looks like a wall, and every model that caught it caught it by comparing the URL with the content rather than by looking for a stub.
One number in the table needs explaining. Sonnet 5.5 was billed a mean of 1,979 input tokens per page and Haiku 4.5 a mean of 1,311, for identical prompts. The medians tell the same story, 448 against 331, a ratio of 1.35. Anthropic's models overview says the current tokenizer fits roughly 555,000 words in a million tokens where earlier models fit about 750,000, which is the same 1.35. Sonnet 5.5 is billed at twice Haiku's per-token price and on 35% more tokens for the same text, so it costs 2.9 times as much per page, not 2.
Nine pages separate Haiku 4.5 from the rest
Haiku 4.5's 125 comes from eight bad pages accepted and one good page rejected, and the eight fall into two shapes.
Six were empty shells: a search page, a category page, a route page and a review directory that each returned between 708 and 2,674 bytes whose only visible text was the site's own name. Haiku 4.5 accepted all six with a probability of 0.85. Every other model rejected all six, and so did Jev. These are the pages a 300-character minimum length rule catches for free, which is why the ../jev-block-page-detection/ pipeline runs the free checks before any model.
Two were homepages served in place of an Austin search on two accommodation sites, 576,101 and 529,633 bytes of real content about nothing in particular. Haiku 4.5 accepted both. So did Jev in the two-question configuration, with match scores of 0.61 and 0.55. Sonnet 5.5, GPT-6 Luna and DeepSeek V4.1 Flash rejected both.
The ninth is the only good page any of the four LLMs threw away: a 525,068-byte product page for a phone, with the product title, the storage and colour variants and the store navigation all present. Haiku 4.5 rejected it with a probability of 0.15. That single decision is worth more attention than the eight accepts, for a reason the next section explains.
A false accept poisons the dataset; a false reject bills the fetch twice
The two kinds of error do not cost the same, and a pipeline should not treat them as one accuracy number.
A false accept is a wall or a wrong page stored as data. The parser runs on it and returns empty fields, or worse, plausible fields: the wrong city's hotel prices become Austin's hotel prices in the table, and nothing downstream can tell. The cost is the fetch, the classification, the parse, the storage, and then every decision made on the row, and the error is silent. Haiku 4.5's eight false accepts on this corpus would have put six empty rows and two sets of homepage text into a dataset with no flag on them.
A false reject is a real page thrown away. The cost is visible and bounded: you fetch it again, so you pay the fetch twice and the classifier twice, and if the second fetch returns the same page the classifier rejects it again. On our API a plain fetch is $0.0002 and a rendered one $0.001, so Haiku 4.5's one false reject costs a fifth of a cent to a tenth of a cent to retry, plus $0.0014 for the second classification. It never reaches the dataset. When you choose a threshold, lean it towards rejecting: a rejected real page costs one refetch, an accepted wall costs a row you will act on.
Read the table with that in mind. Sonnet 5.5's one error is a false accept, the expensive kind, on the hardest page. Jev's six errors under the two-question rule are all false rejects of homepages the URL actually asked for, the cheap kind, and its two false accepts are the same two homepages that fooled Haiku. GPT-6 Luna and DeepSeek V4.1 Flash made neither kind.
The classifier costs less than the fetch, except with Sonnet
The question that decides whether to run a model on every page is what it adds to the cost of the page. Our web scraping API charges from $0.0002 per page over plain HTTP and $0.001 with JavaScript rendering, and bills only pages that come back. Per 1,000 pages:
| Classifier | Model per 1,000 | Plain fetch + model | Rendered fetch + model | Model as a multiple of the plain fetch |
|---|---|---|---|---|
| Jev 1.13 | $0.075 | $0.275 | $1.075 | 0.4 |
| GPT-6 Luna | $0.147 | $0.347 | $1.147 | 0.7 |
| DeepSeek V4.1 Flash | $0.236 | $0.436 | $1.236 | 1.2 |
| Claude Haiku 4.5 | $1.421 | $1.621 | $2.421 | 7.1 |
| Claude Sonnet 5.5 | $4.170 | $4.370 | $5.170 | 20.9 |
With GPT-6 Luna the check adds 70% to a plain fetch and 15% to a rendered one, and it caught every one of the 79 unusable pages. With Sonnet 5.5 the check costs 21 plain fetches per page. For comparison, the same 134 pages pulled through residential proxies at $0.80/GB, counting 1 GB as 10^9 bytes, would cost $0.064 in bandwidth, about $0.47 per 1,000 pages at this corpus's mean weight of 593,815 bytes; a Sonnet 5.5 verdict on each of those pages costs nine times the bytes it judges.
Sonnet 5.5's list price is $2 per million input tokens and $10 per million output, with cache reads at $0.20, the same as Sonnet 5; Anthropic's launch page says it runs over 30% faster than Sonnet 5 and costs up to 30% less for most work because it uses fewer tokens to get there. Neither claim changes this task. A page classification has no reasoning to shorten: Sonnet 5.5 produced a mean of 21 output tokens per page, Haiku 4.5 produced 22, and the bill is the input. GPT-6 Luna at $0.10 per million input tokens and DeepSeek V4.1 Flash at $0.30 through OpenRouter, or $0.15 direct off-peak, are 20 and 7 times cheaper per input token, and on this task they were also more accurate.
Latency is the same one and a half seconds everywhere, until the tail
Median latency does not separate these models: 1.38 seconds for DeepSeek V4.1 Flash, 1.46 for Sonnet 5.5, 1.47 for Haiku 4.5, 1.96 for GPT-6 Luna, all measured from Europe through OpenRouter with two requests in flight. The tail does. DeepSeek's 90th percentile was 3.05 seconds; Sonnet 5.5, Haiku 4.5 and GPT-6 Luna all had a 90th percentile between 9.6 and 10.1 seconds, and the slowest single call in each of those three took 22 to 29 seconds.
For a batch job that runs overnight the tail is irrelevant. For a check that sits inside a request, before a page is stored or handed to an agent, a 10-second 90th percentile means one call in ten holds the pipeline for ten seconds, and it is the reason the earlier test put Jev, at about 320 ms median, in front of every LLM for in-request decisions. If you need the tail bounded and want an LLM, DeepSeek V4.1 Flash was the only one of the four that kept it under four seconds on this run.
Where Jev sits after today
In the ../jev-block-page-detection/ test Jev with one question got 131 of 134, and with the two-question rule used here 126, all of the difference being homepages the second question was written to reject. On this run two LLMs beat both figures. That does not reverse the earlier conclusion, it narrows it: Jev is still the cheapest classifier per page at $0.075 per 1,000 and by far the fastest, and every one of its errors came with a hesitant score. GPT-6 Luna at $0.147 is twice the price, made no errors, and answers in two seconds instead of a third of one.
The honest caveat is prompt sensitivity. The ../jev-vs-llm-benchmark/ post, run the day before with a different prompt shape and each provider's default reasoning, had Haiku 4.5 at 132 of 134 on this same corpus; today's plain JSON prompt with 200 output tokens gave it 125. Seven cases moved on a prompt change, which is more than the gap between the top three models. Sonnet 5.5's 133, and the two 134s, are from this prompt; run your own prompt on your own labelled pages before you commit a budget to any of them.
The setting that works on block-page detection
- Model: GPT-6 Luna, through OpenRouter or OpenAI, at $0.10 per million input tokens: 134 of 134 on this corpus, $0.147 per 1,000 pages, 1.96 s median. DeepSeek V4.1 Flash is the alternative when the latency tail matters: 134 of 134, $0.236 per 1,000, 3.05 s at the 90th percentile.
- Claude Sonnet 5.5: 133 of 134 at $4.17 per 1,000. Use it where the same call also extracts fields or explains the verdict; as a yes-or-no gate it is 28 times the price of a model that made no errors.
- Fetch: our web scraping API over plain HTTP at $0.0002 per page, rendered at $0.001 only for the pages the classifier marks as empty shells. The API retries unusable responses on its own and bills only pages that come back, so the model sees the 4.5% of pages a fetcher cannot judge: the right site answering with the wrong page.
- Threshold: reject below 0.5, and route anything a model scores between 0.3 and 0.8 to a render or a second model; on this run that band held 2 of Sonnet 5.5's 134 answers and 1 of DeepSeek's, with no errors in it.
- Free checks first: a body under 300 characters, a network error, a missing canonical. They remove the six empty shells that fooled Haiku 4.5 before any model is paid.
- When the classifier is not enough: a rejected page on a listing site is usually the wrong page, not a wall; refetch it with the country pinned to the one the URL names, and work through ../proxy-not-working-checklist/ for the fetch-side settings that change the answer.
- Every account gets $2 of free API usage per month: at $0.0002 a page that is 10,000 plain fetches, and the GPT-6 Luna verdicts on all 10,000 cost $1.47 more.