Unbiased listed Pareto 26.10 Preview on OpenRouter on 1 October 2026, and on 2 October we ran it on 134 real scraper responses labelled by hand: it judged 125 correctly, at $0.84 per 1,000 pages billed. All nine mistakes were bad pages passed as good, and six of them were pages whose entire visible text was the site name, accepted with 0.95 confidence. The September Pareto release rejected those six and scored 131 of 134 at $1.36 per 1,000. GPT-6 Luna scored 134 of 134 at $0.147. If a model behind your scraper decides whether a page is the page you asked for, do not use the preview for it: use GPT-6 Luna, or Jev with Luna as the second opinion.
Pareto 26.10 Preview judged 125 of 134 pages right
The corpus is the one we built for ../jev-block-page-detection/: 70 heavily protected sites fetched twice each, once by a plain Python client and once through our web scraping API over residential exits, reduced to 134 responses and labelled by hand. 55 are usable, meaning the page is the listing, article, product or search the URL asks for. 79 are not: an empty shell, a sign-in page, a consent page, an error page, or a real page that is the wrong page.
Each model saw the URL, the title and up to 12,000 characters of visible text, and one question: is this the content the URL asks for? It answered as JSON with a true or false and a probability, at temperature 0, with 200 output tokens allowed, through OpenRouter from Europe, two requests in flight. The two Pareto rows ran on 2 October 2026; the other rows ran on the same rig and prompt on 29 and 30 September. Cost is what OpenRouter billed per call.
| Model | Correct of 134 | Bad pages accepted | Good pages rejected | Cost per 1,000 | Median latency |
|---|---|---|---|---|---|
| Pareto 26.10 Preview | 125 | 9 | 0 | $0.839 | 1.47 s |
| Pareto, September release | 131 | 3 | 0 | $1.360 | 1.51 s |
| GPT-6 Luna | 134 | 0 | 0 | $0.147 | 2.17 s |
| DeepSeek V4.1 Flash | 134 | 0 | 0 | $0.236 | 1.38 s |
| Jev 1.13, one question | 131 | 3 | 0 | $0.073 | 0.32 s |
| Claude Haiku 4.5 | 125 | 8 | 1 | $1.421 | 1.47 s |
Every model answered every page. The preview's listing says it does not support the structured-output parameter, so JSON is not enforced; all 134 of its answers still came back as valid JSON from the instruction alone. The 90th-percentile latency was 2.24 s for the preview and 3.50 s for the September release.
Six pages that were only the site name, passed as real
Six of the preview's nine false accepts are the clearest pages in the corpus: a marketplace search for ceramic mugs, a pet-store dog-food category, a subreddit, two copies of a flight route page, and a local plumber search. Each response carried between 6 and 10 characters of visible text, the site's own name and nothing else, with a title that was the same name. The preview called every one of them the requested content, with a probability of 0.95.
The September Pareto release saw the same six pages with the same prompt and rejected all six, with a probability of 0.99. So did GPT-6 Luna, DeepSeek V4.1 Flash and Jev 1.13. This is a regression in the preview, not a hard case: the same vendor's previous release gets these pages right.
The other three false accepts are the three accommodation searches for Austin, Texas that also fooled Jev and the September Pareto: one travel site answered with a complete hotel listing for a different city, two others with their homepage. A text classifier only catches those by comparing the place in the URL with the place on the page, and on this corpus only the three OpenAI models and DeepSeek V4.1 Flash caught all three.
The preview gives no warning when it is wrong
Jev's three misses came with scores of 0.54, 0.76 and 0.68, which is why we route Jev's middle band to a second model. The preview's nine misses came with probabilities of 0.95 or 1.0, and its correct answers on good pages carry the same 0.95. There is no band to route on: the score does not tell its right answers from its wrong ones.
That leaves a free check as the only fix. A visible body under 300 characters is never the listing the URL asked for: in this corpus 53 pages are that short and every one is unusable. Put that rule in front and the preview scores 131 of 134, the same as Jev and the September release. At that point it is a $0.84 model doing the work a $0.073 model already does at the same score in a fifth of the time.
Cheaper than the September Pareto, and the price you pay depends on repeats
OpenRouter lists the preview at $0.80 per million input tokens and $3.20 per million output, against $2.50 and $7.50 for the September release: about a third of the input price. On our corpus the bill was $0.1124 for the preview and $0.1822 for the September release, $0.839 and $1.360 per 1,000 pages.
Both bills came in under list-price arithmetic: 152,929 input and 2,040 output tokens at the preview's list price come to $0.1289, and the September release's tokens to $0.3991. The corpus fetches every site twice, and identical prompts were read from cache at $0.03 and $0.25 per million. Page text from a live crawl rarely repeats, so budget at list: $0.96 per 1,000 pages for the preview and $2.98 for the September release. GPT-6 Luna and DeepSeek V4.1 Flash still cost less on the bill and judged every page right.
Our web scraping API charges from $0.0002 per page over plain HTTP and $0.001 with JavaScript rendering, and bills only pages that come back. Per 1,000 pages, with the check added:
| Classifier | Model per 1,000 | Plain fetch + model | Rendered fetch + model | Model as a multiple of the plain fetch |
|---|---|---|---|---|
| Jev 1.13 | $0.073 | $0.273 | $1.073 | 0.4 |
| Jev, then GPT-6 Luna on scores 0.3 to 0.8 | $0.087 | $0.287 | $1.087 | 0.4 |
| GPT-6 Luna | $0.147 | $0.347 | $1.147 | 0.7 |
| DeepSeek V4.1 Flash | $0.236 | $0.436 | $1.236 | 1.2 |
| Pareto 26.10 Preview | $0.839 | $1.039 | $1.839 | 4.2 |
| Pareto, September release | $1.360 | $1.560 | $2.360 | 6.8 |
A preview verdict costs 4.2 plain fetches and lets through nine bad pages in 134. A GPT-6 Luna verdict costs 0.7 plain fetches and let none through. A false accept is the expensive error: a wrong page stored as data returns empty or plausible fields and nothing downstream flags it, while a false reject costs one refetch at $0.0002.
A preview can change under the same model name
OpenRouter describes the release as a preview of the next Pareto version that may change without notice, and points to the September release for stable behaviour. A classifier that sits in a data pipeline needs the opposite: the same answer on the same page next week. Pin a dated model ID for any gate that decides what gets stored, and re-run a labelled set the day the vendor moves it. Ours is 134 pages and costs under $0.20 per model to re-run.
Where the preview may earn its price is the work it is sold for: long agent sessions with a 1,048,576-token context and cache reads at $0.03 per million. A yes-or-no gate on one page uses neither. For that job the numbers above settle it: ../gpt-6-1-sol-block-pages/ has the full OpenAI rows, and ../jev-vs-llm-benchmark/ the original comparison.
The setting that works on block-page detection
- Model: Jev 1.13 first, at $0.073 per 1,000 pages and 0.32 s median, with GPT-6 Luna judging every page Jev scores between 0.3 and 0.8. On this corpus: 134 of 134, $0.087 per 1,000. One model only: GPT-6 Luna, 134 of 134 at $0.147, or DeepSeek V4.1 Flash, 134 of 134 at $0.236 and 1.38 s median.
- Pareto 26.10 Preview: 125 of 134 at $0.84 per 1,000, with six bare site-name pages accepted at 0.95 confidence. Keep it out of any gate that decides what gets stored; the September release scored 131 on the same pages.
- Free checks first: a visible body under 300 characters, a network error, a title that is only the site name. On this corpus they remove 6 of the preview's 9 false accepts before any model is paid.
- Fetch: our web scraping API over plain HTTP at $0.0002 per page, rendered at $0.001 only for the pages the check marks as empty shells, over residential proxies from $0.80/GB when you run your own client. The API retries unusable responses and bills only pages that come back.
- When the classifier is not enough: a rejected page on a listing site is usually the wrong page for the place in the URL; refetch it with the country pinned to that place, and work through ../proxy-not-working-checklist/ for the fetch-side settings. For whole listings, the collectors return rows instead of pages to judge.
- Every account gets $2 of free API usage per month: at $0.0002 a page that is 10,000 plain fetches, and the Jev-then-Luna verdicts on all 10,000 cost $0.87 more.