# DeepSeek V4.1 Flash Providers: Pin 1 of 13

> One run of DeepSeek V4.1 Flash on OpenRouter landed on 13 providers. On 134 labelled pages it scored 134, 132, 130 and 131 across four runs. Pin one host.

[Home](https://quanticdata.io/)/[Blog](https://quanticdata.io/blog/)/DeepSeek V4.1 Flash Providers: Pin 1 of 13

# DeepSeek V4.1 Flash Providers: Pin 1 of 13

AI scrapingOct 5, 2026·9 min read·By [Aldo Morese](https://quanticdata.io/about/), founder of QuanticData

DeepSeek V4.1 Flash judged the same 134 scraped pages on OpenRouter: 134 right on 29 September, 130 on 5 October with the router choosing the host, 131 pinned to DeepInfra, 131 pinned to Sail Research, against GPT-6 Luna at 134; the card shows that one unpinned run of 134 calls landed on 13 different providers.

On this page [The release watch caught a model that moved without a new version](/blog/deepseek-v4-1-flash-providers/#the-release-watch-caught-a-model-that-moved-without-a-new-ve) [One run of 134 calls landed on 13 providers](/blog/deepseek-v4-1-flash-providers/#one-run-of-134-calls-landed-on-13-providers) [The hosts disagree on the same four pages](/blog/deepseek-v4-1-flash-providers/#the-hosts-disagree-on-the-same-four-pages) [An empty answer is a reject, not a retry](/blog/deepseek-v4-1-flash-providers/#an-empty-answer-is-a-reject-not-a-retry) [What it costs per 1,000 pages, fetch included](/blog/deepseek-v4-1-flash-providers/#what-it-costs-per-1-000-pages-fetch-included) [How to pin a provider on OpenRouter](/blog/deepseek-v4-1-flash-providers/#how-to-pin-a-provider-on-openrouter) [The setting that works on DeepSeek V4.1 Flash as a page classifier](/blog/deepseek-v4-1-flash-providers/#the-setting-that-works-on-deepseek-v4-1-flash-as-a-page-clas)

DeepSeek V4.1 Flash has 28 providers on OpenRouter, and the provider decides the answer: on the same 134 hand-labelled scraper responses, at temperature 0, the unpinned model ID scored 134 on 29 September, 132 on 2 October, and 130 and 131 on 5 October. One run of 134 calls landed on 13 different hosts. Pinned to one host on 5 October it scored 131 on DeepInfra, 130 on Decart and 131 on Sail Research, and the misses fall on the same four pages every time. If DeepSeek V4.1 Flash decides what your scraper keeps, pin one provider with fallbacks off and treat an empty answer as a reject: 133 of 134 at $0.107 per 1,000 pages. If the gate must give the same answer every week, use GPT-6 Luna: 134 of 134 in all three runs, at $0.147.

## The release watch caught a model that moved without a new version

Every Monday we re-run each model already measured on the corpus from ../jev-block-page-detection/: 134 responses from 70 heavily protected sites, fetched once by a plain Python client and once through our [web scraping API](https://quanticdata.io/web-scraping-api/), labelled by hand. 55 are the page the URL asks for. 79 are not: an empty shell, a sign-in page, a consent page, an error page, or a real page for the wrong thing. Each model gets the URL, the title, up to 12,000 characters of visible text and one question, answered as JSON.

On 5 October seven of the nine models gave the same score as their baseline, within one page. DeepSeek V4.1 Flash went from 134 to 130 under the same model ID and the same dated slug, deepseek-v4.1-flash-20260910. Nothing in the listing said the model had changed. What changes on every call is the host: OpenRouter balances requests across the providers that serve a model, by price, unless you tell it otherwise.

| Run | Who picked the host | Correct of 134 | Bad pages accepted | No answer | Cost per 1,000 | Median latency |
| --- | --- | --- | --- | --- | --- | --- |
| 29 September | OpenRouter | 134 | 0 | 0 | $0.236 | 1.38 s |
| 2 October | OpenRouter | 132 | 1 | 1 | $0.331 | 1.02 s |
| 5 October, morning | OpenRouter | 130 | 2 | 2 | $0.207 | 1.25 s |
| 5 October, afternoon | OpenRouter | 131 | 1 | 2 | $0.175 | 1.24 s |
| 5 October | Pinned: DeepInfra, fp8 | 131 | 3 | 0 | $0.149 | 0.52 s |
| 5 October | Pinned: Decart, fp4 | 130 | 0 | 4 | $0.111 | 1.66 s |
| 5 October | Pinned: Sail Research, fp4 | 131 | 0 | 3 | $0.107 | 3.51 s |

No run rejected a good page that it answered. Cost is what OpenRouter billed per call. The quantization column on the provider list is the host's own declaration: fp4 and fp8 are reduced-precision builds of the same weights, and several hosts do not declare one.

**Try it on your own targets.** Every QuanticData account gets $2 of free API usage per month.

[Start free — $2/month included](https://quanticdata.io/signup/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog-inline&amp;utm_content=deepseek-v4-1-flash-providers)[Explore Web Scraping API](https://quanticdata.io/web-scraping-api/)

## One run of 134 calls landed on 13 providers

For the afternoon run we logged the host that served each call. With no provider preference, 134 calls went to InferenceNet 36 times, DeepInfra 25, CoreWeave 18, AtlasCloud 18, Together 13, Decart 9, Sail Research 8, Baidu 2, and Wafer, Ionstream, Morph, Makora and SiliconFlow once each. Every error in that run came from the 9 calls Decart served: 6 right, 1 bad page accepted, 2 empty answers. The other 125 calls, on 12 hosts, were all right.

That is why the score drifts while the model ID stays put. Your score on a given day is the mix of hosts the router picked that day, weighted by which pages happened to land on which host. The same rig at temperature 0 gives the same answer from the same host far more often than it gives the same answer from the same model ID.

## The hosts disagree on the same four pages

The misses are not spread across the corpus. They sit on four pages, and each host fails them in its own way:

- A travel site's Austin hotel search that came back with a full hotel list for Orangeburg, South Carolina.

- A second travel site's Austin hotels URL that came back with a generic deals page.

- A vacation-rental search for Austin that came back with the site's homepage.

- A furniture retailer's sofa category: a real, usable page with 70,706 characters of text.

Pinned to DeepInfra, the model accepted all three wrong-city pages as the requested content, with probabilities of 0.9 and 0.95. Pinned to Decart or Sail Research, it accepted none of them: on those pages, and on the sofa category, the fp4 hosts returned an empty answer, a 200 response with no text, on all eight attempts. On the morning run one call came back as a refusal in Chinese on a news site's technology section, a page every other run judged usable.

The three wrong-city pages are the ones that also fooled Jev, the September Pareto release and Pareto 26.10 Preview in ../pareto-26-10-preview-block-pages/. A text classifier only catches them by comparing the place in the URL with the place on the page. The fp4 hosts did not get them wrong; they gave no answer, and that is the safer failure.

## An empty answer is a reject, not a retry

The rig retries an empty answer eight times and then records no verdict. A pipeline should not: an empty answer from the classifier means "do not store this page yet". Count it as a reject and send the page back for one refetch at $0.0002, and the scores change:

| Host | As measured | Empty answer counted as reject | Bad pages accepted | Good pages rejected |
| --- | --- | --- | --- | --- |
| Sail Research, pinned | 131 | 133 | 0 | 1 |
| Decart, pinned | 130 | 133 | 0 | 1 |
| OpenRouter picks, 5 Oct afternoon | 131 | 132 | 1 | 1 |
| DeepInfra, pinned | 131 | 131 | 3 | 0 |

The one good page rejected is the sofa category, which costs one refetch. The bad pages accepted are the expensive error: a hotel list for the wrong city stored as data looks complete and nothing downstream flags it. On that measure the two fp4 hosts are the best DeepSeek V4.1 Flash we found on 5 October, and DeepInfra, the fastest at 0.52 s median, is the worst.

We could not pin DeepSeek's own endpoint. OpenRouter removed it from routing for our account because our privacy setting excludes providers that may train on prompts. If you leave that setting off, the first-party host is one of the candidates the router mixes in; if you keep it on, as a pipeline handling client URLs should, it never is.

## What it costs per 1,000 pages, fetch included

The listing price at the top of the OpenRouter page read $0.30 per million input tokens and $1.20 per million output on 5 October. The routed calls billed well under that, because the router prefers cheaper hosts: Sail Research lists $0.08 and $0.40, Decart $0.09 and $0.18, DeepInfra $0.14 and $0.42. DeepInfra also answered in 14 output tokens a page on average, against 73 to 76 on the fp4 hosts, which is most of its speed.

Our web scraping API charges from $0.0002 per page over plain HTTP and $0.001 with JavaScript rendering, and bills only pages that come back. Per 1,000 pages, classifier included:

| Classifier | Correct of 134 | Model per 1,000 | Plain fetch + model | Rendered fetch + model | Model as a multiple of the plain fetch |
| --- | --- | --- | --- | --- | --- |
| Jev 1.13, then GPT-6 Luna on scores 0.3 to 0.8 | 134 | $0.087 | $0.287 | $1.087 | 0.4 |
| DeepSeek V4.1 Flash, Sail Research pinned, empty = reject | 133 | $0.107 | $0.307 | $1.107 | 0.5 |
| DeepSeek V4.1 Flash, Decart pinned, empty = reject | 133 | $0.111 | $0.311 | $1.111 | 0.6 |
| GPT-6 Luna, three runs | 134 | $0.147 | $0.347 | $1.147 | 0.7 |
| DeepSeek V4.1 Flash, DeepInfra pinned | 131 | $0.149 | $0.349 | $1.149 | 0.7 |
| DeepSeek V4.1 Flash, OpenRouter picks | 130 to 134 | $0.175 to $0.331 | $0.375 to $0.531 | $1.175 to $1.331 | 0.9 to 1.7 |

Leaving the host to the router is the most expensive way to run this model on this job and the least repeatable. Pinning makes it cheaper and makes the score a property of your setup instead of a property of the day. GPT-6 Luna scored 134 on 30 September, 2 October and 5 October; the full OpenAI rows are in ../gpt-6-1-sol-block-pages/.

## How to pin a provider on OpenRouter

Add a provider object to the request: the host's name in `order` and `allow_fallbacks` set to false. Without the second field OpenRouter treats the order as a preference and falls back to any other host when yours is busy, which puts you back in the mix.

```
"provider": { "order": ["Sail Research"], "allow_fallbacks": false }
```

The trade is availability: with fallbacks off, a request fails when that host is down. List a second host only if you have measured it on your own labelled set, and log the `provider` field of every response, so a score change can be traced to a host change. Re-run the labelled set weekly; ours is 134 pages and cost $0.0144 on the cheapest pinned host. The method and the first comparison are in ../jev-vs-llm-benchmark/.

## The setting that works on DeepSeek V4.1 Flash as a page classifier

- **Host**: pin one with `order` and `allow_fallbacks: false`. On 5 October: Sail Research or Decart, both fp4, 133 of 134 with an empty answer counted as a reject, $0.107 and $0.111 per 1,000 pages, no bad page accepted. DeepInfra is the fastest at 0.52 s median and accepted 3 wrong-city pages.

- **Empty answer**: a reject and one refetch, never a retry loop. It is how the fp4 hosts handle the pages the others get wrong.

- **Stable instead of cheap**: Jev 1.13 on every page with GPT-6 Luna on scores between 0.3 and 0.8, 134 of 134 at $0.087 per 1,000; or GPT-6 Luna alone, 134 of 134 in three runs at $0.147.

- **Fetch**: our [web scraping API](https://quanticdata.io/web-scraping-api/) over plain HTTP at $0.0002 per page, rendered at $0.001 only for empty shells, over [residential proxies](https://quanticdata.io/residential-proxies/) from $0.80/GB when you run your own client. Pin the exit country to the place in the URL: on this corpus the wrong-city pages are the ones every classifier struggles with, and ../proxy-not-working-checklist/ covers the fetch-side settings.

- **When the classifier is not enough**: for listings you need as rows, the [collectors](https://quanticdata.io/collectors/) return structured records instead of pages to judge.

- Every account gets $2 of free API usage per month: 10,000 plain fetches at $0.0002, and the pinned DeepSeek V4.1 Flash verdicts on all 10,000 cost $1.07 more.

### Sources & further reading

- [OpenRouter model page: DeepSeek V4.1 Flash (providers, pricing)](https://openrouter.ai/deepseek/deepseek-v4.1-flash)

- [OpenRouter documentation: Provider Routing (order, allow_fallbacks, quantizations)](https://openrouter.ai/docs/guides/routing/provider-selection)

- [DeepSeek API news: DeepSeek-V4.1-Flash release (10 September 2026)](https://api-docs.deepseek.com/news/news260910/)

- [Hugging Face model card: deepseek-ai/DeepSeek-V4.1-Flash](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash)

## FAQ

Quick answers on deepseek v4.1 flash providers.

[Something else? Ask us](mailto:hello@quanticdata.io)

### How many providers serve DeepSeek V4.1 Flash on OpenRouter?

OpenRouter lists 28 providers for DeepSeek V4.1 Flash. In one run of 134 calls with no provider preference on 5 October 2026, 13 different hosts served the requests, from InferenceNet with 36 calls to five hosts with one call each.

### Do DeepSeek V4.1 Flash providers give the same answers?

Not on our corpus. On the same 134 labelled pages at temperature 0, pinned on 5 October 2026, DeepInfra scored 131 with 3 bad pages accepted, Decart 130 with 4 empty answers and Sail Research 131 with 3 empty answers. Unpinned, the model ID scored 134, 132, 130 and 131 across four runs.

### How do I pin a provider on OpenRouter?

Send a provider object with the host in order and allow_fallbacks set to false, for example order ["Sail Research"]. Without allow_fallbacks false, OpenRouter falls back to other hosts. Log the provider field of each response; in our unpinned run 134 calls went to 13 hosts.

### Which DeepSeek V4.1 Flash provider is best for classifying scraped pages?

On 5 October 2026, Sail Research and Decart, both fp4: 133 of 134 with an empty answer counted as a reject, no bad page accepted, at $0.107 and $0.111 per 1,000 pages. DeepInfra was the fastest at 0.52 s median but accepted 3 wrong-city pages.

### What does DeepSeek V4.1 Flash cost per 1,000 pages classified?

Between $0.107 and $0.149 per 1,000 pages pinned to one host, and $0.175 to $0.331 when OpenRouter picks the host, billed on our 134-page corpus. The listing price on 5 October 2026 read $0.30 per million input tokens and $1.20 per million output.

### Can I pin DeepSeek first-party on OpenRouter?

Not with a privacy setting that excludes providers that may train on prompts: OpenRouter removed the DeepSeek endpoint from routing for our account on 5 October 2026. With that setting on, the 28 listed providers drop to the third-party hosts.

## Give the classifier pages worth judging

Every page in this test came from our web scraping API: plain HTTP or rendered, residential exits, unusable responses retried and never billed, so the model only sees the cases a fetcher cannot settle. Every account gets $2 of free API usage per month.

[Start free — $2/month included](https://quanticdata.io/signup/?utm_source=blog&amp;utm_medium=website&amp;utm_campaign=blog-end&amp;utm_content=deepseek-v4-1-flash-providers)[Explore Web Scraping API](https://quanticdata.io/web-scraping-api/)

## Related reading

[AI scraping Pareto 26.10 Preview: 125 of 134 Pages Right The day after Unbiased listed Pareto 26.10 Preview on OpenRouter we sent it the 134 hand-labelled scraper responses from our Jev tests. It got 125 right at $0.84 per 1,000 pages, and accepted six pages whose entire visible text was the site name, each with 0.95 confidence. The September Pareto release rejected all six and scored 131. GPT-6 Luna scored 134 at $0.147. Read more](https://quanticdata.io/blog/pareto-26-10-preview-block-pages/) [AI scraping ChatGPT Web Scraping: 20 Live Pages via MCP ChatGPT reads live web pages once you connect an MCP server as a developer-mode app. On 2 October 2026 we sent 20 public pages through the QuanticData scrape tool: all 20 came back, 18 over plain HTTP at a median 5.4 seconds, and nine pages of raw HTML worth about 947,431 tokens arrived as 75,777 tokens of Markdown. The 20 pages cost $0.0056. Read more](https://quanticdata.io/blog/chatgpt-web-scraping/) [AI scraping LlamaIndex Web Scraping: 95% Fewer Tokens We sent 20 pages a LlamaIndex RAG app would ingest through one batch job on 2 October 2026: 19 came back in 49.8 seconds, all over plain HTTP, for $0.0038. Raw HTML would have put 1,904,058 tokens into your Documents; smart Markdown with link targets stripped put 103,333. Here is the code, the chunk sizes and the embedding bill. Read more](https://quanticdata.io/blog/llamaindex-web-scraping/)

---

Source: https://quanticdata.io/blog/deepseek-v4-1-flash-providers/ · Site index for AI: https://quanticdata.io/llms.txt
