The QuanticData Blog
Practical guides on web scraping, SERP data, proxies and web data pipelines for AI — every post backed by real SERP research, not opinions.
Social proxies Lead generation Web scraping Proxies Troubleshooting SEO data Guides Data for AI MCP & agents Browser agents AI scraping Crawling SERP & search Use cases Anti-bot
All guides, newest first
Our weekly re-run caught DeepSeek V4.1 Flash moving from 134 of 134 to 130 under the same model ID. The model did not change: the host did. One run of 134 calls landed on 13 providers, and the hosts disagree on the same four hard pages. Pinned to one fp4 host and with an empty answer treated as a reject, it scores 133 at $0.107 per 1,000 pages.
Oct 5, 2026 · 9 minRead more TroubleshootingSSL CERTIFICATE_VERIFY_FAILED Behind a ProxyA plain CONNECT proxy never touches certificates, so CERTIFICATE_VERIFY_FAILED means something re-signed the site or your trust bundle is wrong. We reproduced the error 42 times across seven clients to show which tail means what, and which fixes really work.
Oct 5, 2026 · 8 minRead more Lead generationGoogle Maps Scraper, Tested: 1,644 PlacesEight keyword-and-city searches through our Google Maps collector on 4 October 2026 delivered 1,644 unique places in 7 minutes 19 seconds, for $1.64 at the list price of $0.001 per place. Name, category, coordinates and place id were on every row; rating on 96.4%, address on 96.0%, phone on 95.3%, website on 90.7%. Run by run, with field coverage and cost.
Oct 4, 2026 · 7 minRead more SERP & searchSERP API Benchmark: Google, Bing, DuckDuckGoWe sent the same 40 queries to Google, Bing and DuckDuckGo through our SERP API on 4 October 2026 and kept every row. 117 of 120 calls returned a full results page on the first attempt, the median call took 2.92 seconds, and the three engines returned different blocks: AI Overviews on 21 of 38 Google pages, inline videos on 21 of 39 Bing pages. 1,000 searches cost $0.50 on Bing or DuckDuckGo and $2 on Google.
Oct 4, 2026 · 7 minRead more Social proxiesMobile Proxies for Instagram & TikTok, TestedWe fetched public Instagram and TikTok profiles through QuanticData mobile sessions on 4 October 2026. All 17 connected sessions exited on a carrier network in the requested country, across six countries, and returned 51 of 51 Instagram profiles with a 1.21-second median time to first byte. On TikTok, 22 of 24 profile requests from the United States, the United Kingdom, Germany and Brazil returned the profile.
Oct 4, 2026 · 10 minRead more Social proxiesUntappd Proxies: What 23 Requests ShowTwenty-three requests to Untappd on 2 October 2026 from the US, Ireland, Germany, the UK and Japan. Every public beer and brewery page loaded on the first attempt over plain HTTP, at about 160 KB, against 685 KB rendered. The "where to find" page changes completely by exit country: 509 venues from the US, 109 from Germany, 11 from Ireland. Untappd's terms exclude robots and data mining, and its official API allows 100 calls an hour per key. Here is what is public, what each route costs and what a proxy is for on Untappd.
Oct 2, 2026 · 9 minRead more Social proxiesTaringa Proxies: What the Domain Serves NowTaringa!, the Argentine social network, shut down on 25 March 2024. On 2 October 2026 we sent eighteen requests to taringa.net and the Wayback Machine from Argentina, Mexico and the United States. Every path, robots.txt included, returns the same 5,524-byte memorial page with status 200, the API host no longer resolves, and the old content lives only in archives. Here is what that means for proxies, link checks and research.
Oct 2, 2026 · 8 minRead more Social proxiesSpoutible Proxies: What 16 Requests ShowSixteen requests to Spoutible on 2 October 2026 from the US, the UK, Germany, Japan and India. Nothing was blocked, yet plain HTTP returned a 5 KB shell with no posts, and only a browser that waited eight seconds saw content, at about 1.84 MB a page. Spoutible's terms also forbid false IP addresses. Here is what is public, what each request costs, and what a proxy is honestly good for there.
Oct 2, 2026 · 8 minRead more Social proxiesNostr Proxies: What 33 Requests ShowThirty-three requests to Nostr surfaces on 2 October 2026 from the US, Germany, Brazil and the UK. Nothing blocked us. The cost of the same profile ranged from 99 bytes to 1.1 MB depending on where you ask, and the real limits are the ones each relay publishes about itself. Here is what is public, what each surface weighs, and where a proxy actually helps.
Oct 2, 2026 · 10 minRead more Social proxiesHive Social Proxies: What 25 Requests ShowTwenty-five requests on 2 October 2026 from eight countries to Hive Social and its App Store listings. The website is a 20 KB brochure with no profiles or posts, identical everywhere; the app has not been updated on iOS since 30 August 2023; and Germany's App Store returns nothing for it. Here is what is public, what Hive's rules say, and the small set of jobs a proxy is actually good for.
Oct 2, 2026 · 9 minRead more Social proxiesGoodreads Proxies: What 19 Requests ShowNineteen requests to Goodreads on 2 October 2026 from six countries. Book, author and list pages loaded over plain HTTP on the first try every time, with the full text and ratings in the HTML. Rendering the same book page cost 2.9 times the traffic for no extra words. The public API has been closed to new keys since December 2020, and the terms exclude robots and data mining. Here is what is public, what it costs, and what a proxy is actually for on Goodreads.
Oct 2, 2026 · 9 minRead more Social proxiesDeviantArt Proxies: What 20 Requests ShowTwenty requests to DeviantArt on 2 October 2026 from the US, Germany, Japan, the UK and Brazil. Every profile and artwork page loaded on the first attempt over plain HTTP, with no challenge. The real story is weight: a single artwork page is 886 KB, the same facts through the official oEmbed endpoint are 3.5 KB. Here is what DeviantArt serves, what changes by country, what the terms say about scraping and AI, and when a proxy is worth it at all.
Oct 2, 2026 · 9 minRead more AI scrapingLlamaIndex Web Scraping: 95% Fewer TokensWe sent 20 pages a LlamaIndex RAG app would ingest through one batch job on 2 October 2026: 19 came back in 49.8 seconds, all over plain HTTP, for $0.0038. Raw HTML would have put 1,904,058 tokens into your Documents; smart Markdown with link targets stripped put 103,333. Here is the code, the chunk sizes and the embedding bill.
Oct 2, 2026 · 10 minRead more AI scrapingChatGPT Web Scraping: 20 Live Pages via MCPChatGPT reads live web pages once you connect an MCP server as a developer-mode app. On 2 October 2026 we sent 20 public pages through the QuanticData scrape tool: all 20 came back, 18 over plain HTTP at a median 5.4 seconds, and nine pages of raw HTML worth about 947,431 tokens arrived as 75,777 tokens of Markdown. The 20 pages cost $0.0056.
Oct 2, 2026 · 8 minRead more AI scrapingPareto 26.10 Preview: 125 of 134 Pages RightThe day after Unbiased listed Pareto 26.10 Preview on OpenRouter we sent it the 134 hand-labelled scraper responses from our Jev tests. It got 125 right at $0.84 per 1,000 pages, and accepted six pages whose entire visible text was the site name, each with 0.95 confidence. The September Pareto release rejected all six and scored 131. GPT-6 Luna scored 134 at $0.147.
Oct 2, 2026 · 8 minRead more Social proxiesClapper Proxies: What 17 Requests ShowSeventeen requests to Clapper on 2 October 2026 from the US, Germany, the UK and Mexico. Profiles and videos never loaded over plain HTTP, only two of six browser renders got through Cloudflare, and a failed render cost more traffic than a successful one. Clapper's terms forbid scraping outright. Here is what is public, what each error means, and what a proxy is actually good for on Clapper.
Oct 2, 2026 · 9 minRead more MCP & agentsLocal LLM Web Search MCP: What Fits in 4KWe measured what an MCP web-search server costs a local model on 2 October 2026. The 28 tool definitions alone are 13,947 tokens, more than three times the 4,096-token context Ollama gives a model on a GPU under 24 GiB. Loading one tool cuts that to 356 tokens, and a bounded search-and-read answer fits beside it with 1,882 tokens to spare.
Oct 2, 2026 · 7 minRead more MCP & agentsGive Ollama Internet Access: 360-Token SetupOllama runs models, it does not browse. To give an Ollama model internet access you put an MCP client in front of it and connect a web-search server. On 2 October 2026 our full tool list measured 13,947 tokens, more than Ollama's 4,096-token default, so the setup below loads one tool for 356 tokens and caps every result.
Oct 2, 2026 · 6 minRead more MCP & agentsLM Studio MCP Web Search: 14K vs 360 TokensLM Studio has been an MCP host since version 0.3.17, so web search is one block in mcp.json. The catch is in LM Studio's own docs: servers built for cloud models can flood a local context. We measured ours on 2 October 2026: 13,947 tokens for all 28 tools, 4,215 for the three web tools, 356 for one.
Oct 2, 2026 · 5 minRead more MCP & agentsQwen Web Search MCP: Qwen3 Online via OllamaA local Qwen3 model gets web search through Qwen-Agent, which speaks MCP. The config has one trap: a server with a URL is treated as SSE unless you set the type to streamable-http. On 2 October 2026 our full tool list measured 13,947 tokens; one tool measured 356, small enough for a 4K Qwen on a laptop.
Oct 2, 2026 · 5 minRead more MCP & agentsOpen WebUI MCP Server: Web Search in 5 StepsOpen WebUI has spoken MCP natively since v0.6.31, so an admin can give every Ollama model in the instance web search from one form. We measured what that costs the model on 2 October 2026: 13,947 tokens of definitions for all 28 of our tools, 4,215 for the three web tools, 356 for one.
Oct 2, 2026 · 5 minRead more Social proxiesGETTR Proxies: What 17 Fetches ShowSeventeen fetches of GETTR on 1 October 2026 from the United States, Germany and Brazil. The HTML sits behind Imperva; the no-JS page is 8,793 bytes of metadata, the rendered one 3.8 MB. The JSON endpoints the app itself uses answered on the first attempt: 1 KB for a profile, 60 KB for twenty posts. Here is the measurement, which proxy it needs, the cost math and the rules GETTR publishes.
Oct 1, 2026 · 9 minRead more AI scrapingGPT-6.1 Sol Web Scraping: 134 of 134 RightThe day after OpenAI released GPT-6.1 Sol we sent it the 134 hand-labelled scraper responses from our Jev tests: is this the page the URL asked for, or an empty shell, a sign-in page, an error, the wrong page? It got all 134 right, for $2.74 per 1,000 pages. So did GPT-6 Sol at the same price, and GPT-6 Luna at $0.147. Jev 1.13 got 131 in a third of a second, and all three of its misses came with a hesitant score.
Sep 30, 2026 · 10 minRead more Social proxiesWhatsApp Proxies: What 19 Fetches ShowNineteen fetches of WhatsApp on 30 September 2026 from the United States, Germany and India. A public channel page returns about 202 KB over plain HTTP with the name, description and follower count, and no posts; the channel directory lists nothing without the app. A German exit prints the exact follower count where English rounds it to 593K. Here is the measurement, the arithmetic, the proxy the app itself expects, and the line WhatsApp’s terms draw.
Sep 30, 2026 · 11 minRead more MCP & agentsMCP Server OAuth 2.1: Connect in 3 StepsThe hosted QuanticData MCP server now supports OAuth 2.1 sign-in. Add https://api.quanticdata.io/mcp to a client with MCP authorization, such as a Claude custom connector, sign in and click Allow access. No API key to copy, one revocable key per connection, API keys still accepted.
Sep 29, 2026 · 7 minRead more CrawlingCloudflare Crawl API: 1 in 54 Pages Needed JSCloudflare's /crawl endpoint launches a headless browser for every page unless you pass render: false. We crawled 77 pages across four sites and compared both modes on 54: plain HTTP returned 35,957 words, the browser 32,908. Only one page gained more than 20%. We also read 78 robots.txt files: 13 disallow its crawler.
Sep 29, 2026 · 8 minRead more MCP & agentsMCP Web Scraper: 11 Pages, 96% Fewer TokensWe fetched 11 public pages four ways on 29 September 2026. Raw HTML came to about 852,229 tokens, full Markdown 97,586, smart Markdown 63,374 and smart Markdown without link targets 37,039. The content mode, not the model, sets the bill for an agent that reads the web.
Sep 29, 2026 · 8 minRead more AI scrapingOpenAI Scraping Lawsuit: 0 of 30 Block BingbotUnsealed filings in the New York Times-led case say Microsoft and OpenAI built training sets from news, including data gathered for Bing. We read the robots.txt of 30 news sites on 29 September 2026: 20 block GPTBot, 27 block ClaudeBot, and not one blocks Bingbot.
Sep 29, 2026 · 9 minRead → AI scrapingClaude Sonnet 5.5 for Web Scraping: 133 of 134Hours after Anthropic released Claude Sonnet 5.5 we sent it the same 134 hand-labelled scraper responses we used for the Jev tests: is this the requested page, or a stub, a sign-in page, an error, the wrong page? It got 133 right for $0.56, about $4.17 per 1,000 pages. GPT-6 Luna and DeepSeek V4.1 Flash got all 134 for $0.147 and $0.236 per 1,000. Claude Haiku 4.5 got 125. The one page Sonnet 5.5 missed was a real one.
Sep 29, 2026 · 12 minRead → Social proxiesXiaohongshu Proxies: 70 KB a Note, No BrowserTwenty-four fetches of Xiaohongshu on 28 September 2026. A public note requested with the token the feed hands out answers a plain HTTP client with 70,499 bytes, the full note text in the meta description and an Article JSON-LD block; rendered, the same note weighs 2,475,522 bytes. The explore feed is 186,518 bytes of server-rendered cards. The exit country changes the feed, not the note.
Sep 29, 2026 · 9 minRead → Social proxiesNaver Proxies: 768 Words Without a BrowserTwenty fetches of Naver on 28 September 2026. The desktop blog is a frame that returns 0 words with or without JavaScript; the same post on the mobile subdomain returns 768 words in 155,493 bytes over plain HTTP, identical from Korea, the United States and Germany. The blog RSS hands over 50 posts in 99,632 bytes, a public cafe serves 8,280 words in a 1990s Korean charset, and the Open API answers a keyless request with 113 bytes and a documented quota of 25,000 calls a day.
Sep 29, 2026 · 10 minRead → Social proxiesKuaishou Proxies: 1,009 Words, Zero BrowserThirteen fetches of Kuaishou on 28 September 2026. A public profile answers a plain HTTP client with 341,397 bytes, 1,009 words and 31 video cards from Singapore, 341,309 bytes and 1,010 words from the United States. Rendered it is 606,175 bytes, 39.7 seconds and the same 1,009 words. The head is generic on every profile, the video page is the one surface that needs a browser, and robots.txt allows fourteen named crawlers, GPTBot included, and nobody else.
Sep 29, 2026 · 8 minRead → Social proxiesDouyin Proxies: 2.1 KB a Profile via the APINineteen fetches of Douyin on 28 September 2026. A public user page answers a plain HTTP client with 72,914 bytes and zero words from every country. The endpoint that does the job needs no key: 2,100 bytes of JSON with follower and like counts, identical from all three exits. That is 3,902 times less than the 8,194,019-byte render.
Sep 29, 2026 · 9 minRead → Social proxiesBilibili Proxies: 158 Bytes Per Follower CountTwenty-one fetches of Bilibili on 28 September 2026. A public video page answers a plain HTTP client with 230,769 bytes, a VideoObject JSON-LD block and a hydration state carrying the exact view, like and coin counts; rendered, it is 1,696,995 bytes, and from the US the render added nothing. Three keyless API endpoints returned JSON from every country, the smallest of them a follower count in 158 bytes, with a header naming the exit region.
Sep 29, 2026 · 9 minRead → Social proxiesZalo Proxies: 120 Words Rendered, 0 WithoutZalo answers a plain HTTP client with 4,036 bytes and zero words of body, and the same 120 rendered words from Vietnam, Singapore and the United States. The head is worth reading anyway: it carries the account name and a description with five social handles. Measured on 28 September 2026, with the two-line robots.txt, the developer portal and the cost of a rendered page.
Sep 29, 2026 · 8 minRead → Social proxiesOdnoklassniki Proxies: 1,434 Words, No BrowserOdnoklassniki hands a plain HTTP client the whole public group page: 1,434 words, the member count in the meta description, a canonical and breadcrumb JSON-LD, from Germany, the United States and Kazakhstan alike. Rendering the same URL costs 2,073,620 bytes instead of 554,224 and returns the same words. Measured on 28 September 2026, with the robots.txt, the API rule and the cost per page.
Sep 29, 2026 · 8 minRead → Social proxiesMeWe Proxies: 17 Words, and Where the Data IsEvery URL on mewe.com, robots.txt included, is the same 15,465-byte shell with 17 words, from the United States, Germany and the United Kingdom, rendered or not. The terms, which only exist after rendering, forbid data mining and robots. So the public data about MeWe is not on mewe.com: an App Store record in 1.0 second, a Google Play record in 4.7, twenty reviews in 1.2. Measured on 28 September 2026.
Sep 29, 2026 · 9 minRead → Social proxiesLINE Proxies: 79 Words in Japan, 112 in the USOne LINE Official Account profile page, three residential exits, 28 September 2026: 79 words from Japan, 80 from Thailand, 112 from the United States, identical with and without JavaScript. The exit country rewrites the language of the page chrome, not the account. Plain HTTP costs 91,422 bytes; rendering costs 1,833,420 for the same words. Here is the measurement, the 70-shard sitemap, the API rule and the price per page.
Sep 29, 2026 · 8 minRead → Social proxiesKakao Proxies: 33 Words Rendered, 4 From JapanA Kakao Talk Channel page hands a plain HTTP client 4,198 bytes and zero words, with an Organization block in the head. Rendered, it returns 33 words from a Korean or an American exit and 4 from a Japanese one, all with status 200. Measured on 28 September 2026: the exit decides whether the render finishes, robots.txt allows everything, and the REST API needs an app key for your own channel.
Sep 29, 2026 · 7 minRead → Social proxiesXING Proxies: 2,514 Words Without a BrowserTwenty-two fetches of XING on 28 September 2026 from Germany, Austria and the United States. A public company page returns 2,514 words to a plain HTTP client, byte-identical from Germany and Austria, with follower count, headcount band, news, twelve job ads and named employees. A member profile returns 679 words and a complete Person JSON-LD block with every role and start date. Rendering the company page costs 1,511,542 bytes for 394 extra words of interface. Here is the whole measurement, what robots.txt closes, and the one line about personal data.
Sep 29, 2026 · 10 minRead → Social proxiesStrava Proxies: 117 Words Without a LoginTwenty-one fetches of Strava on 28 September 2026 from the United States, Germany and Italy. A public segment page returns 117 words to a logged-out client: the segment name, sport and location in the title, and a sign-in prompt in the body, from any country and with or without a browser. The club page is the one public surface with data: 7,284,511 members and twenty dated posts in 963,234 bytes over plain HTTP. The API needs OAuth and documents its own limits. Here is the whole measurement and the sentence in Strava’s terms that governs all of it.
Sep 29, 2026 · 10 minRead → Social proxiesNextdoor Proxies: 2,404 Words, No LoginNineteen fetches of Nextdoor on 28 September 2026 from the United States, Canada and the United Kingdom. A public city page returns 2,404 words to a plain HTTP client, identical from all three countries: resident counts, six neighbour conversations, 1,147 groups, marketplace listings with prices and 44 events. Rendering the same URL costs 1,046,614 bytes and returns 322 fewer words. Here is the whole measurement, the two paths robots.txt leaves open, and where Nextdoor’s own Display API takes over.
Sep 29, 2026 · 10 minRead → Social proxiesLetterboxd Proxies: 2,830 Words, No BrowserNineteen fetches of Letterboxd on 28 September 2026 from the United States, the United Kingdom and Germany. A public film page returns 2,830 words to a plain HTTP client, identical from all three countries, with cast, crew and twelve popular reviews. The average rating is not in that HTML: it arrives in a 5,981-byte fragment the page loads separately. Member profiles are the one surface that answers only a real browser, at 1,270,529 bytes and 32.9 seconds. Here is the whole measurement and the line Letterboxd draws around its API.
Sep 29, 2026 · 10 minRead → Social proxiesFlickr Proxies: 145 Words, No BrowserTwenty-one fetches of Flickr on 28 September 2026 from the United States, Germany and Japan. A public photostream returns 145 words to a plain HTTP client and the same 145 words to a browser, for 663,499 bytes against 2,005,256. A photo page carries the caption, the view count and the licence in plain HTML. The oEmbed endpoint answers with no key in 1,189 bytes. Here is the whole measurement, the 3,600-per-hour limit Flickr documents, and the line where its own guide draws it.
Sep 29, 2026 · 10 minRead → Social proxiesVero Proxies: 2,393 Words From 3 CountriesOne public Vero profile measured through residential exits in the United States, Germany and the United Kingdom on 28 September 2026. Over plain HTTP it returns 2,393 words in 265,121 bytes from all three countries, identical; rendered, 2,421 words for 331,650 bytes. Every response carries a meta robots tag of noindex, noimageindex, noarchive. There is no public API and no oEmbed endpoint. On Vero the setting is plain HTTP through any residential exit, and the honest line is that the platform publishes profiles to the open web while asking that they not be indexed or archived.
Sep 29, 2026 · 8 minRead → Social proxiesTruth Social Proxies: 1,440 Bytes, No LoginTruth Social measured through residential exits in the United States, Germany and the United Kingdom on 28 September 2026. The public profile page hands a plain HTTP client 7 words from every exit, and a browser hydrates it only from the US: 400 words for 937,369 bytes. The Mastodon-compatible API needs no account and returned the identical 1,440-byte JSON from all three countries, with a documented ceiling of 300 requests per five-minute window. On Truth Social the setting is the API, plain HTTP, any exit.
Sep 29, 2026 · 9 minRead → Social proxiesTriller Proxies: 0 Words, Domain for SaleTriller measured through residential exits in the United States, Germany and Brazil on 28 September 2026. The profile URL returns 114 bytes and 0 words over plain HTTP, a script that redirects to a listing where the domain is for sale; rendered, the sale notice arrives in the exit country's language, 92 words in English or Portuguese and 161 in German. The app's online services stopped in 2026. There are no Triller profiles to read through a proxy, and this page says so with the numbers, then shows where the archived pages are: the Wayback Machine collector returns 10 snapshots in 7.3 seconds.
Sep 29, 2026 · 9 minRead → Social proxiesPlurk Proxies: 118 Words Without a BrowserOne public Plurk profile measured through residential exits in Taiwan, Japan and the United States on 28 September 2026. Over plain HTTP it returns 118 words in 32,781 bytes from Taiwan and the US, with the profile header, the description and the counters in the HTML; rendering adds 78 words for 765,843 bytes. From Japan the same page is 103 words plain and 137 rendered, because the interface follows the exit. A single plurk page is server-rendered at 17,858 bytes. The API needs a signed OAuth request, even for public profiles. On Plurk the setting is plain HTTP, a Taiwanese exit, and an app key for timelines.
Sep 29, 2026 · 8 minRead → Social proxiesMinds Proxies: 738 Words From Germany, 38 USOne public Minds channel measured through residential exits in the United States, Germany and the United Kingdom on 28 September 2026. Over plain HTTP the page is a 37-word shell everywhere. Rendered, it hydrates from Germany (738 words) and the United Kingdom (764) but not from the United States (38 words, HTTP 200). The unauthenticated channel endpoint returns 3,485 bytes from all three countries, but robots.txt closes the API to crawlers except discovery. On Minds the setting is a European exit, rendered, and the discovery endpoint for search.
Sep 29, 2026 · 8 minRead → Social proxiesWeibo Proxies: 0 Words From 3 CountriesOne public Weibo profile measured through residential exits in the United States, Germany and Singapore on 28 September 2026, plus the mobile site, the hot-search page, a video page, a topic page and two ajax endpoints. Every web surface hands an anonymous client a 9,540 to 9,930-byte visitor page with 0 words, rendered or not; the ajax endpoints answer 21 bytes. What does return data: the documented Open API with an app key, and the search index, where one collector run returned 10 Weibo rows in 5.9 seconds including a 158-million follower count. Weibo is the target where the proxy is not the product.
Sep 28, 2026 · 8 minRead → Social proxiesVK Proxies: 1 Word From 3 CountriesThe official VK community page measured through proxy exits in Germany, the United States and Kazakhstan on 28 September 2026. Every exit is redirected from vk.com to vk.ru and receives a 179,861 to 188,502-byte page with 1 extractable word, tagged noindex; rendering it in a browser returns 1 to 8 words. The API, called without a token, answers in 220 bytes of JSON naming the missing key. VK is the target where the exit country changes nothing and the token changes everything.
Sep 28, 2026 · 9 minRead → Social proxiesTumblr Proxies: 4,505 Words, No BrowserOne public Tumblr blog measured through residential exits in the United States, Germany, Japan and the United Kingdom on 28 September 2026. From Germany and Japan the plain HTTP fetch returns 4,505 and 4,421 words in about 400 KB and the browser adds exactly zero. From the US the first plain request returns 2 words in 6,869 bytes and the second returns the full 386,280-byte page. The same blog on www.tumblr.com is 1.7 MB, and the exit country rewrites its title. Tumblr is the target where the retry, not the browser, is the setting.
Sep 28, 2026 · 10 minRead → Social proxiesRumble Proxies: 205 Words With No BrowserOne public Rumble channel measured through residential exits in the United States, Germany and Canada on 28 September 2026. Plain HTTP returns 205, 203 and 203 words with the h1, the follower count, the canonical and two JSON-LD blocks in 416,546 bytes. Rendering the same URL costs 1,115,659 bytes and returns 272 words from the US exit and 43 to 44 words from the other two. Rumble is the target where the browser buys nothing and, outside the US, costs you the page.
Sep 28, 2026 · 9 minRead → Social proxiesQuora Proxies: 2,254 Words Need the BrowserOne public Quora question measured through residential exits in the United States, Germany and India on 28 September 2026. Over plain HTTP, German and Indian exits return the head in 233 to 250 KB: title, canonical, the answer count and a 300-character opening of the top answer, 19 to 22 words. The full 2,254-word rendered page arrives only from a US exit, costs 1,246,131 bytes and 18.7 seconds, and carries a noindex tag. Quora is the target where the head and the body come from different countries.
Sep 28, 2026 · 9 minRead → Social proxiesThreads Proxies: 12.7M Followers in the HeadThirteen fetches of Threads on 28 September 2026 from the United States, Germany and Brazil. A public profile answers a plain HTTP client with 1,034,341 bytes, the follower and post counts in the meta description, and the recent posts in the HTML, no login and no JavaScript. Rendering it costs 2,028,277 bytes for an unstable word count. The exit country rewrites the number itself: 12.7M, 12,7 Mio., 12,7 mi. Here is the measurement, the arithmetic, what the Threads API needs, and the sentence in robots.txt that governs all of it.
Sep 28, 2026 · 12 minRead → Social proxiesTelegram Proxies: 1,139 Words, No JavaScriptNineteen fetches of Telegram on 28 September 2026 from the United States, Germany and India. The t.me/s/ web preview of a public channel returns 1,139 words and 127,691 bytes to a plain HTTP client, identical from all three countries, paginated by message id, with no account and no token. Rendering adds 40 words of chrome for 7.6 seconds. The Bot API wants a token for every method, and there is no robots.txt on t.me at all. Here is the measurement, the arithmetic, and the line Telegram’s own terms draw.
Sep 28, 2026 · 12 minRead → Social proxiesSnapchat Proxies: 757K Subscribers, No JSSixteen fetches of Snapchat on 28 September 2026 from the United States, Germany and India. A public profile returns its subscriber count, category, website and a Spotlight list with view counts to a plain HTTP client in 662,597 bytes, no login, no JavaScript. Rendering the same page costs 2,637,627 bytes for the same 94 words. The exit country rewrites the title and the counter label, and Snapchat tells you which country it saw in a response header. Here is the measurement, the arithmetic, and where the Snap Terms stop you.
Sep 28, 2026 · 11 minRead → Social proxiesPinterest Proxies: 1.3 MB Profile, 107 WordsFifteen fetches of Pinterest on 28 September 2026 from the United States, Germany and Brazil. A public profile answers a plain HTTP client with 1,304,978 bytes, an h1, a description, ProfilePage JSON-LD and the board grid. Rendering it costs 2,628,042 bytes and 33 seconds for 151 words, and from Brazil the render never hydrated at all. The exit country picks the domain: www, de. or br. Here is the measurement, the oEmbed endpoint that does exist, the robots allowlist of 330 named bots, and the one line in the Terms.
Sep 28, 2026 · 11 minRead → Social proxiesDiscord Proxies: 2,810 Bytes, No TokenTwenty fetches of Discord on 28 September 2026 from the United States, Germany and Brazil. The public invite page is a 9-word shell over plain HTTP and a 4,939,296-byte shell when rendered. The invite API answers the same question with no token in 2,810 bytes, identical from all three countries: member count, online count, verification level and boost tier. Here is the whole measurement, the per-address limit Discord documents, and where its Developer Policy draws the line.
Sep 28, 2026 · 12 minRead → ProxiesWeb Scraping API vs Proxy: 25 eBay Rows, 19 sWeb scraping API vs proxy is the wrong question on eight of the 53 sites we measured. On Amazon, eBay, Booking.com, Skyscanner, Airbnb, Expedia, Taobao and Spotify the page a proxy fetches holds between 0 and 72 words in either fetch mode. On 28 September 2026 the ebay_search collector returned 25 priced listings in 19.5 seconds and booking_stays returned 20 priced Rome properties in 5.1 seconds. The table, the cost per row and the setting for each site.
Sep 28, 2026 · 13 minRead → TroubleshootingRead the Block: 4 Shapes, 4 Settings That WorkA refused reply has a shape, and the shape tells you the setting. Across 53 sites measured between 26 and 28 September 2026 we met four: a short page with no canonical (Amazon, 2,007 bytes, 0 words), a redirect to the login form (Glassdoor, rendered: canonical /member/profile/login, 54 words), a rate-limit reply (Expedia, rendered: 381,468 bytes, 45 words, no h1) and an empty rendered page (Zillow, 759 words plain, 0 rendered). Each one resolves into a fetch mode, an exit and, where the page never exists, a collector with rows and seconds.
Sep 28, 2026 · 12 minRead → ProxiesHeadless Browser Scraping: 53 Targets MeasuredHeadless browser scraping pays on 15 of the 53 sites we measured between 26 and 28 September 2026, and costs words on 12. Every target was fetched through a residential exit as a plain HTTP client and again fully rendered; the table below has both word counts, the verdict and the one-line setting for all 53. Fifteen need the browser, four gain a little, fourteen gain nothing for up to 15 times the bytes, twelve lose words, and eight return no rows in either mode and want a collector.
Sep 28, 2026 · 17 minRead → ProxiesGeo-Targeted Proxies: 987 Words UK vs 149 DEThe same URL, rendered through residential exits: 987 words from the United Kingdom, 149 from Germany, 2,031 from the United States. Nothing was translated; each exit landed on a different licensed product. On 18 of the 53 sites we measured, geo-targeted proxies change the page itself: the domain on AliExpress, the entity on OKX, the price line on Netflix, the number format on Facebook and LinkedIn, and whether the channel page exists at all on YouTube. Pin the exit per job or you are comparing two different pages.
Sep 28, 2026 · 12 minRead → Proxies12 Sites Where Headless Browsers Lock You OutTwelve of the 53 sites we measured in September 2026 give a headless browser less than they give a plain HTTP client, and seven of them give it nothing. The same twelve URLs return 11,157 words over plain HTTP and 3,032 rendered. Steam, Etsy, Idealista, Reddit, Indeed, Zillow, Tripadvisor, DoorDash, Expedia, DraftKings, Vinted and Glassdoor, with the numbers, the cost of the browser you do not need, and the collector rows for the pages the fetch cannot reach.
Sep 28, 2026 · 10 minRead → ProxiesFollower Counts Are in the Head: 5 PlatformsFive platforms print the follower count before any script runs: Instagram and LinkedIn in the meta description, Facebook in og:description, YouTube and Twitch in JSON-LD. We measured eleven targets with and without a browser through residential exits. On these five the rendered pass adds bytes and no number. TikTok and Spotify are the exceptions, and the tiktok_profile collector returns 2 profiles in 14.4 seconds.
Sep 28, 2026 · 10 minRead → Social proxiesTrustpilot Proxies: 2,553 Words, Browser OnTrustpilot measured through residential exits on 26 and 28 September 2026: the home page hands a plain HTTP client 19 words in 991 bytes and a rendered browser 2,553 words, then 3,588 on a second day. A company review page renders to 4,389 words in 1,631,458 bytes, with a status code that does not describe the body. The reviews collector returned 10 reviews with ratings, dates, verification and the company reply in 14.8 seconds.
Sep 28, 2026 · 11 minRead → Social proxiesReddit Proxies: 574 Words Without a BrowserReddit measured through residential exits on 28 September 2026: a subreddit listing answers a plain HTTP client with 574 words and 566,030 bytes, and the same URL rendered in a browser with 39 words and no canonical. The rate limit is printed in the response headers, 200 requests per window. Comment threads need the collector: 20 comments in 7.9 seconds.
Sep 28, 2026 · 10 minRead → Social proxiesLinkedIn Proxies: 1,993 Words, No BrowserA LinkedIn company page measured through residential exits on 28 September 2026: 1,993 words and 373,575 bytes over plain HTTP, 1,920 words and 825,521 bytes rendered. The follower count sits in the meta description, the Organization schema in the head, and a German exit rewrites the number as 7.122.892 Follower:innen. The collector returns the company row in 4.3 seconds.
Sep 28, 2026 · 11 minRead → Social proxiesIndeed Proxies: 1,026 Words, First RequestIndeed measured through residential exits on 26 and 28 September 2026. A plain HTTP fetch with a browser TLS fingerprint returned 1,026 words on the first request, then 601,347 bytes and 2,394 words in 9.4 seconds; the rendered page cost 1,066,178 bytes and 52 seconds for 683 words. From a UK exit the same URL becomes uk.indeed.com. The jobs collector returned 10 listings with parsed salaries in 7.3 seconds.
Sep 28, 2026 · 10 minRead → Social proxiesGlassdoor Proxies: 624 Words, Skip the BrowserGlassdoor measured through residential exits on 26 and 28 September 2026: the home page hands a plain HTTP client 624 words from the United States and 589 from the United Kingdom, with the h1 and the Organization schema; rendered in a browser it lands on the login page, canonical /member/profile/login, for 705,071 bytes and 54 words. Glassdoor is the target where the browser is the thing that loses you the data.
Sep 28, 2026 · 11 minRead → Use casesTripadvisor Proxies: 713 Words Over Plain HTTPTripadvisor measured on 28 September 2026 through residential exits in the United States and Italy. The home page hands a plain HTTP client 713 words, the h1 "Where to?", a canonical and three kinds of JSON-LD, including 19 LocalBusiness objects with URLs and countries, for 388,397 bytes in 3.6 seconds. From Italy the same URL returns the same English page with 757 words. The browser spends 46 seconds and returns 5,915 bytes. The setting is plain HTTP; the tripadvisor_search collector returned 30 Chicago restaurants in 8.1 seconds.
Sep 28, 2026 · 10 minRead → Use casesSkyscanner Proxies: 0 Words, 15 Fares in 34 sSkyscanner measured on 28 September 2026 through residential exits in the United Kingdom and the United States. The home page returns a title, a description, a canonical and a WebSite JSON-LD block, and zero words of body text, to a plain HTTP client and to a full browser alike, in both countries. No fetch mode changes that. The google_flights collector, asked for London Heathrow to New York JFK on 20 October in pounds, returned 15 itineraries in 33.7 seconds.
Sep 28, 2026 · 9 minRead → Use casesKayak Proxies: 2,451 Words With No BrowserKayak measured on 28 September 2026 through residential exits in the United States and the United Kingdom. The home page hands a plain HTTP client 2,451 words, an h1, a canonical and three JSON-LD blocks including a four-question FAQPage, for 1,656,272 bytes. A full browser render returns 2,455 words for 2,415,047 bytes and 40.7 seconds. Every tutorial on the SERP opens Selenium first; the measurement says close it.
Sep 28, 2026 · 8 minRead → Use casesExpedia Proxies: 72 Words Without a BrowserExpedia measured on 28 September 2026 through residential exits in the United States and the United Kingdom. From the United States a plain HTTP client gets the home page: 72 words, the h1 "The one place you go to go places", a canonical and a description for 507,949 bytes in 4.1 seconds. The browser gets 14 words. From the United Kingdom the plain fetch gets 14 words too. Structure over plain HTTP from a US exit; prices from the hotels and google_flights collectors.
Sep 28, 2026 · 9 minRead → Use casesBooking.com Proxies: 26 Words, 20 Rows in 6 sBooking.com measured on 28 September 2026 through residential exits in Italy and the United States. The home page answers a plain HTTP client with 3,962 bytes and 26 words, and a full browser render with 857,992 bytes and zero words of content. The collector, run for Rome on 20 to 22 October, returned 20 properties with nightly price, struck-through price, review score and distance in 5.9 seconds. That is the setting.
Sep 28, 2026 · 9 minRead → Use casesAirbnb Proxies: 61 Words, 18 Stays in 5.3 sAirbnb measured on 28 September 2026 through residential exits in the United States and Germany. The home page returns 61 words to a plain HTTP client and the same 61 words to a full browser, for 597,053 and 1,561,460 bytes respectively. From Germany the plain fetch is a 3-word redirect notice. The listings live in the page's hydration island, which is what the airbnb_stays collector reads: 18 priced stays for Austin in 5.3 seconds.
Sep 28, 2026 · 9 minRead → Social proxiesYouTube Proxies: 15 Words Carry the Channel IDYouTube measured through residential exits on 28 September 2026. A channel page answers a plain HTTP client with 15 words and 1,746,955 bytes, but the head carries the canonical channel ID, the itemprop identifier and a ProfilePage block with 15,100,000 subscribers. Rendering costs 3,544,622 bytes and 33.7 seconds for 583 words. From Germany the same URL is a consent page. The collector returns 30 videos in 3.84 seconds.
Sep 28, 2026 · 7 minRead → Use casesTwitch Proxies: 2,461,147 Followers in 214 KBTwitch measured through residential exits on 28 September 2026. The home page answers a plain HTTP client with 200 and zero words, and a browser needs 1,718,470 bytes and 45.7 seconds for 101. A channel page over plain HTTP carries the follower count, 2,461,147, and ten videos with view counts in its JSON-LD. The Helix API answers 401 in 72 bytes without a token, and the terms say what a proxy is for.
Sep 28, 2026 · 8 minRead → Use casesSteam Proxies: 2,792 Words Without a BrowserThe Steam store measured through residential exits on 28 September 2026. Plain HTTP returns 2,792 words of storefront; a headless browser returns 1,394 and an error heading. Prices live in a 172-byte JSON endpoint that answers $59.99 from a US exit and 59,99 EUR from a German one. The collector returns 10 apps in 1.83 seconds.
Sep 28, 2026 · 8 minRead → Use casesSpotify Proxies: 0 Words, 164 KB, Use the APISpotify measured through residential exits on 28 September 2026. open.spotify.com answers a plain HTTP client with 200, 164,388 bytes and zero words, and a headless browser with zero words after 2,143,243 bytes. An artist page has no description and no structured data. The Web API answers 401 in 94 bytes without a token. What a proxy still reads: the public Premium price, $12.99 from the US and 12,99 € from Germany.
Sep 28, 2026 · 8 minRead → Use casesRoblox Proxies: 0 Words, 1,802-Byte Games APIRoblox measured through residential exits on 28 September 2026. The home page answers a plain HTTP client with 200, 101,528 bytes and zero words; a headless browser spends 4,440,448 bytes and 19.9 seconds for 128. The games endpoint answers 1,802 bytes with live players, visits and favourites, and every response prints its rate limit in a header.
Sep 28, 2026 · 8 minRead → Use casesNetflix Proxies: 3 Prices, 3 Countries, No JSNetflix measured through residential exits on 28 September 2026. The public home answers a plain HTTP client with 466 words and the plan prices; rendering changes the h1 and nothing else. The same URL says $8.99 to $26.99 from the United States, £7.99 to £20.99 from the United Kingdom and 6,99 € to 21,99 € from Germany. A title page is 711 words in English and 680 in German. That is what a proxy verifies on Netflix, and the terms say what it does not.
Sep 28, 2026 · 8 minRead → Use casesTicketmaster Proxies: 494 Words, No BrowserTicketmaster measured through residential exits on 28 September 2026. The home page answers a plain HTTP client with 494 words and a canonical; the concerts category page adds twenty MusicEvent objects in JSON-LD, with venue, date and availability, for 681,189 bytes and no browser. A UK exit gets the identical page. And the BOTS Act covers circumventing purchase limits, not reading public pages.
Sep 28, 2026 · 9 minRead → Use casesStubHub Proxies: 0 Words Without a BrowserStubHub measured through residential exits on 28 September 2026. The home page hands a plain HTTP client 226,417 bytes and zero words from the US and from the UK. Rendering it costs 3,238,143 bytes and 35 seconds from a US exit, and returns 398 words instead of 4,721 from a UK one. Ticketmaster, measured the same day, serves 494 words and twenty structured events with no browser.
Sep 28, 2026 · 8 minRead → Use casesPop Mart Proxies: 841 Words Over Plain HTTPPop Mart measured through residential exits on 28 September 2026. The US home page answers a plain HTTP client with 841 words for 765,161 bytes; rendering it returns 867 words for 1,972,361 bytes and 21 seconds. A product page is 117,555 bytes, carries a canonical and is served with no-store, which is exactly what a drop monitor wants. The UK storefront at /gb is a different application with 6,904 words.
Sep 28, 2026 · 7 minRead → Use casesPokemon Center Proxies: 1,622 Words Over HTTPPokemon Center measured through residential exits on 28 September 2026. Residential exit, plain HTTP fetch, 1,622 words for 674,534 bytes. Rendering the same URL costs 1,231,142 bytes and 62 seconds and returns 765 words. A UK exit gets the same page plus a header offering the UK store. robots.txt names the price and availability endpoints you should not poll, and the ebay_search collector returns 20 resale listings in 46 seconds.
Sep 28, 2026 · 9 minRead → Use casesDraftKings Proxies: 647 Words, Never RenderDraftKings measured through residential exits on 28 September 2026. The sportsbook home page serves a plain HTTP client 647 words; rendering it costs 2,877,467 bytes and returns 642. The NFL odds page answers plain HTTP with 5,418 words of lines. A Canadian exit gets a different book: 697 words on the home page, 5,611 on NFL. And the terms allow one account per person, which no proxy changes.
Sep 28, 2026 · 8 minRead → Use casesTaobao Proxies: 36 Words at Home, 410 Deeperworld.taobao.com hands a plain HTTP client 36 words and a headless browser 56, from Singapore and from the United States alike; a full render downloads 1,032,793 bytes of category menu with no listing and no price. One level down, a /lang/en-us/ page returns 410 words and four prices for 49,272 bytes without JavaScript. The home is not the page. Here is which pages are, and what they cost.
Sep 28, 2026 · 9 minRead → Use casesShopee Proxies: 1,427 Words, Pin SingaporeShopee.sg answers a plain HTTP client from Singapore with 1,427 words and answers a headless browser with the same 1,427 words, for 1.74 times the bytes and 55 seconds of waiting. From an Indonesian exit the same URL returns a 163,632-byte shell with no title and no words. The whole setting is one line: residential, plain HTTP, Singapore.
Sep 28, 2026 · 9 minRead → Use casesRakuten Proxies: 882 Words From Any CountryRakuten Ichiba serves a plain HTTP client 882 words, a canonical and Corporation plus WebSite JSON-LD in 381,287 bytes, and it serves exactly the same 882 words to a US exit. Rendering adds 268 words of rotating modules for 4.2 times the bytes. The catalogue itself has an official API. Here is what to fetch, from where, and what it costs.
Sep 28, 2026 · 8 minRead → Use casesLazada Proxies: 1,196 Words Without a BrowserLazada Singapore hands a plain HTTP client 1,196 words, a canonical and an h1 in 469,188 bytes. A headless browser gets 1,197 words for 1,841,729 bytes and 48 seconds. A Malaysian exit gets the same Singapore page, because on Lazada the domain picks the market and the exit does not. The setting is plain HTTP on the residential Basic line, and here is the arithmetic.
Sep 28, 2026 · 8 minRead → Use casesAliExpress Proxies: 519 Words, 2 DomainsSend the same aliexpress.com URL through a US residential exit and you land on aliexpress.us with 519 words; send it through a German exit and you land on de.aliexpress.com with 427 words in German. Rendering lifts those to 2,502 and 2,085. The exit does not localise AliExpress, it picks the domain. And for listings, the collector returned 20 items with price, original price and discount in 4.47 seconds.
Sep 28, 2026 · 8 minRead → Use casesZalando Proxies: 266 Words, One False 200Zalando.it measured through residential exits on 28 September 2026. The home answers a plain HTTP client with 200 OK, 570,402 bytes and 266 words; rendered, 1,607,265 bytes and 268 words. The number that matters is the one status code that lies: Zalando can answer 200 with an Akamai interstitial instead of the catalogue, and the collector retried it from a fresh exit and delivered 10 products with EU Omnibus reference prices in 3.6 seconds.
Sep 28, 2026 · 8 minRead → Use casesVinted Proxies: 227 Words in 2 MB of HTMLVinted measured through residential exits on 28 September 2026. The Italian home answers a plain HTTP client with 200 OK, 1,997,674 bytes and 227 words; the rendered page takes 36 seconds and returns 204 words. The country lives in the domain, not in the exit: from Germany, vinted.it is still Italian. And robots.txt says, in its own words, what bots may not do.
Sep 28, 2026 · 9 minRead → Use casesSubito Proxies: 600 Words for 183 KBSubito.it measured through residential exits on 28 September 2026. The home answers a plain HTTP client with 200 OK, 182,983 bytes and 600 words; rendering it costs 1,263,991 bytes and returns the same 600. The subito_search collector returns 10 ads in 2.7 seconds with price, comune, seller type and shipping cost, out of the 9,268 Subito reports for the query.
Sep 28, 2026 · 8 minRead → Use casesKleinanzeigen Proxies: 506 Words, No Geo GateKleinanzeigen.de measured through residential exits on 28 September 2026. The home answers a plain HTTP client with 200 OK, 288,478 bytes and 506 words from Germany, 503 from the United Kingdom; rendered, 1,305,776 bytes and 530 words. Its robots.txt closes pages 6 onward of every listing to crawlers, and the kleinanzeigen_search collector returns 10 ads with price, VB flag and ZIP in 3.5 seconds.
Sep 28, 2026 · 8 minRead → Use casesIdealista Proxies: 1,180 Words, No BrowserIdealista.it measured through residential exits on 28 September 2026. Plain HTTP with a browser TLS profile returns 200 OK, 101,680 bytes and 1,180 words, the full head and the h1, from Italy and from Spain alike; the rendered pass returns zero words after 38.5 seconds, so the browser is the one cost to skip. The idealista_search collector returns 10 Milan listings with price, size and agency in 5.5 seconds.
Sep 28, 2026 · 8 minRead → Use casesAutotrader Proxies: 1,115 Words, Fenced SearchAutotrader.co.uk measured through residential exits on 28 September 2026. The home answers a plain HTTP client with 200 OK, 366,925 bytes and 1,115 words, from the UK and from the US alike; rendered, 835,543 bytes and 1,133 words. Its robots.txt, updated 24 September 2026, closes /car-search, /car-details and the GraphQL gateway to every agent. The autotrader_search collector covers the US site: 10 Honda Civic listings with VIN, mileage and days on site in 15.3 seconds.
Sep 28, 2026 · 7 minRead → Use casesWalmart Proxies: 74% of the Page Without JSwalmart.com measured through residential exits in September 2026. Three quarters of the home page, 2,190 of 2,975 words, arrives before any script runs, with canonical, JSON-LD and an INDEX,FOLLOW robots tag. The remaining quarter costs 7.5 times the bytes and 24 times the seconds. The exit country changes the home you are handed, and the terms, updated on 9 September 2026, changed too.
Sep 28, 2026 · 8 minRead → Use casesTarget Proxies: 2,163 Words Need the Browsertarget.com measured through residential exits in September 2026. It is the retailer where the browser pays: 317 words without JavaScript, 2,163 rendered, 6.8 times more, the opposite of Best Buy. The exit country changes nothing in the page, a British exit and an American one get the same 256 words, but the server-side page state assigns every visitor a store and a ZIP from the IP alone, and that is what a Target proxy actually selects.
Sep 28, 2026 · 8 minRead → Use casesEtsy Proxies: 678 Words Without a Browseretsy.com measured through residential exits on 28 September 2026. A plain HTTP client from the United States receives the full home: 200, 205,911 bytes, 678 words, title, canonical and JSON-LD. The same client from Germany or the United Kingdom receives 8 words, and a headless browser from the United States receives none. The setting is three words long: residential, plain HTTP, US.
Sep 28, 2026 · 8 minRead → Use caseseBay Proxies: 25 Listings in 185 Secondsebay.com measured through residential exits on 28 September 2026. Over plain HTTP the home page returns the same 1,976-byte, 24-word page from the United States and from Germany, and marks it noindex. The ebay_search collector returned 25 road-bike listings with price, condition, format, location and shipping in 185 seconds.
Sep 28, 2026 · 9 minRead → Use casesBest Buy Proxies: 285 Words, Never Renderbestbuy.com measured through residential exits in September 2026. The home gives a plain HTTP client 285 words; a headless browser gets 284. On 28 September the plain fetch weighed 506,484 bytes in 5.1 seconds and the rendered one 4,968,555 bytes in 45.9 seconds: 9.8 times the bytes and 9 times the wait for one word fewer. From Canada the same URL is a 103-word country chooser. The setting: residential, plain HTTP, US, and never the browser.
Sep 28, 2026 · 9 minRead → Use casesAmazon Proxies: 20 Products in 43 SecondsAmazon.com measured through residential exits on 28 September 2026. A plain HTTP fetch of the home page returns 2,007 bytes and not one word of product content; a rendered fetch from the United States returns none either, and the same URL rendered from Germany returns 546 words of a German-language home. The number that matters is the collector: 20 products with ASIN, price, rating and review count in 43 seconds, for two cents.
Sep 28, 2026 · 9 minRead → Use casesZillow Proxies: 747 Words Over Plain HTTPZillow measured on 28 September 2026 through residential exits in the United States and Germany. The homepage answers a plain HTTP client with 200 OK and 747 words from both countries; a listing page answers with 755,913 bytes and its price, beds, baths and square feet in the meta description, no JavaScript run. The browser is the one thing not to buy here. For rows, the Zillow search collector returned 20 listings for Austin, TX in 3.02 seconds and the Redfin collector 50 of 240 for one ZIP in 7.01.
Sep 28, 2026 · 9 minRead → Use casesUber Eats Proxies: 99 Words Until You RenderUber Eats measured on 28 September 2026 through residential exits in the United States and the United Kingdom. The homepage answers a plain HTTP client with 200 OK and 99 words from the US, and redirects a UK exit to /gb with 42 words. Rendered in a browser it returns 206 words for 1,951,573 bytes. Nothing on the storefront exists until JavaScript runs and a delivery address is set, and Uber’s terms reserve commercial use of the data to parties with written permission.
Sep 28, 2026 · 8 minRead → Use casesPayPal Proxies: 1 Static IP for PayflowPayPal proxies, the developer version. Payflow Pro lets a merchant allowlist up to 16 Internet-facing server IPs in PayPal Manager and answers everything else with RESULT=1, User authentication failed. PayPal’s own help says REST and Classic APIs need no allowlisting and recommends DNS over hard-coded ranges. For an app on cloud infrastructure whose egress changes, a dedicated static ISP IP is one address to register once. Measured on 28 September 2026: paypal.com/us/home is 1,693 words over plain HTTP from the US and identical from Germany, so nothing here needs a browser either.
Sep 28, 2026 · 9 minRead → Use casesDoorDash Proxies: 485 Words, No Browser NeededDoorDash measured on 28 September 2026 through residential exits in the United States and Canada. The homepage answers a plain HTTP client with 200 OK and 485 words; a store page answers with 2.95 MB and 1,087 words of menu and prices, no JavaScript run. The browser is the one thing you should not buy here. When you want rows instead of pages, the DoorDash restaurants collector returned 50 restaurants for Austin, TX in 5.98 seconds.
Sep 28, 2026 · 9 minRead → Use casesDeliveroo Proxies: The Render Pays 11.5xDeliveroo measured on 28 September 2026 through residential exits in the UK, the US, France and Italy. The homepage answers a plain HTTP client with 200 OK and 195 words from every country; rendering it in a browser returns 2,247 words for 3,903,291 bytes. The Soho restaurant listing ships 3.38 MB of server-rendered HTML with no JavaScript at all, and a header that tells you the rate limit.
Sep 28, 2026 · 9 minRead → Use casesShopify Proxies: 250 Products Per 1.76 MB CallShopify storefronts measured through residential exits on 28 September 2026. A real store's /products.json?limit=250 answered a plain HTTP client with 200 OK, 1,759,991 bytes, 250 products and 2,720 variants with price and availability in 2.2 seconds, no key and no browser, and returned the same bytes from Germany. shopify.com itself is 749 words without JavaScript and 27 more rendered. The Admin API is limited per app per store, so a proxy changes nothing there.
Sep 28, 2026 · 10 minRead → Use casesOKX Proxies: Same URL, 623 Words DE, 371 USOKX measured through residential exits in Germany, Singapore and the United States on 28 September 2026. The same URL answered a plain HTTP client with 623, 464 and 371 words under three different titles, three h1s and three canonicals, chosen from the exit IP alone. Rendering the German page cost 14.9 times the bytes for 44 more words. The public ticker is 364 bytes, answers every exit, and rate-limits per IP; the API key allowlist takes up to 20 addresses and wants them fixed.
Sep 28, 2026 · 9 minRead → Use casesBinance Proxies: 558 Bytes Per Price TickBinance measured through residential exits in Germany and the United States on 28 September 2026. The homepage hands a plain HTTP client 2,108 bytes and 26 words; rendering it costs 444,189 bytes and 56 seconds. A 24-hour ticker from the public market-data host is 558 bytes, needs no key, and answered both countries identically. The one place a proxy choice matters is the API key allowlist, and there the answer is a static IP, not a pool.
Sep 28, 2026 · 10 minRead → Use casesbet365 Proxies: 987 Words UK, 149 Germanybet365 answers a plain HTTP client with 200 OK, 45,815 bytes and zero words. Rendered, the same URL returns 987 words from a UK exit, 149 from Germany and 2,031 from the United States, because each exit lands on a different licensed product. One setting reads it: residential exit, rendered fetch, country pinned to the market you are monitoring.
Sep 28, 2026 · 10 minRead → Use casesJev on 730 Google Reviews: What 1-Stars SayJev read 730 Google Maps reviews in English and Italian and agreed with the star rating 99% of the time in both languages. Three of its disagreements were reviewers who picked the wrong stars. Then it sorted 304 one- and two-star reviews by what they blame: American dentists get punished over money, Italian dentists over the treatment.
Sep 28, 2026 · 7 minRead → AI scrapingJev vs GPT, Claude, Gemini on Real Web DataSame 317 hand-labelled cases, same question, six models. Accuracy was a tie: 306 to 309 correct. Jev was 2 to 42 times cheaper and answered in a third of a second. The difference that matters is where the mistakes were: all 11 of Jev's errors came with an uncertain score, while the LLMs made theirs sounding sure.
Sep 28, 2026 · 8 minRead → Lead generationJev Lead Scoring on 183 Real BusinessesA lead list scraped from a directory and Google Maps is never clean: in our sample of 183 businesses, 63 were not the buyer's target at all. We asked Jev to sort them with one sentence and no tuning. It got 175 right, one fewer than the keyword rules we had tuned by hand on the same data, for under half a cent.
Sep 28, 2026 · 7 minRead → Anti-botCan Jev Tell a Block Page From Content?We fetched 70 heavily protected sites twice each, labelled 134 responses by hand, and asked Jev one question: is this the page that was requested? It got 131 right; the HTTP status code got 111. Every page Jev scored above 0.8 or below 0.5 was correct, and the six in between are exactly the ones a human would want to look at.
Sep 28, 2026 · 8 minRead → AI scrapingJev for Web Scraping: Will Your Page Fit?Jev, the System One model from TypeSafe AI, reads at most 32,000 tokens of state and bills every input token. We fetched 20 real pages on 28 September 2026: raw HTML fit on 4 of them, with a median of 152,277 tokens. The same pages as Markdown fit on all 20, median 4,512, and 2,060 with links stripped.
Sep 28, 2026 · 10 minRead → Social proxiesKick Proxies: What Kick.com Serves a BotKick.com measured through residential exits on 28 September 2026. The channel JSON behind every page is 8,210 bytes, edge-cached for 15 seconds and answered 72 of 72 requests in a burst. The HTML page is 670,936 bytes with 63 words, the rendered page 2.76 MB. Before any of that: Kick's terms forbid scraping, and it now ships an official API.
Sep 28, 2026 · 10 minRead → Social proxiesFacebook Proxies: What You Can ReadSixteen fetches of Facebook on 25 September 2026. A public Page answers a plain HTTP client with 200 OK, 923,794 bytes and not one word of content. Rendering the same URL costs 2,474,309 bytes and returns 350 words. Marketplace is the one surface that repays the render: 15 listings with prices, from a logged-out residential exit, for a fifth of a cent.
Sep 25, 2026 · 14 minRead → Social proxiesLemmy Proxies: How to Scrape LemmyEleven fetches of Lemmy on 24 September 2026: the community page of lemmy.world refuses a plain HTTP client and costs 690,458 bytes to render, while the same instance hands its public API 10 posts in 37,702 bytes and 0.77 seconds. Ask two instances about one post and the vote counts disagree.
Sep 24, 2026 · 13 minRead → Social proxiesTwitter Proxies: What X Returns to a BotSeven fetches of x.com through residential exits on 23 September 2026: a profile returns five posts and 254 KB over plain HTTP, rendering it costs 18 times the bandwidth for the same five posts, and every failure arrives as HTTP 200.
Sep 23, 2026 · 12 minRead → Social proxiesMastodon Proxies: What 11 Servers ReturnEleven Mastodon servers, one anonymous request each: seven open, four closed, and a single cached response replayed on four continents. What that means for proxy choice, rate limits and cost.
Sep 21, 2026 · 13 minRead more TroubleshootingIPv6 Proxy Not Working? The Five ChecksFive checks that settle almost every "IPv6 proxy not working" ticket: whether the target publishes an AAAA record at all (57 of 102 major sites do not), DNS resolved at the exit with socks5h, credentials the client can actually pass, the error you really got, and pacing by /64 when the site answers but blocks.
Sep 20, 2026 · 9 minRead more Data for AIIPv6 Proxies for AI Training Data and AgentsA DNS probe of the sources AI pipelines pull from and of the model API hosts agents call: which ones an IPv6 proxy reaches, which are IPv4-only, the routing rule that sends each fetch to the cheapest network that works, and the cost of a corpus at $0.20/GB versus residential.
Sep 20, 2026 · 8 minRead more SEO dataIPv6 Proxies for SEO: Engines That Accept ThemWhich search engines and SEO tools publish IPv6 addresses (DNS census of 102 domains), what a Google result page weighs when fetched through our own SERP layer (48.5–59.2 KB), the cost per thousand SERPs at each pack, and how to pace by /64 so the savings survive.
Sep 20, 2026 · 10 minRead more Social proxiesTikTok Proxies: Measured From Four CountriesOne public TikTok profile, four exit countries, measured on 20 September 2026. Three exits returned around 400 kilobytes containing zero readable words, one returned a 1.4-kilobyte firewall interstitial under an HTTP 200 status, and the exit country decided which of TikTok’s regional backends answered at all.
Sep 20, 2026 · 10 minRead more Social proxiesBluesky Proxies: What the API ReturnsWe read Bluesky through five proxy exits on 18 September 2026. The public API answered without authentication and returned identical bytes from every country, while the web app served a 47-word shell to any client without JavaScript. The one thing the exit country changed was which national moderation labeler the response names.
Sep 18, 2026 · 11 minRead more Social proxiesInstagram Proxies: What the Site ReturnsInstagram publishes a profile's follower, following and post counts in its meta tags and almost nothing in the DOM. Eight measured passes across four surfaces and four exit countries, the two error shapes that lie to you, and the cost math per profile.
Sep 17, 2026 · 12 minRead more ProxiesIPv6 Proxies in 2026: Which Sites Accept Them, Tested102 major scraping, SEO and social domains checked for IPv6 on 14 September 2026, from three resolvers: 45 accept IPv6 on the host the homepage lands on, 57 are IPv4-only. Google, YouTube, Bing and Yahoo yes; TikTok, X, Reddit, eBay, GitHub, Booking and DuckDuckGo no. Plus why QuanticData ranks #1 on IPv6: $0.20/GB down to $0.083/GB at 3 TB, /48 and /64 subnets, rotating or static, no KYC. Raw CSV attached.
Sep 14, 2026 · 17 minRead more ProxiesGolang HTTP Client Proxy: Auth, SOCKS5How to attach a proxy to an http.Client in Go, authenticate over CONNECT, use SOCKS5 from the standard library and rotate exits per request, with a measurement of how far a no-JavaScript client actually gets.
Sep 14, 2026 · 11 minRead more ProxiesBest Residential Proxies in 2026, TestedQuanticData against the major residential providers, priced from their own pages on 8 September 2026: the lowest entry rate, $1.00/GB, down to $0.80/GB at 1,000 GB with no subscription. Plus 170 requests through our own residential pool with the raw rows attached: 50/50 country match, 41 ASNs, 1.26 s median TTFB, 20 of 20 sticky sessions holding one IP and a median fraud score of 2.
Sep 13, 2026 · 25 minRead more ProxiesBest Mobile Proxies in 2026, TestedQuanticData against the major mobile providers, priced from their own pages on 13 September 2026: three of the biggest publish no mobile price at all, and QuanticData publishes the whole ladder from $5 for 1 GB to $2.30/GB. Plus 98 requests through our own 4G/5G network with the raw rows attached: 40/40 success, 100% country match, 23 carrier ASNs, 1.74 s median TTFB.
Sep 13, 2026 · 19 minRead more ProxiesShadowrocket Proxy Setup: The iPhone GuideShadowrocket is a routing engine, not a VPN subscription. The add-server screen takes two minutes; the rules decide whether it is worth the money. Setup, config syntax, and measured exits.
Sep 12, 2026 · 9 minRead more ProxiesPlaywright Proxy: Setup, Auth and BandwidthPlaywright takes a proxy as an object, not a URL, and a browser pays for the whole page rather than the HTML. Per-context exits, the auth rules, the net errors, and measured bytes for seven real pages.
Sep 11, 2026 · 8 minRead more Proxieshttpx Proxy: Setup, Auth and the proxies= FixThe httpx proxy argument changed name, and most tutorials still teach the removed one. A version-by-version matrix, the exact exception text, and what the proxy sees on the wire.
Sep 10, 2026 · 8 minRead more Proxiescurl Proxy: Setup, Auth, SOCKS5 and ErrorsOne flag sets a proxy in curl. The rest of the work is authentication, SOCKS5 versus SOCKS5h, HTTPS tunnelling, and reading the exit code when it fails.
Sep 9, 2026 · 10 minRead more Troubleshootingrequests.exceptions.ProxyError: How to FixProxyError says the connection to the proxy failed, never that the site blocked you. We ran ten isolated failure modes against a controlled proxy on three requests/urllib3 stacks and mapped each cause to the exception class and the exact text it prints — including two causes that stopped being ProxyError in urllib3 2.x.
Sep 8, 2026 · 9 minRead more TroubleshootingERR_TUNNEL_CONNECTION_FAILED: How to FixChromium error -111 is raised when a CONNECT request to your proxy does not come back as a usable tunnel. Here is how to tell which of the seven causes you have, in the browser and in curl, Playwright, Puppeteer and Selenium.
Sep 7, 2026 · 10 minRead more Lead generationHow to Find B2B Clients From Public DataBusinesses publish their category, location, phone and website on maps, and their email addresses on their own sites. Collecting that into a targeted prospect list is the cheap part: our measurement puts a usable business record at $0.0143. The workflow, the yield you should expect at each step from 458 measured records, and the part most guides skip: which lawful basis lets you actually email the list in Italy, Germany, Spain, France and the UK, with the regulator page for each.
Sep 4, 2026 · 9 minRead more Lead generationHow Much Does a Lead Cost? Price by SourceThe word "lead" covers a scraped business record and a booked sales meeting, and they differ in price by four orders of magnitude. Published per-record prices from named list vendors, agency retainer and pay-per-lead rates, paid-search cost per lead, blended cost-per-lead benchmarks by industry, and our own first-hand measurement of $0.0101 to $0.0143 per usable business record from public data. What each number does and does not include.
Sep 4, 2026 · 8 minRead more Lead generationGoogle Maps Leads: What 458 Rows ContainEvery Google Maps scraper page lists the fields you get. None of them publish how often those fields are actually filled. We collected 458 dentist listings across Milan, Berlin, Madrid, Paris and London on one afternoon and counted: phone 99.3%, website 96.5%, rating 99.6%, review count 12.7%, an email on the site pass 60.3%. Cost per usable lead, the Paris booking-platform problem, the method, the limitations and the published dataset.
Sep 4, 2026 · 8 minRead more Lead generationIs Buying Leads Worth It? Math by Lead Type“Buying leads” covers five different products, from real-time shared leads to a spreadsheet of contacts, and they are worth it at very different prices. The break-even formula that turns your margin and close rate into a maximum price per lead, benchmark costs by industry, why bought leads underperform and how to test a vendor with a small sample, where the law actually bites, and the public-data route that produces a contactable business record for cents.
Sep 3, 2026 · 9 minRead more Web scrapingWeb Scraping vs API: Which One Should You Use?Web scraping versus API is a false binary: there are three options, and most real pipelines use two of them. When an official API is the right answer and the four ways it stops being one; what scraping costs in maintenance and risk; where scraping APIs and ready-made collectors fit; a seven-question decision tree, a worked cost example, and the hybrid pattern that holds up.
Sep 3, 2026 · 8 minRead more ProxiesResidential vs Datacenter Proxy: Which to BuyResidential and datacenter proxies differ in one fact the target can look up in a millisecond: the network the IP belongs to. Everything else, speed, price, block rate, follows from that. What the label means, a table that compares them honestly, a 100-request test that tells you which one your target requires, the cost-per-successful-page math, and when ISP or mobile is the right third answer.
Sep 3, 2026 · 8 minRead more ProxiesHow Much Proxy Data Do I Need? Real NumbersProxy plans are sold by the gigabyte and nobody tells you how many pages that is. We fetched 11 major pages and recorded what they cost on the wire: 47 KB to 302 KB compressed for the HTML, against 2.5 to 2.9 MB for a full browser load. The formula that turns pages per month into GB, three worked examples, the break-even where per-page pricing beats per-GB, and the settings that cut usage by five to twenty times.
Sep 3, 2026 · 8 minRead more TroubleshootingProxy Not Working? A 10-Step Checklist“Proxy not working” is four different problems wearing one name: you cannot reach the proxy, the proxy refuses you, the proxy reaches the site but the site refuses it, or it works and is slow. Each layer has its own error strings and its own fix. A table that maps the message to the layer, curl timings that separate slow from broken, and a 10-step checklist in the order that saves the most time.
Sep 3, 2026 · 8 minRead more TroubleshootingHow to Fix 429 Too Many Requests When ScrapingA 429 is the one block that tells the truth: you crossed a rate limit. What it does not tell you is what the limit is keyed on, and that decides whether proxies help at all. A two-request test to find the key, backoff code that honours Retry-After, the concurrency formula that turns pages per hour into IPs in flight, and a per-host token bucket.
Sep 3, 2026 · 7 minRead more TroubleshootingWeb Scraping 403 Forbidden: Causes and FixesThe browser loads the page and your script gets 403. The response headers usually name the system that refused you, and that decides the fix: a user agent, the full header set, the IP type, or the TLS fingerprint. A decision tree you can run in five minutes, Python fixes for each branch, and what we saw fetching 16 major sites with a pure HTTP client behind US residential IPs.
Sep 3, 2026 · 7 minRead more TroubleshootingHow to Fix 407 Proxy Authentication RequiredA 407 is the proxy refusing to forward your request until it sees credentials it accepts, so rotating headers or slowing down cannot fix it. On HTTPS it does not even arrive as a status code. How to read the challenge, tell the five causes apart, and fix each one in curl, Python requests, httpx, Node and Scrapy.
Sep 3, 2026 · 8 minRead more SEO datallms.txt vs robots.txtrobots.txt answers whether an agent may fetch a page and is honoured by every major AI operator. llms.txt answers what is worth reading and is documented as read by none of them. The formats, the resolution rules, what neither file can do, and measured adoption for both across the same 355 domains.
Sep 2, 2026 · 6 minRead more SEO dataHow Many Sites Use llms.txt? New Data64 of 355 top domains publish a real llms.txt. Count by status code instead and you get 39.7%, because 77 sites answer 200 with their homepage. Includes what the real files contain, and the cross-tab nobody has run: publishers of llms.txt block AI crawlers four to seven times less, and none of them blocks a retrieval crawler.
Sep 2, 2026 · 8 minRead more SEO dataAI Crawler User Agent ListA reference table of every AI user agent: which operator runs it, whether its job is training, retrieval or a user-triggered fetch, what blocking it actually costs you, and the share of 326 top domains that block it today. Plus the robots.txt mechanics that make these rules mean something other than intended.
Sep 2, 2026 · 8 minRead more SEO dataShould I Block AI Crawlers? New DataWe parsed the robots.txt of 326 of the web’s top domains. AI crawlers are blocked seven times more often than Googlebot, training bots twice as often as the search bots from the same company — and a large share of the blocking lands on crawlers that would have cited the site and never trained on it.
Sep 2, 2026 · 9 minRead more Lead generationHow to Find Business Email AddressesAn address is published, inferred or guessed, and the three behave completely differently once you press send. Where business addresses actually live, why role mailboxes are a category rather than a fallback, what verification can and cannot prove, and the obligations that attach the moment you find one.
Sep 2, 2026 · 9 minRead more Lead generationHow to Build a B2B Lead ListFour passes: express the ICP as predicates a machine can execute, source companies from public surfaces that answer different questions, enrich each row into a contactable record with provenance, and verify before anything is sent. Includes decay maths derived from official turnover statistics rather than a vendor deck.
Sep 2, 2026 · 9 minRead more Lead generationIs Scraping Google Maps Legal?The question hides four: the CFAA, contract, data protection and copyright. hiQ won the CFAA argument and lost on breach of contract; Google’s terms define prohibited automated access by pointing at robots.txt, which explicitly allows /maps/search/ and /maps/place/; and GDPR applies regardless of whether the data was public.
Sep 2, 2026 · 9 minRead more Lead generationHow to Scrape Leads From Google MapsA Maps listing gives you name, category, address, phone, rating, coordinates and a domain — but no email address. Building a usable lead list therefore means tiling the city, de-duplicating on place id, crawling each domain for a published contact, and knowing which of those businesses you are actually allowed to email.
Sep 2, 2026 · 10 minRead more SEO dataDo AI Crawlers Render JavaScript? A Data StudyWe audited 800 top domains with JS on and off: 1 in 15 is invisible to AI crawlers, 1 in 6 loses half its words — and the biggest sites fare worst.
Aug 25, 2026 · 5 minRead more GuidesHow to Use scrapy-playwrightInstall the download handler, flag requests to render with Playwright, wait for content, add proxies — and understand the throughput cost of a browser per page.
Jul 30, 2026 · 5 minRead more GuidesHow to Use a Proxy in Node.jsNative fetch needs an undici ProxyAgent; axios takes a proxy config or agent. The setup for both, authentication, HTTPS tunneling, and rotation for scraping.
Jul 30, 2026 · 5 minRead more GuidesHow to Use undetected-chromedriverInstall and launch the patched driver that evades Selenium detection, add a proxy, and understand the ceiling — IP reputation and behavior it can't fix.
Jul 30, 2026 · 5 minRead more GuidesHow to Use a Proxy in PuppeteerSet the proxy in launch args, authenticate with page.authenticate, rotate per page via browser contexts, and dodge the mistakes that leak your real IP.
Jul 30, 2026 · 5 minRead more GuidesHow to Use MCP in CursorWhat MCP gives Cursor's agent, how to add a server in mcp.json, project vs global scope, approving tool calls, and connecting a web-data server for live scraping.
Jul 30, 2026 · 5 minRead more GuidesHow to Use a Proxy with Python RequestsThe proxies dict, authenticated and SOCKS5 proxies, session reuse, rotating per request, and the mistakes — HTTPS key, verify, env vars — that silently break it.
Jul 30, 2026 · 5 minRead more GuidesPlaywright Stealth in PythonInstall playwright-stealth, understand which automation tells it patches, why sophisticated detectors still win, and the IP-plus-fingerprint combination that lasts.
Jul 30, 2026 · 5 minRead more GuidesHow to Use a Proxy in n8nTwo ways to route n8n through a proxy — per-node in the HTTP Request node or globally with env vars — the gotchas that trip people up, and rotation for scraping.
Jul 30, 2026 · 5 minRead more Data for AIHow to Feed Data to an LLMContext window, RAG, tool calls or fine-tuning — the four ways to give an LLM your data, how each works, when to pick it, and how to keep the source fresh.
Jul 30, 2026 · 5 minRead more MCP & agentsHow to Create an MCP ServerPick the SDK, define tools with clear schemas, choose a transport, test in a client — plus the tool-description rules that decide whether a model actually calls it.
Jul 30, 2026 · 5 minRead more Web scrapingIs Web Scraping Legal in Europe?The EU has no anti-scraping law, but four layers draw the lines: the GDPR, the database right, the DSM text-and-data-mining exception, and unfair-competition rules.
Jul 30, 2026 · 5 minRead more ProxiesWhat Is a Rotating Proxy?A rotating proxy hands out a fresh IP per request from a pool behind one endpoint. How that works, how it differs from static and sticky, and what it's for.
Jul 30, 2026 · 5 minRead more Data for AIHow Do Data Pipelines Work?The four stages every pipeline shares, ETL vs ELT, batch vs streaming, how orchestration ties it together — and where web data feeds in at the ingest step.
Jul 30, 2026 · 5 minRead more Data for AIHow to Create an LLM DatasetChoose the format for your goal, source raw text from the web, clean and deduplicate, structure the examples, and quality-check — the pipeline that decides model quality.
Jul 30, 2026 · 5 minRead more Browser agentsHow to Use Playwright for ScrapingInstall and launch, wait for content that loads late, extract with locators, add proxies and stealth — and the point where a browser per page stops being worth it.
Jul 30, 2026 · 5 minRead more ProxiesHow Can Residential Proxies Be Legal?The legality of a residential proxy comes down to one thing — did the person whose IP you're using agree? The ethical/illegal line, and how to check which pool you're buying.
Jul 30, 2026 · 5 minRead more AI scrapingHow to Stop Web Scraping (Honestly)A defender's honest guide: the tactics that stop casual scrapers, the ones that only add friction, and why protecting logins and PII beats blocking public prices.
Jul 30, 2026 · 5 minRead more AI scrapingHow to Build an AI Web ScraperWhere the LLM actually belongs in a scraping pipeline, how prompt-based extraction survives redesigns, and how to keep token costs from eating the project.
Jul 30, 2026 · 5 minRead more ProxiesHow to Set Up Rotating ProxiesHow the single rotating endpoint works, config in curl/Python/Node, sticky vs per-request sessions, and the settings that decide whether you get banned.
Jul 30, 2026 · 5 minRead more MCP & agentsHow MCP Servers Work: ArchitectureHost, client, server; JSON-RPC and transports; tools, resources and prompts — and a single tool call traced from user question to structured result.
Jul 30, 2026 · 5 minRead more CrawlingWhat Is Web Crawling in Python?The frontier algorithm behind every crawler, a minimal Python example, how Scrapy and Crawlee fit in, and the scaling wall where DIY stops paying.
Jul 30, 2026 · 5 minRead more SERP & searchHow a SERP API Works, End to EndThe full path from query to JSON: localization parameters, the identity layer, HTTP vs rendered fetches, parsing rich blocks, verticals and pricing.
Jul 30, 2026 · 5 minRead more Web scrapingIs Web Scraping Legal in Germany?Germany has no anti-scraping law — but GDPR, database rights, the TDM exception and unfair-competition rules draw real lines. Here is where they sit.
Jul 30, 2026 · 5 minRead more AI scrapingCan AI Work Without Data?Two different questions hide in one query: AI runs offline just fine, but AI without data is a contradiction — and stale data is a slower version of none.
Jul 30, 2026 · 5 minRead more Browser agentsHow to Build an AI Browser AgentThe observe-decide-act loop, the browser-control stack, prompt-injection defenses and honest cost math — everything a working browser agent actually needs.
Jul 30, 2026 · 5 minRead more SEO dataHow to Perform an SEO AuditA practical six-step SEO audit process with a checklist, the crawler-vs-user diff most audits skip, and how to run the whole thing programmatically.
Jul 30, 2026 · 6 minRead more Use casesHow to Price Watch on AmazonThree ways to price watch on Amazon — native price history and alerts, third-party trackers, or your own API watcher — with honest cost math for each.
Jul 29, 2026 · 11 minRead more Anti-botHow to Browser Fingerprint: MethodsThe signals, hashing and detection logic behind browser fingerprinting — plus how automation stacks get caught, and what it costs to avoid the problem entirely.
Jul 29, 2026 · 9 minRead more ProxiesHow to Rotate Proxy in Selenium PythonChrome fixes its proxy at launch. Here are the three real ways to rotate IPs in Selenium Python, with code, auth handling and honest bandwidth math.
Jul 29, 2026 · 9 minRead more Data for AIHow to Get Data for AIFive real sources of AI training data, how to judge them, and the cost math nobody publishes — plus working API calls for live web data.
Jul 29, 2026 · 10 minRead more MCP & agentsHow to Use the Claude API: Key to AgentCreate an API key, send your first message, then layer tool use and MCP on top — a practical Claude API walkthrough with real cost math.
Jul 29, 2026 · 9 minRead more SERP & searchHow to Get a SERP API in MinutesA practical guide to getting a SERP API — sign-up, key generation, first request, free tier limits and the cost math that decides which provider fits.
Jul 29, 2026 · 11 minRead more Web scrapingHow to Web Scrape Using PythonA practical Python scraping walkthrough: fetch, parse, paginate, handle JavaScript pages — and the cost math for when to stop maintaining your own stack.
Jul 29, 2026 · 9 minRead more AI scrapingIs Web Scraping Legal in the UK?Web scraping is not banned in the UK, but four separate legal layers decide whether your specific job is lawful. Here is how each one works in practice.
Jul 29, 2026 · 10 minRead more Browser agentsWhat Is AI Automation?AI automation puts a model in the middle of a workflow so it can read messy input, decide, and act. Definition, examples, tooling criteria and real cost math.
Jul 29, 2026 · 10 minRead more SEO dataHow to Site Audit in Semrush, Step by StepA practical walkthrough of Semrush Site Audit setup and reports — plus the JS-rendering gap a dashboard crawl hides, and per-URL API cost math.
Jul 29, 2026 · 10 minRead more Use casesHow to Price Monitor: Track Any MonitorTwo readings of "how to price monitor": what a monitor should cost by spec class, and how to build an automated price tracker with real cost math.
Jul 29, 2026 · 9 minRead more Anti-botIs Device Fingerprinting Legal?Device fingerprinting is conditionally legal: ePrivacy consent and GDPR balancing in the EU, notice and opt-out in the US. A practical, layer-by-layer breakdown.
Jul 29, 2026 · 10 minRead more ProxiesHow to Use Rotating ProxiesGateway setup, rotation cadence, sticky sessions, retry logic and the real cost math behind rotating proxies — with Python and curl you can run today.
Jul 29, 2026 · 9 minRead more ProxiesHow to use a residential proxyGateway credentials, username targeting flags, rotating vs sticky sessions, an exit-IP check, and the real per-GB cost math for residential bandwidth.
Jul 29, 2026 · 9 minRead more Data for AIHow to Use Data for AITraining, grounding and analysis need different data. A practical guide to sourcing, cleaning and serving data for AI — with real per-page cost math.
Jul 29, 2026 · 10 minRead more CrawlingIs Web Crawling Legal? Rules by LayerCrawling public pages is generally lawful. Legality turns on four separate layers — access, contract, content rights and privacy — plus how much load you create.
Jul 29, 2026 · 10 minRead more SERP & searchHow to Create a SERP API KeyA practical walkthrough for creating a SERP API key: signup and verification, key generation, safe storage, a first authenticated request, and quota math.
Jul 29, 2026 · 10 minRead more Web scrapingDoes Web Scraping Use an API?Web scraping and APIs are not opposites. Most modern scrapers hit a JSON endpoint, a public API, or a third-party scraping API — here is how to pick.
Jul 29, 2026 · 10 minRead more AI scrapingIs AI Web Scraping Legal?AI web scraping is not one legal question but three: how you access, what you collect, and what your model does with it. Here is the test, the case law and the pipeline.
Jul 29, 2026 · 10 minRead more Browser agentsWhat Is Browser Automation?Browser automation drives a real browser with code or an AI agent. Here is how it works, when it beats a plain HTTP request, and what it actually costs.
Jul 29, 2026 · 12 minRead more SEO dataHow to SEO Audit a WebsiteA step-by-step SEO audit process — crawlability, indexation, rendering, on-page and links — plus how to script the whole thing per URL instead of paying per dashboard.
Jul 29, 2026 · 11 minRead more Use casesIs Lead Generation Legal? Rules by LayerLead generation is legal — collection, contact channel and sector rules are what get people in trouble. A layered breakdown, vendor due diligence and real cost math.
Jul 29, 2026 · 9 minRead more Anti-botIs Browser Fingerprinting Legal?Browser fingerprinting is legal but regulated: consent rules for tracking, more room for fraud detection, and separate questions for anyone resisting it in automation.
Jul 29, 2026 · 9 minRead more ProxiesHow to Rotate Proxies in PythonRotate proxies in Python three ways — a cycled list, an async pool with cooldowns, or a rotating gateway — plus the cost math nobody publishes.
Jul 29, 2026 · 9 minRead more ProxiesHow to Detect Residential ProxiesA practical guide to detecting residential proxy traffic: IP attribution, TCP/TLS round-trip gaps, telemetry mismatches, behavioural signals — and the honest false-positive math.
Jul 29, 2026 · 10 minRead more Data for AIWhat is web data? Types, examples and usesA precise definition of web data, its types and examples, how it is collected, and what it actually costs to acquire at scale for analytics and AI.
Jul 29, 2026 · 9 minRead more MCP & agentsIs an MCP Server Like an API?An MCP server is an API in the broad sense but not a REST API. Here is what differs on the wire, with examples, a decision table and honest cost math.
Jul 29, 2026 · 10 minRead more CrawlingHow to Web Crawl in PythonA working Python crawl loop in 40 lines, the same job in Scrapy, and when to swap the loop for a crawl API — with real cost math per 10,000 pages.
Jul 29, 2026 · 10 minRead more SERP & searchHow to Use a SERP APIA working guide to using a SERP API — first request, the parameters that change results, parsing organic and PAA blocks, and what a real workload costs.
Jul 29, 2026 · 10 minRead more Web scrapingWhat Is a Web Scraper API?One endpoint, one URL in, structured data out. How web scraper APIs work under the hood, what they cost per page, and how to judge one.
Jul 29, 2026 · 10 minRead more AI scrapingIs Web Scraping Legal in the US?Scraping publicly available pages is generally lawful in the US. What creates liability is how you access, what you collect and how you reuse it.
Jul 29, 2026 · 10 minRead more