Documentation Python quickstart Blog Free tools hello@quanticdata.ioLog in

Read the Block: 4 Shapes, 4 Settings That Work

Plain HTTP against rendered word counts for the four reply shapes, residential exits, 28 September 2026: Amazon 0 and 0 words on a 2,007-byte page, Glassdoor 624 plain and 54 on a login redirect, Expedia 72 plain and 45 on a rate-limit reply, Zillow 759 plain and 0 on an empty render
Plain HTTP against rendered word counts for the four reply shapes, residential exits, 28 September 2026: Amazon 0 and 0 words on a 2,007-byte page, Glassdoor 624 plain and 54 on a login redirect, Expedia 72 plain and 45 on a rate-limit reply, Zillow 759 plain and 0 on an empty render

When a site refuses a scraper, read the shape of the reply, because the shape is the setting. Across 53 sites measured through residential exits between 26 and 28 September 2026 we met four shapes. A short page with no canonical (Amazon, 2,007 bytes, 0 words in both modes). A redirect to the login form (Glassdoor rendered: canonical /member/profile/login, 54 words, where plain HTTP returns 624). A rate-limit reply (Expedia rendered: 381,468 bytes, 45 words, no h1, where plain HTTP returns 72 words with the h1 in 4.1 seconds). An empty rendered page (Zillow: 759 words plain, 0 rendered). Three of the four resolve to plain HTTP with no browser; the fourth resolves to a collector with rows and seconds.

The guides list fourteen tips; none of them show a refused reply

The first page for this search is four vendor guides with ten, twelve and fourteen tips each, a Reddit link to one of them, two YouTube tutorials and a Quora answer that says add retry and backoff. The longest guide runs to 6,004 words and covers rotation, headers, fingerprints, request rate, honeypots and browsers. Every technique in it is real. What none of the guides do is show what a refusal looks like on the wire, so the reader applies all fourteen tips to every refusal, which on most of the sites below means opening a browser on a page the plain client was already reading.

We have the replies. Every site in the table was fetched twice from the same residential exit, as a plain HTTP client and fully rendered, and the audit reports the status, the title, the canonical, the h1 and the word count of each view. Four shapes cover everything we saw, and each has one setting that works, with numbers.

Thirteen targets, four shapes, one table

TargetExitPlain HTTP, wordsRendered, wordsShape servedThe setting
AmazonUS0 (2,007 bytes)0Short page, no canonicalamazon_search: 20 products in 43 s
eBayUS24 (1,976 bytes)24Short page, no canonical, noindexebay_search: 25 listings in 19.5 s
GlassdoorUS62454Login redirect when renderedPlain HTTP, never render; salaries via indeed_jobs
ExpediaUS7214Rate-limit reply when renderedPlain HTTP, US exit; hotels 20 rows in 160 s, google_flights 15 fares in 19 s
ZillowUS7590Empty renderPlain HTTP, never render; zillow_search 20 rows in 3.0 s
EtsyUS6780Empty renderPlain HTTP with a browser TLS profile, never render, never an EU exit
IdealistaIT1,1800Empty renderPlain HTTP in 3.8 s, never render; idealista_search 10 rows in 5.5 s
RedditUS5740Empty renderPlain HTTP for listings; reddit_comments 20 rows in 7.9 s
DoorDashUS48426Empty renderPlain HTTP, never render; doordash_restaurants 50 rows in 6.0 s
TripadvisorUS7130Empty renderPlain HTTP, 19 LocalBusiness JSON-LD blocks; tripadvisor_search 30 rows in 8.1 s
IndeedUS2,394683Render returns lessPlain HTTP with a browser TLS profile, 601 KB first request; indeed_jobs 10 rows in 7.3 s
SteamUS2,8021,396Render breaks the pagePlain HTTP; prices from the appdetails JSON with cc=
DraftKingsUS647642Render returns lessPlain HTTP; league pages 5,418 words

Count the settings. Eleven of thirteen say plain HTTP and no browser. Two say collector. None say render, and none say a better IP, because on every row the plain client from a Basic residential exit at $0.80/GB read the page or the page did not exist to be read.

Shape 1: a short page with no canonical

Amazon hands a plain client from a United States exit 2,007 bytes with a 200 status and 0 words: no title worth the name, no canonical, no description, no JSON-LD. The rendered fetch returns 0 words as well. eBay, weighed again at 21:22 UTC on 28 September, is 1,976 bytes, 24 words, no canonical, marked noindex, and the same 24 words arrive from Germany and from a browser. The tell is not the status and not the size. It is that the head describes nothing: a real Amazon or eBay page carries a canonical and a description, and these replies carry neither.

What works: stop fetching the page, because no fetch mode returns a catalogue that is not served to an anonymous first request. The amazon_search collector returns 20 products with ASIN, price, rating and review count in 43 seconds at $0.001 per product. The ebay_search collector returned 25 priced listings in 19.5 seconds today, every row with item id, price, format and seller location, and wrote into its notes which exit it served the run through. Retrying a short page with a browser, a slower rate or a new IP is the most common waste we see, and it is wasted on the cheapest bytes: a gigabyte of the eBay reply is 506,000 fetches of nothing.

Shape 2: a redirect to the login form

We rendered glassdoor.com/ from a United States residential exit at 21:22 UTC today. The browser landed on a page titled "Log In | Glassdoor" with a canonical of https://www.glassdoor.com/member/profile/login, 54 words, no h1, in 22.9 seconds. The Glassdoor post got 624 words with JSON-LD from a plain United States client and 589 from the United Kingdom on the 28 September afternoon pass. The site does not refuse the browser; it routes it to a form, and the form has a canonical that says so.

What works: classify on the canonical. A canonical ending in /login means the surface you rendered is gated, not that the item is missing, and the plain client is the one reading the public page. Setting: residential, engine: tls, never a render on this site. For salaries and postings as rows, the indeed_jobs collector returns 10 rows in 7.3 seconds. The same shape ended mbasic.facebook.com, which now redirects to /login/?next=...&refsrc=deprecated with a canonical of facebook.com/login; any tutorial that parses it is parsing a page that no longer answers. Glassdoor's robots.txt is explicit about what it wants read: it disallows /member/, /profile/ and /search/ for every agent and reserves the login and join pages as the only allowed paths under them.

Shape 3: a rate-limit reply

Expedia is the site where the plain fetch is the good one and the browser gets the refusal. At 21:22 UTC today the plain audit from a United States exit returned 72 words, the h1 "The one place you go to go places", a self-referencing canonical, a description and a robots meta of index,follow. The rendered fetch of the same URL returned a rate-limit reply: 381,468 bytes, 45 words, no canonical, no h1, no description, in 28.0 seconds. The Expedia post measured the same pair on the afternoon of 28 September: 72 words plain in 4.1 seconds, 14 rendered in 50.9.

What works: keep the plain fetch, keep the United States exit, and treat the absence of the h1 as the health check, because the refused reply has none. From a United Kingdom exit the plain fetch drops to 14 words with no canonical, so the exit is part of the setting. When a reply of this shape carries a Retry-After header, honour it per exit rather than per pool: MDN documents the header as either a number of seconds or an HTTP date, and the correct client behaviour is to wait that long before the next request from that address. Ours did not carry one, so the rule is the one in our proxy not working checklist: back off the exit that got the reply, not the whole pool. Rates never come from expedia.com pages anyway. The search URLs are the ones robots.txt fences off (disallow: /Hotel-Search, disallow: /Flights-Search, disallow: /*chkin=*), so prices come from the hotels collector, 20 priced properties in 160 seconds, and from google_flights, 15 fares in 19 seconds.

Shape 4: an empty rendered page

This is the shape that costs the most, because the plain client was already reading the page. Zillow: 759 words plain, with price, beds and square feet in the meta description of listing pages, and 0 rendered. Etsy: 678 words in 205,911 bytes plain, 0 rendered. Idealista: 1,180 words in 3.8 seconds plain, 0 words in 38 seconds rendered. Tripadvisor: 713 words and 19 LocalBusiness JSON-LD objects plain, 0 rendered. Reddit listings: 574 words plain, 0 rendered. DoorDash: 484 words plain, 26 rendered. Indeed: 2,394 words in 601 KB on the first plain request, 683 rendered. Steam: 2,802 plain, 1,396 rendered with the page's own error message as the h1.

What works: residential, engine: tls, with a browser TLS profile on Etsy and Indeed, and no render, ever, on any of the eight. For rows, the collector named in the table: zillow_search 20 in 3.0 seconds, idealista_search 10 in 5.5, reddit_comments 20 in 7.9, doordash_restaurants 50 in 6.0, tripadvisor_search 30 in 8.1, indeed_jobs 10 in 7.3. The guides' first instinct on a thin-looking page is to open a browser. On these eight the browser is the thing that produces the empty reply, and the fix is to not send it.

The classifier: canonical, h1, words, never the status

Every shape above can arrive with a 200, and a real page can arrive with a status that looks wrong: Booking.com's 26-word shell is a 202 and Trustpilot's 3,588-word rendered home reports a non-200 status with the correct title and h1. A retry loop keyed on the status code is blind on this family of sites. Key it on the head:

def shape(r):
    # r: status, bytes, words, canonical, h1, jsonld, retry_after
    if r.canonical and r.canonical.endswith("/login"):
        return "login_redirect"     # gated surface: read it plain, never rendered
    if r.retry_after or (r.words < 50 and not r.h1 and not r.canonical and r.bytes > 100_000):
        return "rate_limit"         # back off this exit; keep the plain mode that worked
    if r.bytes < 5_000 and not r.canonical:
        return "short_page"         # no page exists to fetch: collector
    if r.words == 0 and r.rendered:
        return "empty_render"       # the plain client reads it: engine tls, no browser
    return "page"

Run it on the plain reply first and on the rendered reply only when the plain one is a shell with a canonical and an h1, which is the rule in 53 targets measured: browser or not. On the thirteen rows above that order never opens a browser on Zillow, Etsy, Idealista, Reddit, DoorDash, Tripadvisor, Indeed, Steam or Glassdoor, and never fetches Amazon or eBay twice.

What getting the shape wrong costs

Residential Basic at $0.80/GB, counting a gigabyte as 10^9 bytes. The middle column is the fetch the guides recommend; the right column is the one that works.

TargetRendered fetch: bytes, words, secondsCost of that fetchPlain fetch that worksCost
Expedia381,468 bytes, 45 words, 28.0 s$0.00031507,949 bytes, 72 words, h1 and canonical, 4.1 s$0.00041
Glassdoor705,071 bytes, 54 words on the login form, 60.8 s$0.00056624 words with JSON-LDnot weighed
Indeed1,066,178 bytes, 683 words$0.00085601,347 bytes, 2,394 words$0.00048
Steam1,359,498 bytes, 1,396 words, error h1$0.001091,073,636 bytes, 2,802 words$0.00086
eBay24 words, any mode$0.0000016ebay_search, 25 listings in 19.5 s$0.025
Amazon0 words, any mode$0.0000016amazon_search, 20 products in 43 s$0.02

The bandwidth numbers are small; the time is not. Expedia's rendered refusal took 28 seconds and Glassdoor's took 60.8 to reach a form, against 4.1 seconds for the plain Expedia page. On a pool, a shape-4 mistake costs a browser slot for a minute per page and returns nothing, which is the throughput the guides are spending on their tips.

Where we stop

Everything above is about reading what a site serves an anonymous visitor: public listings, salaries, prices, availability, restaurant lists. The robots files we fetched draw the lines and we keep to them. Reddit's is one rule for every unnamed agent, Disallow: /, with a pointer to its Public Content Policy, which is why Reddit data is a collector job under that policy and not a crawl. Zillow's opens by pointing to its Terms of Use and allows only the top of a few listing trees. Indeed's allows the first ten result pages for everyone and disallows job views. Glassdoor's disallows member, profile and search paths. None of this is a reason to render harder or rotate faster; it is the map of what is public, and the settings above stay inside it. We do not help with evading a login, an account restriction or a ban.

The setting that works on these targets

  • Network: residential, Basic line at $0.80/GB. On all thirteen rows the plain client from a Basic exit read the page or the page did not exist; Premium at $2.20/GB and mobile at $2.30/GB changed nothing we could count.
  • Fetch mode: engine: tls on eleven of thirteen, with the browser TLS profile on Etsy (678 words) and Indeed (2,394 words); no render on any. The rendered fetch produced the refusal on Glassdoor (54 words), Expedia (45), Zillow, Etsy, Idealista, Tripadvisor and Reddit (0).
  • Country: the United States on Expedia (72 words; the United Kingdom gets 14 and no canonical), Etsy (an EU exit gets 8 words), Glassdoor (624 US, 589 GB); Italy on Idealista. The exit that received a rate-limit reply backs off alone; the pool does not.
  • When the proxy is not enough: the short-page shape, where no fetch mode returns rows: amazon_search, 20 products in 43 s, and ebay_search, 25 listings in 19.5 s, both at $0.001 per row. For rows on the empty-render sites, zillow_search 20 in 3.0 s, idealista_search 10 in 5.5 s, reddit_comments 20 in 7.9 s, doordash_restaurants 50 in 6.0 s. For any page as Markdown with the fetch mode chosen for you, the web scraping API from $0.0002 per page, billed on success.
  • Failed requests are never billed, and every account gets $2 of free API usage per month, which is 10,000 plain fetches or 2,000 collector rows before the first charge.

Sources & further reading

FAQ

Quick answers on web scraping blocked what to do.

Something else? Ask us →

My scraper gets an empty page from a headless browser but not from plain requests. What do I do?

Stop rendering that site. On 8 of the 53 sites we measured the plain client reads the page and the browser returns nothing: Zillow 759 words against 0, Etsy 678 against 0, Idealista 1,180 against 0, Tripadvisor 713 against 0, Reddit listings 574 against 0, DoorDash 484 against 26, Indeed 2,394 against 683, Steam 2,802 against 1,396. The setting on all eight is residential with engine tls, a browser TLS profile on Etsy and Indeed, and no render.

How do I tell a login redirect from a missing page?

Read the canonical, not the status. Glassdoor rendered on 28 September 2026 returned a 200 with the title "Log In | Glassdoor" and a canonical of /member/profile/login, 54 words, in 22.9 seconds; the plain client got 624 words with JSON-LD. A canonical ending in /login means the surface you rendered is gated and the item still exists on the public page. mbasic.facebook.com shows the same shape with refsrc=deprecated in the redirect, which means the surface itself is gone.

What should a scraper do with a rate-limit reply?

Keep the fetch mode that worked and back off the exit that got the reply. Expedia returned 72 words with the h1 and canonical to a plain United States client in 4.1 seconds and a rate-limit reply of 381,468 bytes and 45 words to the browser in 28 seconds, so the fix is not to render. If the reply carries Retry-After, MDN defines it as seconds or an HTTP date; wait that long on that address. Prices come from the hotels collector (20 rows in 160 s) and google_flights (15 fares in 19 s), not from the search pages robots.txt fences off.

Does a short 200 with no content mean I need a better proxy?

No. Amazon returned 2,007 bytes and 0 words to a plain United States client and 0 words to a browser; eBay returned 1,976 bytes and 24 words from the United States, from Germany and from a browser alike. Nothing scored the address, so a Premium or mobile exit changes the price and not the page. The setting is the collector: amazon_search, 20 products in 43 seconds; ebay_search, 25 listings in 19.5 seconds on 28 September 2026, at $0.001 per row.

Why not key retries on the HTTP status code?

Because every shape above can arrive as a 200 and a real page can arrive with a status that looks wrong: Booking.com serves its 26-word shell with a 202, Trustpilot serves a 3,588-word rendered home with a non-200 status and the correct h1, and Glassdoor serves its login form with a 200. On the 13 targets in this post the reliable fields are the canonical, the h1, the word count and the byte size; a classifier on those four never opens a browser on the empty-render sites and never fetches Amazon or eBay twice.

See the shape before you change the setting

Every reply in this post came from one tool: fetch a URL plain and rendered from the same exit, read the canonical, h1 and word count of each, and pick the mode that returned the page. Every account gets $2 of free API usage per month, and failed requests are never billed.

Related reading