On 12 of the 53 sites we measured in September 2026, a headless browser returns fewer words than a plain HTTP client, and on 7 of them it returns nothing usable at all. Etsy, Idealista, Zillow, Tripadvisor and Reddit hand a rendered request 0 words; DoorDash 26; Expedia 14. The same twelve URLs, fetched over plain HTTP with a current browser TLS fingerprint through residential exits, return 11,157 words. Rendered, 3,032. The setting for all twelve is residential, engine: tls, never rendered, and it is also the cheaper one.
Every guide says "if you are stuck, render". On these 12 sites that is backwards
Search for headless browser web scraping and the first page tells you what a headless browser is, lists eight of them, and shows you Puppeteer. The most upvoted human voice on that page is a Reddit thread whose author stopped using one: high memory, high latency, and an HTTP client plus the site's own JSON endpoints did the job. The vendor explainers are careful to say a browser "does not guarantee access". None of them says which sites punish it, or by how much.
We can, because we fetched 53 targets twice each: once as a pure HTTP client with a browser TLS profile, once fully rendered, both through residential proxies from the country the site serves. The verdict column of that table has six values. Two of them, "less" and "nothing", cover the twelve sites in this post. This is the group the tutorials never mention, because on these sites the standard advice makes the result worse and the bill bigger at the same time.
Twelve sites, 11,157 words over plain HTTP and 3,032 rendered
Every row below is a home page or a public listing page, fetched logged out, no cookies, no Accept-Language, through a residential exit in the country shown. Word counts are extractable text; bytes are the decoded body of the plain fetch.
| Site | Exit | Plain HTTP words | Rendered words | Plain fetch bytes | Setting |
|---|---|---|---|---|---|
| Steam | US | 2,802 | 1,396, with an error heading | 1,073,636 | plain HTTP, prices via appdetails JSON |
| Indeed | US | 2,394 | 683 | 601,347 | plain HTTP with a browser TLS profile |
| Idealista | IT | 1,180 | 0 | 101,680 | plain HTTP; collector for listings |
| Zillow | US | 759 | 0 | 423,697 | plain HTTP; collector for search |
| Tripadvisor | US | 713 | 0 | 388,397 | plain HTTP, 19 LocalBusiness objects in JSON-LD |
| Etsy | US | 678 | 0 | 205,911 | plain HTTP, US exit only |
| DraftKings | US | 647 | 642 | 1,581,100 | plain HTTP; league pages plain too |
| Glassdoor | US | 624 | 54, on the login form | n/a | plain HTTP, one exit per edition |
| US | 574 | 0 | 566,030 | plain HTTP; collector for comments | |
| DoorDash | US | 484 | 26 | 2,025,194 | plain HTTP; collector for restaurants |
| Vinted | IT | 230 | 217 | 1,997,674 | plain HTTP, country of the domain |
| Expedia | US | 72 | 14 | 507,949 | plain HTTP; collectors for dated prices |
Add the two columns and the shape of the problem is plain: 11,157 words without a browser, 3,032 with one, a factor of 3.7 in favour of the cheaper fetch. Five sites return exactly zero words to the browser. Two more return under thirty. On the remaining five the browser returns a smaller page than the raw HTML did, and on one of them it returns a page that is actively wrong.
Three shapes of harm: fewer words, no words, and a wrong page
Fewer words. DraftKings serves 647 words to a plain client and 642 to a browser, for 2,877,467 bytes instead of 1,581,100. Vinted serves 230 and 217. Steam serves 2,802 and 1,396: the storefront's server-rendered HTML carries the whole front page, and the rendered pass loses half of it. Re-audited today from a US exit, Steam returned 2,785 words plain and 1,396 rendered in a 10.5-second audit, the same shape two days later.
A wrong page. Steam is also the case where the browser does not merely lose text but replaces it. The rendered pass has an h1 that the plain HTML does not have, and it reads "Something went wrong while displaying this content. Refresh". A parser keyed on the h1 will file that as the page title. Glassdoor is the same shape with a different destination: 624 words plain, and the rendered pass lands on the member login form with 54 words and a canonical of /member/profile/login.
No words. Etsy is the cleanest example, because it is stable. On 26 September the plain fetch returned 707 words and the render returned nothing; on 28 September, 678 and nothing; today, 678 words, a canonical, WebSite and Organization JSON-LD over plain HTTP, and again nothing from the browser, in a 9.9-second audit. Idealista, Zillow, Tripadvisor and Reddit behave the same way: the raw HTML is the page, and the browser is handed a document with no content in it. Reddit is the extreme: 566,030 bytes and 574 words plain, 10,827 bytes and 0 words rendered.
Notice what is not in those three paragraphs: a status code, a vendor name, a screenshot of the page the browser got instead. That page is not the finding. The finding is that on these sites the plain fetch is the one that works, and the instruction is to use it.
The fingerprint decides, not the address
Indeed is the row that explains the mechanism. On 26 September a plain fetch returned 1,026 words on the first request. On 28 September the plain fetch, sent with a current Firefox TLS profile, returned 601,347 bytes and 2,394 words in 9.4 seconds. The same address rendered returned 683 words after 52 seconds and 1,066,178 bytes. What the site scores is the shape of the client, and a plain HTTP client presenting a current browser TLS handshake through a household IP is the shape it serves in full.
That reorders the shopping list for all twelve. Nothing in this group rewarded a more expensive address: Premium residential at $2.20/GB, mobile at $2.30/GB and ISP at $2.50 per IP per month at volume were not the variable that changed the outcome. Residential Basic at $0.80/GB was, every time, with the TLS profile doing the work. Pin the country the site serves, because that does change the page: Etsy returned 678 words from the United States and 8 from Germany and the United Kingdom; Indeed redirected a UK exit to uk.indeed.com and returned 216 words; Glassdoor returned 624 from the US and 589 from the UK edition.
What the browser costs, and what you save by not buying it
Bytes and seconds, counting a gigabyte as 10^9 bytes and residential Basic at $0.80/GB. The rendered bytes are the browser's decoded document; the seconds are the wall time of the pass.
| Site | Plain bytes / seconds | Rendered bytes / seconds | Plain cost per fetch | Rendered cost per fetch | Words you gain by rendering |
|---|---|---|---|---|---|
| Steam | 1,073,636 / 9.1 s | 1,359,498 / 9.0 s | $0.00086 | $0.00109 | -1,406 |
| Indeed | 601,347 / 9.4 s | 1,066,178 / 52.1 s | $0.00048 | $0.00085 | -1,711 |
| DraftKings | 1,581,100 | 2,877,467 / 34.8 s | $0.00126 | $0.00230 | -5 |
| Vinted | 1,997,674 / 3.5 s | 1,915,011 / 36.4 s | $0.00160 | $0.00153 | -13 |
| Idealista | 101,680 / 3.8 s | 5,703 / 38.5 s | $0.00008 | $0.000005 | -1,180 |
| 566,030 / 6.4 s | 10,827 / 41.0 s | $0.00045 | $0.000009 | -574 | |
| Glassdoor | n/a | 705,071 / 60.8 s | n/a | $0.00056 | -570 |
Read the Idealista and Reddit rows twice. The rendered pass is cheap in bandwidth because it returns almost nothing, and it takes ten times longer to do it. A budget that counts bytes will call those rows a saving; a budget that counts words per second will not. Through the web scraping API the same difference is priced explicitly: $0.0002 per page over plain HTTP, $0.001 with rendering. On these twelve sites the five-times-dearer option buys a smaller page, and on seven of them it buys an empty one.
Where the page is not the product, the collector returns rows
Seven of the twelve have a ready-made collector, and on each of them the plain fetch gives you the page while the collector gives you the table. Runs from 28 September, billed per delivered row and never for failures:
- idealista_search: 10 Milan listings in 5.51 seconds with price, size and agency.
- indeed_jobs: 10 listings for "data engineer" in New York in 7.3 seconds, salary ranges parsed.
- zillow_search: 20 listings for Austin, TX in 3.02 seconds.
- tripadvisor_search: 30 Chicago restaurants in 8.1 seconds with phone, rating and review count.
- doordash_restaurants: 50 Austin restaurants in 5.98 seconds.
- reddit_comments: 20 comments in 7.9 seconds, the surface that no logged-out fetch returns.
- steam: 10 apps for "portal" in 1.83 seconds.
Expedia's dated prices come through the hotels collector, 20 priced New York properties in 160 seconds, and DraftKings and Vinted have no collector, because their terms rule one out. On Steam, DraftKings and the two marketplaces there is a further line to state plainly: each site's terms forbid automated access by account holders, and where a site says that, we read it logged out and stop there.
Robots files: read them, they are short
Steam's robots.txt is 303 bytes, ten Disallow lines, last modified 26 May 2026, and it closes only account and sharing paths. Zillow's is 10,524 bytes, last modified 22 September 2026, with 15 sitemaps. Etsy's is 52,810 bytes and lists 1,681 disallow lines across three user-agent groups. A robots file is a technical statement of what the site wants crawled; it is not a permission to do anything else, and the plain-HTTP rule in this post is about what the site serves, not about what you may keep.
The setting that works on these 12 sites
- Network: residential proxies, Basic line at $0.80/GB. Every plain fetch in the table came through it on the first attempt; no site in this group rewarded Premium at $2.20/GB, mobile at $2.30/GB or ISP at $2.50 per IP per month.
- Fetch mode:
engine: tls, plain HTTP with a current browser TLS profile, never rendered. 11,157 words across the twelve against 3,032 rendered; Steam 2,802 against 1,396, Indeed 2,394 against 683, Etsy 678 against 0. - Country to pin: the market the site serves. Etsy gives a US exit 678 words and a German one 8; Indeed sends a UK exit to
uk.indeed.comfor 216 words instead of 2,394; Idealista and Vinted are pinned to the country of the domain. - When the proxy is not enough: the collectors above, from zillow_search at 20 rows in 3.02 seconds to doordash_restaurants at 50 rows in 5.98 seconds; for the raw page as Markdown, the web scraping API at $0.0002 per page with rendering switched off.
- Free tier: every account gets $2 of free API usage per month, which is 10,000 plain-fetch pages before anything is billed.