# 12 Sites Where Headless Browsers Lock You Out

> We measured 53 sites with and without a headless browser. On 12 the browser returns fewer words, on 7 none. Plain HTTP gets 11,157 words: use engine tls.

[Home](https://quanticdata.io/)/[Blog](https://quanticdata.io/blog/)/12 Sites Where Headless Browsers Lock You Out

# 12 Sites Where Headless Browsers Lock You Out

ProxiesSep 28, 2026·10 min read·By [Aldo Morese](https://quanticdata.io/about/), founder of QuanticData

Words returned by plain HTTP versus a headless browser on six sites measured through residential proxies in September 2026: Steam 2,802 against 1,396, Indeed 2,394 against 683, Idealista 1,180 against 0, Etsy 678 against 0, Reddit 574 against 0, Glassdoor 624 against 54

On this page [Every guide says "if you are stuck, render". On these 12 sites that is backwards](/blog/headless-browser-locks-you-out-plain-http/#every-guide-says-if-you-are-stuck-render-on-these-12-sites-t) [Twelve sites, 11,157 words over plain HTTP and 3,032 rendered](/blog/headless-browser-locks-you-out-plain-http/#twelve-sites-11-157-words-over-plain-http-and-3-032-rendered) [Three shapes of harm: fewer words, no words, and a wrong page](/blog/headless-browser-locks-you-out-plain-http/#three-shapes-of-harm-fewer-words-no-words-and-a-wrong-page) [The fingerprint decides, not the address](/blog/headless-browser-locks-you-out-plain-http/#the-fingerprint-decides-not-the-address) [What the browser costs, and what you save by not buying it](/blog/headless-browser-locks-you-out-plain-http/#what-the-browser-costs-and-what-you-save-by-not-buying-it) [Where the page is not the product, the collector returns rows](/blog/headless-browser-locks-you-out-plain-http/#where-the-page-is-not-the-product-the-collector-returns-rows) [Robots files: read them, they are short](/blog/headless-browser-locks-you-out-plain-http/#robots-files-read-them-they-are-short) [The setting that works on these 12 sites](/blog/headless-browser-locks-you-out-plain-http/#the-setting-that-works-on-these-12-sites)

On 12 of the 53 sites we measured in September 2026, a headless browser returns fewer words than a plain HTTP client, and on 7 of them it returns nothing usable at all. Etsy, Idealista, Zillow, Tripadvisor and Reddit hand a rendered request 0 words; DoorDash 26; Expedia 14. The same twelve URLs, fetched over plain HTTP with a current browser TLS fingerprint through residential exits, return 11,157 words. Rendered, 3,032. The setting for all twelve is residential, `engine: tls`, never rendered, and it is also the cheaper one.

## Every guide says "if you are stuck, render". On these 12 sites that is backwards

Search for headless browser web scraping and the first page tells you what a headless browser is, lists eight of them, and shows you Puppeteer. The most upvoted human voice on that page is a Reddit thread whose author stopped using one: high memory, high latency, and an HTTP client plus the site's own JSON endpoints did the job. The vendor explainers are careful to say a browser "does not guarantee access". None of them says which sites punish it, or by how much.

We can, because we fetched 53 targets twice each: once as a pure HTTP client with a browser TLS profile, once fully rendered, both through [residential proxies](https://quanticdata.io/residential-proxies/) from the country the site serves. The verdict column of that table has six values. Two of them, "less" and "nothing", cover the twelve sites in this post. This is the group the tutorials never mention, because on these sites the standard advice makes the result worse and the bill bigger at the same time.

## Twelve sites, 11,157 words over plain HTTP and 3,032 rendered

Every row below is a home page or a public listing page, fetched logged out, no cookies, no `Accept-Language`, through a residential exit in the country shown. Word counts are extractable text; bytes are the decoded body of the plain fetch.

| Site | Exit | Plain HTTP words | Rendered words | Plain fetch bytes | Setting |
| --- | --- | --- | --- | --- | --- |
| Steam | US | 2,802 | 1,396, with an error heading | 1,073,636 | plain HTTP, prices via appdetails JSON |
| Indeed | US | 2,394 | 683 | 601,347 | plain HTTP with a browser TLS profile |
| Idealista | IT | 1,180 | 0 | 101,680 | plain HTTP; collector for listings |
| Zillow | US | 759 | 0 | 423,697 | plain HTTP; collector for search |
| Tripadvisor | US | 713 | 0 | 388,397 | plain HTTP, 19 LocalBusiness objects in JSON-LD |
| Etsy | US | 678 | 0 | 205,911 | plain HTTP, US exit only |
| DraftKings | US | 647 | 642 | 1,581,100 | plain HTTP; league pages plain too |
| Glassdoor | US | 624 | 54, on the login form | n/a | plain HTTP, one exit per edition |
| Reddit | US | 574 | 0 | 566,030 | plain HTTP; collector for comments |
| DoorDash | US | 484 | 26 | 2,025,194 | plain HTTP; collector for restaurants |
| Vinted | IT | 230 | 217 | 1,997,674 | plain HTTP, country of the domain |
| Expedia | US | 72 | 14 | 507,949 | plain HTTP; collectors for dated prices |

Add the two columns and the shape of the problem is plain: 11,157 words without a browser, 3,032 with one, a factor of 3.7 in favour of the cheaper fetch. Five sites return exactly zero words to the browser. Two more return under thirty. On the remaining five the browser returns a smaller page than the raw HTML did, and on one of them it returns a page that is actively wrong.

## Three shapes of harm: fewer words, no words, and a wrong page

**Fewer words.** DraftKings serves 647 words to a plain client and 642 to a browser, for 2,877,467 bytes instead of 1,581,100. Vinted serves 230 and 217. Steam serves 2,802 and 1,396: the storefront's server-rendered HTML carries the whole front page, and the rendered pass loses half of it. Re-audited today from a US exit, Steam returned 2,785 words plain and 1,396 rendered in a 10.5-second audit, the same shape two days later.

**A wrong page.** Steam is also the case where the browser does not merely lose text but replaces it. The rendered pass has an h1 that the plain HTML does not have, and it reads "Something went wrong while displaying this content. Refresh". A parser keyed on the h1 will file that as the page title. Glassdoor is the same shape with a different destination: 624 words plain, and the rendered pass lands on the member login form with 54 words and a canonical of `/member/profile/login`.

**No words.** Etsy is the cleanest example, because it is stable. On 26 September the plain fetch returned 707 words and the render returned nothing; on 28 September, 678 and nothing; today, 678 words, a canonical, WebSite and Organization JSON-LD over plain HTTP, and again nothing from the browser, in a 9.9-second audit. Idealista, Zillow, Tripadvisor and Reddit behave the same way: the raw HTML is the page, and the browser is handed a document with no content in it. Reddit is the extreme: 566,030 bytes and 574 words plain, 10,827 bytes and 0 words rendered.

Notice what is not in those three paragraphs: a status code, a vendor name, a screenshot of the page the browser got instead. That page is not the finding. The finding is that on these sites the plain fetch is the one that works, and the instruction is to use it.

## The fingerprint decides, not the address

Indeed is the row that explains the mechanism. On 26 September a plain fetch returned 1,026 words on the first request. On 28 September the plain fetch, sent with a current Firefox TLS profile, returned 601,347 bytes and 2,394 words in 9.4 seconds. The same address rendered returned 683 words after 52 seconds and 1,066,178 bytes. What the site scores is the shape of the client, and a plain HTTP client presenting a current browser TLS handshake through a household IP is the shape it serves in full.

That reorders the shopping list for all twelve. Nothing in this group rewarded a more expensive address: Premium residential at $2.20/GB, mobile at $2.30/GB and ISP at $2.50 per IP per month at volume were not the variable that changed the outcome. Residential Basic at $0.80/GB was, every time, with the TLS profile doing the work. Pin the country the site serves, because that does change the page: Etsy returned 678 words from the United States and 8 from Germany and the United Kingdom; Indeed redirected a UK exit to `uk.indeed.com` and returned 216 words; Glassdoor returned 624 from the US and 589 from the UK edition.

## What the browser costs, and what you save by not buying it

Bytes and seconds, counting a gigabyte as 10^9 bytes and residential Basic at $0.80/GB. The rendered bytes are the browser's decoded document; the seconds are the wall time of the pass.

| Site | Plain bytes / seconds | Rendered bytes / seconds | Plain cost per fetch | Rendered cost per fetch | Words you gain by rendering |
| --- | --- | --- | --- | --- | --- |
| Steam | 1,073,636 / 9.1 s | 1,359,498 / 9.0 s | $0.00086 | $0.00109 | -1,406 |
| Indeed | 601,347 / 9.4 s | 1,066,178 / 52.1 s | $0.00048 | $0.00085 | -1,711 |
| DraftKings | 1,581,100 | 2,877,467 / 34.8 s | $0.00126 | $0.00230 | -5 |
| Vinted | 1,997,674 / 3.5 s | 1,915,011 / 36.4 s | $0.00160 | $0.00153 | -13 |
| Idealista | 101,680 / 3.8 s | 5,703 / 38.5 s | $0.00008 | $0.000005 | -1,180 |
| Reddit | 566,030 / 6.4 s | 10,827 / 41.0 s | $0.00045 | $0.000009 | -574 |
| Glassdoor | n/a | 705,071 / 60.8 s | n/a | $0.00056 | -570 |

Read the Idealista and Reddit rows twice. The rendered pass is cheap in bandwidth because it returns almost nothing, and it takes ten times longer to do it. A budget that counts bytes will call those rows a saving; a budget that counts words per second will not. Through the [web scraping API](https://quanticdata.io/web-scraping-api/) the same difference is priced explicitly: $0.0002 per page over plain HTTP, $0.001 with rendering. On these twelve sites the five-times-dearer option buys a smaller page, and on seven of them it buys an empty one.

## Where the page is not the product, the collector returns rows

Seven of the twelve have a ready-made collector, and on each of them the plain fetch gives you the page while the collector gives you the table. Runs from 28 September, billed per delivered row and never for failures:

- [idealista_search](https://quanticdata.io/collectors/idealista-scraper-api/): 10 Milan listings in 5.51 seconds with price, size and agency.

- [indeed_jobs](https://quanticdata.io/collectors/indeed-jobs-api/): 10 listings for "data engineer" in New York in 7.3 seconds, salary ranges parsed.

- [zillow_search](https://quanticdata.io/collectors/zillow-scraper-api/): 20 listings for Austin, TX in 3.02 seconds.

- [tripadvisor_search](https://quanticdata.io/collectors/tripadvisor-scraper-api/): 30 Chicago restaurants in 8.1 seconds with phone, rating and review count.

- [doordash_restaurants](https://quanticdata.io/collectors/doordash-scraper-api/): 50 Austin restaurants in 5.98 seconds.

- [reddit_comments](https://quanticdata.io/collectors/reddit-comments-api/): 20 comments in 7.9 seconds, the surface that no logged-out fetch returns.

- [steam](https://quanticdata.io/collectors/steam-store-api/): 10 apps for "portal" in 1.83 seconds.

Expedia's dated prices come through the [hotels collector](https://quanticdata.io/collectors/google-hotels-api/), 20 priced New York properties in 160 seconds, and DraftKings and Vinted have no collector, because their terms rule one out. On Steam, DraftKings and the two marketplaces there is a further line to state plainly: each site's terms forbid automated access by account holders, and where a site says that, we read it logged out and stop there.

## Robots files: read them, they are short

Steam's `robots.txt` is 303 bytes, ten `Disallow` lines, last modified 26 May 2026, and it closes only account and sharing paths. Zillow's is 10,524 bytes, last modified 22 September 2026, with 15 sitemaps. Etsy's is 52,810 bytes and lists 1,681 disallow lines across three user-agent groups. A robots file is a technical statement of what the site wants crawled; it is not a permission to do anything else, and the plain-HTTP rule in this post is about what the site serves, not about what you may keep.

## The setting that works on these 12 sites

- **Network**: [residential proxies](https://quanticdata.io/residential-proxies/), Basic line at $0.80/GB. Every plain fetch in the table came through it on the first attempt; no site in this group rewarded Premium at $2.20/GB, mobile at $2.30/GB or ISP at $2.50 per IP per month.

- **Fetch mode**: `engine: tls`, plain HTTP with a current browser TLS profile, never rendered. 11,157 words across the twelve against 3,032 rendered; Steam 2,802 against 1,396, Indeed 2,394 against 683, Etsy 678 against 0.

- **Country to pin**: the market the site serves. Etsy gives a US exit 678 words and a German one 8; Indeed sends a UK exit to `uk.indeed.com` for 216 words instead of 2,394; Idealista and Vinted are pinned to the country of the domain.

- **When the proxy is not enough**: the collectors above, from [zillow_search](https://quanticdata.io/collectors/zillow-scraper-api/) at 20 rows in 3.02 seconds to [doordash_restaurants](https://quanticdata.io/collectors/doordash-scraper-api/) at 50 rows in 5.98 seconds; for the raw page as Markdown, the [web scraping API](https://quanticdata.io/web-scraping-api/) at $0.0002 per page with rendering switched off.

- **Free tier**: every account gets $2 of free API usage per month, which is 10,000 plain-fetch pages before anything is billed.

### Sources & further reading

- [store.steampowered.com/robots.txt (fetched 28 September 2026, 303 bytes)](https://store.steampowered.com/robots.txt)

- [etsy.com/robots.txt (fetched 28 September 2026, 52,810 bytes)](https://www.etsy.com/robots.txt)

- [zillow.com/robots.txt (fetched 28 September 2026, last modified 22 September 2026)](https://www.zillow.com/robots.txt)

- [Your preferred method to scrape? Headless browser or private APIs, r/webscraping](https://www.reddit.com/r/webscraping/comments/1hjuan9/your_preferred_method_to_scrape_headless_browser/)

## FAQ

Quick answers on headless browser.

[Something else? Ask us →](mailto:hello@quanticdata.io)

### Does a headless browser get you further on sites that limit scrapers?

Not on these twelve. Across 53 measured targets, 12 give a rendered request fewer words than a plain one, and on 7 of those the browser receives nothing usable: Etsy, Idealista, Zillow, Tripadvisor and Reddit return 0 words, DoorDash 26, Expedia 14. The same URLs return 11,157 words over plain HTTP against 3,032 rendered. The instruction is to fetch them with engine: tls and a current browser TLS profile.

### Which proxy type do these sites need?

Residential Basic at $0.80/GB, with the fetch mode doing the work. Indeed returned 2,394 words to a plain client presenting a current Firefox TLS handshake and 683 to a browser; nothing in the group changed its answer for Premium at $2.20/GB, mobile at $2.30/GB or ISP at $2.50 per IP per month. Pin the exit country, because it changes the page: Etsy served 678 words to a US exit and 8 to a German one.

### How much does skipping the browser save?

Both time and money. Indeed: 9.4 seconds and 601,347 bytes plain against 52.1 seconds and 1,066,178 bytes rendered, $0.00048 against $0.00085 at $0.80/GB with a gigabyte counted as 10^9 bytes. Idealista: 3.8 seconds against 38.5. Glassdoor: the rendered pass alone takes 60.8 seconds and lands on a login form. Through the web scraping API the difference is $0.0002 per page against $0.001 with rendering, for a smaller page.

### What does the browser actually return on Steam?

Half the page and a wrong heading. The Steam storefront served 2,802 words to a plain client and 1,396 to a browser on 28 September 2026, and 2,785 against 1,396 when re-audited today. The rendered pass carries an h1 the raw HTML does not have, "Something went wrong while displaying this content. Refresh", so a parser keyed on the h1 files an error message as the page title. Prices come from the appdetails JSON endpoint in 172 bytes with cc=us.

### When do I still need a collector rather than the plain fetch?

When the product is a table, not a page. On 28 September the collectors returned 10 Idealista listings in 5.51 seconds, 10 Indeed jobs in 7.3, 20 Zillow listings in 3.02, 30 Tripadvisor restaurants in 8.1, 50 DoorDash restaurants in 5.98, 20 Reddit comments in 7.9 and 10 Steam apps in 1.83, each billed per delivered row and never for a failure. The plain fetch gives you the page; the collector gives you the rows, and neither needs a browser.

### Is it against the rules to read these sites without a browser?

Reading a public page logged out is what the sites serve to anyone, and the robots files are short: Steam closes 10 paths in 303 bytes, Zillow publishes 15 sitemaps in 10,524 bytes. What the terms forbid is different: Steam, DraftKings, Etsy and Vinted all forbid automated access by account holders, and where a site says that, we read logged out and stop there. No setting in this post logs in, evades a limit or automates an account.

## Measure your target before you pay for a browser

One seo_audit call fetches any URL twice, plain and rendered, and returns both word counts. Every account gets $2 of free API usage per month, and failed requests are never billed.

[Start free — $2/month included](https://quanticdata.io/signup/)[Explore Residential Proxies from $0.80/GB](https://quanticdata.io/residential-proxies/)

## Related reading

[Proxies Golang HTTP Client Proxy: Auth, SOCKS5 How to attach a proxy to an http.Client in Go, authenticate over CONNECT, use SOCKS5 from the standard library and rotate exits per request, with a measurement of how far a no-JavaScript client actually gets. Read →](https://quanticdata.io/blog/golang-http-client-proxy/) [Proxies Follower Counts Are in the Head: 5 Platforms Five platforms print the follower count before any script runs: Instagram and LinkedIn in the meta description, Facebook in og:description, YouTube and Twitch in JSON-LD. We measured eleven targets with and without a browser through residential exits. On these five the rendered pass adds bytes and no number. TikTok and Spotify are the exceptions, and the tiktok_profile collector returns 2 profiles in 14.4 seconds. Read →](https://quanticdata.io/blog/follower-counts-in-the-head/) [Proxies Geo-Targeted Proxies: 987 Words UK vs 149 DE The same URL, rendered through residential exits: 987 words from the United Kingdom, 149 from Germany, 2,031 from the United States. Nothing was translated; each exit landed on a different licensed product. On 18 of the 53 sites we measured, geo-targeted proxies change the page itself: the domain on AliExpress, the entity on OKX, the price line on Netflix, the number format on Facebook and LinkedIn, and whether the channel page exists at all on YouTube. Pin the exit per job or you are comparing two different pages. Read →](https://quanticdata.io/blog/same-url-different-country-different-page/)

## Also on this site

Quantic**Data**

Residential proxies & web data APIs for AI.

#### Proxies

- [Residential Basic](https://quanticdata.io/residential-proxies/#basic)

- [Residential Premium](https://quanticdata.io/residential-proxies/#plans)

- [Cheap Residential](https://quanticdata.io/cheap-residential-proxies/)

- [Mobile Proxies](https://quanticdata.io/mobile-proxies/)

- [Datacenter Proxies](https://quanticdata.io/datacenter-proxies/)

- [ISP Proxies](https://quanticdata.io/isp-proxies/)

- [Rotating Proxies](https://quanticdata.io/rotating-proxies/)

- [Sneaker Proxies](https://quanticdata.io/sneaker-proxies/)

- [SOCKS5 Proxies](https://quanticdata.io/socks5-proxies/)

- [IPv6 Proxies](https://quanticdata.io/ipv6-proxies/)

- [Proxy locations](https://quanticdata.io/proxies/)

#### Data APIs

- [MCP Server](https://quanticdata.io/mcp-server/)

- [Web Scraper API](https://quanticdata.io/web-scraping-api/)

- [SERP API](https://quanticdata.io/serp-api/)

- [Collectors](https://quanticdata.io/collectors/)

- [Web Data for AI](https://quanticdata.io/web-data-api-for-ai/)

- [Quantic AI](https://quanticdata.io/ai-web-scraping-service/)

- [Crawl & Map](https://quanticdata.io/crawl-map/)

- [SEO Audit](https://quanticdata.io/seo-audit/)

#### Use cases

- [Company data](https://quanticdata.io/scrape-company-data/)

- [Price monitoring](https://quanticdata.io/competitor-price-monitoring/)

- [Market research](https://quanticdata.io/market-research-data/)

- [Real estate data](https://quanticdata.io/real-estate-data-scraping/)

- [Scrape job postings](https://quanticdata.io/scrape-job-postings/)

#### Company

- [Documentation](https://quanticdata.io/docs/)

- [Blog](https://quanticdata.io/blog/)

- [Free tools](https://quanticdata.io/tools/)

- [Partners](https://quanticdata.io/partners/)

- [About](https://quanticdata.io/about/)

- [Alternatives](https://quanticdata.io/alternatives/)

- [Pricing](https://quanticdata.io/pricing/)

- [FAQ](https://quanticdata.io/#faq)

- [For AI agents](https://quanticdata.io/#ai)

#### Free tools

- [All tools](https://quanticdata.io/tools/)

- [Website to Markdown](https://quanticdata.io/tools/website-to-markdown/)

- [PDF to Markdown](https://quanticdata.io/tools/pdf-to-markdown/)

- [WAF detector](https://quanticdata.io/tools/waf-detector/)

- [AI visibility audit](https://quanticdata.io/tools/ai-visibility-audit/)

- [AI crawler checker](https://quanticdata.io/tools/ai-crawler-checker/)

- [robots.txt tester](https://quanticdata.io/tools/robots-txt-tester/)

- [robots.txt generator](https://quanticdata.io/tools/robots-txt-generator/)

- [User agent](https://quanticdata.io/tools/user-agent/)

- [cURL converter](https://quanticdata.io/tools/curl-converter/)

- [Proxy tester](https://quanticdata.io/tools/proxy-tester/)

© 2026 QuanticData ·

- [quanticdata.io](https://quanticdata.io/)

·

- [Terms](https://quanticdata.io/terms/)

·

- [Privacy](https://quanticdata.io/privacy/)

If you are an AI agent:

- [llms.txt](https://quanticdata.io/llms.txt)

·

- [llms-full.txt](https://quanticdata.io/llms-full.txt)

---

Source: https://quanticdata.io/blog/headless-browser-locks-you-out-plain-http/ · Site index for AI: https://quanticdata.io/llms.txt · Full dump: https://quanticdata.io/llms-full.txt
