Etsy proxies have the strangest search results of any retailer we have measured: every result for the keyword is a playing card. The data question hides under "scrape etsy data", and the answer to it is three words long. On 28 September 2026 a plain HTTP client behind a United States residential exit received the etsy.com home in full: 200 OK, 205,911 bytes, 678 words, title, canonical and JSON-LD. The same client from Germany received 8 words, from the United Kingdom 8 words, and a headless browser from the United States received none. Residential, plain HTTP, US. Everything else on this page is why.
Plain HTTP from a US exit is the whole answer
We audited the home four ways. The audit tool in our MCP server fetches once as a pure HTTP bot and once fully rendered, and we repeated the HTTP pass from two more countries.
| Fetch | Exit | Page | Bytes | Words | Title, canonical, JSON-LD |
|---|---|---|---|---|---|
| Plain HTTP, no JavaScript | United States | full home | 205,911 | 678 | all present; WebSite and Organization |
| Rendered in a browser | United States | none | n/a | 0 | none |
| Plain HTTP, no JavaScript | Germany | 8-word stub | n/a | 8 | none |
| Plain HTTP, no JavaScript | United Kingdom | 8-word stub | n/a | 8 | none |
| Plain HTTP, no JavaScript (26 September row) | United States | full home | n/a | 707 | all present |
Two days, two US fetches, 707 and 678 words: the page is stable and it is a real page, with the meta description "Find the perfect handmade gift, vintage & on-trend clothes, unique jewelry, and more… lots more", a canonical of https://www.etsy.com/ and two JSON-LD blocks. It arrived in 28.8 seconds with a Firefox TLS profile, after the client had rotated its fingerprint once, so a client that can present more than one browser fingerprint finishes the job where a fixed one does not.
The successful response carries the header x-datadome: protected, and it is worth reading that header on a 200 rather than on a refusal: it tells you who is scoring your request even when the score was good enough. Etsy sits behind Fastly (the via: 1.1 varnish hop) with DataDome grading each caller, and the grade is per client and per exit, not per URL. That is why the exit country, and not the path, is the variable in the table above.
The browser is the one client Etsy turns away
Row two is the row to save money on. A headless browser from the same US exit that got 678 words over plain HTTP got none, on 26 September and again on 28 September. This is the pattern we call harmful rendering: the fetch mode that costs the most is the one the site refuses. Of the 52 targets we measured in September, Etsy is one of seven where the browser makes the result worse rather than better, and it is the cleanest case, because the HTTP page is complete on its own.
The practical rule follows: set the fetch mode to engine: tls and never let an auto-escalation reach for the browser after a thin-looking 200. The HTTP page is not thin. It has the category grid, the seasonal blocks and the structured data. Listing and shop pages, which the vendor tutorials on the second SERP agree embed schema.org Product JSON-LD, are the same shape: server-rendered HTML that a plain client can parse without executing anything. If your parser needs a DOM, it needs it for your convenience, not for Etsy's markup.
The exit country decides whether you get a page at all
The other two rows are the geo point, and on Etsy it is not about language or currency. From Germany and from the United Kingdom the identical plain HTTP client received an 8-word page with no title, no canonical and no structured data. From the United States it received 678 words. Same URL, same TLS profile, same headers, three exits, one page. Etsy's marketplace is global, and every listing has a currency that follows the visitor, but the front door is graded per exit, and in our measurement only the US exit was let through.
So pin the country. country: "us" is not a preference here, it is the difference between 678 words and 8. If your job is to see prices in euros as a German shopper would, that is a second-order problem you solve after the fetch, by reading the currency that Etsy prints; the top-ranking scraping guide for this platform found 19 currencies in one sample and warns that an unfiltered average of that column was 227 times too high. Get the page from the US exit first, then normalise the currency you find in it.
Etsy's robots.txt has 1,681 Disallow lines and no sitemap
etsy.com/robots.txt is 52,811 bytes across 1,822 lines, and it is the longest robots file of the six retailers in this series. It declares three user-agent groups: Spinn3r is disallowed everywhere, AdsBot-Google-Mobile is kept off /people in every locale, and the wildcard gets 1,681 Disallow lines. There is no Sitemap directive at all, and no AI crawler is named. The wildcard rules matter for a data team: keyword search is closed (/search?*q= and its locale variants), as are price-bucket and attribute filters, sold histories under /shop/*/sold, favoriters, and the internal API. Listing and shop pages are not in the file.
The Terms of Use are stricter than the file, and section 6C says it in one sentence: you agree not to crawl, scrape or spider any page of the Services without Etsy's express permission, and if you want to use the API, follow the API Terms of Use. That is the line, and we leave it there. What Etsy licenses is the Open API v3, which is keyed per application with a Queries Per Day and a Queries Per Second limit set in the Developer Portal, enforced on a sliding 24-hour window, with a stated prohibition on creating additional keys to raise the limit. Its documentation now ships a developer MCP server of its own. For your own shop, and for the 31 public GET operations the tutorials count, that is the sanctioned route.
What a gigabyte buys on Etsy
Prices from our pricing page: residential Basic $0.80/GB, and the web scraping API from $0.0002 per page, or $0.001 with rendering. A gigabyte here is 10^9 bytes.
| Approach | Bytes per page | Pages per GB | Words per page | Cost per 1,000 pages |
|---|---|---|---|---|
| Plain HTTP, residential US, $0.80/GB | 205,911 | 4,856 | 678 | $0.16 |
| Rendered, residential US | any | any | 0 | undefined: no page |
| Plain HTTP, residential DE or GB | any | any | 8 | undefined: no page |
| Web scraping API, engine tls, country us | n/a | n/a | 678 | $0.20 |
| Web scraping API with rendering | n/a | n/a | 0 | $1.00 for nothing |
Sixteen cents per thousand home pages over raw residential bandwidth, twenty cents through the API, and the API's twenty cents buys the fingerprint rotation that the first row needed. The bottom row is the one to strike from every budget: paying five times the price to render is paying five times the price for zero words. If you are sizing the pool for a catalogue crawl, how much proxy data you need does the same arithmetic in the other direction, and the neighbouring marketplaces behave very differently: Amazon and eBay return nothing to a raw client and everything to a collector.
There is no Etsy collector in our catalogue. For a cross-marketplace price check on the same product, google_shopping returns priced listings per query, and for the Etsy page itself the web scraping API with engine: tls and country: us is the job.
The setting that works on Etsy
- Network: residential, Basic line at $0.80/GB. Every successful fetch in this post went through it, at 205,911 bytes per home page.
- Fetch mode:
engine: tls, never rendered. 678 words over plain HTTP against 0 words rendered from the same exit, on two different days. - Country:
country: "us". From Germany and the United Kingdom the identical client received 8 words; from the United States, 678. - When the proxy is not enough: use the web scraping API with
engine: tls, which retries with a second browser TLS profile the way our 28.8-second fetch did, at $0.0002 per page. Do not add rendering: it costs $0.001 and returns nothing on this site. - Free tier: every account gets $2 of free API usage per month, which is 10,000 Etsy pages over plain HTTP at the API price.