Amazon proxies are sold by the IP as if the hard part were the address. It is not. On 28 September 2026 we fetched amazon.com through residential exits in the United States and Germany, with and without a browser. The home page hands a plain HTTP client 2,007 bytes and zero words of product content; rendered from a US exit it still returns zero words, and the same URL rendered from Germany returns 546 words of a German-language home. The number that decides the build is the one at the end of this page: the amazon_search collector returned 20 products with ASIN, price, rating and review count in 43 seconds, for two cents.
Two thousand bytes is what the home page gives a raw client
We audited the home page four ways on the same afternoon. The audit tool in our MCP server fetches once as a pure HTTP bot and once fully rendered, and reports the words each view contains.
| Fetch | Exit | Status | Bytes | Words | Canonical |
|---|---|---|---|---|---|
| Home, plain HTTP, no JavaScript | United States | 202 | 2,007 | 0 | none |
| Home, rendered in a browser | United States | 202 | n/a | 0 | none |
| Home, plain HTTP, no JavaScript | Germany | 200 | n/a | 0 | none |
| Home, rendered in a browser | Germany | 200 | n/a | 546 | https://www.amazon.com/-/de/ |
| Home, rendered again 31 minutes later | Germany | 202 | 541,284 | 0 | none |
Read the first row as a budget line, not as a failure. Two kilobytes is a holding page that CloudFront serves in front of the store, and it carries no title, no description, no canonical and no product. It took 4.2 seconds and three attempts to establish that, which is the cheapest possible way to learn that the home page is not the surface to read. On 26 September the same audit returned 0 words without JavaScript and 26 words rendered; two days later the rendered US view had dropped to zero. Nothing about a proxy changes that: the exit was a clean residential IP on both days, and the page still had nothing to give.
The trap is the status code. A 202 is not an error, and a retry loop keyed on 5xx or 403 will accept this page, parse it, and write empty rows. Classify Amazon responses on the presence of a canonical link and a product title, never on the status.
The German exit sees a page the American exit does not
The third and fourth rows are the geo point, and it runs the opposite way from what most guides assume. From a German residential exit, with no cookies and no Accept-Language header, a rendered fetch of amazon.com returned a full home: title "Amazon.com. Spend less. Smile more.", a 350-character meta description, 546 words, and a canonical of https://www.amazon.com/-/de/, which is Amazon's German-language view of the US store. From the United States the same rendered fetch returned nothing.
So the exit country changes the page itself: a German exit is handed the /-/de/ view, in German, and any parser keyed on English labels reads it wrong. It does not turn the home into a data surface. A second rendered fetch from the same German pool 31 minutes later weighed 541,284 bytes and contained zero words, so the rendered home is at best intermittent and at worst half a megabyte of nothing. If your job is to see what a German shopper sees on amazon.com, pin a German exit and expect the /-/de/ canonical. If your job is US prices in US dollars, the home page is not where they are, and the next section is.
Twenty products in 43 seconds is the number that matters
We ran the amazon_search collector once, with the input {"query":"wireless earbuds","country":"us"}. It started at 16:57:20 UTC and finished at 16:58:04: 20 products in 43.3 seconds, delivered as rows, not as HTML.
| Field | Coverage in the run | Example |
|---|---|---|
| asin, url, rank, page | 20 of 20 | B0FQFB8FMG, rank 6, page 1 |
| price, price_value | 18 of 20 | $234.00 → 234 |
| list_price | 15 of 20 | $249.00 on the $234.00 row |
| rating, reviews | 20 of 20 | 4.4 stars, 15,600 reviews |
| sponsored | 5 of 20 flagged true | ranks 1, 2, 7, 12, 17 |
| image | 20 of 20 | m.media-amazon.com thumbnail |
Three things in that output are worth more than the fact that it worked. First, the two rows without a price (ranks 8 and 19) are real products whose price is behind a variant picker on the product page, so a null here means "open the product", not "missing". Second, the list_price column is marketing, not history: rank 4 shows a list price of $299.99 against a price of $23.98, and rank 12 shows a list price of $10.00 against a price of $19.99, which is lower than the price. Any discount metric computed from that column is fiction unless you filter list prices below the sale price. Third, the first two results are sponsored and the fifth sponsored row sits at rank 17; if you are measuring organic rank, drop the flagged rows before you count.
The prices ran from $13.99 to $234.00 and the review counts from 3 to 117,500 on the same page, which is the spread you get when a search term mixes accessory sellers with Apple. For one product's full record, the amazon_product collector reads the detail page, and amazon_reviews reads the reviews behind the count.
Amazon's robots.txt names 100 bots and blocks every one
amazon.com/robots.txt is 7,887 bytes and was last modified on 15 May 2026. The wildcard block is long but ordinary: sign-in, cart, wishlist, the buy box, review submission, offer listings. After it come 100 named user-agent blocks, each with a single line, Disallow: /. GPTBot, ClaudeBot (listed twice), PerplexityBot, CCBot, Google-Extended, Bytespider, Scrapy, Crawl4AI, DeepSeekBot, GrokBot and Claude-User are all there. The file allows Googlebot the store and denies every AI and scraping agent it can name.
The Conditions of Use, last updated 14 August 2026, are the contract behind that file. The licence to use Amazon Services excludes "any collection and use of any product listings, descriptions, or prices" and "any use of data mining, robots, or similar data gathering and extraction tools", and adds that AI-generated content from the services may not be used to train models. That is one plain line and we will leave it at that: Amazon does not permit unlicensed automated collection of its listings.
What Amazon does license is worth knowing before you build. The Product Advertising API 5.0 is now deprecated; its documentation URL redirects to a notice that applications still calling it receive a 403 with an AccessDeniedException, and points to the Creators API as the successor for publishers and affiliates. Sellers and their tools use the Selling Partner API. Both are authenticated, rate-limited and tied to a programme membership, and both are the sanctioned route to catalogue data. For the reading of public prices that the courts in the United States have not treated as intrusion, our note on whether web scraping is legal in the US covers the ground; the contract layer above is Amazon's, not ours.
What a gigabyte buys on Amazon
Prices from our pricing page: residential Basic $0.80/GB, the web scraping API from $0.0002 per page ($0.001 with rendering), and the amazon_search collector at $0.001 per delivered product. A gigabyte here is 10^9 bytes.
| Approach | Bytes per request | Requests per GB | What you get | Cost per 1,000 products |
|---|---|---|---|---|
| Home, plain HTTP, residential US | 2,007 | 498,256 | 0 products | undefined: no products |
| Home, rendered, residential DE | 541,284 | 1,847 | 546 words once, 0 words once, 0 prices | undefined: no products |
| amazon_search collector | n/a | n/a | 20 products per run | $1.00 |
| amazon_product collector, one detail page each | n/a | n/a | full record | $8.00 |
The first row is the one that empties budgets. Half a million fetches per gigabyte sounds like a bargain until you notice the column that says zero products. On Amazon the cost driver is not bandwidth, it is delivered rows, and a run that delivers nothing is billed nothing on our side: failed requests are never charged. The 20-product run above cost $0.02. If you are budgeting a price-watch across a catalogue, how to price watch on Amazon works the arithmetic per SKU, and eBay and Walmart are the neighbouring measurements if you are comparing marketplaces.
The setting that works on Amazon
- Network: residential, Basic line at $0.80/GB. Every fetch in this post went through it, and the exit was never the problem; the surface was.
- Fetch mode: neither plain HTTP nor a browser on the home page (0 and 0 words from the United States). Read search and product pages through the collector, which returns rows.
- Country: pin
country: "us"for USD prices and English labels. A German exit returns a different page at /-/de/ with 546 words in German, and a parser built for one silently misreads the other. - When the proxy is not enough: it is not enough here. Use amazon_search: 20 products in 43 seconds, $0.001 per delivered product, 18 of 20 with a price, 5 of 20 flagged sponsored. For the full record of one item, amazon_product at $0.008.
- Free tier: every account gets $2 of free API usage per month, which is 2,000 delivered search rows at this collector's price.