Web scraping API vs proxy is a question with a measured answer on eight of the 53 sites we audited, and the answer is neither fetch mode. On Amazon, eBay, Booking.com, Skyscanner, Airbnb, Expedia, Taobao and Spotify the home page a residential exit fetches holds between 0 and 72 words, with or without a browser. On 28 September 2026 the ebay_search collector returned 25 priced road-bike listings in 19.5 seconds and booking_stays returned 20 priced Rome properties in 5.1 seconds. On these eight sites the collector is the setting; the proxy is what it runs on.
The search results compare products; nobody weighs a response
The first page for "web scraping api vs proxy" in the United States is a Reddit thread asking the question, six vendor posts answering it in the abstract, and one Medium essay about hidden proxy costs. The vendor pieces agree on the frame: proxies give you IP addresses and you build the rest, a scraping API gives you data and bills per success. That frame is right. What none of them do is fetch a page and count what came back, so the reader leaves knowing the pricing models and not knowing which sites need which.
We have that number for 53 sites, measured between 26 and 28 September 2026 through residential proxies, each one fetched as a plain HTTP client and again fully rendered. The full table is in 53 targets measured: browser or not. This post takes the eight rows where both fetch modes return close to nothing and shows what the collector returns instead, with today's runs.
Eight sites, two fetch modes, and the rows a collector delivers
Words are the word count of extractable body text reported by the audit. Bytes are the on-wire weight of the plain fetch. Collector rows and seconds are from runs we made ourselves; the eBay and Booking.com runs are from 28 September 2026 at 21:21 and 21:23 UTC, the rest from the per-target posts linked in the first column.
| Target | Exit | Plain HTTP, words | Rendered, words | Verdict | The setting |
|---|---|---|---|---|---|
| eBay | US | 24 (1,976 bytes) | 24 | Collector | ebay_search: 25 listings in 19.5 s, $0.001 per listing |
| Booking.com | IT | 26 (3,962 bytes) | 0 | Collector | booking_stays: 20 priced properties in 5.1 s, $0.02 per property |
| Amazon | US | 0 (2,007 bytes) | 0 | Collector | amazon_search: 20 products in 43 s, $0.001 per product |
| Skyscanner | GB | 0 (265,657 bytes) | 0 | Collector | google_flights: 15 LHR to JFK fares in 33.7 s, $0.003 per itinerary |
| Airbnb | US | 59 (597,053 bytes) | 47 | Collector | airbnb_stays: 18 stays with nightly price in 5.3 s, $0.001 per stay |
| Expedia | US | 72 (507,949 bytes) | 14 | Plain HTTP for structure, collector for rates | hotels: 20 priced properties in 160 s, $0.002 each; google_flights: 15 fares in 19 s |
| Taobao | SG | 36 (157,657 bytes) | 56 | No collector | Plain HTTP on item and guide pages: 410 words in 49,272 bytes through the web scraping API |
| Spotify | US | 0 (164,388 bytes) | 14 | No collector | The Spotify Web API with a client credentials token; plain HTTP only on spotify.com price pages |
| TikTok | US | 5 | 12 | Collector | The TikTok profile collector; nothing usable in the head or the DOM |
| StubHub | US | 0 (226,417 bytes) | 157 | Render through the API | Web scraping API with rendering at $0.001 per page; no collector by design |
| Trustpilot | US | 19 (991 bytes) | 3,588 | Browser, then collector for reviews | Rendered with a 6 s wait; trustpilot_reviews: 10 reviews in 14.8 s, $0.002 each |
| Binance | DE | 26 (2,108 bytes) | 1,196 | Public API | data-api.binance.vision over plain HTTP, 558 bytes per ticker; coingecko_coins: 50 coins in 14.8 s |
Read the first three rows together. eBay, Booking.com and Amazon hand a plain client a page under 4,000 bytes, and the browser does not improve on it: 24, 0 and 0 words. These are not thin sites. They are sites that do not serve their catalogue to an anonymous first request in any mode we can buy bandwidth for. The rows in the last column are the only rows that exist.
eBay: 24 words from any country, 25 listings in 19.5 seconds
We weighed ebay.com/ again at 21:22 UTC on 28 September through a United States residential exit: 1,976 bytes, the title "Error Page | eBay", no canonical, no description, 24 words. The eBay post got the identical 24 words from Germany and from a rendered fetch two days earlier. The exit country changes nothing on this page, which is the useful signal: a better IP is not the missing piece.
One minute earlier we ran the ebay_search collector with the input {"query":"road bike"}. It started at 21:21:25 UTC and finished at 21:21:45: 25 listings in 19.5 seconds, at $0.001 per delivered listing. Every row carried an item id, a URL, a price, a format, a seller location and an image; 24 of 25 carried a condition, 3 carried a struck-through list price, and the prices ran from $259.00 to $11,500.00 across sellers in seven countries. All 25 rows carried the sponsored flag on this run, which is a fact about the first page of an eBay search for a generic term and a reason to filter on that field before you average anything.
The run's notes field came back with one line, the same one the 28 September afternoon run wrote:
US exit pool throttled; served through a DE exit - same marketplace and listings, but eBay prices a visitor in the exit's currency and shipping hints follow it
That sentence is the difference between a proxy and a collector in one place. A proxy gives you the exit. The collector chose a working exit, delivered the rows, and told you the side effect so you can normalise the currency. The afternoon run took 185 seconds for the same 25 rows; this one took 19.5. Both are billed on 25 delivered listings, so the slow run and the fast run cost the same two and a half cents.
Booking.com: 26 words in both modes, 20 priced properties in 5.1 seconds
booking.com/ over plain HTTP from an Italian residential exit at 21:22 UTC: 3,962 bytes, no title, no canonical, 26 words. The Booking.com post rendered the same URL for 857,992 bytes and got the head of the page and zero words of body. Two hundred and sixteen times the bytes buys a title tag.
The booking_stays collector, run at 21:23 UTC with the input {"location":"Roma, Italia","check_in_date":"2026-10-12","check_out_date":"2026-10-14","currency":"EUR","country":"it"}, resolved the place to 41.89671, 12.48220 and returned 20 properties in 5.1 seconds. Each row has a nightly price for those exact dates in the currency asked for, the struck-through price where one exists, the discount it implies, review score and count, star class, room type, coordinates, distance from the searched point and a photo. On this run: prices from 402 to 3,103 euros for the two nights, 12 of 20 with a struck-through price, discounts up to 34 percent, 4 of 20 with free cancellation, review counts from 4 to 1,939, distances from 20 to 250 metres.
The two inputs that matter are the ones a raw fetch cannot pin. Booking.com decides currency and language from the exit address, and a search page is a different URL for every date pair. The collector takes country and currency explicitly, so a comp set built for a hotel on Via del Corso is priced in euros for Italian visitors on purpose, not because the pool happened to land in Milan.
Amazon, Skyscanner, Airbnb and Expedia are the same shape at four sizes
Amazon is the smallest: 2,007 bytes and 0 words over plain HTTP from the United States, 0 words rendered. The amazon_search collector returned 20 products with ASIN, price, rating and review count in 43 seconds, 18 of 20 priced and 5 flagged sponsored, at $0.001 per product.
Skyscanner is the largest empty page: 265,657 bytes of head, title, description and a WebSite JSON-LD block, and not one word of body, in either mode, from the United Kingdom or the United States. The google_flights collector returned 15 London Heathrow to New York JFK itineraries priced in pounds in 33.7 seconds.
Airbnb gives a plain client a readable head and 59 words, and the render gives 47 for 1,561,460 bytes. The airbnb_stays collector returned 18 stays with a nightly price in 5.3 seconds. Expedia is the one where plain HTTP is worth keeping: 72 words with the h1 and the canonical in 4.1 seconds from a United States exit, which is a structure check, not a price. Rates come from the hotels collector, 20 priced properties in 160 seconds, and from google_flights, 15 fares in 19 seconds, because the search URLs where prices live are exactly the ones Expedia's robots.txt fences off: disallow: /Hotel-Search, disallow: /Flights-Search, disallow: /*chkin=*.
Taobao and Spotify: no collector, and what wins instead
Two of the eight have no collector in our catalogue, and the honest setting is different for each. On Taobao the home page is 36 words plain and 56 rendered, but the item and guide pages under /lang/ return 410 words and four prices in 49,272 bytes over plain HTTP. There the web scraping API without rendering is the tool: $0.0002 per page, Markdown out, no browser.
On Spotify, open.spotify.com returns 0 words plain and 14 rendered for 2,143,243 bytes, and no proxy setting changes that. The catalogue is served by the Spotify Web API under a client credentials token, with followers, popularity and genres per artist, and the Developer Terms rule a scraper out. The proxy earns its keep only on the public spotify.com price pages, 217 words over plain HTTP, one exit per market.
What a gigabyte buys on these sites, and what a collector costs
Every fetch above went through residential Basic at $0.80/GB, counting a gigabyte as 10^9 bytes. The left half of the table is what the bandwidth buys; the right half is what the rows cost.
| Target | Raw fetch, bytes | Cost of that fetch | Words it returns | Collector | Cost for the rows above |
|---|---|---|---|---|---|
| eBay home, plain | 1,976 | $0.0000016 | 24 | ebay_search, 25 listings | $0.025 |
| Booking.com home, rendered | 857,992 | $0.00069 | 0 | booking_stays, 20 properties | $0.40 |
| Amazon home, plain | 2,007 | $0.0000016 | 0 | amazon_search, 20 products | $0.02 |
| Skyscanner home, rendered | 408,075 | $0.00033 | 0 | google_flights, 15 fares | $0.045 |
| Airbnb home, rendered | 1,561,460 | $0.00125 | 47 | airbnb_stays, 18 stays | $0.018 |
| Expedia home, plain | 507,949 | $0.00041 | 72 | hotels, 20 properties | $0.04 |
| Spotify home, rendered | 2,143,243 | $0.00171 | 14 | Web API, client credentials | no proxy cost |
| Taobao item page, plain | 49,272 | $0.00004 | 410 | none needed | $0.0002 per page via the API |
The first row is the trap in numbers. A gigabyte of eBay home pages is 506,000 fetches and 0 listings; it is the cheapest bandwidth on the internet and the most expensive data, because the price per row is infinite. Twenty-five cents buys 250 priced eBay listings from the collector, or 12 Booking.com properties with dates and discounts, or 250 Amazon products with ratings. Budget by what a request yields, not by what it weighs; the same arithmetic in the other direction is in how much proxy data you need.
Where we stop
Everything above is public listing, price and availability data: the rows a logged-out visitor sees, read for price monitoring, comp sets, market research and catalogue matching. Three of the eight publish rules worth quoting. eBay's robots.txt (24,164 bytes, version 30.2, August 2026) states that automated access without eBay's express permission is prohibited and adds a clause aimed at agents: checkouts are strictly for human users, and buy-for-me or LLM-driven flows that place orders are not permitted. Spotify's Developer Terms cover catalogue access through the Web API and nothing else. Booking.com routes partners to its developers portal. We do not build around any of those lines, and the collectors read search results, not accounts, carts or checkouts.
The setting that works on these eight targets
- Network: residential, Basic line at $0.80/GB, underneath every collector run and every plain fetch in this post. Nothing in these measurements scored the address; Premium at $2.20/GB and mobile at $2.30/GB buy nothing we could count here.
- Fetch mode: neither, for the home and search pages of eBay, Booking.com, Amazon, Skyscanner, Airbnb, Spotify and TikTok, which return 0 to 59 words in both modes.
engine: tlson Expedia (72 words, h1 and canonical) and on Taobao item pages (410 words). Never a render on any of the eight: the best it returned was 157 words on StubHub for 3,238,143 bytes. - Country: the market you are pricing, passed to the collector as
countryandcurrency. eBay prices the visitor in the exit's currency and its run notes tell you which exit served you; Booking.com picks language and currency from the address; Amazon from Germany serves a different page at /-/de/. - When the proxy is not enough: it is not enough on these sites, which is the point. ebay_search, 25 listings in 19.5 s; booking_stays, 20 properties in 5.1 s; amazon_search, 20 products in 43 s; google_flights, 15 fares in 33.7 s; airbnb_stays, 18 stays in 5.3 s; hotels, 20 properties in 160 s. Where no collector exists, the web scraping API without rendering at $0.0002 per page, or the platform's own API.
- Failed requests are never billed, and every account gets $2 of free API usage per month, which is 2,000 eBay listings or 100 Booking.com properties before anything is charged.
The other 45 rows, including the 15 sites that do want a browser and the 26 that do worse with one, are in 53 targets measured: browser or not. For the four reply shapes that make people reach for a browser when they should not, read read the block and what to do.