Expedia is the site where the browser returns less than the plain fetch. On 28 September 2026 we fetched the home page through a United States residential exit as a pure HTTP client and received 72 words, the h1 "The one place you go to go places", a self-referencing canonical, a description and a robots meta of index,follow, for 507,949 bytes in 4.1 seconds. The fully rendered fetch of the same URL returned 14 words. From a United Kingdom exit, the plain fetch returned 14 words as well, with no canonical and no h1. The setting for Expedia proxies is residential, engine: tls, country pinned to the United States for page structure, and the hotels and flights collectors for rates, because the search URLs where prices live are the ones robots.txt fences off.
What the SERP teaches, and the one thing it misses
Nobody types "expedia proxies" into Google; autocomplete returns nothing for it and nothing for "proxies for expedia". The demand is under "expedia data scraping" and, from buyers, "expedia hotel price tracking", "price tracker", "price alert" and "price history". The first page for the scraping term is vendor pages, two tutorials, a GitHub scraper published by a vendor and a Reddit thread reporting that listing items do not render in headless mode.
The top tutorial is instructive in a way its author did not intend. It opens Selenium, loads a hotel page, finds that the prices are missing, and fixes that by adding request headers: User-Agent, Accept, Accept-Encoding, Referer. It never tries the obvious control, which is to send those headers without the browser. We did, and it is the whole post.
72 words over plain HTTP, 14 words rendered
We audited expedia.com/ from a United States residential exit as a pure HTTP client and then fully rendered, and weighed each response.
| Fetch | Exit | Bytes | Seconds | Words | h1 | Canonical |
|---|---|---|---|---|---|---|
| Plain HTTP | United States | 507,949 | 4.1 | 72 | The one place you go to go places | expedia.com/ |
| Rendered | United States | 383,086 | 50.9 | 14 | none | none |
| Plain HTTP | United Kingdom | n/a | n/a | 14 | none | none |
| Rendered | United Kingdom | n/a | n/a | 14 | none | none |
The first row is the working configuration and it is complete: title, description, canonical, h1, the navigation and the promotional copy of the day, which on our fetch was a members' sale banner. Seventy-two words is thin because the home page is a search box, not because anything was withheld; the words that exist are all there, in 4.1 seconds, and the response carries a content-language: en-US header that tells you which storefront you got.
The second row is the one to stop paying for. Rendering the same URL took 50.9 seconds and returned a 14-word page with no h1, no canonical and no description. On Expedia the headless browser is not a slower path to the same data; it is a path to less data. Every tutorial on the SERP starts there. Do not.
There is no JSON-LD on the home page in either view, which matters for the pipeline: unlike Kayak, which ships three structured blocks in the source, Expedia gives you markup and headers only. Parse the h1 and the canonical, and treat the presence of the h1 as your health check, because the 14-word responses have none.
Pin the exit to the United States: the same URL returns 14 words from the UK
We repeated the audit from a United Kingdom residential exit, sending no Accept-Language header. Over plain HTTP the same URL returned 14 words, no canonical and no h1; rendered, the same 14 words. The United States exit received the page; the United Kingdom exit received a stub, in both modes.
This is the number that changes when you do not pin the country, and it is not a localisation story like Kayak's redirect to kayak.co.uk. It is a whole page versus almost none. expedia.com is the United States storefront and it serves United States visitors; the other markets have their own domains, and the robots file even disallows /en-au/ and /en-nz/ paths on this one. A pool that rotates through mixed countries will return a working page from some exits and a 14-word stub from others, and a retry loop keyed on status alone will not distinguish them. Key it on the h1. With country: us on every request, the h1 was present on every plain fetch we made.
robots.txt fences off exactly the pages with prices
expedia.com/robots.txt is small, 2,210 bytes and 109 lines, with four user-agent blocks and no Sitemap line at all. The wildcard block is 52 disallow lines and reads like an inventory of where the money is: /Hotel-Search, /hotel-search, /search?, /*/search?, /*chkin=*, /*chkout=*, /Flights-Search, /Flight-SearchResults, /carsearch, /things-to-do/search?, plus every checkout and confirmation path. Any URL with check-in and check-out dates in the query string is disallowed for a generic client. The hotel and destination pages that carry no dates are not listed.
The other three blocks are short. Google's hotel-ads and ads verifiers get 6 disallow lines and an allow. SemrushBot gets Disallow: /. And a block naming OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-User and Claude-SearchBot gets 6 disallow lines of its own, which is Expedia deciding, by name, how far the AI search crawlers may go. We have written about that pattern in the AI crawler user-agent list.
The terms say the rest in one list. Section 2 of the Terms of Service asks you not to access, monitor or copy any content on the service using any robot, spider, scraper or other automated means or any manual process, not to violate the restrictions in any robot exclusion headers, not to impose an unreasonable load, and not to deep link. That is the line, and it is why the rest of this post routes rates through collectors that do not touch a disallowed Expedia URL. For a licensed integration, the Expedia Group Developer Hub documents the Rapid API for lodging, a White Label Travel Platform, a Travel Redirect API and an analytics product, all for approved partners.
Rates come from the collectors, not from the search page
Price monitoring on Expedia means dated searches, and dated searches are the disallowed URLs. So for the prices we ran two collectors on the same afternoon, both through United States exits, neither of which fetches Expedia at all.
The hotels collector, input New York, NY, check-in 20 October, check-out 22 October, two adults, currency USD, up to 20 results, returned 20 properties in 160 seconds: nightly rates from $82 to $263, two-night totals from $195 to $611, rating and review count on every row, and a link to the stay. That is the metasearch view of the same market Expedia sells into, with the same dates, and every property row costs $0.02.
The google_flights collector, input JFK to LAX on 20 October, currency USD, returned 15 itineraries in 19.4 seconds: prices from $204 to $294, 14 nonstop, four airlines, with departure and arrival times, duration and a CO2 estimate per row, at $0.003 per itinerary. For fare tracking that is the whole job, and it does not depend on any one seller's markup.
What a gigabyte buys on Expedia
Every number above came through residential proxies on the Basic line at $0.80/GB. Counting a gigabyte as 10^9 bytes:
| Request | Bytes | Requests per GB | Cost per request | Words |
|---|---|---|---|---|
| Home page, plain HTTP, US exit | 507,949 | 1,968 | $0.00041 | 72 |
| Home page, rendered, US exit | 383,086 | 2,610 | $0.00031 | 14 |
| hotels collector, one property | billed per row | n/a | $0.02 | price, total, rating, reviews |
| google_flights, one itinerary | billed per row | n/a | $0.003 | price, times, stops, airline |
The render row is cheaper per request and worthless, which is the trap: a budget that counts bytes will prefer it. Count words. The plain fetch is $0.00041 for a whole page; the collectors are priced per delivered row and failed rows are never billed, so the 20-property New York run cost $0.40 and the 15-itinerary run cost $0.045. Mobile exits at $2.30/GB change nothing here: the difference between a page and a stub was the country, not the network. Sizing a monthly plan from bytes per request is covered in how much proxy data you need, and the cadence in how to price monitor.
The setting that works on Expedia
Network: residential, Basic line at $0.80/GB. Fetch mode: engine: tls, plain HTTP, which returned 72 words with h1, canonical and description in 4.1 seconds; the browser returned 14 words in 50.9 seconds, so it is off. Country to pin: country: us, because the same URL over plain HTTP returned 72 words from a United States exit and 14 from a United Kingdom one. When the proxy is not enough: for every dated price, which robots.txt disallows on Expedia itself; the hotels collector returned 20 priced New York properties in 160 seconds and the google_flights collector 15 JFK to LAX itineraries in 19.4 seconds, and for undated hotel or destination pages the web scraping API over plain HTTP returns the page as Markdown, and every account gets $2 of free API usage per month.
Read the two lodging measurements next to this one: Booking.com gives a plain client 26 words and a browser none, so it is collector-only, while Tripadvisor gives a plain client 713 words and 19 LocalBusiness objects. Expedia sits with Tripadvisor on the fetch and with Booking.com on the prices.