Documentation Python quickstart Blog Free tools hello@quanticdata.ioLog in

Kayak Proxies: 2,451 Words With No Browser

Kayak measured on 28 September 2026 through residential proxies: 1,656,272 bytes and 2,451 words over plain HTTP with Website, Organization and FAQPage JSON-LD, 2,415,047 bytes and 2,455 words rendered, and 2,124 words on kayak.co.uk from a United Kingdom exit
Kayak measured on 28 September 2026 through residential proxies: 1,656,272 bytes and 2,451 words over plain HTTP with Website, Organization and FAQPage JSON-LD, 2,415,047 bytes and 2,455 words rendered, and 2,124 words on kayak.co.uk from a United Kingdom exit

Kayak is the travel site every tutorial tells you to open in Selenium, and it is the one that needs a browser least. On 28 September 2026 we fetched the home page through a United States residential exit as a plain HTTP client and got 2,451 words, an h1, a self-referencing canonical, a robots meta of index,follow and three JSON-LD blocks: Website, Organization and a FAQPage with four questions. Rendering the same URL took 40.7 seconds, added 758,775 bytes, and returned 2,455 words. The setting for Kayak proxies is residential, engine: tls, country pinned to the market: no browser, and 46% fewer bytes for the same page.

The keyword is half boats, and the other half wants Selenium

Google autocomplete cannot tell the site from the vessel. "Kayak proxies" completes to kayak problem, prototype, propellers; "kayak scraper" completes to hull care and one real query, "kayak scraping"; "proxies for kayak" and "kayak flight data api" return nothing. The SERP for "kayak scraping" is a mix of sanding tutorials and scraping vendors, with no People Also Ask, no AI Overview and no related searches: a thin term with a clear buyer underneath it.

The two guides that rank open the same way. One recommends Selenium because Kayak often uses JavaScript for data loading; the other opens a browser before it fetches anything. Neither measures the HTML, which is the entire finding of this post: the HTML is the page. If you copy those guides, you will pay for a browser to receive four extra words.

2,451 words over plain HTTP, and the JSON-LD is in the source

We audited kayak.com/ from a United States residential exit twice, as a pure HTTP client and fully rendered, and weighed both responses.

FetchExitBytesSecondsWordsh1JSON-LD
Plain HTTPUnited States1,656,2728.72,451Search Flights, Hotels & Rental CarsWebsite, Organization, FAQPage
RenderedUnited States2,415,04740.72,455Search Flights, Hotels & Rental CarsWebsite, Organization, FAQPage
Plain HTTPUnited Kingdomn/an/a2,124Search Flights, Hotels & Hire CarsWebsite, Organization, FAQPage
RenderedUnited Kingdomn/an/a2,367Search Flights, Hotels & Hire CarsWebsite, Organization, FAQPage

The seo_audit diff for the United States pair is empty on every axis: title unchanged, description unchanged, no h1 only in the render, no canonical missing without JavaScript, no content only in JavaScript. The four words the browser added are the difference between a server-rendered page and the same page after its scripts adjusted a label. That is not a data source; that is noise.

The structured data is where Kayak is unusually generous. We pulled every script type="application/ld+json" from the plain HTML with a CSS extraction and got three blocks. The Website block names the site and its URL. The Organization block carries nine sameAs links, from Facebook to Wikipedia to Crunchbase. The FAQPage block holds four questions and their full answers in the source: how to find deals on Kayak, what makes the app useful, how Trips manages bookings, and what Price Alerts are. If you want the site's own description of its price alert feature, it is 1,656,272 bytes away over plain HTTP and you do not need to render a thing. The three h2 headings are in the HTML too.

The exit country changes the domain, the h1 and 327 words

We repeated the audit from a United Kingdom residential exit, no Accept-Language header, nothing else changed. The request to kayak.com/ landed on kayak.co.uk/?ispredir=true, the canonical became kayak.co.uk/, the h1 became "Search Flights, Hotels & Hire Cars" and the plain HTTP word count dropped from 2,451 to 2,124. The title changed from "Rental Cars" to "Car Hire".

Kayak routes by exit IP to a country storefront, and that storefront is a different page: different copy, different currency in the search widget, a different set of promoted routes. Nothing in the geo path is a wall; it is localisation, and it is the reason to pin the exit. A price tracker for the United States that rotates through a mixed pool will silently collect kayak.co.uk pages in pounds and kayak.com pages in dollars and average them. Pin country: us and the h1 stays "Rental Cars" on every request.

robots.txt is rebuilt every night, and the terms are one sentence

kayak.com/robots.txt is 42,050 bytes and 1,271 lines, and it tells you when it was made: the header reads "Build version: R836d" and "Generated on: Mon Sep 28 01:00:00 EDT 2026", the morning of our measurement. Twenty-three user-agent blocks follow. The wildcard gets 23 allow lines and 106 disallow lines, and the allow list is unusually specific: /api/search/V8/hotel/, /i/api/search/v1/hotels/poll, /flights/$, /hotels/$, /hotels/sitemap. Eight agents get a bare Disallow: /, all of them ad-verification and semantic-analysis bots. ChatGPT-User has its own block with 19 disallow lines. Because the file is regenerated nightly, cache it per day, not per deploy, and read the build version before you trust yesterday's copy.

The Terms of Use are shorter and clearer than most. Section 4 asks you not to scrape, harvest, deep-link to, train AI on, or otherwise use automated tools on the site without written permission, and not to bypass any security feature, access restriction or robot exclusion header. That is one line and it is the one that gets cited: reading Kayak with a script is something its terms prohibit without permission. The permitted route is the KAYAK Affiliate Network, which offers deeplinks, widgets, a whitelabel and an API to approved partners and pays on clicks and bookings.

What a gigabyte buys on Kayak

Every number above came through residential proxies on the Basic line at $0.80/GB. Counting a gigabyte as 10^9 bytes:

RequestBytesRequests per GBCost per requestWords
Home page, plain HTTP1,656,272603$0.001332,451
Home page, rendered2,415,047414$0.001932,455

Kayak pages are heavy either way; 1.66 MB for a home page is the price of a metasearch site that inlines its scripts. The render adds 46% to that and four words, and it adds 32 seconds. Over a thousand pages, plain HTTP is $1.33 and the browser is $1.93 plus nine hours of browser time. The decision is not close. Mobile exits at $2.30/GB would cost $3.81 per thousand plain fetches for a site that never scored the address in four fetches; keep the cheap line. Our note on how much proxy data you need shows how to turn bytes per request into a monthly plan, and how to rotate proxies in Python shows the country-pinned session that keeps the storefront fixed.

When you want fares, not the home page

The home page is a shell around a search, and the search results are what most readers actually want. A results URL on Kayak is a rendered, polled experience: the page loads, then it asks its own API repeatedly while fares arrive. That is the one place a browser is legitimate, and the web scraping API with render and a wait for the results container handles it, at the rendered cost above. It is also where the terms and the robots file both point elsewhere: the search endpoints are disallowed for the wildcard and the terms exclude automated use without permission.

For fare monitoring that does not depend on one metasearch site's markup, we ran the Google Flights collector once on 28 September 2026 for JFK to LAX on 20 October, currency USD, exit country United States. It returned 15 itineraries in 19.4 seconds: prices from $204 to $294, 14 nonstop and one with a stop, four airlines, each row with departure and arrival time, duration and a CO2 estimate, at $0.003 per itinerary. The same route pair, the same date, no page to parse.

The setting that works on Kayak

Network: residential, Basic line at $0.80/GB; nothing in four fetches from two countries scored the address. Fetch mode: engine: tls, plain HTTP, which returns 2,451 words, the h1, the canonical and all three JSON-LD blocks for 1,656,272 bytes; the browser returns 2,455 words for 2,415,047 bytes and 40.7 seconds, so leave it off. Country to pin: the storefront you are measuring, because a United Kingdom exit is redirected to kayak.co.uk and the page drops to 2,124 words with a different h1 and currency. When the proxy is not enough: for live fares, the google_flights collector, which returned 15 itineraries for JFK to LAX in 19.4 seconds, or the web scraping API with render for a Kayak results page, and every account gets $2 of free API usage per month.

Among the travel metasearch sites we measured, Kayak is the readable one. Skyscanner hands a plain client zero words in every country, and Expedia hands it 72. Kayak hands it 2,451 and a FAQ, and asks only that you do not open a browser you do not need.

Sources & further reading

FAQ

Quick answers on kayak proxies.

Something else? Ask us →

Do I need Selenium to scrape Kayak?

Not for the pages we measured. On 28 September 2026 a plain HTTP client through a United States residential exit received 2,451 words, the h1, the canonical and three JSON-LD blocks from the home page for 1,656,272 bytes. A full browser render returned 2,455 words for 2,415,047 bytes and 40.7 seconds. The seo_audit diff was empty: nothing exists only in JavaScript.

What structured data does Kayak put in its HTML?

3 JSON-LD blocks in the plain source: a Website block, an Organization block with 9 sameAs links, and a FAQPage with 4 questions and full answers about deals, the app, Trips and Price Alerts. We extracted them with a CSS selector on script type application/ld+json over plain HTTP, no rendering.

Why does Kayak return a different page from the UK?

Because it routes by exit IP to a country storefront. From a United Kingdom residential exit the request to kayak.com landed on kayak.co.uk, the canonical changed, the h1 became "Search Flights, Hotels & Hire Cars" and the plain HTTP word count dropped from 2,451 to 2,124. Pin the exit country or your tracker will mix storefronts and currencies.

Does Kayak allow scraping?

No, not without written permission. Section 4 of the Terms of Use asks users not to scrape, harvest, deep-link to, train AI on or otherwise use automated tools on the site, and not to bypass robot exclusion headers. robots.txt is regenerated nightly (42,050 bytes, 23 user-agent blocks, 106 disallow lines for the wildcard). The permitted route is the KAYAK Affiliate Network, which includes an API for approved partners.

How much does a Kayak page cost through a residential proxy?

$0.00133 over plain HTTP and $0.00193 rendered on the Basic line at $0.80/GB, counting 1 GB as 10^9 bytes, so a gigabyte buys 603 plain pages or 414 rendered ones. Over a thousand pages the browser costs $0.60 more and about nine extra hours of browser time for four extra words per page.

How do I get live flight prices without parsing Kayak?

Run the google_flights collector with a route and a date. Our run for JFK to LAX on 20 October 2026 returned 15 itineraries in 19.4 seconds, priced from $204 to $294, 14 of them nonstop, with airline, times, duration and CO2 per row, at $0.003 per itinerary. It answers the same question a Kayak results page answers, with no markup to maintain.

Fetch the page, not the browser

Kayak returned 2,451 words to a plain HTTP fetch and 2,455 to a 40-second render; audit any URL both ways in one call and see the diff before you pay for a browser. Every account gets $2 of free API usage each month, and failed requests are never billed.

Related reading

Use casesExpedia Proxies: 72 Words Without a Browser

Expedia measured on 28 September 2026 through residential exits in the United States and the United Kingdom. From the United States a plain HTTP client gets the home page: 72 words, the h1 "The one place you go to go places", a canonical and a description for 507,949 bytes in 4.1 seconds. The browser gets 14 words. From the United Kingdom the plain fetch gets 14 words too. Structure over plain HTTP from a US exit; prices from the hotels and google_flights collectors.

Read →
Use casesSkyscanner Proxies: 0 Words, 15 Fares in 34 s

Skyscanner measured on 28 September 2026 through residential exits in the United Kingdom and the United States. The home page returns a title, a description, a canonical and a WebSite JSON-LD block, and zero words of body text, to a plain HTTP client and to a full browser alike, in both countries. No fetch mode changes that. The google_flights collector, asked for London Heathrow to New York JFK on 20 October in pounds, returned 15 itineraries in 33.7 seconds.

Read →
Use casesTripadvisor Proxies: 713 Words Over Plain HTTP

Tripadvisor measured on 28 September 2026 through residential exits in the United States and Italy. The home page hands a plain HTTP client 713 words, the h1 "Where to?", a canonical and three kinds of JSON-LD, including 19 LocalBusiness objects with URLs and countries, for 388,397 bytes in 3.6 seconds. From Italy the same URL returns the same English page with 757 words. The browser spends 46 seconds and returns 5,915 bytes. The setting is plain HTTP; the tripadvisor_search collector returned 30 Chicago restaurants in 8.1 seconds.

Read →