Airbnb gives a proxied client exactly the same page with a browser as without one, and charges 2.6 times the bandwidth for the browser. On 28 September 2026 we fetched the home page through a United States residential exit: 61 words for 597,053 bytes over plain HTTP, 61 words for 1,561,460 bytes fully rendered. The listings are not in either view. They sit in the page's hydration island, which is what the Airbnb stays collector reads: 18 priced stays for Austin, Texas in 5.3 seconds, no browser. The setting for Airbnb proxies is plain HTTP plus the collector, with the exit pinned to the market you are pricing.
Nobody searches for Airbnb proxies, everybody searches for the scraper
Google autocomplete returns nothing for "airbnb proxies" and nothing for "proxies for airbnb". It returns fifteen completions for "airbnb scraper": github, api, apify, python, free, chrome extension, reddit, listing scraper, review scraper, email scraper. The other question people type is "does airbnb have an api", and the honest answer is that Airbnb publishes no public data API; the developer surfaces it does have are for hosts and software partners, not for reading listings.
The first page for "airbnb scraper api" is hosted actors, scraping-API vendor pages, a GitHub project and a Reddit thread, and the AI Overview on that SERP recommends a single vendor by name. The longest tutorial on the page runs 5,846 words, renders the search page in a headless browser, waits for the listing cards and parses them with BeautifulSoup. Asked how many pages you can scrape safely it answers that there is no fixed number. We can do better than that on both counts, because we weighed the page and we did not need the browser.
Same 61 words, 2.6 times the bytes
We audited airbnb.com/ from a United States residential exit as a pure HTTP client and then fully rendered, and weighed both responses separately.
| Fetch | Exit | Bytes | Seconds | Words | h1 | Canonical |
|---|---|---|---|---|---|---|
| Plain HTTP | United States | 597,053 | 8.8 | 61 | Airbnb homepage | airbnb.com/ |
| Rendered | United States | 1,561,460 | 64.0 | 61 | Airbnb homepage | airbnb.com/ |
| Plain HTTP | Germany | n/a | n/a | 3 | none | none |
| Rendered | Germany | n/a | n/a | 54 | Airbnb-Homepage | airbnb.de/ |
Read the first two rows together. The title, description, canonical, h1 and the single WebSite JSON-LD block are all present in the plain HTTP response, and the rendered response adds no words to them. The seo_audit diff reports no content only in JavaScript, no title change, no canonical missing. The browser ran for 64 seconds, pulled 964,407 more bytes than the plain fetch, and produced a page that says the same 61 things. On the home page the render is pure cost.
The home page is also not where the data is, and that is the second finding. Airbnb ships its search results inside a hydration island in the HTML: a JSON blob the client-side app reads to draw the cards. The DOM of a plain fetch of a search page contains no listings, and the DOM of a rendered fetch contains cards dressed in class names that change on every deploy. The JSON contains everything: listing id, name, price breakdown, rating, review count, coordinates, badges. Read the island and you get typed fields without a browser; that is what the collector below does, and it is why its rows carry coordinates that the cards never show.
From Germany the plain fetch is three words
We repeated the audit from a German residential exit, sending no Accept-Language header and changing nothing else. The plain HTTP response was a 3-word page titled "Weiterleitung auf www.airbnb.de", with no canonical and no h1: a redirect notice to the German storefront. The rendered response followed the redirect and landed on airbnb.de/ with a German title, the h1 "Airbnb-Homepage" and 54 words.
Two consequences for a pipeline. First, a plain fetch from the wrong country does not return an English page or a German page; it returns a stub whose word count is 3, and a parser that expects the home page markup finds nothing. Follow the redirect explicitly or pin the exit to the market. Second, Airbnb prices and ranks stays per market, and the collector's own notes say the rank is re-ranked between requests; the country you exit from is part of the result, not a delivery detail. For price monitoring in the United States, exit from the United States.
The collector returns 18 stays in 5.3 seconds
We ran the airbnb_stays collector once on 28 September 2026 with the input Austin, Texas, check-in 20 October, check-out 23 October, up to 18 results. It finished in 5.3 seconds and delivered 18 rows, none partial, each with listing id, URL, card title and property type, nightly and total price, rating and review count, beds, bedrooms and baths, badges, coordinates and the first photo.
| Field | What the 18 rows contained |
|---|---|
| Nightly price | $90.43 to $387.40 |
| Three-night total | $272 to $1,163 on 17 rows |
| Rating | 4.48 to 4.99 on 17 rows; 1 row with no rating yet |
| Reviews | 9 to 745 |
| Guest favorite badge | 7 of 18 |
| Superhost badge | 6 of 18 |
| Free cancellation | 5 of 18 |
| Property type | Condo, Apartment, Home, Guesthouse, Bungalow; 1 "Featured hotel" row with a nightly rate but no total |
Three things in that output would have cost you a week to discover from the cards. The one hotel row arrives with a nightly rate and a null total, so a pipeline that divides total by nights crashes on row 11 of 18. The rank field is Airbnb's order in that response and the collector says so in its notes, so do not diff rank between runs and call it a trend. And the coordinates are Airbnb's approximate map point, not the address, which is exactly right for a supply map and exactly wrong for a mailing list.
Pass check-in and check-out and every listing is priced for that stay; leave them out and Airbnb prices each listing for a stay of its own choosing, which the rows report per listing. The collector paginates through Airbnb's own search cursor, 18 stays per full page, up to 90 per run, de-duplicated by listing id. For a whole city, shape the job as neighbourhoods.
What robots.txt says, and who is on the list
airbnb.com/robots.txt is 20,930 bytes and 780 lines, with one Sitemap line and eighteen user-agent blocks. The wildcard gets 61 disallow lines and 2 allow lines: account pages, API paths, skeleton routes, affiliate click trackers, nothing that touches a public listing. The interesting part is the roll call. OAI-Searchbot, ChatGPT-User, meta-externalagent, PerplexityBot, anthropic-ai and cohere-ai are each named and given the same 61 to 66 rules as Googlebot and Bingbot. Seven agents get a bare Disallow: /: GPTBot, ClaudeBot, Applebot-Extended, Webzio-Extended, AI2Bot, MistralAI-Training and ImagesiftBot. Airbnb lets the AI search crawlers in and shuts the AI training crawlers out, by name, and we have written about which way that cuts in should I block AI crawlers.
The terms are the other half. Airbnb's Terms of Service, in the general rules, say plainly: do not use bots, crawlers, scrapers or other automated means to access or collect data or other content from the platform, and do not scrape, hack or reverse engineer it. That applies to everyone who accepts the terms, which means every account holder. It is one line and we are not going to soften it. Readability is a technical fact; permission is a contractual one; on Airbnb they do not line up, and a proxy changes only the first.
What a gigabyte buys on Airbnb
Every number above came through residential proxies on the Basic line at $0.80/GB. Counting a gigabyte as 10^9 bytes:
| Request | Bytes | Requests per GB | Cost per request |
|---|---|---|---|
| Home page, plain HTTP | 597,053 | 1,674 | $0.00048 |
| Home page, rendered | 1,561,460 | 640 | $0.00125 |
| airbnb_stays, one stay | billed per row | n/a | $0.001 |
The render row is the one to delete from your budget: $0.00125 for the same 61 words the plain fetch returns for $0.00048, plus 64 seconds. The collector is billed per delivered stay, failed rows are never charged, and the 18-row Austin run cost $0.018. Mobile exits at $2.30/GB would nearly triple every row in that table without changing a word count; nothing in these fetches challenged the address, and the German stub was localisation, not a block. Our note on how much proxy data you need covers sizing the pool once you know the bytes per request, and how to price monitor covers the cadence.
The setting that works on Airbnb
Network: residential, Basic line at $0.80/GB; nothing in eight fetches scored the address. Fetch mode: engine: tls, plain HTTP, which returns the full 61-word home page with title, canonical and JSON-LD for 597,053 bytes; rendered returns the identical 61 words for 1,561,460 bytes and 64 seconds, so skip it. Country to pin: the market you are pricing, because a plain fetch from Germany returns a 3-word redirect stub instead of a page and the storefront, currency and ranking follow the exit. When the proxy is not enough: for listings, always; the airbnb_stays collector reads the hydration island and returned 18 priced stays in 5.3 seconds at $0.001 per stay, and for any other page on the domain the web scraping API with app_state mines the same island for you, and every account gets $2 of free API usage per month, which is 2,000 stays.
Airbnb is the middle case of the three big lodging sites. Booking.com hands a plain client 26 words and a browser none, so only the collector works there; Tripadvisor hands a plain client 713 words and 19 LocalBusiness objects and the browser adds nothing. Airbnb sits between them: a readable head, a hidden island, and a collector that knows where it is.