Tripadvisor hands a plain HTTP client more than any browser will get from it. On 28 September 2026 we fetched the home page through a United States residential exit as a pure HTTP client and received 713 words, the h1 "Where to?", a self-referencing canonical, seven h2 headings and 21 JSON-LD objects, 19 of them LocalBusiness entries with a name, a URL and a country, for 388,397 bytes in 3.6 seconds. From an Italian exit the same URL returned the same English page with 757 words and the same canonical. The rendered fetch spent 46.2 seconds and returned 5,915 bytes with nothing to parse. The setting for Tripadvisor proxies is residential, engine: tls, any exit country, no browser, and the Tripadvisor collector when you want a whole destination's listings: 30 Chicago restaurants in 8.1 seconds.
The SERP already knows the API is not the answer
Autocomplete has nothing for "tripadvisor proxies" or "proxies for tripadvisor" and thirteen completions for "tripadvisor scraper": github, api, python, extension, free, reddit, apify, review scraper. The review phrasing adds a revealing pair, "can tripadvisor reviews be traced" and "scraping reviews from tripadvisor python". People want the reviews and they want to know what it costs them.
The first page for "tripadvisor scraper api" holds a GitHub project, a hosted actor, two vendor tutorials, two scraping-API vendors and, at position four, a thread on Tripadvisor's own support forum from 2015 titled "API access or scraping?", where the answer from another user is that without written authorisation, no, and the content is copyrighted. The best tutorial on the page is 4,670 words and technically right: it replicates the site's GraphQL search endpoint and parses the JSON hidden in script variables, and its FAQ notes that the official API returned only 3 reviews per location. What none of them do is fetch the page plain and count.
713 words, 21 JSON-LD objects, no browser
We audited tripadvisor.com/ from a United States residential exit as a pure HTTP client and then fully rendered, then pulled every script type="application/ld+json" block out of the plain HTML with a CSS extraction.
| Fetch | Exit | Bytes | Seconds | Words | h1 | Canonical | JSON-LD |
|---|---|---|---|---|---|---|---|
| Plain HTTP | United States | 388,397 | 3.6 | 713 | Where to? | tripadvisor.com/ | Organization, WebSite, 19 LocalBusiness |
| Rendered | United States | 5,915 | 46.2 | n/a | n/a | n/a | none |
| Plain HTTP | Italy | n/a | n/a | 757 | Where to? | tripadvisor.com/ | Organization, WebSite, LocalBusiness |
The plain response is a complete server-rendered page: title, a 214-character description, canonical, Open Graph tags, a content-language: en header, seven h2 sections from "Find things to do by interest" to "Tripadvisor: join the largest travel community", and 713 words of copy around them. The first JSON-LD block is an @graph with the Organization, seven sameAs links and a WebSite with a SearchAction pointing at /Search?q=. The next two blocks are arrays of LocalBusiness objects, ten and nine, one per promoted tour: a Florence storyteller tour, a Ninh Binh day trip, a Blue Cave boat tour from Dubrovnik, a London pub tour, nineteen in all, each with a name, the AttractionProductReview URL, an image and a PostalAddress whose country is filled in.
That is a structured feed of what Tripadvisor is promoting today, delivered in the source of the home page to a client that runs no JavaScript. It changes as the promotions change, which makes it a cheap daily signal for anyone tracking the experiences market, and it costs 388,397 bytes.
The rendered fetch is the other row. The browser ran for 46.2 seconds and came back with 5,915 bytes and no page to parse. On Tripadvisor the plain fetch returns 713 words and the browser returns none; that is the whole render decision, and it is the same finding the 4,670-word tutorial reached by a longer road when it chose httpx over a browser.
The exit country does not change the page
We repeated the plain audit from an Italian residential exit with no Accept-Language header. The response was the same English page: h1 "Where to?", canonical tripadvisor.com/, content-language: en, the same three JSON-LD types, 757 words. The 44-word difference is the promoted-tour carousel, which rotates between requests, not a localisation.
This is the opposite of what Kayak and Airbnb do, which is redirect a European exit to a country domain. Tripadvisor keeps tripadvisor.com as the en_US storefront regardless of where the request comes from, and serves its other markets on their own domains, which is why the sitemap lines in robots.txt are all suffixed en_US. For the English site, any residential exit will do and the country is a cost decision, not a correctness one. Pin the country only when you fetch a localised domain on purpose.
robots.txt names sixteen crawlers and shuts every one out
tripadvisor.com/robots.txt is 24,267 bytes and 790 lines, and it begins with a recruiting note for its SEO team before listing eight Sitemap indexes. The rules are eight blocks. The first names sixteen agents in one block and gives them a bare Disallow: /: Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Cohere-ai, GPTBot, Google-CloudVertexBot and the rest of the training crawlers. Google-Extended gets its own Disallow: / further down. The wildcard block is shared with ChatGPT-User, ChatGPT-User/2.0, Gemini-Deep-Research and OAI-SearchBot and carries 668 disallow lines and 8 allow lines, a path-by-path list of account, forum, edit and internal endpoints. PerplexityBot, bingbot and Baiduspider each get a short block of their own. The training crawlers are out entirely; the AI search crawlers get the same long list as everyone else; the review and listing pages are not on it. We have written about that split in should I block AI crawlers.
The Terms of Use are the line that gets cited. The prohibited activities list says you may not access, monitor, reproduce or otherwise exploit any content using any robot, spider, scraper or other automated means or any manual process without express written permission, may not violate robot exclusion headers, may not deep link, and, in clause (x), may not use or enable any robot, spider, artificial intelligence system or other automated device to access, retrieve, copy, scrape, aggregate or index any portion of the services except as expressly permitted in writing. That is one sentence and we will not argue with it: reading Tripadvisor with a script is something its terms prohibit. The licensed route used to be the Content API; its reference page now carries a deprecation notice pointing to a successor platform called Terra, so a business that wants licensed reviews applies there.
The collector returns 30 restaurants in 8.1 seconds
For destination-level public data, we ran the tripadvisor_search collector once on 28 September 2026 with the input Chicago, Illinois, type restaurants, exit country United States, up to 30 results. It resolved the destination to Tripadvisor's own geo id 35805, finished in 8.1 seconds and delivered 30 rows, none partial, read from the schema.org ItemList the listing page publishes for search engines.
| Field | What the 30 rows contained |
|---|---|
| Rating | 4.1 to 4.7 on all 30 |
| Review count | 209 to 9,907 |
| Phone | present on all 30 |
| Address and coordinates | present on all 30 |
| Price level | $ on 3, $$ - $$$ on 21, $$$$ on 6 |
| Cuisines | one or two per row, from Italian and Pizza to French and Steakhouse |
| Detail URL | the Restaurant_Review page for every row |
Two details matter for a pipeline. Three names appear twice in the 30 rows, because a chain has two locations in the city; they have different ids, different addresses and different phone numbers, so de-duplicate on id, never on name. And the collector's note says the rank reflects Tripadvisor's order in that response and is re-ranked between requests, so do not store rank as a fact. It paginates 30 per page up to 300 per run, hotels or restaurants, and the rows carry the detail URL, which is the page you fetch plain when you want the full listing.
What a gigabyte buys on Tripadvisor
Every number above came through residential proxies on the Basic line at $0.80/GB. Counting a gigabyte as 10^9 bytes:
| Request | Bytes | Requests per GB | Cost per request | Words |
|---|---|---|---|---|
| Home page, plain HTTP | 388,397 | 2,574 | $0.00031 | 713 |
| Home page, rendered | 5,915 | 169,061 | $0.0000047 | none to parse |
| tripadvisor_search, one place | billed per row | n/a | $0.001 | name, rating, reviews, phone, coordinates |
The render row is nearly free per request and worth nothing, and it costs 46 seconds each; a thousand renders is more than twelve hours of browser time for no words. A thousand plain fetches is $0.31 and about an hour. The collector is billed per delivered place and failed rows are never charged, so the 30-row Chicago run cost $0.03. Mobile exits at $2.30/GB would raise the plain row to $0.00089 for a page that returned the same 713 words to every exit we tried; keep the Basic line. Sizing a monthly plan from bytes per request is covered in how much proxy data you need.
The setting that works on Tripadvisor
Network: residential, Basic line at $0.80/GB; nothing in the plain fetches from two countries scored the address. Fetch mode: engine: tls, plain HTTP, which returned 713 words, the h1, the canonical and 21 JSON-LD objects in 3.6 seconds; the browser returned 5,915 bytes in 46.2 seconds with nothing to parse, so it is off. Country to pin: none required for tripadvisor.com, which served the same English page with the same canonical from the United States and from Italy, 713 and 757 words; pin a country only when you fetch a localised domain on purpose. When the proxy is not enough: for a whole destination, the tripadvisor_search collector returned 30 Chicago restaurants with phone, rating, review count and coordinates in 8.1 seconds at $0.001 per place, and for any single listing page the web scraping API over plain HTTP returns it as Markdown with the JSON-LD intact, and every account gets $2 of free API usage per month, which is 2,000 places.
Of the three lodging and review sites we measured, Tripadvisor is the open one on the fetch. Booking.com hands a plain client 26 words and a browser none; Expedia hands it 72 from the right country and 14 from the wrong one. Tripadvisor hands it 713 words and nineteen businesses in JSON-LD from anywhere, and its terms, not its servers, are where the limit is.