Glassdoor proxies get bought for a headless browser that the measurement says you should never launch. On 26 September 2026 the Glassdoor home page answered a plain HTTP client through a US residential exit with 200 OK and 624 words; rendered in a browser, the same URL returned 54 words and a canonical of /member/profile/login. On 28 September the UK edition gave a plain client 589 words with the h1 and the Organization schema, while the rendered fetch landed on the login page for 705,071 bytes and 61 seconds. The setting is residential, plain HTTP, country pinned to the edition, and no browser at all; for listings with salaries as numbers, the jobs collector of the same operator returns ten rows in 7.3 seconds.
The keyword is a company called Proxy
Google has no autocomplete for "glassdoor proxies", and the first page for the term is a homonym: seven of ten results are Glassdoor's own pages about companies named Proxy and Proxy Live Solutions, 3.9 stars from 27 reviews, plus a search page for 5,020 jobs with proxy in the title. The related searches, "glassdoor proxies legit" and "glassdoor proxies reviews", are people wondering whether to work there. The two results that are not Glassdoor are vendor guides, and the top one teaches a Python and Playwright scraper, on the reasoning that a simple HTML parser will miss dynamically loaded data.
The real demand is one word over. "Glassdoor scraper" autocompletes to selenium, github, api, extension, reddit, python, review scraper, job scraper, interview scraper and salary scraper; "scrape glassdoor" to reviews, interview questions, data, jobs and the question "can you scrape glassdoor". Those searchers want reviews, salaries and listings as rows for compensation benchmarking, employer-brand monitoring and labour-market research. That is what this post measures, and the first finding is that the tool the top guide recommends is the one that returns the least.
Plain HTTP gets the page, the browser gets the login form
We audited glassdoor.com from a United States residential exit on 26 September with our SEO audit, which fetches once as a pure HTTP client and once rendered, repeated it from a United Kingdom exit on 28 September against the UK edition, and weighed the rendered fetch separately.
| Fetch | Status | Words | Title or h1 | Canonical |
|---|---|---|---|---|
| Home, plain HTTP, US, 26 September | 200 | 624 | Glassdoor home | glassdoor.com |
| Home, rendered, US, 26 September | 200 | 54 | Log In | Glassdoor | /member/profile/login |
| Home, plain HTTP, UK, 28 September | 200 | 589 | You deserve a job that loves you back | glassdoor.co.uk/index.htm |
| Home, rendered, UK, 28 September | n/a | 600 | 600 words, no canonical | absent |
| Home, rendered and weighed, US, 28 September | n/a | 44 | Log In | Glassdoor | /member/profile/login |
Read the two US rows together. Without JavaScript the server hands over the home page: navigation, search, the headline, the promotional copy, the footer, 624 words of it. With JavaScript the page's own scripts decide that a browser without a member session belongs on the login form, and the canonical link on what arrives is https://www.glassdoor.com/member/profile/login. The page was delivered, and then the page sent the browser away. That is what a harmful render looks like, and it is the reason the verdict on this target is not "useless" but "harmful": rendering does not just cost more, it replaces the data with a form.
The UK rows say the same thing in a different shape. Over plain HTTP the UK home page arrived with a canonical, an h1, a meta description and two JSON-LD blocks, WebSite and Organization, for 589 words. Rendered, it returned 600 words with no canonical, no h1 from the page and no JSON-LD: the same notice repeated in seven languages. A word count alone would call that a success; the missing canonical is what tells you it is not. On Glassdoor, classify on the canonical link and the h1, and treat a canonical ending in /member/profile/login as "rendered the wrong page", not as "missing item".
The weighed render from the US puts a price on the mistake: 705,071 bytes and 60.8 seconds of browser time to arrive at a login form. Glassdoor puts the reason in the query string of that URL, and the reason is the browser.
The exit country picks the edition
Glassdoor runs national editions on their own domains: glassdoor.com, glassdoor.co.uk, glassdoor.de and others. The exit country does not silently move you between them the way Indeed does; you choose the domain, and the edition sets the currency, the salary format and the review pool. What the exit does change is whether a residential address looks like it belongs to that edition's market, and the two editions we measured agreed on the shape of the answer:
| Edition | Exit | Words, plain HTTP | Words, rendered | Rendered document |
|---|---|---|---|---|
| glassdoor.com | United States | 624 | 54 | Login form |
| glassdoor.co.uk | United Kingdom | 589 | 600 | 600 words, no canonical |
Pin the exit to the edition: country=us for glassdoor.com, country=gb for glassdoor.co.uk. A US address reading the UK edition is a mismatch the site can score, and there is no reason to give it one. And whatever the edition, do not render. The 35-word difference between the two plain fetches is copy; the 570-word difference between plain and rendered on the US edition is the data.
robots.txt says how deep the site wants you to go
glassdoor.com/robots.txt is 6,250 bytes and it is an honest map of what the site is willing to serve. Every user agent is disallowed from /api/, /graph, /search/, /member/ and /profile/, and from pagination: /Reviews/*_P*.htm, /Jobs/*_P*.htm and /Interview/*_P*.htm are all disallowed, with page 2 of a reviews search and pages 2 to 5 of a salaries listing allowed back in by name. The login page, /member/profile/login, is explicitly allowed, which is why the browser's destination is crawlable and the data behind it is not.
Then a block headed "Partially Block High-Quality AI/LLM Bots" names GPTBot, Google-Extended, Amazonbot, anthropic-ai, ClaudeBot, Perplexity, Cohere, Applebot-Extended and Google-CloudVertexBot and disallows them from everything except /blog/, /Award/, /About/ and /employers/; a second block fully blocks CCBot, Bytespider, Diffbot, FacebookBot and others. Reviews, salaries and interviews are off limits to every AI crawler by name, and the pagination rules cap what anyone else can page through. The file opens with a recruiting joke for humans reading it, and it closes the door on machines reading past page two of reviews and page five of salaries.
Whose terms these are: Indeed, Inc.
Glassdoor's Terms of Use, revised 1 July 2026, open with the operator: Indeed, Inc. provides the services Glassdoor.com and Fishbowlapp.com. The House Rules require use solely for lawful purposes consistent with the Terms and incorporate the Community Guidelines. Glassdoor's own security notice says the rest in one line: Glassdoor has been built on the contributions of real employees and job seekers, and it uses advanced security systems to prevent misuse or unauthorized access. The same operator's Site Rules for Indeed say do not access the site through any means other than its public interfaces, and do not access any data, especially personal data, by automated means without permission.
So be accurate about what a proxied fetch is. The home page is served to any visitor, and reading what a server hands the public is a technical fact. Reviews and salaries are contributed by identifiable employees, and Glassdoor's model is that you contribute to read; the login wall the browser lands on is that model enforced. Permission to collect and reuse those contributions is contractual, and Glassdoor has not published a data API for it. Anyone telling you a residential IP settles that question is selling you an IP. Where the honest answer is that the terms forbid it, the answer is that the terms forbid it, and the next section is about what you can build instead.
What to spend on instead of the browser
There is no Glassdoor collector in our catalogue, and given the terms and the login wall we are not building one. The jobs are a different matter, because the same operator publishes them on Indeed, and we ran the Indeed jobs collector the same afternoon: ten "data engineer" listings in New York in 7.3 seconds, each with the salary parsed into salary_min, salary_max and a period, the company's rating and review count, and a sponsored flag. Three of the ten were sponsored, and the ratings came from the same review pool that Glassdoor gates. For a second source with the same input shape, the Google Jobs collector returns listings aggregated across boards. And for the page as it is, the web scraping API over plain HTTP returns the 624 words as Markdown at $0.0002 a page and never renders unless you ask it to. We measured the operator's own board in Indeed through residential proxies, and it points the same way: plain HTTP wins, the browser loses words.
Cost: what the browser bill buys on Glassdoor
Prices from our pricing page: residential Basic $0.80/GB, the web scraping API from $0.0002 per page and $0.001 rendered, the jobs collector $0.001 per delivered listing. A gigabyte is counted as 10^9 bytes.
| Approach | Bytes per page | Pages per GB | Cost per 1,000 | What you get |
|---|---|---|---|---|
| Rendered, residential Basic | 705,071 | 1,418 | $0.56 | 54 words and a login form, 61 seconds each |
| Web scraping API, rendered | n/a | n/a | $1.00 | The same login form as Markdown |
| Web scraping API, plain fetch | n/a | n/a | $0.20 | 624 words, h1, canonical, JSON-LD |
| Indeed jobs collector | n/a | n/a | $1.00 | 1,000 listings with parsed salaries, same operator |
Every rendered row in that table is money spent to reach a form. The plain fetch is the only row that returns the page, and it is also the cheapest. Byte figures are decoded bodies and exclude TLS overhead. If you are sizing a job before you buy bandwidth, how much proxy data you need does the arithmetic from the other side; on Glassdoor the answer is "less than the browser guide told you".
Where we stop
Everything above is about the public home page and public listings, for employer-brand monitoring, compensation research at the level Glassdoor publishes to visitors, and checking how an edition reads from its own market. It does not cover accounts. We will not help with creating accounts to pass the contribute-to-read wall, submitting reviews or salaries you did not experience, or automating a logged-in session. Glassdoor's Community Guidelines govern contributions, its operator's rules forbid fake accounts and automated account creation, and reviews are the words of identifiable employees, which makes them personal data wherever your users live. A proxy does not change who wrote a review or whether you were allowed to keep it.
The setting that works on Glassdoor
- Network: residential proxies, Basic line at $0.80/GB, one exit per edition. The plain fetch returned 624 words from a US exit and 589 from a UK one; nothing in the plain path rewarded a mobile address at $2.30/GB or an ISP address at $2.50 per IP per month.
- Fetch mode:
engine: tls, plain HTTP, never rendered. 624 words with h1, canonical and Organization JSON-LD; rendered, 54 words and a canonical of /member/profile/login for 705,071 bytes. - Country: pin the exit to the edition,
country=usfor glassdoor.com andcountry=gbfor glassdoor.co.uk. The plain word count moves by 35 between editions; the rendered one loses 570. - When the proxy is not enough: there is no Glassdoor collector. For listings with salaries as numbers use the operator's board through the indeed_jobs collector, 10 rows in 7.3 seconds at $0.001 per listing; for the page as Markdown, the web scraping API at $0.0002 per page, plain fetch.
- Free tier: every account gets $2 of free API usage per month, which is 10,000 plain-fetch pages or 2,000 parsed listings before you pay anything.