Shopee is the rare large marketplace where the browser buys nothing at all. On 28 September 2026 we fetched shopee.sg through residential proxies: a plain HTTP client from Singapore received 1,427 words, and a fully rendered browser received the same 1,427 words for 1.74 times the bytes and 55 seconds of waiting. From an Indonesian exit the identical URL returned a 163,632-byte shell with no title and no words. Shopee proxies are a one-line setting, and the line is: residential, plain HTTP, Singapore.
Search for Shopee proxies and Google mostly answers a shopping question
The Singapore results page for the head term is three different markets sharing one phrase. Proxy vendors hold the top slot with a Shopee landing page, followed by a Chrome-extension walkthrough and a generic explainer on why a proxy helps. Then an r/internationalshopper thread asks for a proxy purchasing service, a Shopee listing sells a "Japan proxy service" for buying from Mercari, and another sells MTG proxy cards. The related searches point at "Shopee Indonesia proxy" and "Shopee Taiwan proxy", which is the buyer telling you that the country matters before anyone has measured why.
The data intent lives under other words. The suggest box for "shopee scraper" fills with "shopee scraper api", "shopee scraper python", "scrape shopee product data", "scrape shopee api v4" and "scrape shopee reviews". We read the two top organic guides: one configures a browser extension, the other lists proxy types. Neither fetches a single page. So we did.
Without JavaScript Shopee gives you the whole page, and the render gives you the same page
We audited the Shopee Singapore home from a Singapore residential exit twice in one call: once as a pure HTTP client, once in a headless browser.
| Fetch | Status | Bytes | Words | Title | Time |
|---|---|---|---|---|---|
| Singapore exit, plain HTTP | 200 | 670,938 | 1,427 | Shopee Singapore | Cheaper, Faster On Shopee | 3.9 to 4.8 s |
| Singapore exit, rendered | 200 | 1,170,177 | 1,427 | same | 55.1 s |
| Malaysia exit, plain HTTP | 200 | 670,836 | full page | same | 4.6 s |
| Indonesia exit, plain HTTP | 200 | 163,632 | 0 | none | 1.6 to 5.7 s |
| Indonesia exit, rendered | 200 | shell | 0 | none | audit 79.3 s |
The first two rows are the reason this page exists. Of the forty-plus sites we have measured this way, Shopee is the one where the two views match word for word. The audit's diff reports no content only in JavaScript, no changed title, no changed description. The browser downloaded 499,239 extra bytes of scripts and fonts, waited fifty-five seconds, and produced a document with exactly the words the server had already sent in four seconds.
What the plain HTML contains is the entire home: category grid, flash deals, the daily discover block and the footer. Three things it does not contain matter for a parser. There is no link rel="canonical", there is no JSON-LD of any type, and the robots meta tag reads noindex. Shopee ships its home with an instruction to search engines not to index it. That is a hint about how the site sees its home: a shell for the app, not a document. Category and product pages are the pages to point the same plain client at.
One fetch in four from Singapore is the empty shell, and the tell is the title
We fetched the home nine times over plain HTTP to see how stable the full page is. From Singapore three of four fetches returned the full page at 670,686 to 671,418 bytes; the fourth returned a 163,632-byte document with no title, no description and no words, still with status 200. From Malaysia both fetches were full pages. From Indonesia all three fetches were the shell, and the full audit from Indonesia found zero words in both passes.
Two rules fall out of that, and they are the whole operational setting:
- Classify on the title, never on the status. Both documents are 200 OK. The full page carries "Shopee Singapore | Cheaper, Faster On Shopee" in its head; the shell carries nothing. If
titleis empty, the request did not reach the page you wanted; in our runs the next fetch from the same country returned the full page. - Pin the exit to the market's own country. shopee.sg wants a Singapore or Malaysia exit. An Indonesian address did not get a redirect to shopee.co.id and did not get an error; it got 163,632 bytes of nothing, which is the most expensive kind of failure because it looks like success until the parser runs. Every Shopee market is a separate domain, and the same rule holds per domain: the exit country belongs to the domain you are reading.
def shopee_page(resp):
# 200 in every case; the shell has no head at all
if not resp.title:
return "shell_retry" # 163,632 bytes, zero words
if resp.bytes < 300_000:
return "shell_retry" # a full home is ~671 KB
return "page" # 1,427 words, parse it
The byte check is a belt to the title's braces: in nine fetches the shell was always 163,632 bytes exactly and the full page never dipped below 670,686.
What the rendered fetch costs, and what it saves to skip it
Every number here came through residential proxies on the Basic line at $0.80/GB. Counting a gigabyte as 10^9 bytes, the home costs this:
| Fetch | Bytes | Fetches per GB | Cost per fetch | Words returned |
|---|---|---|---|---|
| Plain HTTP, Singapore | 670,938 | 1,490 | $0.00054 | 1,427 |
| Rendered, Singapore | 1,170,177 | 854 | $0.00094 | 1,427 |
| Plain HTTP, Indonesia (shell) | 163,632 | 6,111 | $0.00013 | 0 |
The render costs 1.74 times the bandwidth and 13 times the wall-clock, and it returns nothing the plain fetch did not. On a thousand home fetches a day that is 0.5 GB you did not need to buy and about fourteen hours of browser time you did not need to wait. On mobile proxies at $2.30/GB the same plain fetch is $0.0015, and nothing in nine requests scored the network, so the residential Basic line is the right one. If you are sizing a pool before you buy, how much proxy data you need runs the same arithmetic the other way round.
The Indonesian row is the trap. It is the cheapest fetch in the table and it is worth exactly nothing, and a budget that counts requests instead of words will report it as a success.
What robots.txt permits, and the crawl delay it asks for
shopee.sg/robots.txt is 1,384 bytes and has two blocks. The first names Googlebot, Googlebot-Mobile and Bingbot, gives them a crawl delay of 0.1 seconds and lists 33 disallowed paths. The second is for everyone else:
User-Agent:*
Crawl-delay:1
Disallow: /cart/
Disallow: /checkout/
Disallow: /buyer/login/otp
Disallow: /user/
Disallow: /me/
Disallow: /order/
Disallow: /daily_discover
Disallow: /mall/just-for-you/
Disallow: *-i.*/similar
Disallow: /find_similar_products/
Disallow: /top_products
Disallow: /search*searchPrefill
Disallow: /index.html
Disallow: *?utm_source
Read it for what it does not say. Product pages, shop pages and category pages are not disallowed for an unnamed client, and plain keyword search is not disallowed either; only the prefilled, hashtag, brand-filter and shop-scoped search variants are listed, and most of those only in the Googlebot block. What the wildcard block asks for is a one-second crawl delay, which is a rate, not a wall. A collector that sends one request per second per exit is doing what the file asks.
The account surfaces are the other half. Cart, checkout, login OTP, orders and the personalised feeds are disallowed for every agent, and they are also the surfaces that require a session. Shopee's terms of service, published through its Help Centre, govern account holders, and automated access with an account is the line we do not help anyone cross. Everything measured here was logged out. For the general shape of these files across platforms, llms.txt vs robots.txt covers what each one can and cannot promise.
Where the structured data is, and where it is not
The home has no JSON-LD, and it is marked noindex, so the home is not the page to parse for anything but the category list. The Shopee Open Platform, at open.shopee.com, is Shopee's official developer platform; it is a JavaScript application end to end, three words of text without a browser, and it is the route for a seller's own listings. For public price and availability reads of other sellers' listings there is no official feed, which is why the plain-HTTP product page is the surface, and why this page measures the cheapest way to get it.
One note from the fetches that matters for a job spanning markets. The home served to Malaysia was the Singapore home with the same title and the same byte count within a few hundred bytes, which tells you the domain, not the exit, decides the market; the exit only decides whether you get the page at all. That is the opposite of what we measured on AliExpress, where the exit rewrites the domain, and it is the same shape as Lazada, Shopee's neighbour in the same city.
The setting that works on Shopee
- Network: residential, Basic line at $0.80/GB through residential proxies. No request in nine scored the IP; the cost driver is bytes, and Basic is the cheapest byte.
- Fetch mode:
engine: tls, plain HTTP. 1,427 words from the home, the same 1,427 the browser returns, for 670,938 bytes instead of 1,170,177 and four seconds instead of fifty-five. - Country: pin
sgfor shopee.sg, ormy, which served the identical page. An Indonesian exit received a 163,632-byte shell with zero words. Classify on the title and retry the shell once. - When the proxy is not enough: there is no Shopee collector in our catalogue. Drive the web scraping API with
engine: tlsagainst product and category URLs, and reach for rendering only on a page whose plain fetch has a title and still no price. For the same job on a marketplace we do ship a collector for, AliExpress search returned 20 priced items in 4.5 seconds on the same day. - Failed requests are never billed, and every account gets $2 of free API usage per month, which is about 3,700 plain fetches of the Shopee home.