Rakuten Ichiba serves the same page to everyone. On 28 September 2026 we fetched rakuten.co.jp through residential proxies: a plain HTTP client from Japan received 882 words, a self-referencing canonical and Corporation plus WebSite JSON-LD in 381,287 bytes, and a client from the United States received the identical 882 words. The render added 268 words of rotating modules for 4.2 times the bytes. Rakuten proxies are a plain-HTTP job on the residential Basic line, and for the catalogue itself there is an official API that needs no proxy at all.
The word proxy means forwarding on this SERP, and the data question hides under scraper
Search for Rakuten proxies from the United States and Google answers a shipping question. The first result is Rakuten's own Global Express help page for its "proxy payment service"; the rest of the page is forwarding agents, a CDJapan buying guide, Buyee explaining Rakuma, and an r/japanesestreetwear thread about proxying an order. Three proxy-network vendor pages sit between them. Related searches are ZenMarket, Rakuma and "does Rakuten Rakuma ship to USA". A proxy here is a person in Japan who buys the parcel for you.
Autocomplete makes the split explicit: "rakuten jp proxy", "rakuten rakuma proxy" and "proxy for rakuten" on one side, "rakuten data scraper" on the other, and no suggestions at all for the bare head term. We read the two vendor pages that rank. One sells Japanese IPs by prefecture for shop management and price tracking; the other describes the US cashback site rakuten.com as if it were the Japanese marketplace. Neither fetches a page from rakuten.co.jp. This post is about the marketplace, Rakuten Ichiba, and about reading its public pages for price and availability monitoring and market research.
882 words over plain HTTP, with the structured data already in the head
We audited the Rakuten Ichiba home from a Japanese residential exit in one call that fetches twice, once without JavaScript and once in a headless browser, and then repeated it from a US exit.
| Fetch | Status | Bytes | Words | Canonical | JSON-LD | Time |
|---|---|---|---|---|---|---|
| Japan exit, plain HTTP | 200 | 381,287 | 882 | https://www.rakuten.co.jp/ | Corporation, WebSite | 4.1 s |
| Japan exit, rendered | 200 | 1,617,601 | 1,150 | same | same | 19.1 s |
| US exit, plain HTTP | 200 | 381 KB class | 882 | same | same | 7.6 s |
| US exit, rendered | 200 | 1.6 MB class | 1,076 | same | same | audit 60.4 s |
| Japan exit, plain HTTP, ?lang=en | 200 | 381,233 | same page | same | same | 4.1 s |
The plain document carries the title, the meta description, the canonical and two JSON-LD blocks, Corporation and WebSite, which is more structure than most marketplace homes ship at all. It has no h1, and its robots meta reads NOYDIR, a directive from the Yahoo Directory era that no current crawler acts on. The audit's diff reports no title, description or canonical change between the two views and no h1 that appears only in the browser.
The rendered view is larger, and it is worth being precise about what the extra 268 words are, because a naive reading says "the render is 30% richer". They are ranking widgets, campaign banners and personalised recommendation modules, which is why the same rendered page returned 1,150 words from Japan and 1,076 from the United States while the plain view returned 882 from both. The render adds content that changes per request; the plain view carries the content that is the same for everyone. For monitoring, the second kind is the one you can diff.
The exit country does not change Rakuten, and neither does lang=en
Most of the marketplaces we have measured this month rewrite something when the exit moves: AliExpress changes the domain, Shopee withholds the page, bet365 changes the licensed offer. Rakuten Ichiba does none of that. A US residential exit received the same title, the same canonical, the same 882 words in Japanese and the same JSON-LD as a Tokyo exit. There is no geo-fence on the public site and no localisation off the IP.
We also tested the one parameter robots.txt hints at. The file tells Googlebot not to crawl ?lang= URLs, so we fetched the home with ?lang=en over plain HTTP from Japan. The response was the same Japanese page, 381,233 bytes against 381,287, same title, same description. The parameter is a crawl-budget exclusion, not an English edition.
The consequence for a parser is that the localisation work is yours regardless of exit. Prices are yen, titles and categories are Japanese, and no exit country will change that. Pin Japan anyway, because shipping estimates and campaign eligibility on item pages are computed for the requesting country, and because it matches what a Japanese buyer sees; but do not expect it to translate anything.
What the render costs, and why the catalogue is cheaper through the API than through any proxy
Every fetch went through residential proxies on the Basic line at $0.80/GB. Counting a gigabyte as 10^9 bytes:
| Fetch | Bytes | Fetches per GB | Cost per fetch | Words |
|---|---|---|---|---|
| Home, plain HTTP | 381,287 | 2,622 | $0.00031 | 882 |
| Home, rendered | 1,617,601 | 618 | $0.00129 | 1,150 |
The render costs 4.2 times the bandwidth and 4.6 times the wall-clock, and the extra words are the ones that change per request. Skip it on the home and on category pages. On mobile proxies at $2.30/GB the plain fetch is $0.00088, and nothing in seven fetches from two countries scored the network, so the Basic line is the correct one.
For the catalogue, the cheapest route is not a proxy at all. Rakuten Web Service publishes the Rakuten Ichiba Item Search API, version 2026-07-01, which returns items by keyword, shop or genre with the fields a price monitor wants, and it asks for an application registered with a Rakuten account rather than an IP. If you are tracking a defined set of items, register the application and call the endpoint; keep the proxy for what the API excludes, which by its own documentation is co-listed items, auctions, flea-market and customer-to-customer listings, and for the rendered view of a page as a Japanese shopper sees it.
What robots.txt says, in 490 bytes
rakuten.co.jp/robots.txt was last modified on 2 March 2026 and is short:
User-Agent: *
Disallow: /com/
Disallow: /images/
Disallow: /backup/
Disallow: /cgi-bin/
Disallow: /shops/
Disallow: /shop/ascosing/
Disallow: /shop/egs-g/
Disallow: /shop/tripleo-shop/
Disallow: /shop/fooody/
Disallow: /shop/kira-con02/
Disallow: /*/lightbox_*.html
Disallow: /aboutus/am/maint/
User-Agent: AdsBot-Google
Disallow: /com/
User-agent: Googlebot
Disallow: /*?lang=
Disallow: /*&lang=
Static asset directories, a legacy shops path, five individually named shops and the lightbox pages are excluded for everyone; there is no crawl delay, no search or item exclusion and no sitemap line. The site's public pages are open to an unnamed client. What is not covered by that file is the member area, the Rakuten points programme and shop management, all of which need a login, and the Rakuten terms for members govern those. Everything measured here was logged out.
One corporate note that matters for anyone deciding what to read. Rakuten Group's corporate site lists the group's businesses, and rakuten.com in the United States is the cashback affiliate service, a different product with a different site. A "rakuten proxies login" search, which appears in the related searches, is someone trying to reach the cashback account from abroad; that is an account question and not one a proxy should answer.
Where the data is on Rakuten Ichiba, and what to fetch
The home is a category and campaign page; its value to a parser is the JSON-LD and the genre links. Item pages and shop pages are where the price, the points rate and stock live; fetch them the same way, plain first, and check that the plain view carries the price before you consider rendering. Search results are paginated, and they are the pages where the official API is cheaper than any fetch.
There is no Rakuten collector in our catalogue. The web scraping API with engine: tls is the tool for public pages, and the cadence and diff logic are the same as for any marketplace, which we set out in how to price monitor. When a marketplace does have a collector the unit changes from bytes to rows: on the same day, AliExpress search returned 20 items with price, original price and discount in 4.5 seconds.
The setting that works on Rakuten
- Network: residential, Basic line at $0.80/GB through residential proxies. Seven fetches from two countries, none scored the IP; the cost is bytes.
- Fetch mode:
engine: tls, plain HTTP. 882 words, canonical and Corporation plus WebSite JSON-LD in 381,287 bytes and 4.1 seconds; the render returns 1,150 words of which 268 rotate per request, for 1,617,601 bytes. - Country: pin
jp. The text is identical from a US exit, 882 words either way, but shipping and campaign eligibility on item pages are computed for the requesting country. Do not expect?lang=ento change anything; it returned the same Japanese page. - When the proxy is not enough: for the catalogue use the official Rakuten Ichiba Item Search API, version 2026-07-01, registered with a Rakuten account. For public pages the web scraping API with
engine: tls; no Rakuten collector exists in our catalogue today. - Failed requests are never billed, and every account gets $2 of free API usage per month, which is about 6,500 plain fetches of the Rakuten home.