Letterboxd is the film diary that became the most cited source of audience ratings on the web, and it hands a logged-out client almost everything except the rating. On 28 September 2026 we fetched it nineteen times through residential exits in the United States, the United Kingdom and Germany. A public film page returned 2,830 words to a plain HTTP client, identical from all three countries, with the full cast, crew, runtime and twelve popular reviews. The 4.5 average and the star distribution arrived in a separate 5,981-byte fragment. Member profiles are the one surface that answers only a real browser: 1,270,529 bytes and 32.9 seconds for a page that lists 2,584 films and 36,306 followers.
On this keyword, "proxy" is a film from 2013
Search for Letterboxd proxies and Google returns Letterboxd itself: the pages for Proxy (2013), Proxy (2015), Proxy (2017) and PROXY (2026), the site’s own search results for the word, an r/movies thread and Wikipedia. The autocomplete box is empty for the head term. The intent that exists is under two other phrases. "Letterboxd scraper" completes eight ways, and the suffixes tell you what people build: list scraper, review scraper, watchlist scraper, data scraper, github. "Letterboxd api" completes fourteen ways, and the suffixes tell you what they run into: reddit, github, python, free, pricing, down, cost.
Both questions have numeric answers, and neither is on that results page. This post is about reading what Letterboxd serves to anyone, for film research, audience analytics and review monitoring, and it ends with what the platform says about its API, which is the paragraph most people on that SERP are looking for.
A film page is 2,830 words without JavaScript, from any country
We audited letterboxd.com/film/parasite-2019/ from three countries, each time as a pure HTTP client and as a full browser.
| Exit | Plain HTTP words | Rendered words | Plain HTTP bytes | Canonical | JSON-LD |
|---|---|---|---|---|---|
| United States | 2,830 | 2,679 | 334,446 | absent | none |
| United Kingdom | 2,830 | 3,018 | 334,390 | absent | none |
| Germany | 2,830 | 2,793 | 334,390 | absent | none |
Six passes, one page. The plain response is the same 2,830 words from all three exits, with 56 bytes of variance. The rendered counts wobble by a few hundred words in both directions because the browser loads ad slots and a "similar films" carousel that differ per request; nothing in that wobble is film data. What the plain HTML carries is the whole card: title, year, original title, director, tagline, synopsis, a 133-minute runtime, the full cast list (30 names before the "show all" fold), twelve popular reviews with their like counts, and the "Mentioned by" stories.
Two things a parser should know before it starts. The page has no canonical link and no structured data of any kind, so there is no JSON-LD shortcut here as there is on Flickr or Snapchat; you read the DOM. And the h1 is "Letterboxd — Your life in film", the site’s slogan, on every page: the film title lives in the <title> and in the response header x-letterboxd-type: Film with an x-letterboxd-identifier, which is the stable id you want to key on.
The rating lives in a 5,981-byte fragment, not in the page
This is the finding that saves the most bandwidth. The average rating and the star distribution are not in the film HTML at all. The page loads them afterwards from a fragment endpoint, and that endpoint answers a plain HTTP client directly:
GET https://letterboxd.com/csi/film/parasite-2019/rating-histogram/
Rating Distribution
half-star 5,593 (0%) | 1 star 13.9K (0%) | 1.5 6,528 (0%) | 2 stars 42.1K (1%)
2.5 37.4K (1%) | 3 stars 240K (4%) | 3.5 283K (5%) | 4 stars 1.3M (22%)
4.5 1M (17%) | 5 stars 2.9M (50%)
Average: 4.5
5,981 bytes, 2.3 seconds, no cookie, no token. It is 56 times smaller than the film page and it carries the number every "letterboxd data scraper" query is really after: ten buckets, their counts and percentages, and the average. If ratings are the job, this fragment is the job, and the film page is optional. Note the counts are abbreviated ("2.9M", "13.9K"); parse the suffix or take the percentages, which are exact to the unit.
Member pages answer the browser, and nothing else
Film pages are served to everyone. Member pages are served to browsers. Over plain HTTP, letterboxd.com/dave/ does not return a profile; rendered, it returns a complete one in 1,270,529 bytes and 32.9 seconds: 2,584 films, 101 this year, 158 lists, 77 following, 36,306 followers, four favourite films, the recent diary with dates and star ratings, pinned and recent reviews with like counts, the member’s own rating distribution, a 620-film watchlist and the activity stream. The page even carries a full month calendar in the DOM, which is where most of the 32.9 seconds go.
So the platform has two surfaces with two settings. Film pages: plain HTTP, 334,446 bytes, 2,830 words. Member pages: a real browser, 1,270,529 bytes, half a minute each. Do not point the browser at film pages to keep one code path: the rendered film page cost 1,471,790 bytes and 13.4 seconds for 151 fewer words than the plain fetch. And be aware that the developer documentation site behaves like the member pages, answering only a browser; the API access page on the main site, by contrast, is plain HTML.
The API is by request, and the request page says who is refused
Letterboxd’s API page is short and unusually specific. Access is available by request only, by email, with the project named in the subject line; applications are read but not individually answered, and access is not guaranteed. Then the sentence that matters for anyone arriving from the "letterboxd api free" query: at this time they are not granting access for data-analysis, visualization or recommendation projects, for LLM or GPT-related use, for private or personal projects, or for anything that recreates features of the paid tiers. For film metadata they point you to TMDB; for your own data, every member profile has an RSS feed of diary entries, reviews and lists, and the account settings offer an export.
Read that as the platform’s statement of what it wants read by machines: your own diary via RSS and export, film metadata via TMDB, and public film pages via the site itself. Audience research on public film pages and the rating fragment sits inside that; harvesting member profiles at scale does not, and a member’s diary, ratings and lists are personal data wherever your users live.
What robots.txt closes, and why the posters matter
letterboxd.com/robots.txt is 3,402 bytes, generated with a public robots builder, and it opens with 24 named AI data scrapers refused everywhere: AI2Bot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Diffbot, GPTBot, Google-Extended, Meta-ExternalAgent, PetalBot and the rest, plus the Algolia crawler. Everyone else is allowed in, with twelve path patterns closed:
User-agent: *
Disallow: /*/by/* # sorting options
Disallow: /*/genre/* # Films by genre
Disallow: /*/decade/* # Films by decade
Disallow: /films/*/size/large/* # Films with large posters (and therefore stats)
Disallow: /*/friends/* # stuff grouped for users' friends
The comments are the platform explaining itself. Sort orders, genre, country, language, decade and year listings are closed because they multiply the crawl space; the large-poster views are closed because, in Letterboxd’s own words, they carry stats. Film pages, review pages and lists are not on the list. That is a cleaner signal than most platforms give: the canonical film page is meant to be read, the combinatorial listings are not, and the stats views are the ones they care about.
What a gigabyte buys on Letterboxd
Nothing in the film-page fetches looked at the address: 200 from a rotating residential exit on the first attempt, from three countries, identical bodies. The cost driver is bytes on film pages and browser time on member pages, and the arithmetic below uses residential proxies on the Basic line at $0.80/GB, counting a gigabyte as 10^9 bytes.
| Request | Bytes | Requests per GB | Cost per request |
|---|---|---|---|
| Film page, plain HTTP | 334,446 | 2,990 | $0.00027 |
| Rating histogram fragment, plain HTTP | 5,981 | 167,196 | $0.000005 |
| Film page, rendered | 1,471,790 | 679 | $0.00118 |
| Member page, rendered | 1,270,529 | 787 | $0.00102 |
Put a real job through it. A catalogue of 50,000 films with their rating distributions is 50,000 film pages plus 50,000 fragments: 17.0 GB, about $13.62, and no browser. The same catalogue through rendered film pages is 73.6 GB and about $59, for fewer words. Ten thousand member pages, which is the smallest useful sample for audience research, are 12.7 GB and about $10 in bandwidth, but at 32.9 seconds each they are also 91 browser-hours, which is the real cost line and the reason to keep that sample small. If you are sizing a pool before you buy, our note on how much proxy data you need does the same arithmetic in the other direction.
The setting that works on Letterboxd
- Network: residential proxies, Basic line, $0.80/GB, rotating for film pages, sticky for the member sessions you render. Film pages met no address-level obstacle from three countries; the pool is for throughput. Mobile at $2.30/GB and ISP at $2.50/IP per month buy nothing we could measure here.
- Fetch mode: two settings. Film pages and the rating fragment:
engine: tls, plain HTTP, 2,830 words in 334,446 bytes plus 5,981 bytes for the histogram. Member pages: rendered, 1,270,529 bytes and 32.9 seconds each, because the plain client does not get the page. - Country: none required. The film page returned 334,446, 334,390 and 334,390 bytes and the same 2,830 words from the United States, the United Kingdom and Germany. Spread the pool for throughput.
- When the proxy is not enough: there is no Letterboxd collector in our catalogue, and film pages do not need one. The web scraping API returns each film page as clean Markdown or JSON from $0.0002 per page without rendering and $0.001 with, retries and rotation included, and failed requests are never billed. For critic and audience scores from another source, the Rotten Tomatoes collector returns Tomatometer and audience scores per delivered result. For your own diary, use the RSS feed and export Letterboxd provides.
- Free tier: every account gets $2 of free API usage per month, which is 10,000 film pages through the scraping API before you pay anything.
Read film pages over plain HTTP, take the rating from the histogram fragment, render only the member pages you have a reason to read, and stay inside what the API page says the platform wants read by machines. That is the whole configuration. For the platform in this series where a public page also carried the data and the API did not, see Facebook; for the one where the open API made the browser pointless, see Mastodon.