Flickr is the largest public archive of licensed photographs on the web, and it is one of the few platforms in this series where the browser buys nothing at all. On 28 September 2026 we fetched it twenty-one times through residential exits in the United States, Germany and Japan. A public photostream returned 145 words to a plain HTTP client and exactly 145 words to a full browser, for 663,499 bytes against 2,005,256. A single photo page carried the caption, the view count and the licence in plain HTML for 536,306 bytes. And the oEmbed endpoint answered with no key at all in 1,189 bytes.
Nobody searches for this, and that is why nobody measures it
Search for Flickr proxies and Google has nothing to suggest: the autocomplete box is empty for the head term and for "proxies for flickr". The results page is not about proxies either. It is a Flickr group literally named "proxy" with 227 photos in it, a forum thread about an image proxy breaking Flickr embeds, two photostreams whose owners happen to be called "high proxies" and "Printing Proxies", a 2007 blog post about reaching Flickr from China, and two vendor pages, one of which redirects to a generic social-media page and the other of which sells mobile lines for managing accounts.
The intent that does exist sits under other words. "Flickr api" returns fifteen suggestions, led by "flickr api key", "flickr api pricing" and "flickr api rate limit". "Flickr scraper" returns three, including "flickr ai scraping". So the people arriving here want two things: to read public photo data at volume, and to know what Flickr allows. Both have numeric answers, and neither appears on that results page.
A photostream is 145 words with or without a browser
We audited flickr.com/photos/nasahqphoto/, the public photostream of NASA HQ Photo with 33,188 pictures, from three countries, each time once as a pure HTTP client and once fully rendered.
| Exit | Plain HTTP words | Rendered words | Plain HTTP bytes | Meta description |
|---|---|---|---|---|
| United States | 145 | 145 | 663,499 | Explore NASA HQ PHOTO’s 33,188 photos on Flickr! |
| Germany | 142 | 142 | 663,703 | Entdecke NASA HQ PHOTOs 33.188 Fotos auf Flickr! |
| Japan | 145 | 145 | 663,464 | Explore NASA HQ PHOTO’s 33,188 photos on Flickr! |
Six passes, one shape. The title, the h1, the canonical link and the description are server-rendered and identical in both views, and the audit found four JSON-LD blocks in the plain HTML: WebSite, Organization, Person and BlogPosting. The page also carries robots: noarchive, which tells search engines not to keep a cached copy but says nothing about reading it.
What the browser adds is the photo grid, and it adds it expensively. A full render from the US cost 2,005,256 bytes and 10.5 seconds; the audit still counted 145 words, and even a generous text extraction that keeps every injected photo title reached 933 words. That is three times the bytes for a list of titles that each photo page hands you anyway, together with the fields the grid never shows.
The photo page carries the caption, the counters and the licence in plain HTML
The unit of value on Flickr is the photo page, not the stream. We fetched one over plain HTTP, /photos/nasahqphoto/55541281672/, and it answered in 536,306 bytes and 2.3 seconds with everything a catalogue needs: the full caption (a 96-word press description with the photographer credit), the counters (3,088 views, 2 faves, 0 comments), the upload date and the capture date, and the licence, both as the words "Some rights reserved" and as structured data:
{ "@type": "ImageObject",
"contentUrl": "https://live.staticflickr.com/65535/55541281672_37da73d6be_b.jpg",
"license": "https://creativecommons.org/licenses/by-nc-nd/4.0/",
"acquireLicensePage": "https://www.flickr.com/photos/nasahqphoto/55541281672",
"author": { "@type": "Person", "name": "NASA HQ PHOTO" } }
That block is the reason this platform is worth reading at all. The licence URL is machine-readable, it is in the first response, and it is the field that decides whether you may use the picture. A BreadcrumbList sits next to it with the owner and the photo title, so a parser that reads two JSON-LD objects has the owner, the title, the licence and the full-size URL without touching the DOM.
oEmbed answers in 1,189 bytes with no key
Flickr's oEmbed endpoint is open. We called it with the photostream URL and no credentials:
GET https://www.flickr.com/services/oembed/?url=https%3A%2F%2Fwww.flickr.com%2Fphotos%2Fnasahqphoto%2F&format=json
{ "type": "rich", "flickr_type": "photostream",
"title": "NASA HQ PHOTO", "author_name": "NASA HQ PHOTO",
"thumbnail_url": "https://live.staticflickr.com/65535/55541281672_37da73d6be_b.jpg",
"license_url": "https://creativecommons.org/licenses/by-nc-nd/4.0/",
"cache_age": 3600, "provider_name": "Flickr" }
1,189 bytes, 1.6 seconds, no key. For a photostream it returns the owner, the canonical URL, the licence and the most recent photo, which is how we found the photo page above. For a photo URL it returns the title, the author and the embed. What it does not return is any counter: no views, no faves, no comments. If you need those, the photo page is the source, at 451 times the bytes.
The REST API is the opposite case. Without a key, flickr.people.getInfo answers in 79 bytes: {"stat":"fail","code":100,"message":"Invalid API Key"}. With a key, Flickr's Developer Guide states the ceiling in one sentence: stay under 3,600 queries per hour across the whole key, aggregated over every user of your integration, and results may be cached for up to 24 hours. That is one query per second per key, which is the number to size against.
The exit country rewrites the number format, not the data
Three exits, three near-identical pages, one difference that breaks parsers. From the German exit Flickr localised the description off the IP alone, with no Accept-Language header sent: "Explore NASA HQ PHOTO’s 33,188 photos" became "Entdecke NASA HQ PHOTOs 33.188 Fotos". The count is the same; the thousands separator is a full stop. From Japan the page stayed in English with the comma. Word counts moved by three, from 145 to 142, because the German chrome is shorter.
int(re.sub(r"[^\d]", "", "33.188")) # 33188, correct
int(re.match(r"[\d,]+", "33.188").group()) # 33 <-- silent, wrong on every DE exit
Two things follow. First, on Flickr a rotating pool that mixes countries produces mixed number formats inside one job, and the failure is silent. Second, this is a formatting effect, not a content effect: the photos, the licences and the counters are identical from all three countries, so geo targeting on this platform is a parsing decision rather than a data decision. Pin one exit country per job, or strip every separator before converting, and never match on translated labels.
What robots.txt and the Developer Guide say
flickr.com/robots.txt is 2,210 bytes and it is a whitelist. Forty-five named user agents share one block with nineteen path exclusions (search, lightboxes, a few specific accounts). The list is worth reading for who is on it: alongside Googlebot, bingbot and the link unfurlers sit ChatGPT-User, OAI-SearchBot, Claude-User, Claude-SearchBot and PerplexityBot. The training crawlers GPTBot and ClaudeBot are not named at all, so they fall under the final rule, which is also the one that governs your client:
User-agent: *
Disallow: /
Then six sitemap indexes: users, tags, sets, photos, groups and cameras. Flickr publishes a map of its public content, admits the assistants that fetch on a user's behalf, and tells unnamed clients to stay out.
The Developer Guide is blunter than the robots file. Its best-practices list says screen scraping flickr.com is not the way, that the API is the scalable route, and that clients doing it get cut off. The API Terms of Use add the conditions that matter for a data product: display no more than 30 user photos per page, cache photos only for reasonable periods, remove anything the owner asks to remove within 24 hours, and never use the API to support surveillance of Flickr users. The photographs belong to the photographers, and commercial use of one requires a Creative Commons licence that allows it. Read public photostreams for research, brand and licence monitoring and attribution checks; for anything at volume, get a key and stay under 3,600 an hour.
What a gigabyte buys, and where the API wins
Nothing in twenty-one fetches looked at the address. Every response was a 200 from a rotating residential exit on the first attempt, from three countries, with no difference between them beyond the number format above. The cost driver on Flickr is bytes, and the arithmetic below uses residential proxies on the Basic line at $0.80/GB, counting a gigabyte as 10^9 bytes.
| Request | Bytes | Requests per GB | Cost per request |
|---|---|---|---|
| Photostream, plain HTTP | 663,499 | 1,507 | $0.00053 |
| Photostream, rendered | 2,005,256 | 499 | $0.00160 |
| Photo page, plain HTTP | 536,306 | 1,865 | $0.00043 |
| oEmbed, no key | 1,189 | 841,043 | $0.000001 |
Put a real job through it. Reading 100,000 photo pages a day for a month moves 3,000,000 pages and 1,609 GB, about $1,287 on residential Basic, and it returns captions, counters and licences. The same 3,000,000 lookups through oEmbed move 3.6 GB and cost about $2.85, and they return titles, authors and licences but no counters. The rendered route is the one to strike out: it costs three times the photostream bytes for the same 145 words and adds nothing the photo pages do not already carry. On Mastodon the same comparison came out at $0.76 against $555 for a month of profiles; the shape of the answer is the same here. Read the page that has the data, not the page that lists it.
The setting that works on Flickr
- Network: residential proxies, Basic line, $0.80/GB, rotating. Twenty-one fetches from three countries met no address-level obstacle; the pool is for throughput and for the one-query-per-second-per-key ceiling on the API, not for trust. Mobile at $2.30/GB and ISP at $2.50/IP per month buy nothing we could measure here.
- Fetch mode:
engine: tls, plain HTTP, never rendered. The photostream returns 145 words either way; the photo page returns the caption, the counters and the ImageObject licence block in 536,306 bytes; oEmbed returns the owner and the licence in 1,189 bytes. - Country: pin one, any one, per job. The data is identical from the United States, Germany and Japan; only the number format changes, and a German exit writes 33,188 as "33.188". Mixing exits inside a job mixes formats.
- When the proxy is not enough: there is no Flickr collector in our catalogue, and the pages do not need one. The web scraping API without rendering returns each photo page as clean Markdown or JSON from $0.0002 per page, retries and rotation included, and failed requests are never billed. For discovery across the open web rather than one photostream, the image search collector returns Google Images results with the source page and host per row. For volume on Flickr itself, the licensed route is a key at 3,600 queries an hour.
- Free tier: every account gets $2 of free API usage per month, which is 10,000 photo pages through the scraping API before you pay anything.
Read the photo pages over plain HTTP, take the licence from the JSON-LD, use oEmbed when the owner and the licence are all you need, and skip the browser. That is the whole configuration. For the platform in this series where the public API also made the browser unnecessary, see Bluesky; for the one where rendering was the only way to read anything, see TikTok.