The Cloudflare crawl API launches a headless browser for every page unless you pass render: false, and on the sites we tested that browser added almost nothing. On 29 September 2026 we crawled 77 pages across four sites and compared both modes on 54: plain HTTP returned 35,957 words, rendering 32,908. One page in 54 gained more than 20%.
The /crawl endpoint went into open beta on 10 March 2026 and its documentation was updated on 26 September. It follows links from a start URL, returns HTML, Markdown or JSON, and bills rendered crawls as browser time. This post measures the two things its docs leave to you: whether your pages need the browser, and whether the sites you target let its crawler in. It closes with the setting we use and when a crawl API with country exits is the better tool.
How the Cloudflare crawl API works and what it bills
A crawl is two calls. A POST to /accounts/{account_id}/browser-run/crawl with a url returns a job id; a GET on the job returns status and records, paginated in 10 MB pages with a cursor. Jobs run for up to seven days and results are kept for 14 days. The documented defaults are what decide the bill:
- limit defaults to 10 pages (maximum 100,000) and depth to 100,000.
- render defaults to
true: every page is loaded in a headless browser and billed under Browser Run pricing. Withrender: falsethe crawl runs as a plain HTML fetch on Workers, not billed during the beta and billed under Workers pricing afterwards. - source defaults to
all, sitemaps plus links found on pages. - crawlPurposes defaults to search, ai-input and ai-train, which matters in the robots section below.
Browser time costs $0.09 per hour beyond the 10 hours included in the Workers Paid plan, and the free plan is capped at 10 minutes of browser use per day. The docs' own example crawl of 50 pages reports 134.7 browser seconds, 2.7 seconds a page, so the free cap covers roughly 220 rendered pages a day. There is no country parameter in the documented options, and the user agent, CloudflareBrowserRenderingCrawler/1.0, cannot be changed.
53 of 54 pages read the same without a browser
We picked four site types a crawl is usually pointed at: a documentation site (the Workers docs, the endpoint's own example), a Shopify store, GOV.UK from a British exit, and a Next.js marketing site (Notion). Each was crawled to 20 pages with the QuanticData crawl endpoint; 77 pages came back. We then fetched the same 77 URLs twice through the batch API, once over plain HTTP and once in a headless browser, and compared word counts on the 54 pages where both passes returned a page.
| Site | Pages paired | Words, plain HTTP | Words, rendered | Pages gaining over 20% rendered |
|---|---|---|---|---|
| Docs site | 13 | 13,994 | 12,334 | 0 |
| Shopify store | 9 | 4,725 | 3,658 | 1 |
| GOV.UK | 20 | 9,441 | 9,104 | 0 |
| Notion (Next.js) | 12 | 7,797 | 7,812 | 0 |
| Total | 54 | 35,957 | 32,908 | 1 |
On 48 of the 54 pages the plain fetch returned as many words as the browser or more. The browser often returns fewer, because the server-rendered HTML carries text that scripts later collapse into tabs, carousels and menus. The one page that gained was the store's home page, 173 words plain and 250 rendered, where product tiles fill in after load. Even the Next.js site, the stack people assume needs a browser, matched within 15 words across 12 pages: its pages are rendered on the server.
Across all 77 crawled pages the plain fetch returned 55,365 words, with a median per page of 485 on the docs, 620 on the store, 178 on GOV.UK and 699 on Notion. Our wider study of 800 top domains found the same shape: most sites serve their text in the first response, and 6.4% serve crawlers nothing readable without scripts (do AI crawlers render JavaScript?).
Set render to false, and render only the pages that need it
The instruction that follows is simple: start every crawl with render: false. Check the output of a handful of pages against the live site. Turn rendering on only for the paths that come back thin, using include patterns so the browser runs on those paths and nowhere else.
The time cost is the other half. Our 77 URLs took 200 seconds as one plain-HTTP batch and 741 seconds rendered, 3.7 times longer, and rendered pages move far more bytes: the Notion home page alone weighed 1,105,226 bytes in the browser. On the /crawl endpoint the render default is also the only part that is billed during the beta, so leaving it on buys browser hours for text you already had.
13 of 78 sites disallow the crawler in robots.txt
The /crawl endpoint obeys robots.txt, including crawl-delay, and waits 0.5 seconds between requests to a domain when none is set. Disallowed URLs come back with the status disallowed. So before choosing it we read the robots.txt of 78 sites through our batch API: 48 of the retail, travel, social and marketplace targets we measure, and 30 US and UK news sites. All 78 answered 200.
- 3 name it and disallow it: usatoday.com, bbc.com and amazon.com each have a
User-agent: CloudflareBrowserRenderingCrawlergroup withDisallow: /. - 10 disallow every unnamed client, which includes it: linkedin.com, facebook.com, instagram.com, reddit.com, netflix.com, trustpilot.com, bet365.com, world.taobao.com, wsj.com and reuters.com.
- 1 more rejects the default request: vinted.it sets
Content-Signal: ai-train=nofor all agents. Because the defaultcrawlPurposesinclude ai-train, the documented result is a 400 until you declare only the purposes you actually have, for example["search"].
That is 14 of 78 sites, 18%, where a default call returns nothing useful. Content Signals are still rare: only 2 of the 78 files use the directive at all, and theatlantic.com applies it only to named search bots. The honest reading is that robots.txt is the rule for any crawler, this one included. When a site disallows automated clients, the answer is its API, its feeds or a licence, not a different crawler. Our robots.txt tester shows the deciding line for any URL and agent before you start, and the AI crawler user agent list maps who is who.
When a different crawl API fits better
For a site that allows it, with content in the first HTML response and no need to be seen from a particular country, the /crawl endpoint with render: false is hard to beat on price while the beta lasts. Three cases point elsewhere.
- The page depends on the visitor's country. Prices, stock, currency, legal notices and whole catalogues change with the exit IP. The /crawl options have no country parameter; our crawl takes
countryand runs through residential exits in that country. GOV.UK in this test was crawled from a British exit. - You want the crawl inside a wider job. One API key covers
mapto list a site's URLs first, search results, batch fetches of known URLs, collectors that return structured rows, and the same tools over MCP for an agent. - You do not want a platform plan. Our crawl is billed per page returned, and pages we could not fetch are refunded.
Whichever tool you use, the legal layer does not change: read the terms, respect robots.txt and rate limits, and keep personal data out unless you have a lawful basis. Is web crawling legal? sets out the four layers, and how to web crawl in Python shows the same loop by hand.
The setting that works for a crawl API
- Rendering: off by default. 53 of 54 pages we compared read the same without a browser, and the plain pass finished 3.7 times faster.
- Network: residential exits, pinned to the country whose version of the site you need; the residential proxies Basic line is $0.80/GB when you run your own crawler.
- Endpoint: the crawl and map API, an async breadth-first crawl up to 500 pages and depth 10 at $0.0003 per page, unfetched pages refunded;
maplists a site's URLs first for $0.0005. Each of our four 20-page crawls cost $0.006 or less. - Before the first request: run the seed through the robots.txt tester and drop the sites that disallow automated clients.
- When pages are not enough: for single URLs use the web scraping API at $0.0002 per page, $0.001 with rendering, and for structured targets a collector, such as the Google News collector at $0.0005 per delivered article.
- Free tier: every account gets $2 of free API usage per month.