Documentation Python quickstart Blog Free tools Enterprise solutions hello@quanticdata.ioLog in

Cloudflare Crawl API: 1 in 54 Pages Needed JS

Plain HTTP against rendered words on 54 crawled pages, 29 September 2026: docs 13,994 vs 12,334, store 4,725 vs 3,658, GOV.UK 9,441 vs 9,104, Notion 7,797 vs 7,812
Plain HTTP against rendered words on 54 crawled pages, 29 September 2026: docs 13,994 vs 12,334, store 4,725 vs 3,658, GOV.UK 9,441 vs 9,104, Notion 7,797 vs 7,812

The Cloudflare crawl API launches a headless browser for every page unless you pass render: false, and on the sites we tested that browser added almost nothing. On 29 September 2026 we crawled 77 pages across four sites and compared both modes on 54: plain HTTP returned 35,957 words, rendering 32,908. One page in 54 gained more than 20%.

The /crawl endpoint went into open beta on 10 March 2026 and its documentation was updated on 26 September. It follows links from a start URL, returns HTML, Markdown or JSON, and bills rendered crawls as browser time. This post measures the two things its docs leave to you: whether your pages need the browser, and whether the sites you target let its crawler in. It closes with the setting we use and when a crawl API with country exits is the better tool.

How the Cloudflare crawl API works and what it bills

A crawl is two calls. A POST to /accounts/{account_id}/browser-run/crawl with a url returns a job id; a GET on the job returns status and records, paginated in 10 MB pages with a cursor. Jobs run for up to seven days and results are kept for 14 days. The documented defaults are what decide the bill:

  • limit defaults to 10 pages (maximum 100,000) and depth to 100,000.
  • render defaults to true: every page is loaded in a headless browser and billed under Browser Run pricing. With render: false the crawl runs as a plain HTML fetch on Workers, not billed during the beta and billed under Workers pricing afterwards.
  • source defaults to all, sitemaps plus links found on pages.
  • crawlPurposes defaults to search, ai-input and ai-train, which matters in the robots section below.

Browser time costs $0.09 per hour beyond the 10 hours included in the Workers Paid plan, and the free plan is capped at 10 minutes of browser use per day. The docs' own example crawl of 50 pages reports 134.7 browser seconds, 2.7 seconds a page, so the free cap covers roughly 220 rendered pages a day. There is no country parameter in the documented options, and the user agent, CloudflareBrowserRenderingCrawler/1.0, cannot be changed.

53 of 54 pages read the same without a browser

We picked four site types a crawl is usually pointed at: a documentation site (the Workers docs, the endpoint's own example), a Shopify store, GOV.UK from a British exit, and a Next.js marketing site (Notion). Each was crawled to 20 pages with the QuanticData crawl endpoint; 77 pages came back. We then fetched the same 77 URLs twice through the batch API, once over plain HTTP and once in a headless browser, and compared word counts on the 54 pages where both passes returned a page.

SitePages pairedWords, plain HTTPWords, renderedPages gaining over 20% rendered
Docs site1313,99412,3340
Shopify store94,7253,6581
GOV.UK209,4419,1040
Notion (Next.js)127,7977,8120
Total5435,95732,9081

On 48 of the 54 pages the plain fetch returned as many words as the browser or more. The browser often returns fewer, because the server-rendered HTML carries text that scripts later collapse into tabs, carousels and menus. The one page that gained was the store's home page, 173 words plain and 250 rendered, where product tiles fill in after load. Even the Next.js site, the stack people assume needs a browser, matched within 15 words across 12 pages: its pages are rendered on the server.

Across all 77 crawled pages the plain fetch returned 55,365 words, with a median per page of 485 on the docs, 620 on the store, 178 on GOV.UK and 699 on Notion. Our wider study of 800 top domains found the same shape: most sites serve their text in the first response, and 6.4% serve crawlers nothing readable without scripts (do AI crawlers render JavaScript?).

Set render to false, and render only the pages that need it

The instruction that follows is simple: start every crawl with render: false. Check the output of a handful of pages against the live site. Turn rendering on only for the paths that come back thin, using include patterns so the browser runs on those paths and nowhere else.

The time cost is the other half. Our 77 URLs took 200 seconds as one plain-HTTP batch and 741 seconds rendered, 3.7 times longer, and rendered pages move far more bytes: the Notion home page alone weighed 1,105,226 bytes in the browser. On the /crawl endpoint the render default is also the only part that is billed during the beta, so leaving it on buys browser hours for text you already had.

13 of 78 sites disallow the crawler in robots.txt

The /crawl endpoint obeys robots.txt, including crawl-delay, and waits 0.5 seconds between requests to a domain when none is set. Disallowed URLs come back with the status disallowed. So before choosing it we read the robots.txt of 78 sites through our batch API: 48 of the retail, travel, social and marketplace targets we measure, and 30 US and UK news sites. All 78 answered 200.

  • 3 name it and disallow it: usatoday.com, bbc.com and amazon.com each have a User-agent: CloudflareBrowserRenderingCrawler group with Disallow: /.
  • 10 disallow every unnamed client, which includes it: linkedin.com, facebook.com, instagram.com, reddit.com, netflix.com, trustpilot.com, bet365.com, world.taobao.com, wsj.com and reuters.com.
  • 1 more rejects the default request: vinted.it sets Content-Signal: ai-train=no for all agents. Because the default crawlPurposes include ai-train, the documented result is a 400 until you declare only the purposes you actually have, for example ["search"].

That is 14 of 78 sites, 18%, where a default call returns nothing useful. Content Signals are still rare: only 2 of the 78 files use the directive at all, and theatlantic.com applies it only to named search bots. The honest reading is that robots.txt is the rule for any crawler, this one included. When a site disallows automated clients, the answer is its API, its feeds or a licence, not a different crawler. Our robots.txt tester shows the deciding line for any URL and agent before you start, and the AI crawler user agent list maps who is who.

When a different crawl API fits better

For a site that allows it, with content in the first HTML response and no need to be seen from a particular country, the /crawl endpoint with render: false is hard to beat on price while the beta lasts. Three cases point elsewhere.

  1. The page depends on the visitor's country. Prices, stock, currency, legal notices and whole catalogues change with the exit IP. The /crawl options have no country parameter; our crawl takes country and runs through residential exits in that country. GOV.UK in this test was crawled from a British exit.
  2. You want the crawl inside a wider job. One API key covers map to list a site's URLs first, search results, batch fetches of known URLs, collectors that return structured rows, and the same tools over MCP for an agent.
  3. You do not want a platform plan. Our crawl is billed per page returned, and pages we could not fetch are refunded.

Whichever tool you use, the legal layer does not change: read the terms, respect robots.txt and rate limits, and keep personal data out unless you have a lawful basis. Is web crawling legal? sets out the four layers, and how to web crawl in Python shows the same loop by hand.

The setting that works for a crawl API

  • Rendering: off by default. 53 of 54 pages we compared read the same without a browser, and the plain pass finished 3.7 times faster.
  • Network: residential exits, pinned to the country whose version of the site you need; the residential proxies Basic line is $0.80/GB when you run your own crawler.
  • Endpoint: the crawl and map API, an async breadth-first crawl up to 500 pages and depth 10 at $0.0003 per page, unfetched pages refunded; map lists a site's URLs first for $0.0005. Each of our four 20-page crawls cost $0.006 or less.
  • Before the first request: run the seed through the robots.txt tester and drop the sites that disallow automated clients.
  • When pages are not enough: for single URLs use the web scraping API at $0.0002 per page, $0.001 with rendering, and for structured targets a collector, such as the Google News collector at $0.0005 per delivered article.
  • Free tier: every account gets $2 of free API usage per month.

Sources & further reading

FAQ

Quick answers on cloudflare crawl api.

Something else? Ask us

What is the Cloudflare crawl API?

It is the /crawl endpoint of Cloudflare's Browser Run service, in open beta since 10 March 2026. One POST starts a job that follows links from a start URL, by default up to 10 pages, and a GET returns the pages as HTML, Markdown or JSON. By default every page is rendered in a headless browser.

How much does the Cloudflare crawl API cost?

Rendered crawls are billed as browser time: $0.09 per hour beyond the 10 hours included in Workers Paid, and the free plan allows 10 minutes a day. The docs' example used 134.7 browser seconds for 50 pages. Crawls with render set to false run on Workers and are not billed during the beta.

Should I set render to false on the /crawl endpoint?

Yes, as the default. On 54 pages from four sites that we fetched both ways on 29 September 2026, plain HTTP returned 35,957 words and the browser 32,908, and only one page gained more than 20% when rendered. Turn rendering on only for the paths that come back thin.

Which sites block the Cloudflare crawler?

In the robots.txt of 78 sites we read on 29 September 2026, 3 disallow CloudflareBrowserRenderingCrawler by name (usatoday.com, bbc.com, amazon.com) and 10 more disallow every unnamed client. One more, vinted.it, rejects the default request through a Content-Signal ai-train=no line. That is 14 of 78.

Can I choose the country the Cloudflare crawler uses?

The documented /crawl options have no country parameter, and the user agent is fixed. When the page changes by country, use a crawl API that takes a country: ours runs through residential exits in the country you name, at $0.0003 per crawled page.

What is Content-Signal in robots.txt?

A robots.txt line in which a site states whether its content may be used for search, ai-input or ai-train. The /crawl endpoint rejects a job with a 400 when a signal says no to a purpose you declared, and it declares all three by default. Only 2 of the 78 files we read used the directive.

Crawl without a browser first

Our crawl returned 77 pages from four sites over plain HTTP, and 53 of 54 matched their rendered text. Every account gets $2 of free API usage each month, and pages we cannot fetch are refunded.

Related reading