# Cloudflare Crawl API: 1 in 54 Pages Needed JS

> The Cloudflare crawl API renders every page in a browser by default. We crawled 77 pages on 29 September 2026: 53 of 54 read the same without JavaScript.

[Home](https://quanticdata.io/)/[Blog](https://quanticdata.io/blog/)/Cloudflare Crawl API: 1 in 54 Pages Needed JS

# Cloudflare Crawl API: 1 in 54 Pages Needed JS

CrawlingSep 29, 2026·8 min read·By [Aldo Morese](https://quanticdata.io/about/), founder of QuanticData

Plain HTTP against rendered words on 54 crawled pages, 29 September 2026: docs 13,994 vs 12,334, store 4,725 vs 3,658, GOV.UK 9,441 vs 9,104, Notion 7,797 vs 7,812

On this page [How the Cloudflare crawl API works and what it bills](/blog/cloudflare-crawl-api/#how-the-cloudflare-crawl-api-works-and-what-it-bills) [53 of 54 pages read the same without a browser](/blog/cloudflare-crawl-api/#53-of-54-pages-read-the-same-without-a-browser) [Set render to false, and render only the pages that need it](/blog/cloudflare-crawl-api/#set-render-to-false-and-render-only-the-pages-that-need-it) [13 of 78 sites disallow the crawler in robots.txt](/blog/cloudflare-crawl-api/#13-of-78-sites-disallow-the-crawler-in-robots-txt) [When a different crawl API fits better](/blog/cloudflare-crawl-api/#when-a-different-crawl-api-fits-better) [The setting that works for a crawl API](/blog/cloudflare-crawl-api/#the-setting-that-works-for-a-crawl-api)

The Cloudflare crawl API launches a headless browser for every page unless you pass `render: false`, and on the sites we tested that browser added almost nothing. On 29 September 2026 we crawled 77 pages across four sites and compared both modes on 54: plain HTTP returned 35,957 words, rendering 32,908. One page in 54 gained more than 20%.

The /crawl endpoint went into open beta on 10 March 2026 and its documentation was updated on 26 September. It follows links from a start URL, returns HTML, Markdown or JSON, and bills rendered crawls as browser time. This post measures the two things its docs leave to you: whether your pages need the browser, and whether the sites you target let its crawler in. It closes with the setting we use and when a crawl API with country exits is the better tool.

## How the Cloudflare crawl API works and what it bills

A crawl is two calls. A `POST` to `/accounts/{account_id}/browser-run/crawl` with a `url` returns a job id; a `GET` on the job returns status and records, paginated in 10 MB pages with a cursor. Jobs run for up to seven days and results are kept for 14 days. The documented defaults are what decide the bill:

- **limit** defaults to 10 pages (maximum 100,000) and **depth** to 100,000.

- **render** defaults to `true`: every page is loaded in a headless browser and billed under Browser Run pricing. With `render: false` the crawl runs as a plain HTML fetch on Workers, not billed during the beta and billed under Workers pricing afterwards.

- **source** defaults to `all`, sitemaps plus links found on pages.

- **crawlPurposes** defaults to search, ai-input and ai-train, which matters in the robots section below.

Browser time costs $0.09 per hour beyond the 10 hours included in the Workers Paid plan, and the free plan is capped at 10 minutes of browser use per day. The docs' own example crawl of 50 pages reports 134.7 browser seconds, 2.7 seconds a page, so the free cap covers roughly 220 rendered pages a day. There is no country parameter in the documented options, and the user agent, `CloudflareBrowserRenderingCrawler/1.0`, cannot be changed.

## 53 of 54 pages read the same without a browser

We picked four site types a crawl is usually pointed at: a documentation site (the Workers docs, the endpoint's own example), a Shopify store, GOV.UK from a British exit, and a Next.js marketing site (Notion). Each was crawled to 20 pages with the [QuanticData crawl endpoint](https://quanticdata.io/crawl-map/); 77 pages came back. We then fetched the same 77 URLs twice through the batch API, once over plain HTTP and once in a headless browser, and compared word counts on the 54 pages where both passes returned a page.

| Site | Pages paired | Words, plain HTTP | Words, rendered | Pages gaining over 20% rendered |
| --- | --- | --- | --- | --- |
| Docs site | 13 | 13,994 | 12,334 | 0 |
| Shopify store | 9 | 4,725 | 3,658 | 1 |
| GOV.UK | 20 | 9,441 | 9,104 | 0 |
| Notion (Next.js) | 12 | 7,797 | 7,812 | 0 |
| **Total** | **54** | **35,957** | **32,908** | **1** |

On 48 of the 54 pages the plain fetch returned as many words as the browser or more. The browser often returns fewer, because the server-rendered HTML carries text that scripts later collapse into tabs, carousels and menus. The one page that gained was the store's home page, 173 words plain and 250 rendered, where product tiles fill in after load. Even the Next.js site, the stack people assume needs a browser, matched within 15 words across 12 pages: its pages are rendered on the server.

Across all 77 crawled pages the plain fetch returned 55,365 words, with a median per page of 485 on the docs, 620 on the store, 178 on GOV.UK and 699 on Notion. Our wider study of 800 top domains found the same shape: most sites serve their text in the first response, and 6.4% serve crawlers nothing readable without scripts ([do AI crawlers render JavaScript?](https://quanticdata.io/blog/do-ai-crawlers-render-javascript/)).

## Set render to false, and render only the pages that need it

The instruction that follows is simple: start every crawl with `render: false`. Check the output of a handful of pages against the live site. Turn rendering on only for the paths that come back thin, using include patterns so the browser runs on those paths and nowhere else.

The time cost is the other half. Our 77 URLs took 200 seconds as one plain-HTTP batch and 741 seconds rendered, 3.7 times longer, and rendered pages move far more bytes: the Notion home page alone weighed 1,105,226 bytes in the browser. On the /crawl endpoint the render default is also the only part that is billed during the beta, so leaving it on buys browser hours for text you already had.

## 13 of 78 sites disallow the crawler in robots.txt

The /crawl endpoint obeys robots.txt, including `crawl-delay`, and waits 0.5 seconds between requests to a domain when none is set. Disallowed URLs come back with the status `disallowed`. So before choosing it we read the robots.txt of 78 sites through our batch API: 48 of the retail, travel, social and marketplace targets we measure, and 30 US and UK news sites. All 78 answered 200.

- **3 name it and disallow it**: usatoday.com, bbc.com and amazon.com each have a `User-agent: CloudflareBrowserRenderingCrawler` group with `Disallow: /`.

- **10 disallow every unnamed client**, which includes it: linkedin.com, facebook.com, instagram.com, reddit.com, netflix.com, trustpilot.com, bet365.com, world.taobao.com, wsj.com and reuters.com.

- **1 more rejects the default request**: vinted.it sets `Content-Signal: ai-train=no` for all agents. Because the default `crawlPurposes` include ai-train, the documented result is a 400 until you declare only the purposes you actually have, for example `["search"]`.

That is 14 of 78 sites, 18%, where a default call returns nothing useful. Content Signals are still rare: only 2 of the 78 files use the directive at all, and theatlantic.com applies it only to named search bots. The honest reading is that robots.txt is the rule for any crawler, this one included. When a site disallows automated clients, the answer is its API, its feeds or a licence, not a different crawler. Our [robots.txt tester](https://quanticdata.io/tools/robots-txt-tester/) shows the deciding line for any URL and agent before you start, and the [AI crawler user agent list](https://quanticdata.io/blog/ai-crawler-user-agent-list/) maps who is who.

## When a different crawl API fits better

For a site that allows it, with content in the first HTML response and no need to be seen from a particular country, the /crawl endpoint with `render: false` is hard to beat on price while the beta lasts. Three cases point elsewhere.

1. **The page depends on the visitor's country.** Prices, stock, currency, legal notices and whole catalogues change with the exit IP. The /crawl options have no country parameter; our crawl takes `country` and runs through residential exits in that country. GOV.UK in this test was crawled from a British exit.

2. **You want the crawl inside a wider job.** One API key covers `map` to list a site's URLs first, search results, batch fetches of known URLs, collectors that return structured rows, and the same tools over MCP for an agent.

3. **You do not want a platform plan.** Our crawl is billed per page returned, and pages we could not fetch are refunded.

Whichever tool you use, the legal layer does not change: read the terms, respect robots.txt and rate limits, and keep personal data out unless you have a lawful basis. [Is web crawling legal?](https://quanticdata.io/blog/is-web-crawling-legal/) sets out the four layers, and [how to web crawl in Python](https://quanticdata.io/blog/how-to-web-crawl-python/) shows the same loop by hand.

## The setting that works for a crawl API

- **Rendering:** off by default. 53 of 54 pages we compared read the same without a browser, and the plain pass finished 3.7 times faster.

- **Network:** residential exits, pinned to the country whose version of the site you need; the [residential proxies](https://quanticdata.io/residential-proxies/) Basic line is $0.80/GB when you run your own crawler.

- **Endpoint:** the [crawl and map API](https://quanticdata.io/crawl-map/), an async breadth-first crawl up to 500 pages and depth 10 at $0.0003 per page, unfetched pages refunded; `map` lists a site's URLs first for $0.0005. Each of our four 20-page crawls cost $0.006 or less.

- **Before the first request:** run the seed through the robots.txt tester and drop the sites that disallow automated clients.

- **When pages are not enough:** for single URLs use the [web scraping API](https://quanticdata.io/web-scraping-api/) at $0.0002 per page, $0.001 with rendering, and for structured targets a collector, such as the [Google News collector](https://quanticdata.io/collectors/google-news-api/) at $0.0005 per delivered article.

- **Free tier:** every account gets $2 of free API usage per month.

### Sources & further reading

- [RFC 9309: Robots Exclusion Protocol, IETF](https://www.rfc-editor.org/rfc/rfc9309.html)

- [bbc.com/robots.txt (fetched 29 September 2026)](https://www.bbc.com/robots.txt)

- [usatoday.com/robots.txt (fetched 29 September 2026)](https://www.usatoday.com/robots.txt)

- [amazon.com/robots.txt (fetched 29 September 2026)](https://www.amazon.com/robots.txt)

- [vinted.it/robots.txt (fetched 29 September 2026)](https://www.vinted.it/robots.txt)

## FAQ

Quick answers on cloudflare crawl api.

[Something else? Ask us](mailto:hello@quanticdata.io)

### What is the Cloudflare crawl API?

It is the /crawl endpoint of Cloudflare's Browser Run service, in open beta since 10 March 2026. One POST starts a job that follows links from a start URL, by default up to 10 pages, and a GET returns the pages as HTML, Markdown or JSON. By default every page is rendered in a headless browser.

### How much does the Cloudflare crawl API cost?

Rendered crawls are billed as browser time: $0.09 per hour beyond the 10 hours included in Workers Paid, and the free plan allows 10 minutes a day. The docs' example used 134.7 browser seconds for 50 pages. Crawls with render set to false run on Workers and are not billed during the beta.

### Should I set render to false on the /crawl endpoint?

Yes, as the default. On 54 pages from four sites that we fetched both ways on 29 September 2026, plain HTTP returned 35,957 words and the browser 32,908, and only one page gained more than 20% when rendered. Turn rendering on only for the paths that come back thin.

### Which sites block the Cloudflare crawler?

In the robots.txt of 78 sites we read on 29 September 2026, 3 disallow CloudflareBrowserRenderingCrawler by name (usatoday.com, bbc.com, amazon.com) and 10 more disallow every unnamed client. One more, vinted.it, rejects the default request through a Content-Signal ai-train=no line. That is 14 of 78.

### Can I choose the country the Cloudflare crawler uses?

The documented /crawl options have no country parameter, and the user agent is fixed. When the page changes by country, use a crawl API that takes a country: ours runs through residential exits in the country you name, at $0.0003 per crawled page.

### What is Content-Signal in robots.txt?

A robots.txt line in which a site states whether its content may be used for search, ai-input or ai-train. The /crawl endpoint rejects a job with a 400 when a signal says no to a purpose you declared, and it declares all three by default. Only 2 of the 78 files we read used the directive.

## Crawl without a browser first

Our crawl returned 77 pages from four sites over plain HTTP, and 53 of 54 matched their rendered text. Every account gets $2 of free API usage each month, and pages we cannot fetch are refunded.

[Start free — $2/month included](https://quanticdata.io/signup/)[Explore Website Crawler API](https://quanticdata.io/crawl-map/)

## Related reading

[Crawling How to Web Crawl in Python A working Python crawl loop in 40 lines, the same job in Scrapy, and when to swap the loop for a crawl API — with real cost math per 10,000 pages. Read more](https://quanticdata.io/blog/how-to-web-crawl-python/) [Crawling Is Web Crawling Legal? Rules by Layer Crawling public pages is generally lawful. Legality turns on four separate layers — access, contract, content rights and privacy — plus how much load you create. Read more](https://quanticdata.io/blog/is-web-crawling-legal/) [Crawling What Is Web Crawling in Python? The frontier algorithm behind every crawler, a minimal Python example, how Scrapy and Crawlee fit in, and the scaling wall where DIY stops paying. Read more](https://quanticdata.io/blog/what-is-web-crawling-in-python/)

---

Source: https://quanticdata.io/blog/cloudflare-crawl-api/ · Site index for AI: https://quanticdata.io/llms.txt
