# ChatGPT Web Scraping: 20 Live Pages via MCP

> ChatGPT web scraping through an MCP app: we read 20 live pages on 2 October 2026. All 20 came back, median 5.4 s, 92% fewer tokens than raw HTML, $0.0056 total.

[Home](https://quanticdata.io/)/[Blog](https://quanticdata.io/blog/)/ChatGPT Web Scraping: 20 Live Pages via MCP

# ChatGPT Web Scraping: 20 Live Pages via MCP

AI scrapingOct 2, 2026·8 min read·By [Aldo Morese](https://quanticdata.io/about/), founder of QuanticData

ChatGPT web scraping measured on 20 live pages: raw HTML of nine pages against the Markdown the scrape tool returned, and how many pages came back over plain HTTP versus rendered from a US exit

On this page [ChatGPT web scraping works through an MCP app, set up in 5 steps](/blog/chatgpt-web-scraping/#chatgpt-web-scraping-works-through-an-mcp-app-set-up-in-5-st) [We read 20 live pages through the scrape tool: all 20 came back](/blog/chatgpt-web-scraping/#we-read-20-live-pages-through-the-scrape-tool-all-20-came-ba) [The Markdown is 92% smaller than the raw HTML of the same pages](/blog/chatgpt-web-scraping/#the-markdown-is-92-smaller-than-the-raw-html-of-the-same-pag) [Search first, then read: 3 Google searches in 7.6 to 15.7 seconds](/blog/chatgpt-web-scraping/#search-first-then-read-3-google-searches-in-7-6-to-15-7-seco) [20 pages cost $0.0056, and $2 a month covers 10,000](/blog/chatgpt-web-scraping/#20-pages-cost-0-0056-and-2-a-month-covers-10-000) [How the app differs from ChatGPT's built-in browsing](/blog/chatgpt-web-scraping/#how-the-app-differs-from-chatgpt-s-built-in-browsing) [The setting that works on ChatGPT](/blog/chatgpt-web-scraping/#the-setting-that-works-on-chatgpt)

ChatGPT web scraping works once you connect an MCP server as a developer-mode app: ChatGPT then calls the server's scrape and search tools and reads the live page instead of guessing. On 2 October 2026 we sent 20 public pages through the QuanticData scrape tool with its default settings. All 20 came back: 18 over plain HTTP on the first call at a median 5.4 seconds, and 2 rendered from a US exit. Nine pages of raw HTML, about 947,431 tokens, arrived as 75,777 tokens of Markdown, 92% fewer. The 20 pages cost $0.0056.

## ChatGPT web scraping works through an MCP app, set up in 5 steps

ChatGPT on its own answers from its training data, its built-in search, or a page you paste. To make it read any public URL you name, OpenAI's developer mode lets you add a remote MCP server as an app. OpenAI lists developer mode for Pro, Plus, Business, Enterprise and Education accounts on the web, and describes it as full MCP client support for read and write tools.

The steps below follow OpenAI's developer-mode guide and our [MCP server](https://quanticdata.io/mcp-server/) page:

1. In ChatGPT, open **Settings**, then **Security and login**, and turn on **Developer mode**.

2. Go to ChatGPT Plugins, select the plus button, and give the app a name and a description, for example QuanticData and "live web pages and Google results".

3. Under **Connection**, enter `https://api.quanticdata.io/mcp`, including the `/mcp` path, and create the connection. No key field is needed.

4. ChatGPT opens the QuanticData sign-in page. Sign in and click **Allow access**. The tools ChatGPT discovers include `scrape`, `search`, `search_and_read`, `map`, `crawl`, `batch` and `seo_audit`.

5. In a new chat, choose **Developer mode** from the plus menu and select the QuanticData app for the conversation.

The sign-in is OAuth 2.1, so there is no API key to copy: ChatGPT receives a 1-hour access token with a rotating 30-day refresh token, valid only for the MCP endpoint, and a dedicated key named after the client appears on your dashboard. ../mcp-server-oauth/ covers the flow in detail, and Disconnect on the dashboard MCP page ends it.

## We read 20 live pages through the scrape tool: all 20 came back

We picked the kinds of pages people ask ChatGPT to read: encyclopedia articles, news sections, a forum front page, a subreddit, GitHub READMEs, reference docs, pricing pages, product pages, a movie page, a Q&A thread, a paper abstract and a package page. Each went through the scrape endpoint exactly as the MCP tool sends it: Markdown, smart content mode, engine auto, two requests in flight, from Europe, on 2 October 2026.

| Page | Fetch | Seconds | Words | Markdown tokens |
| --- | --- | --- | --- | --- |
| Wikipedia, Web scraping | plain HTTP | 2.2 | 7,012 | 12,222 |
| Wikipedia, Model Context Protocol | plain HTTP | 18.7 | 2,927 | 5,748 |
| BBC Technology | plain HTTP | 8.4 | 2,949 | 4,767 |
| The Guardian, Technology | plain HTTP | 5.6 | 2,467 | 4,310 |
| Reuters Technology | plain HTTP | 22.0 | 758 | 1,727 |
| Hacker News front page | plain HTTP | 5.1 | 2,059 | 3,703 |
| Reddit r/webscraping | plain HTTP | 5.9 | 966 | 1,535 |
| GitHub, MCP servers README | plain HTTP | 4.6 | 1,608 | 3,256 |
| GitHub, openai-python README | plain HTTP | 6.4 | 6,356 | 11,893 |
| Python docs, urllib.request | plain HTTP | 3.8 | 10,403 | 18,601 |
| MDN, HTTP status reference page | plain HTTP | 3.1 | 631 | 1,066 |
| OpenAI API pricing | plain HTTP | 5.3 | 3,023 | 5,693 |
| Anthropic pricing | plain HTTP | 11.3 | 967 | 1,585 |
| Stripe pricing | plain HTTP | 4.6 | 5,020 | 10,036 |
| Amazon product page | plain HTTP | 4.5 | 3,121 | 4,838 |
| Apple AirPods Pro page | plain HTTP | 7.0 | 4,739 | 8,261 |
| IMDb title page | rendered, US exit | 10.9 | 4,298 | 7,037 |
| Stack Overflow question | rendered, US exit | 31.3 | 3,752 | 6,478 |
| arXiv abstract | plain HTTP | 3.1 | 1,224 | 2,045 |
| PyPI, requests | plain HTTP | 6.2 | 6,773 | 10,621 |

Eighteen pages came back on the first call over plain HTTP, with no browser, at a median of 5.4 seconds and a slowest of 22.0 seconds. The movie page and the Q&A thread came back on a second call with the exit pinned to the United States and the page rendered: 4,298 and 3,752 words. That is the one instruction to give ChatGPT when a page arrives nearly empty: call scrape again with country set to us.

## The Markdown is 92% smaller than the raw HTML of the same pages

For nine of the pages we also asked for the raw HTML. We count HTML tokens as characters divided by 4, which is an estimate; the Markdown counts are the tool's own token figures. The scrape tool's smart mode keeps the page body with its tables and links and drops navigation, footers and cookie banners.

| Page | Raw HTML, tokens | Markdown, tokens | Fewer |
| --- | --- | --- | --- |
| The Guardian, Technology | 175,715 | 4,310 | 97.5% |
| GitHub, MCP servers README | 88,113 | 3,256 | 96.3% |
| BBC Technology | 121,494 | 4,767 | 96.1% |
| Stripe pricing | 242,596 | 10,036 | 95.9% |
| Apple AirPods Pro page | 130,618 | 8,261 | 93.7% |
| PyPI, requests | 67,429 | 10,621 | 84.2% |
| Wikipedia, Web scraping | 58,868 | 12,222 | 79.2% |
| Python docs, urllib.request | 54,065 | 18,601 | 65.6% |
| Hacker News front page | 8,535 | 3,703 | 56.6% |
| **Nine pages** | **947,431** | **75,777** | **92.0%** |

Across all 20 pages the median page came to 5,266 Markdown tokens, and the largest, the Python reference page, to 18,601. When you only need one fact from a long page, tell ChatGPT to pass a `query` to scrape, which keeps only the matching sections, or a `max_tokens` cap, which cuts at a section boundary. ../mcp-web-scraper/ compares every content mode on 11 pages.

## Search first, then read: 3 Google searches in 7.6 to 15.7 seconds

Most questions start without a URL. The `search` tool returns Google, Bing or DuckDuckGo results as structured JSON, and `search_and_read` runs the search, fetches the top pages and returns numbered sources with one token-bounded context. Our three Google searches each returned 10 organic results, in 8.8, 15.7 and 7.6 seconds.

OpenAI's guide says ChatGPT treats any tool without the `readOnlyHint` annotation as a write action and asks you to confirm it. In the QuanticData app, `search`, `search_and_read`, `map` and `seo_audit` are marked read-only and run without a prompt. `scrape` is not, because it can also click and type on a page when you ask it to, so ChatGPT asks before each call; you can approve it once and have ChatGPT remember that choice for the rest of the conversation.

OpenAI also recommends naming the app and the tool when several could answer. A prompt that works:

```
Use the QuanticData app's search_and_read tool to find the three
most recent pages about the EU AI Act timeline, then summarize them
with numbered citations. Do not use built-in browsing.
```

## 20 pages cost $0.0056, and $2 a month covers 10,000

The [web scraping API](https://quanticdata.io/web-scraping-api/) behind the scrape tool charges per page, not per byte: $0.0002 for a plain HTTP page and $0.001 for a rendered one. Search runs through the [SERP API](https://quanticdata.io/serp-api/) from $0.0005 per search. Failed requests are never billed. Usage through ChatGPT is charged to your account balance like any API key.

| What ChatGPT asked for | Calls | Price each | Cost |
| --- | --- | --- | --- |
| Pages over plain HTTP | 18 | $0.0002 | $0.0036 |
| Pages rendered, US exit | 2 | $0.001 | $0.0020 |
| Google searches | 3 | $0.0005 | $0.0015 |
| **Total** | **23** |  | **$0.0071** |

The 20 pages moved 13.3 MB of HTML, and the heaviest page, at 2.5 MB, cost the same $0.001 as any rendered page. At these prices, $2 buys 10,000 plain pages, 2,000 rendered pages or 4,000 searches.

## How the app differs from ChatGPT's built-in browsing

OpenAI documents that when a user asks ChatGPT a question it may visit a page with the ChatGPT-User agent, and that this agent does not crawl automatically. The MCP app is a different path: the request leaves from our residential exits, in the country you pin, and the page comes back as Markdown, HTML or text with its title, word count and token count, so you and ChatGPT both see what was read.

Read only public pages, and respect each site's terms of use. The app's tool set in ChatGPT leaves out collectors that return people's contact details, and its raw request tool fetches public pages only. OpenAI calls developer mode powerful but dangerous and asks users to review every write action before approving it: keep that habit with any MCP app, ours included.

## The setting that works on ChatGPT

- **Connection**: a developer-mode app on a Plus, Pro, Business, Enterprise or Education account, URL `https://api.quanticdata.io/mcp`, OAuth sign-in, no key to paste.

- **Tool**: `search_and_read` for questions without a URL, at 10 organic results per search; `scrape` for a link you have, at a median 5.4 seconds on our 20 pages; `map` or `crawl` for a whole site.

- **Fetch mode**: leave engine on auto. 18 of 20 pages came back over plain HTTP at $0.0002, with no browser to pay for.

- **Country**: pin `us` when a page comes back nearly empty. The two pages that needed it returned 4,298 and 3,752 words rendered, at $0.001 each.

- **Token budget**: smart Markdown, 92% fewer tokens than raw HTML on nine pages; add `query` or `max_tokens` on long reference pages.

- **When a page is not enough**: for listings, places, jobs or prices across many pages, ask ChatGPT to call `list_collectors` and run a [collector](https://quanticdata.io/collectors/), which returns rows instead of pages, priced per delivered result.

- **Price**: $0.0002 per plain page, $0.001 rendered, $0.0005 per search, pay per success, and every account gets $2 of free API usage per month.

### Sources & further reading

- [ChatGPT Developer mode, OpenAI Developers (eligibility, setup, OAuth, tool confirmation; fetched 2 October 2026)](https://developers.openai.com/api/docs/guides/developer-mode)

- [Connect and test your plugin, OpenAI Developers (adding an MCP server in ChatGPT; fetched 2 October 2026)](https://developers.openai.com/plugins/deploy/connect-chatgpt)

- [Overview of OpenAI Crawlers, OpenAI Developers (ChatGPT-User; fetched 2 October 2026)](https://developers.openai.com/api/docs/bots)

- [How is ChatGPT able to extract webpages so quickly?, OpenAI Developer Community](https://community.openai.com/t/how-is-chatgpt-able-to-extract-webpages-so-quickly/857980)

## FAQ

Quick answers on chatgpt web scraping.

[Something else? Ask us](mailto:hello@quanticdata.io)

### Can ChatGPT scrape websites?

Yes, once an MCP server with a scrape tool is connected as a developer-mode app. Through the QuanticData app we read 20 public pages on 2 October 2026 and all 20 came back: 18 over plain HTTP at a median 5.4 seconds, 2 rendered from a US exit. Read only public pages and respect each site's terms.

### Which ChatGPT plans can connect an MCP server?

OpenAI lists developer mode for Pro, Plus, Business, Enterprise and Education accounts on the web. Turn it on under Settings, Security and login, then add the server from ChatGPT Plugins with the plus button: 5 steps in all, sign-in included.

### Do I need an API key to use QuanticData in ChatGPT?

No. The hosted endpoint https://api.quanticdata.io/mcp supports OAuth 2.1: you sign in and click Allow access. ChatGPT gets a 1-hour access token and a rotating 30-day refresh token bound to the MCP endpoint, and a dedicated key named after the client appears on your dashboard.

### How much does ChatGPT web scraping cost with QuanticData?

$0.0002 per page over plain HTTP, $0.001 per rendered page and $0.0005 per Google search, billed only on success. Our 20 pages and 3 searches cost $0.0071. Every account gets $2 of free API usage per month, which covers 10,000 plain pages.

### Why does ChatGPT ask me to confirm the scrape tool?

OpenAI treats any tool without the readOnlyHint annotation as a write action. Scrape can also click and type on a page, so it is not marked read-only; search, search_and_read, map and seo_audit are, and run without a prompt. You can approve scrape once and let ChatGPT remember it for the conversation.

### What should I do when a page comes back nearly empty in ChatGPT?

Ask ChatGPT to call scrape again with country set to us. In our test 2 of 20 pages needed this, and they returned 4,298 and 3,752 words rendered, in 10.9 and 31.3 seconds, at $0.001 each.

## Give ChatGPT the live web

All 20 pages in this test came back through the QuanticData app, at a median 5.4 seconds and 92% fewer tokens than raw HTML. Connect https://api.quanticdata.io/mcp with one sign-in, and every account gets $2 of free API usage per month.

[Start free — $2/month included](https://quanticdata.io/signup/)[Explore Web Scraping MCP Server for AI Agents](https://quanticdata.io/mcp-server/)

## Related reading

[AI scraping OpenAI Scraping Lawsuit: 0 of 30 Block Bingbot Unsealed filings in the New York Times-led case say Microsoft and OpenAI built training sets from news, including data gathered for Bing. We read the robots.txt of 30 news sites on 29 September 2026: 20 block GPTBot, 27 block ClaudeBot, and not one blocks Bingbot. Read more](https://quanticdata.io/blog/openai-scraping-lawsuit/) [AI scraping GPT-6.1 Sol Web Scraping: 134 of 134 Right The day after OpenAI released GPT-6.1 Sol we sent it the 134 hand-labelled scraper responses from our Jev tests: is this the page the URL asked for, or an empty shell, a sign-in page, an error, the wrong page? It got all 134 right, for $2.74 per 1,000 pages. So did GPT-6 Sol at the same price, and GPT-6 Luna at $0.147. Jev 1.13 got 131 in a third of a second, and all three of its misses came with a hesitant score. Read more](https://quanticdata.io/blog/gpt-6-1-sol-block-pages/) [AI scraping Pareto 26.10 Preview: 125 of 134 Pages Right The day after Unbiased listed Pareto 26.10 Preview on OpenRouter we sent it the 134 hand-labelled scraper responses from our Jev tests. It got 125 right at $0.84 per 1,000 pages, and accepted six pages whose entire visible text was the site name, each with 0.95 confidence. The September Pareto release rejected all six and scored 131. GPT-6 Luna scored 134 at $0.147. Read more](https://quanticdata.io/blog/pareto-26-10-preview-block-pages/)

---

Source: https://quanticdata.io/blog/chatgpt-web-scraping/ · Site index for AI: https://quanticdata.io/llms.txt
