Documentation Python quickstart Blog Free tools Enterprise solutions hello@quanticdata.ioLog in

ChatGPT Web Scraping: 20 Live Pages via MCP

Twenty live pages read into ChatGPT through the QuanticData MCP scrape tool on 2 October 2026: nine pages of raw HTML at 947,431 tokens against 75,777 tokens of Markdown, 92% fewer; the card shows all 20 pages back, 18 over plain HTTP at a 5.4 s median, for $0.0056.
ChatGPT web scraping measured on 20 live pages: raw HTML of nine pages against the Markdown the scrape tool returned, and how many pages came back over plain HTTP versus rendered from a US exit

ChatGPT web scraping works once you connect an MCP server as a developer-mode app: ChatGPT then calls the server's scrape and search tools and reads the live page instead of guessing. On 2 October 2026 we sent 20 public pages through the QuanticData scrape tool with its default settings. All 20 came back: 18 over plain HTTP on the first call at a median 5.4 seconds, and 2 rendered from a US exit. Nine pages of raw HTML, about 947,431 tokens, arrived as 75,777 tokens of Markdown, 92% fewer. The 20 pages cost $0.0056.

ChatGPT web scraping works through an MCP app, set up in 5 steps

ChatGPT on its own answers from its training data, its built-in search, or a page you paste. To make it read any public URL you name, OpenAI's developer mode lets you add a remote MCP server as an app. OpenAI lists developer mode for Pro, Plus, Business, Enterprise and Education accounts on the web, and describes it as full MCP client support for read and write tools.

The steps below follow OpenAI's developer-mode guide and our MCP server page:

  1. In ChatGPT, open Settings, then Security and login, and turn on Developer mode.
  2. Go to ChatGPT Plugins, select the plus button, and give the app a name and a description, for example QuanticData and "live web pages and Google results".
  3. Under Connection, enter https://api.quanticdata.io/mcp, including the /mcp path, and create the connection. No key field is needed.
  4. ChatGPT opens the QuanticData sign-in page. Sign in and click Allow access. The tools ChatGPT discovers include scrape, search, search_and_read, map, crawl, batch and seo_audit.
  5. In a new chat, choose Developer mode from the plus menu and select the QuanticData app for the conversation.

The sign-in is OAuth 2.1, so there is no API key to copy: ChatGPT receives a 1-hour access token with a rotating 30-day refresh token, valid only for the MCP endpoint, and a dedicated key named after the client appears on your dashboard. ../mcp-server-oauth/ covers the flow in detail, and Disconnect on the dashboard MCP page ends it.

We read 20 live pages through the scrape tool: all 20 came back

We picked the kinds of pages people ask ChatGPT to read: encyclopedia articles, news sections, a forum front page, a subreddit, GitHub READMEs, reference docs, pricing pages, product pages, a movie page, a Q&A thread, a paper abstract and a package page. Each went through the scrape endpoint exactly as the MCP tool sends it: Markdown, smart content mode, engine auto, two requests in flight, from Europe, on 2 October 2026.

PageFetchSecondsWordsMarkdown tokens
Wikipedia, Web scrapingplain HTTP2.27,01212,222
Wikipedia, Model Context Protocolplain HTTP18.72,9275,748
BBC Technologyplain HTTP8.42,9494,767
The Guardian, Technologyplain HTTP5.62,4674,310
Reuters Technologyplain HTTP22.07581,727
Hacker News front pageplain HTTP5.12,0593,703
Reddit r/webscrapingplain HTTP5.99661,535
GitHub, MCP servers READMEplain HTTP4.61,6083,256
GitHub, openai-python READMEplain HTTP6.46,35611,893
Python docs, urllib.requestplain HTTP3.810,40318,601
MDN, HTTP status reference pageplain HTTP3.16311,066
OpenAI API pricingplain HTTP5.33,0235,693
Anthropic pricingplain HTTP11.39671,585
Stripe pricingplain HTTP4.65,02010,036
Amazon product pageplain HTTP4.53,1214,838
Apple AirPods Pro pageplain HTTP7.04,7398,261
IMDb title pagerendered, US exit10.94,2987,037
Stack Overflow questionrendered, US exit31.33,7526,478
arXiv abstractplain HTTP3.11,2242,045
PyPI, requestsplain HTTP6.26,77310,621

Eighteen pages came back on the first call over plain HTTP, with no browser, at a median of 5.4 seconds and a slowest of 22.0 seconds. The movie page and the Q&A thread came back on a second call with the exit pinned to the United States and the page rendered: 4,298 and 3,752 words. That is the one instruction to give ChatGPT when a page arrives nearly empty: call scrape again with country set to us.

The Markdown is 92% smaller than the raw HTML of the same pages

For nine of the pages we also asked for the raw HTML. We count HTML tokens as characters divided by 4, which is an estimate; the Markdown counts are the tool's own token figures. The scrape tool's smart mode keeps the page body with its tables and links and drops navigation, footers and cookie banners.

PageRaw HTML, tokensMarkdown, tokensFewer
The Guardian, Technology175,7154,31097.5%
GitHub, MCP servers README88,1133,25696.3%
BBC Technology121,4944,76796.1%
Stripe pricing242,59610,03695.9%
Apple AirPods Pro page130,6188,26193.7%
PyPI, requests67,42910,62184.2%
Wikipedia, Web scraping58,86812,22279.2%
Python docs, urllib.request54,06518,60165.6%
Hacker News front page8,5353,70356.6%
Nine pages947,43175,77792.0%

Across all 20 pages the median page came to 5,266 Markdown tokens, and the largest, the Python reference page, to 18,601. When you only need one fact from a long page, tell ChatGPT to pass a query to scrape, which keeps only the matching sections, or a max_tokens cap, which cuts at a section boundary. ../mcp-web-scraper/ compares every content mode on 11 pages.

Search first, then read: 3 Google searches in 7.6 to 15.7 seconds

Most questions start without a URL. The search tool returns Google, Bing or DuckDuckGo results as structured JSON, and search_and_read runs the search, fetches the top pages and returns numbered sources with one token-bounded context. Our three Google searches each returned 10 organic results, in 8.8, 15.7 and 7.6 seconds.

OpenAI's guide says ChatGPT treats any tool without the readOnlyHint annotation as a write action and asks you to confirm it. In the QuanticData app, search, search_and_read, map and seo_audit are marked read-only and run without a prompt. scrape is not, because it can also click and type on a page when you ask it to, so ChatGPT asks before each call; you can approve it once and have ChatGPT remember that choice for the rest of the conversation.

OpenAI also recommends naming the app and the tool when several could answer. A prompt that works:

Use the QuanticData app's search_and_read tool to find the three
most recent pages about the EU AI Act timeline, then summarize them
with numbered citations. Do not use built-in browsing.

20 pages cost $0.0056, and $2 a month covers 10,000

The web scraping API behind the scrape tool charges per page, not per byte: $0.0002 for a plain HTTP page and $0.001 for a rendered one. Search runs through the SERP API from $0.0005 per search. Failed requests are never billed. Usage through ChatGPT is charged to your account balance like any API key.

What ChatGPT asked forCallsPrice eachCost
Pages over plain HTTP18$0.0002$0.0036
Pages rendered, US exit2$0.001$0.0020
Google searches3$0.0005$0.0015
Total23$0.0071

The 20 pages moved 13.3 MB of HTML, and the heaviest page, at 2.5 MB, cost the same $0.001 as any rendered page. At these prices, $2 buys 10,000 plain pages, 2,000 rendered pages or 4,000 searches.

How the app differs from ChatGPT's built-in browsing

OpenAI documents that when a user asks ChatGPT a question it may visit a page with the ChatGPT-User agent, and that this agent does not crawl automatically. The MCP app is a different path: the request leaves from our residential exits, in the country you pin, and the page comes back as Markdown, HTML or text with its title, word count and token count, so you and ChatGPT both see what was read.

Read only public pages, and respect each site's terms of use. The app's tool set in ChatGPT leaves out collectors that return people's contact details, and its raw request tool fetches public pages only. OpenAI calls developer mode powerful but dangerous and asks users to review every write action before approving it: keep that habit with any MCP app, ours included.

The setting that works on ChatGPT

  • Connection: a developer-mode app on a Plus, Pro, Business, Enterprise or Education account, URL https://api.quanticdata.io/mcp, OAuth sign-in, no key to paste.
  • Tool: search_and_read for questions without a URL, at 10 organic results per search; scrape for a link you have, at a median 5.4 seconds on our 20 pages; map or crawl for a whole site.
  • Fetch mode: leave engine on auto. 18 of 20 pages came back over plain HTTP at $0.0002, with no browser to pay for.
  • Country: pin us when a page comes back nearly empty. The two pages that needed it returned 4,298 and 3,752 words rendered, at $0.001 each.
  • Token budget: smart Markdown, 92% fewer tokens than raw HTML on nine pages; add query or max_tokens on long reference pages.
  • When a page is not enough: for listings, places, jobs or prices across many pages, ask ChatGPT to call list_collectors and run a collector, which returns rows instead of pages, priced per delivered result.
  • Price: $0.0002 per plain page, $0.001 rendered, $0.0005 per search, pay per success, and every account gets $2 of free API usage per month.

Sources & further reading

FAQ

Quick answers on chatgpt web scraping.

Something else? Ask us

Can ChatGPT scrape websites?

Yes, once an MCP server with a scrape tool is connected as a developer-mode app. Through the QuanticData app we read 20 public pages on 2 October 2026 and all 20 came back: 18 over plain HTTP at a median 5.4 seconds, 2 rendered from a US exit. Read only public pages and respect each site's terms.

Which ChatGPT plans can connect an MCP server?

OpenAI lists developer mode for Pro, Plus, Business, Enterprise and Education accounts on the web. Turn it on under Settings, Security and login, then add the server from ChatGPT Plugins with the plus button: 5 steps in all, sign-in included.

Do I need an API key to use QuanticData in ChatGPT?

No. The hosted endpoint https://api.quanticdata.io/mcp supports OAuth 2.1: you sign in and click Allow access. ChatGPT gets a 1-hour access token and a rotating 30-day refresh token bound to the MCP endpoint, and a dedicated key named after the client appears on your dashboard.

How much does ChatGPT web scraping cost with QuanticData?

$0.0002 per page over plain HTTP, $0.001 per rendered page and $0.0005 per Google search, billed only on success. Our 20 pages and 3 searches cost $0.0071. Every account gets $2 of free API usage per month, which covers 10,000 plain pages.

Why does ChatGPT ask me to confirm the scrape tool?

OpenAI treats any tool without the readOnlyHint annotation as a write action. Scrape can also click and type on a page, so it is not marked read-only; search, search_and_read, map and seo_audit are, and run without a prompt. You can approve scrape once and let ChatGPT remember it for the conversation.

What should I do when a page comes back nearly empty in ChatGPT?

Ask ChatGPT to call scrape again with country set to us. In our test 2 of 20 pages needed this, and they returned 4,298 and 3,752 words rendered, in 10.9 and 31.3 seconds, at $0.001 each.

Give ChatGPT the live web

All 20 pages in this test came back through the QuanticData app, at a median 5.4 seconds and 92% fewer tokens than raw HTML. Connect https://api.quanticdata.io/mcp with one sign-in, and every account gets $2 of free API usage per month.

Related reading