ChatGPT web scraping works once you connect an MCP server as a developer-mode app: ChatGPT then calls the server's scrape and search tools and reads the live page instead of guessing. On 2 October 2026 we sent 20 public pages through the QuanticData scrape tool with its default settings. All 20 came back: 18 over plain HTTP on the first call at a median 5.4 seconds, and 2 rendered from a US exit. Nine pages of raw HTML, about 947,431 tokens, arrived as 75,777 tokens of Markdown, 92% fewer. The 20 pages cost $0.0056.
ChatGPT web scraping works through an MCP app, set up in 5 steps
ChatGPT on its own answers from its training data, its built-in search, or a page you paste. To make it read any public URL you name, OpenAI's developer mode lets you add a remote MCP server as an app. OpenAI lists developer mode for Pro, Plus, Business, Enterprise and Education accounts on the web, and describes it as full MCP client support for read and write tools.
The steps below follow OpenAI's developer-mode guide and our MCP server page:
- In ChatGPT, open Settings, then Security and login, and turn on Developer mode.
- Go to ChatGPT Plugins, select the plus button, and give the app a name and a description, for example QuanticData and "live web pages and Google results".
- Under Connection, enter
https://api.quanticdata.io/mcp, including the/mcppath, and create the connection. No key field is needed. - ChatGPT opens the QuanticData sign-in page. Sign in and click Allow access. The tools ChatGPT discovers include
scrape,search,search_and_read,map,crawl,batchandseo_audit. - In a new chat, choose Developer mode from the plus menu and select the QuanticData app for the conversation.
The sign-in is OAuth 2.1, so there is no API key to copy: ChatGPT receives a 1-hour access token with a rotating 30-day refresh token, valid only for the MCP endpoint, and a dedicated key named after the client appears on your dashboard. ../mcp-server-oauth/ covers the flow in detail, and Disconnect on the dashboard MCP page ends it.
We read 20 live pages through the scrape tool: all 20 came back
We picked the kinds of pages people ask ChatGPT to read: encyclopedia articles, news sections, a forum front page, a subreddit, GitHub READMEs, reference docs, pricing pages, product pages, a movie page, a Q&A thread, a paper abstract and a package page. Each went through the scrape endpoint exactly as the MCP tool sends it: Markdown, smart content mode, engine auto, two requests in flight, from Europe, on 2 October 2026.
| Page | Fetch | Seconds | Words | Markdown tokens |
|---|---|---|---|---|
| Wikipedia, Web scraping | plain HTTP | 2.2 | 7,012 | 12,222 |
| Wikipedia, Model Context Protocol | plain HTTP | 18.7 | 2,927 | 5,748 |
| BBC Technology | plain HTTP | 8.4 | 2,949 | 4,767 |
| The Guardian, Technology | plain HTTP | 5.6 | 2,467 | 4,310 |
| Reuters Technology | plain HTTP | 22.0 | 758 | 1,727 |
| Hacker News front page | plain HTTP | 5.1 | 2,059 | 3,703 |
| Reddit r/webscraping | plain HTTP | 5.9 | 966 | 1,535 |
| GitHub, MCP servers README | plain HTTP | 4.6 | 1,608 | 3,256 |
| GitHub, openai-python README | plain HTTP | 6.4 | 6,356 | 11,893 |
| Python docs, urllib.request | plain HTTP | 3.8 | 10,403 | 18,601 |
| MDN, HTTP status reference page | plain HTTP | 3.1 | 631 | 1,066 |
| OpenAI API pricing | plain HTTP | 5.3 | 3,023 | 5,693 |
| Anthropic pricing | plain HTTP | 11.3 | 967 | 1,585 |
| Stripe pricing | plain HTTP | 4.6 | 5,020 | 10,036 |
| Amazon product page | plain HTTP | 4.5 | 3,121 | 4,838 |
| Apple AirPods Pro page | plain HTTP | 7.0 | 4,739 | 8,261 |
| IMDb title page | rendered, US exit | 10.9 | 4,298 | 7,037 |
| Stack Overflow question | rendered, US exit | 31.3 | 3,752 | 6,478 |
| arXiv abstract | plain HTTP | 3.1 | 1,224 | 2,045 |
| PyPI, requests | plain HTTP | 6.2 | 6,773 | 10,621 |
Eighteen pages came back on the first call over plain HTTP, with no browser, at a median of 5.4 seconds and a slowest of 22.0 seconds. The movie page and the Q&A thread came back on a second call with the exit pinned to the United States and the page rendered: 4,298 and 3,752 words. That is the one instruction to give ChatGPT when a page arrives nearly empty: call scrape again with country set to us.
The Markdown is 92% smaller than the raw HTML of the same pages
For nine of the pages we also asked for the raw HTML. We count HTML tokens as characters divided by 4, which is an estimate; the Markdown counts are the tool's own token figures. The scrape tool's smart mode keeps the page body with its tables and links and drops navigation, footers and cookie banners.
| Page | Raw HTML, tokens | Markdown, tokens | Fewer |
|---|---|---|---|
| The Guardian, Technology | 175,715 | 4,310 | 97.5% |
| GitHub, MCP servers README | 88,113 | 3,256 | 96.3% |
| BBC Technology | 121,494 | 4,767 | 96.1% |
| Stripe pricing | 242,596 | 10,036 | 95.9% |
| Apple AirPods Pro page | 130,618 | 8,261 | 93.7% |
| PyPI, requests | 67,429 | 10,621 | 84.2% |
| Wikipedia, Web scraping | 58,868 | 12,222 | 79.2% |
| Python docs, urllib.request | 54,065 | 18,601 | 65.6% |
| Hacker News front page | 8,535 | 3,703 | 56.6% |
| Nine pages | 947,431 | 75,777 | 92.0% |
Across all 20 pages the median page came to 5,266 Markdown tokens, and the largest, the Python reference page, to 18,601. When you only need one fact from a long page, tell ChatGPT to pass a query to scrape, which keeps only the matching sections, or a max_tokens cap, which cuts at a section boundary. ../mcp-web-scraper/ compares every content mode on 11 pages.
Search first, then read: 3 Google searches in 7.6 to 15.7 seconds
Most questions start without a URL. The search tool returns Google, Bing or DuckDuckGo results as structured JSON, and search_and_read runs the search, fetches the top pages and returns numbered sources with one token-bounded context. Our three Google searches each returned 10 organic results, in 8.8, 15.7 and 7.6 seconds.
OpenAI's guide says ChatGPT treats any tool without the readOnlyHint annotation as a write action and asks you to confirm it. In the QuanticData app, search, search_and_read, map and seo_audit are marked read-only and run without a prompt. scrape is not, because it can also click and type on a page when you ask it to, so ChatGPT asks before each call; you can approve it once and have ChatGPT remember that choice for the rest of the conversation.
OpenAI also recommends naming the app and the tool when several could answer. A prompt that works:
Use the QuanticData app's search_and_read tool to find the three
most recent pages about the EU AI Act timeline, then summarize them
with numbered citations. Do not use built-in browsing.
20 pages cost $0.0056, and $2 a month covers 10,000
The web scraping API behind the scrape tool charges per page, not per byte: $0.0002 for a plain HTTP page and $0.001 for a rendered one. Search runs through the SERP API from $0.0005 per search. Failed requests are never billed. Usage through ChatGPT is charged to your account balance like any API key.
| What ChatGPT asked for | Calls | Price each | Cost |
|---|---|---|---|
| Pages over plain HTTP | 18 | $0.0002 | $0.0036 |
| Pages rendered, US exit | 2 | $0.001 | $0.0020 |
| Google searches | 3 | $0.0005 | $0.0015 |
| Total | 23 | $0.0071 |
The 20 pages moved 13.3 MB of HTML, and the heaviest page, at 2.5 MB, cost the same $0.001 as any rendered page. At these prices, $2 buys 10,000 plain pages, 2,000 rendered pages or 4,000 searches.
How the app differs from ChatGPT's built-in browsing
OpenAI documents that when a user asks ChatGPT a question it may visit a page with the ChatGPT-User agent, and that this agent does not crawl automatically. The MCP app is a different path: the request leaves from our residential exits, in the country you pin, and the page comes back as Markdown, HTML or text with its title, word count and token count, so you and ChatGPT both see what was read.
Read only public pages, and respect each site's terms of use. The app's tool set in ChatGPT leaves out collectors that return people's contact details, and its raw request tool fetches public pages only. OpenAI calls developer mode powerful but dangerous and asks users to review every write action before approving it: keep that habit with any MCP app, ours included.
The setting that works on ChatGPT
- Connection: a developer-mode app on a Plus, Pro, Business, Enterprise or Education account, URL
https://api.quanticdata.io/mcp, OAuth sign-in, no key to paste. - Tool:
search_and_readfor questions without a URL, at 10 organic results per search;scrapefor a link you have, at a median 5.4 seconds on our 20 pages;maporcrawlfor a whole site. - Fetch mode: leave engine on auto. 18 of 20 pages came back over plain HTTP at $0.0002, with no browser to pay for.
- Country: pin
uswhen a page comes back nearly empty. The two pages that needed it returned 4,298 and 3,752 words rendered, at $0.001 each. - Token budget: smart Markdown, 92% fewer tokens than raw HTML on nine pages; add
queryormax_tokenson long reference pages. - When a page is not enough: for listings, places, jobs or prices across many pages, ask ChatGPT to call
list_collectorsand run a collector, which returns rows instead of pages, priced per delivered result. - Price: $0.0002 per plain page, $0.001 rendered, $0.0005 per search, pay per success, and every account gets $2 of free API usage per month.
Sources & further reading
- ChatGPT Developer mode, OpenAI Developers (eligibility, setup, OAuth, tool confirmation; fetched 2 October 2026)
- Connect and test your plugin, OpenAI Developers (adding an MCP server in ChatGPT; fetched 2 October 2026)
- Overview of OpenAI Crawlers, OpenAI Developers (ChatGPT-User; fetched 2 October 2026)
- How is ChatGPT able to extract webpages so quickly?, OpenAI Developer Community