Documentation Python quickstart Blog Free tools Enterprise solutions hello@quanticdata.ioLog in

OpenAI Scraping Lawsuit: 0 of 30 Block Bingbot

How many of 30 news sites block each crawler in robots.txt, measured 29 September 2026: ClaudeBot 27, CCBot 26, GPTBot 20, Google-Extended 20, ChatGPT-User 15, OAI-SearchBot 13, Bingbot 0
How many of 30 news sites block each crawler in robots.txt, measured 29 September 2026: ClaudeBot 27, CCBot 26, GPTBot 20, Google-Extended 20, ChatGPT-User 15, OAI-SearchBot 13, Bingbot 0

The OpenAI scraping lawsuit turned on a new page on 17 September 2026, when a summary-judgment motion from the news plaintiffs led by The New York Times was unsealed and quoted a Microsoft director calling the scraping of news for AI training perhaps "the largest theft of labor in human history". The filings also say data gathered for Bing ended up as training material. So on 29 September 2026 we read the robots.txt of 30 US and UK news sites, the seven plaintiff domains included: 20 block GPTBot, 27 block ClaudeBot, and not one blocks Bingbot.

That last number is the whole story in one cell. A publisher can refuse an AI training crawler at no cost. It cannot refuse a search engine crawler without disappearing from search. This post sets out what the filings allege, what the 30 files actually say, and the rules that follow for anyone collecting web data today, whether for a model, a dashboard or a price feed. It is not legal advice; it is a reading of public documents and a measurement.

What the unsealed OpenAI scraping lawsuit filings say

The case began on 27 December 2023, when The New York Times sued OpenAI and Microsoft over the use of its articles to train and ground models. The New York Daily News and seven sister papers followed, as did the Center for Investigative Reporting, and the suits now run together in New York. The September 2026 filing is the plaintiffs' own brief; as TechCrunch notes, most underlying exhibits are still sealed, so the quotes below are the plaintiffs' characterisation.

Four allegations matter for anyone who scrapes:

  • Scale. OpenAI's mid-training datasets are said to contain more than 91,692 copies of works from the three plaintiff groups, and a Common Crawl-derived set more than 2 million documents from nytimes.com alone.
  • Paywalls. The brief quotes an internal message about a "hack to get around nytimes paywall" and a two-word reply from OpenAI's president: "ah nice". Microsoft's CEO testified that anything paywalled should be licensed by whoever wants it for grounding or training.
  • Repurposed search data. The plaintiffs say Microsoft sold a dataset purchased for Bing to OpenAI as training data, and that OpenAI used a New York Times dataset of 1.8 million articles obtained from a third party bound to non-commercial use.
  • Stripped notices. Researchers allegedly removed copyright notices from training data so the model would not output them.

Microsoft's response, per Ars Technica, is that the director's memos reflect one employee's view and that its products are transformative fair use. Courts have so far been receptive to fair-use arguments for training, and the US government filed a brief supporting OpenAI earlier in September. None of this is decided. What is already clear is which behaviours the plaintiffs chose to put in front of the judge: going around paywalls, reusing data collected under one kind of consent for another purpose, and removing attribution.

Twenty of thirty news sites block GPTBot in robots.txt

We fetched each site's robots.txt through the QuanticData batch endpoint over plain HTTP from a US residential exit. All 30 returned 200, 163,199 bytes of text in total, with a median of 47 user-agent lines per file; USA Today's names 288. We then applied standard group matching: an agent is "blocked" when the group that applies to it disallows the root path and nothing re-allows it.

CrawlerJobBlocked, all 30Blocked, 7 plaintiffs
ClaudeBotTraining (Anthropic)277
CCBotCommon Crawl archive266
PerplexityBotSearch index (Perplexity)236
GPTBotTraining (OpenAI)207
Google-ExtendedGemini training token206
ChatGPT-UserUser-triggered fetch (OpenAI)156
OAI-SearchBotChatGPT search index136
GooglebotGoogle Search00
BingbotBing Search00

Twenty of 30 is 67%. In our earlier study of 326 top domains across all sectors, 17.8% blocked GPTBot (see Should I block AI crawlers?). News publishers refuse OpenAI's training crawler at almost four times the web's rate, and 29 of the 30 files block at least one AI agent; only foxnews.com has no AI-specific rule at all.

Look at CCBot too. The filings say a Common Crawl-derived set held over 2 million nytimes.com documents. Today CCBot is blocked on 26 of the 30 sites and on 6 of the 7 plaintiffs. Blocking now does not recall an archive built years ago, which is why the lawsuits exist.

Bingbot: the crawler no publisher can refuse

Googlebot and Bingbot are blocked on zero of 30 sites. That is not an oversight; it is the price of being found. It is also why the "Bing dataset" allegation stands out. If data collected by a search crawler is later used to train a model, robots.txt could never have stopped it, because the publisher's only lever would have cost it its search traffic.

Google split the problem in 2023 by publishing Google-Extended, a robots.txt token that controls use for Gemini training without touching Search; 20 of our 30 sites use it. Microsoft took a different route: Bing reads page-level NOCACHE and NOARCHIVE meta tags to limit how content appears in its chat answers, rather than a separate user-agent. A page-level tag is invisible to a robots.txt audit like ours, so our count of Microsoft-specific opt-outs is structurally zero, not measured zero.

The practical lesson for anyone running a crawler is narrow and firm: one user agent, one declared purpose. If your bot collects prices, call it that and use the data for prices. Reusing a crawl gathered for one purpose for another is precisely the conduct the plaintiffs are asking a court to punish.

OpenAI runs three agents with three jobs: GPTBot for training, OAI-SearchBot for the ChatGPT search index, and ChatGPT-User for pages a user asks about. Our 30 files treat them differently. Thirteen sites block all three. Seven block GPTBot but leave OAI-SearchBot open: revealnews.org, latimes.com, reuters.com, apnews.com, techcrunch.com, forbes.com and cbsnews.com. The message is "cite us, do not train on us", and the file is how they say it. Our AI crawler user agent list maps every token to its job.

A few named GPTBot and left it open, among them wsj.com, theverge.com and theatlantic.com. Two sites, wsj.com and reuters.com, run default-deny: User-agent: * with Disallow: /, then an allowlist of named crawlers. Any client not on that list, including yours, is refused everything by the file.

Here is every site, grouped:

SiteGPTBotOAI-SearchBotClaudeBotGoogle-ExtendedBingbot
Plaintiffs in the New York case (7)
The New York Times (nytimes.com)blockedblockedblockedblockedopen
New York Daily News (nydailynews.com)blockedblockedblockedblockedopen
Chicago Tribune (chicagotribune.com)blockedblockedblockedblockedopen
The Denver Post (denverpost.com)blockedblockedblockedblockedopen
The Mercury News (mercurynews.com)blockedblockedblockedblockedopen
Orange County Register (ocregister.com)blockedblockedblockedblockedopen
Reveal (CIR) (revealnews.org)blockedopenblockedopenopen
Other US and UK news sites (23)
apnews.comblockedopenblockedopenopen
arstechnica.comopenopenblockedblockedopen
axios.comopenopenopenopenopen
bbc.comblockedblockedblockedblockedopen
bloomberg.comblockedblockedblockedblockedopen
businessinsider.comopenopenblockedopenopen
cbsnews.comblockedopenopenopenopen
cnn.comblockedblockedblockedblockedopen
forbes.comblockedopenblockedopenopen
foxnews.comopenopenopenopenopen
latimes.comblockedopenblockedopenopen
nbcnews.comblockedblockedblockedblockedopen
npr.orgblockedblockedblockedblockedopen
politico.comblockedblockedblockedblockedopen
reuters.comblockedopenblockedblockedopen
techcrunch.comblockedopenblockedblockedopen
theatlantic.comopenopenblockedblockedopen
theguardian.comopenopenblockedopenopen
theverge.comopenopenblockedblockedopen
usatoday.comblockedblockedblockedblockedopen
washingtonpost.comopenopenblockedopenopen
wired.comopenopenblockedblockedopen
wsj.comopenopenblockedblockedopen

What the case means for your scraper

The lawsuits target the copying of expressive work to build a substitute product. Collecting public facts, such as prices, availability, rankings, headlines and links, sits on different ground, and our pieces on whether AI web scraping is legal and web scraping law in the US cover the doctrine. Four rules come straight from what the filings chose to highlight:

  1. Never go around a paywall. Nadella's own testimony draws the line at paywalled content: license it. A proxy changes where a request comes from; it does not change what you are entitled to read, and no setting we sell is for that.
  2. Read robots.txt and the terms before the first request. Eleven of the 30 files open with a written legal notice in comments, six of the seven plaintiffs among them; the New York Times file prohibits automated collection without written permission. Our free robots.txt tester shows the exact rule that allows or blocks a URL.
  3. Keep purpose and provenance. Log what each crawl was for and what the file said on the day. The filings show a dataset outliving its consent; your logs are how you prove yours did not.
  4. Keep attribution. Store the source URL and publisher with every record. Stripping notices is one of the allegations.

The setting that works on news sites

For monitoring what publishers allow, or collecting headlines and links rather than article bodies, this is what our measurement supports:

  • Network: residential Basic from $0.80/GB. The 30 robots.txt files came to 163,199 bytes of text, about $0.00013 of bandwidth at that rate, counting 1 GB as 10^9 bytes.
  • Fetch mode: plain HTTP (engine: tls). All 30 files answered 200 without a browser; rendering robots.txt buys nothing. Through the web scraping API the 30 fetches cost $0.006 at $0.0002 per page.
  • Country: a US exit for US publishers. Robots rules are the same worldwide, but paywall and consent layers on the pages themselves vary by country.
  • When you need the news itself: the Google News collector returns dated articles with title, source, snippet and link for $0.0005 per delivered article, which keeps you on metadata and links rather than copied text. To check a single site quickly, the free AI crawler checker reads its live robots.txt.
  • Free tier: $2 of free API usage per month, and failed requests are never billed.

Sources & further reading

FAQ

Quick answers on openai scraping lawsuit.

Something else? Ask us →

What is the OpenAI scraping lawsuit about?

News publishers led by The New York Times, which sued on 27 December 2023, allege that OpenAI and Microsoft copied millions of articles to train and ground AI models without a licence. Filings unsealed on 17 September 2026 add allegations about paywall workarounds, a dataset bought for Bing reused as training data, and more than 2 million nytimes.com documents in one Common Crawl-derived set.

How many news sites block GPTBot?

In our reading of 30 US and UK news sites on 29 September 2026, 20 block GPTBot, 27 block ClaudeBot, 26 block CCBot and 13 block OAI-SearchBot. All seven plaintiff domains we identified block GPTBot. Across 326 top domains from all sectors, the rate we measured earlier was 17.8%.

Why do no news sites block Bingbot?

Because blocking Bingbot removes a site from Bing Search, and 0 of the 30 sites we checked accept that cost; Googlebot is also blocked on 0 of 30. Google offers a separate Google-Extended token for AI training, used by 20 of the 30 sites. Bing relies on page-level NOCACHE and NOARCHIVE meta tags instead, which a robots.txt audit cannot see.

Did the court decide the NYT vs OpenAI case?

Not as of 29 September 2026. The September filing is a motion for summary judgment by the news plaintiffs, asking the court to rule on articles where outputs show extensive verbatim overlap. Microsoft says its products are transformative fair use, and courts have so far leaned toward fair use for training. Treat the quotes as the plaintiffs' allegations until the court decides.

Is it legal to scrape news websites?

It depends on what you take and how, so here is the working rule: facts, headlines and links are low-risk; copying full articles, bypassing paywalls or ignoring written prohibitions is where the claims sit. 11 of the 30 robots.txt files we read open with a written legal notice, and the New York Times file prohibits automated collection without permission. When in doubt, license the content.

Can a proxy get around a news paywall?

No, and we do not sell it for that. A proxy changes the IP address a request comes from; it does not grant access to paid content, and circumventing a paywall is one of the specific behaviours the filings highlight. For news data, use headlines and links, for example the Google News collector at $0.0005 per article, or a licence from the publisher.

Read what a site allows before you collect

All 30 robots.txt files in this study came back over plain HTTP for less than a cent through one batch call. Every account gets $2 of free API usage each month, and failed requests are never billed.

Related reading