# OpenAI Scraping Lawsuit: 0 of 30 Block Bingbot

> The OpenAI scraping lawsuit filings, measured: 20 of 30 news sites block GPTBot in robots.txt on 29 September 2026, and none blocks Bingbot. What it means.

[Home](https://quanticdata.io/)/[Blog](https://quanticdata.io/blog/)/OpenAI Scraping Lawsuit: 0 of 30 Block Bingbot

# OpenAI Scraping Lawsuit: 0 of 30 Block Bingbot

AI scrapingSep 29, 2026·9 min read·By [Aldo Morese](https://quanticdata.io/about/), founder of QuanticData

How many of 30 news sites block each crawler in robots.txt, measured 29 September 2026: ClaudeBot 27, CCBot 26, GPTBot 20, Google-Extended 20, ChatGPT-User 15, OAI-SearchBot 13, Bingbot 0

On this page [What the unsealed OpenAI scraping lawsuit filings say](/blog/openai-scraping-lawsuit/#what-the-unsealed-openai-scraping-lawsuit-filings-say) [Twenty of thirty news sites block GPTBot in robots.txt](/blog/openai-scraping-lawsuit/#twenty-of-thirty-news-sites-block-gptbot-in-robots-txt) [Bingbot: the crawler no publisher can refuse](/blog/openai-scraping-lawsuit/#bingbot-the-crawler-no-publisher-can-refuse) [Publishers now separate training from search](/blog/openai-scraping-lawsuit/#publishers-now-separate-training-from-search) [What the case means for your scraper](/blog/openai-scraping-lawsuit/#what-the-case-means-for-your-scraper) [The setting that works on news sites](/blog/openai-scraping-lawsuit/#the-setting-that-works-on-news-sites)

The OpenAI scraping lawsuit turned on a new page on 17 September 2026, when a summary-judgment motion from the news plaintiffs led by The New York Times was unsealed and quoted a Microsoft director calling the scraping of news for AI training perhaps "the largest theft of labor in human history". The filings also say data gathered for Bing ended up as training material. So on 29 September 2026 we read the robots.txt of 30 US and UK news sites, the seven plaintiff domains included: 20 block GPTBot, 27 block ClaudeBot, and not one blocks Bingbot.

That last number is the whole story in one cell. A publisher can refuse an AI training crawler at no cost. It cannot refuse a search engine crawler without disappearing from search. This post sets out what the filings allege, what the 30 files actually say, and the rules that follow for anyone collecting web data today, whether for a model, a dashboard or a price feed. It is not legal advice; it is a reading of public documents and a measurement.

## What the unsealed OpenAI scraping lawsuit filings say

The case began on 27 December 2023, when The New York Times sued OpenAI and Microsoft over the use of its articles to train and ground models. The New York Daily News and seven sister papers followed, as did the Center for Investigative Reporting, and the suits now run together in New York. The September 2026 filing is the plaintiffs' own brief; as TechCrunch notes, most underlying exhibits are still sealed, so the quotes below are the plaintiffs' characterisation.

Four allegations matter for anyone who scrapes:

- **Scale.** OpenAI's mid-training datasets are said to contain more than 91,692 copies of works from the three plaintiff groups, and a Common Crawl-derived set more than 2 million documents from nytimes.com alone.

- **Paywalls.** The brief quotes an internal message about a "hack to get around nytimes paywall" and a two-word reply from OpenAI's president: "ah nice". Microsoft's CEO testified that anything paywalled should be licensed by whoever wants it for grounding or training.

- **Repurposed search data.** The plaintiffs say Microsoft sold a dataset purchased for Bing to OpenAI as training data, and that OpenAI used a New York Times dataset of 1.8 million articles obtained from a third party bound to non-commercial use.

- **Stripped notices.** Researchers allegedly removed copyright notices from training data so the model would not output them.

Microsoft's response, per Ars Technica, is that the director's memos reflect one employee's view and that its products are transformative fair use. Courts have so far been receptive to fair-use arguments for training, and the US government filed a brief supporting OpenAI earlier in September. None of this is decided. What is already clear is which behaviours the plaintiffs chose to put in front of the judge: going around paywalls, reusing data collected under one kind of consent for another purpose, and removing attribution.

## Twenty of thirty news sites block GPTBot in robots.txt

We fetched each site's robots.txt through the QuanticData batch endpoint over plain HTTP from a US residential exit. All 30 returned 200, 163,199 bytes of text in total, with a median of 47 user-agent lines per file; USA Today's names 288. We then applied standard group matching: an agent is "blocked" when the group that applies to it disallows the root path and nothing re-allows it.

| Crawler | Job | Blocked, all 30 | Blocked, 7 plaintiffs |
| --- | --- | --- | --- |
| ClaudeBot | Training (Anthropic) | 27 | 7 |
| CCBot | Common Crawl archive | 26 | 6 |
| PerplexityBot | Search index (Perplexity) | 23 | 6 |
| GPTBot | Training (OpenAI) | 20 | 7 |
| Google-Extended | Gemini training token | 20 | 6 |
| ChatGPT-User | User-triggered fetch (OpenAI) | 15 | 6 |
| OAI-SearchBot | ChatGPT search index | 13 | 6 |
| Googlebot | Google Search | 0 | 0 |
| Bingbot | Bing Search | 0 | 0 |

Twenty of 30 is 67%. In our earlier study of 326 top domains across all sectors, 17.8% blocked GPTBot (see [Should I block AI crawlers?](https://quanticdata.io/blog/should-i-block-ai-crawlers/)). News publishers refuse OpenAI's training crawler at almost four times the web's rate, and 29 of the 30 files block at least one AI agent; only foxnews.com has no AI-specific rule at all.

Look at CCBot too. The filings say a Common Crawl-derived set held over 2 million nytimes.com documents. Today CCBot is blocked on 26 of the 30 sites and on 6 of the 7 plaintiffs. Blocking now does not recall an archive built years ago, which is why the lawsuits exist.

## Bingbot: the crawler no publisher can refuse

Googlebot and Bingbot are blocked on zero of 30 sites. That is not an oversight; it is the price of being found. It is also why the "Bing dataset" allegation stands out. If data collected by a search crawler is later used to train a model, robots.txt could never have stopped it, because the publisher's only lever would have cost it its search traffic.

Google split the problem in 2023 by publishing Google-Extended, a robots.txt token that controls use for Gemini training without touching Search; 20 of our 30 sites use it. Microsoft took a different route: Bing reads page-level NOCACHE and NOARCHIVE meta tags to limit how content appears in its chat answers, rather than a separate user-agent. A page-level tag is invisible to a robots.txt audit like ours, so our count of Microsoft-specific opt-outs is structurally zero, not measured zero.

The practical lesson for anyone running a crawler is narrow and firm: one user agent, one declared purpose. If your bot collects prices, call it that and use the data for prices. Reusing a crawl gathered for one purpose for another is precisely the conduct the plaintiffs are asking a court to punish.

## Publishers now separate training from search

OpenAI runs three agents with three jobs: GPTBot for training, OAI-SearchBot for the ChatGPT search index, and ChatGPT-User for pages a user asks about. Our 30 files treat them differently. Thirteen sites block all three. Seven block GPTBot but leave OAI-SearchBot open: revealnews.org, latimes.com, reuters.com, apnews.com, techcrunch.com, forbes.com and cbsnews.com. The message is "cite us, do not train on us", and the file is how they say it. Our [AI crawler user agent list](https://quanticdata.io/blog/ai-crawler-user-agent-list/) maps every token to its job.

A few named GPTBot and left it open, among them wsj.com, theverge.com and theatlantic.com. Two sites, wsj.com and reuters.com, run default-deny: `User-agent: *` with `Disallow: /`, then an allowlist of named crawlers. Any client not on that list, including yours, is refused everything by the file.

Here is every site, grouped:

| Site | GPTBot | OAI-SearchBot | ClaudeBot | Google-Extended | Bingbot |
| --- | --- | --- | --- | --- | --- |
| **Plaintiffs in the New York case (7)** |  |  |  |  |  |
| The New York Times (nytimes.com) | blocked | blocked | blocked | blocked | open |
| New York Daily News (nydailynews.com) | blocked | blocked | blocked | blocked | open |
| Chicago Tribune (chicagotribune.com) | blocked | blocked | blocked | blocked | open |
| The Denver Post (denverpost.com) | blocked | blocked | blocked | blocked | open |
| The Mercury News (mercurynews.com) | blocked | blocked | blocked | blocked | open |
| Orange County Register (ocregister.com) | blocked | blocked | blocked | blocked | open |
| Reveal (CIR) (revealnews.org) | blocked | open | blocked | open | open |
| **Other US and UK news sites (23)** |  |  |  |  |  |
| apnews.com | blocked | open | blocked | open | open |
| arstechnica.com | open | open | blocked | blocked | open |
| axios.com | open | open | open | open | open |
| bbc.com | blocked | blocked | blocked | blocked | open |
| bloomberg.com | blocked | blocked | blocked | blocked | open |
| businessinsider.com | open | open | blocked | open | open |
| cbsnews.com | blocked | open | open | open | open |
| cnn.com | blocked | blocked | blocked | blocked | open |
| forbes.com | blocked | open | blocked | open | open |
| foxnews.com | open | open | open | open | open |
| latimes.com | blocked | open | blocked | open | open |
| nbcnews.com | blocked | blocked | blocked | blocked | open |
| npr.org | blocked | blocked | blocked | blocked | open |
| politico.com | blocked | blocked | blocked | blocked | open |
| reuters.com | blocked | open | blocked | blocked | open |
| techcrunch.com | blocked | open | blocked | blocked | open |
| theatlantic.com | open | open | blocked | blocked | open |
| theguardian.com | open | open | blocked | open | open |
| theverge.com | open | open | blocked | blocked | open |
| usatoday.com | blocked | blocked | blocked | blocked | open |
| washingtonpost.com | open | open | blocked | open | open |
| wired.com | open | open | blocked | blocked | open |
| wsj.com | open | open | blocked | blocked | open |

## What the case means for your scraper

The lawsuits target the copying of expressive work to build a substitute product. Collecting public facts, such as prices, availability, rankings, headlines and links, sits on different ground, and our pieces on [whether AI web scraping is legal](https://quanticdata.io/blog/is-ai-web-scraping-legal/) and [web scraping law in the US](https://quanticdata.io/blog/is-web-scraping-legal-in-us/) cover the doctrine. Four rules come straight from what the filings chose to highlight:

1. **Never go around a paywall.** Nadella's own testimony draws the line at paywalled content: license it. A proxy changes where a request comes from; it does not change what you are entitled to read, and no setting we sell is for that.

2. **Read robots.txt and the terms before the first request.** Eleven of the 30 files open with a written legal notice in comments, six of the seven plaintiffs among them; the New York Times file prohibits automated collection without written permission. Our free [robots.txt tester](https://quanticdata.io/tools/robots-txt-tester/) shows the exact rule that allows or blocks a URL.

3. **Keep purpose and provenance.** Log what each crawl was for and what the file said on the day. The filings show a dataset outliving its consent; your logs are how you prove yours did not.

4. **Keep attribution.** Store the source URL and publisher with every record. Stripping notices is one of the allegations.

## The setting that works on news sites

For monitoring what publishers allow, or collecting headlines and links rather than article bodies, this is what our measurement supports:

- **Network:** residential Basic from $0.80/GB. The 30 robots.txt files came to 163,199 bytes of text, about $0.00013 of bandwidth at that rate, counting 1 GB as 10^9 bytes.

- **Fetch mode:** plain HTTP (`engine: tls`). All 30 files answered 200 without a browser; rendering robots.txt buys nothing. Through the [web scraping API](https://quanticdata.io/web-scraping-api/) the 30 fetches cost $0.006 at $0.0002 per page.

- **Country:** a US exit for US publishers. Robots rules are the same worldwide, but paywall and consent layers on the pages themselves vary by country.

- **When you need the news itself:** the [Google News collector](https://quanticdata.io/collectors/google-news-api/) returns dated articles with title, source, snippet and link for $0.0005 per delivered article, which keeps you on metadata and links rather than copied text. To check a single site quickly, the free [AI crawler checker](https://quanticdata.io/tools/ai-crawler-checker/) reads its live robots.txt.

- **Free tier:** $2 of free API usage per month, and failed requests are never billed.

### Sources & further reading

- [Microsoft exec called AI scraping the "largest theft of labor in human history", Ars Technica, 17 September 2026](https://arstechnica.com/tech-policy/2026/09/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history/)

- [Microsoft exec called AI scraping "the largest theft of labor in human history," new unredacted filings reveal, TechCrunch, 17 September 2026](https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/)

- [OpenAI, Microsoft Sued by Publishers for Scraping Articles, Bloomberg Law](https://news.bloomberglaw.com/litigation/publishers-sue-microsoft-openai-over-unauthorized-content-use)

- [Announcing new options for webmasters to control usage of their content in Bing Chat, Bing Webmaster Blog](https://blogs.bing.com/webmaster/2023/9/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat/)

- [nytimes.com/robots.txt (fetched 29 September 2026)](https://www.nytimes.com/robots.txt)

- [wsj.com/robots.txt (fetched 29 September 2026)](https://www.wsj.com/robots.txt)

## FAQ

Quick answers on openai scraping lawsuit.

[Something else? Ask us →](mailto:hello@quanticdata.io)

### What is the OpenAI scraping lawsuit about?

News publishers led by The New York Times, which sued on 27 December 2023, allege that OpenAI and Microsoft copied millions of articles to train and ground AI models without a licence. Filings unsealed on 17 September 2026 add allegations about paywall workarounds, a dataset bought for Bing reused as training data, and more than 2 million nytimes.com documents in one Common Crawl-derived set.

### How many news sites block GPTBot?

In our reading of 30 US and UK news sites on 29 September 2026, 20 block GPTBot, 27 block ClaudeBot, 26 block CCBot and 13 block OAI-SearchBot. All seven plaintiff domains we identified block GPTBot. Across 326 top domains from all sectors, the rate we measured earlier was 17.8%.

### Why do no news sites block Bingbot?

Because blocking Bingbot removes a site from Bing Search, and 0 of the 30 sites we checked accept that cost; Googlebot is also blocked on 0 of 30. Google offers a separate Google-Extended token for AI training, used by 20 of the 30 sites. Bing relies on page-level NOCACHE and NOARCHIVE meta tags instead, which a robots.txt audit cannot see.

### Did the court decide the NYT vs OpenAI case?

Not as of 29 September 2026. The September filing is a motion for summary judgment by the news plaintiffs, asking the court to rule on articles where outputs show extensive verbatim overlap. Microsoft says its products are transformative fair use, and courts have so far leaned toward fair use for training. Treat the quotes as the plaintiffs' allegations until the court decides.

### Is it legal to scrape news websites?

It depends on what you take and how, so here is the working rule: facts, headlines and links are low-risk; copying full articles, bypassing paywalls or ignoring written prohibitions is where the claims sit. 11 of the 30 robots.txt files we read open with a written legal notice, and the New York Times file prohibits automated collection without permission. When in doubt, license the content.

### Can a proxy get around a news paywall?

No, and we do not sell it for that. A proxy changes the IP address a request comes from; it does not grant access to paid content, and circumventing a paywall is one of the specific behaviours the filings highlight. For news data, use headlines and links, for example the Google News collector at $0.0005 per article, or a licence from the publisher.

## Read what a site allows before you collect

All 30 robots.txt files in this study came back over plain HTTP for less than a cent through one batch call. Every account gets $2 of free API usage each month, and failed requests are never billed.

[Start free — $2/month included](https://quanticdata.io/signup/)[Explore Web Scraping API](https://quanticdata.io/web-scraping-api/)

## Related reading

[AI scraping Jev for Web Scraping: Will Your Page Fit? Jev, the System One model from TypeSafe AI, reads at most 32,000 tokens of state and bills every input token. We fetched 20 real pages on 28 September 2026: raw HTML fit on 4 of them, with a median of 152,277 tokens. The same pages as Markdown fit on all 20, median 4,512, and 2,060 with links stripped. Read →](https://quanticdata.io/blog/jev-web-scraping/) [AI scraping Jev vs GPT, Claude, Gemini on Real Web Data Same 317 hand-labelled cases, same question, six models. Accuracy was a tie: 306 to 309 correct. Jev was 2 to 42 times cheaper and answered in a third of a second. The difference that matters is where the mistakes were: all 11 of Jev's errors came with an uncertain score, while the LLMs made theirs sounding sure. Read →](https://quanticdata.io/blog/jev-vs-llm-benchmark/) [AI scraping Claude Sonnet 5.5 for Web Scraping: 133 of 134 Hours after Anthropic released Claude Sonnet 5.5 we sent it the same 134 hand-labelled scraper responses we used for the Jev tests: is this the requested page, or a stub, a sign-in page, an error, the wrong page? It got 133 right for $0.56, about $4.17 per 1,000 pages. GPT-6 Luna and DeepSeek V4.1 Flash got all 134 for $0.147 and $0.236 per 1,000. Claude Haiku 4.5 got 125. The one page Sonnet 5.5 missed was a real one. Read →](https://quanticdata.io/blog/claude-sonnet-5-5-block-pages/)

---

Source: https://quanticdata.io/blog/openai-scraping-lawsuit/ · Site index for AI: https://quanticdata.io/llms.txt
