The OpenAI scraping lawsuit turned on a new page on 17 September 2026, when a summary-judgment motion from the news plaintiffs led by The New York Times was unsealed and quoted a Microsoft director calling the scraping of news for AI training perhaps "the largest theft of labor in human history". The filings also say data gathered for Bing ended up as training material. So on 29 September 2026 we read the robots.txt of 30 US and UK news sites, the seven plaintiff domains included: 20 block GPTBot, 27 block ClaudeBot, and not one blocks Bingbot.
That last number is the whole story in one cell. A publisher can refuse an AI training crawler at no cost. It cannot refuse a search engine crawler without disappearing from search. This post sets out what the filings allege, what the 30 files actually say, and the rules that follow for anyone collecting web data today, whether for a model, a dashboard or a price feed. It is not legal advice; it is a reading of public documents and a measurement.
What the unsealed OpenAI scraping lawsuit filings say
The case began on 27 December 2023, when The New York Times sued OpenAI and Microsoft over the use of its articles to train and ground models. The New York Daily News and seven sister papers followed, as did the Center for Investigative Reporting, and the suits now run together in New York. The September 2026 filing is the plaintiffs' own brief; as TechCrunch notes, most underlying exhibits are still sealed, so the quotes below are the plaintiffs' characterisation.
Four allegations matter for anyone who scrapes:
- Scale. OpenAI's mid-training datasets are said to contain more than 91,692 copies of works from the three plaintiff groups, and a Common Crawl-derived set more than 2 million documents from nytimes.com alone.
- Paywalls. The brief quotes an internal message about a "hack to get around nytimes paywall" and a two-word reply from OpenAI's president: "ah nice". Microsoft's CEO testified that anything paywalled should be licensed by whoever wants it for grounding or training.
- Repurposed search data. The plaintiffs say Microsoft sold a dataset purchased for Bing to OpenAI as training data, and that OpenAI used a New York Times dataset of 1.8 million articles obtained from a third party bound to non-commercial use.
- Stripped notices. Researchers allegedly removed copyright notices from training data so the model would not output them.
Microsoft's response, per Ars Technica, is that the director's memos reflect one employee's view and that its products are transformative fair use. Courts have so far been receptive to fair-use arguments for training, and the US government filed a brief supporting OpenAI earlier in September. None of this is decided. What is already clear is which behaviours the plaintiffs chose to put in front of the judge: going around paywalls, reusing data collected under one kind of consent for another purpose, and removing attribution.
Twenty of thirty news sites block GPTBot in robots.txt
We fetched each site's robots.txt through the QuanticData batch endpoint over plain HTTP from a US residential exit. All 30 returned 200, 163,199 bytes of text in total, with a median of 47 user-agent lines per file; USA Today's names 288. We then applied standard group matching: an agent is "blocked" when the group that applies to it disallows the root path and nothing re-allows it.
| Crawler | Job | Blocked, all 30 | Blocked, 7 plaintiffs |
|---|---|---|---|
| ClaudeBot | Training (Anthropic) | 27 | 7 |
| CCBot | Common Crawl archive | 26 | 6 |
| PerplexityBot | Search index (Perplexity) | 23 | 6 |
| GPTBot | Training (OpenAI) | 20 | 7 |
| Google-Extended | Gemini training token | 20 | 6 |
| ChatGPT-User | User-triggered fetch (OpenAI) | 15 | 6 |
| OAI-SearchBot | ChatGPT search index | 13 | 6 |
| Googlebot | Google Search | 0 | 0 |
| Bingbot | Bing Search | 0 | 0 |
Twenty of 30 is 67%. In our earlier study of 326 top domains across all sectors, 17.8% blocked GPTBot (see Should I block AI crawlers?). News publishers refuse OpenAI's training crawler at almost four times the web's rate, and 29 of the 30 files block at least one AI agent; only foxnews.com has no AI-specific rule at all.
Look at CCBot too. The filings say a Common Crawl-derived set held over 2 million nytimes.com documents. Today CCBot is blocked on 26 of the 30 sites and on 6 of the 7 plaintiffs. Blocking now does not recall an archive built years ago, which is why the lawsuits exist.
Bingbot: the crawler no publisher can refuse
Googlebot and Bingbot are blocked on zero of 30 sites. That is not an oversight; it is the price of being found. It is also why the "Bing dataset" allegation stands out. If data collected by a search crawler is later used to train a model, robots.txt could never have stopped it, because the publisher's only lever would have cost it its search traffic.
Google split the problem in 2023 by publishing Google-Extended, a robots.txt token that controls use for Gemini training without touching Search; 20 of our 30 sites use it. Microsoft took a different route: Bing reads page-level NOCACHE and NOARCHIVE meta tags to limit how content appears in its chat answers, rather than a separate user-agent. A page-level tag is invisible to a robots.txt audit like ours, so our count of Microsoft-specific opt-outs is structurally zero, not measured zero.
The practical lesson for anyone running a crawler is narrow and firm: one user agent, one declared purpose. If your bot collects prices, call it that and use the data for prices. Reusing a crawl gathered for one purpose for another is precisely the conduct the plaintiffs are asking a court to punish.
Publishers now separate training from search
OpenAI runs three agents with three jobs: GPTBot for training, OAI-SearchBot for the ChatGPT search index, and ChatGPT-User for pages a user asks about. Our 30 files treat them differently. Thirteen sites block all three. Seven block GPTBot but leave OAI-SearchBot open: revealnews.org, latimes.com, reuters.com, apnews.com, techcrunch.com, forbes.com and cbsnews.com. The message is "cite us, do not train on us", and the file is how they say it. Our AI crawler user agent list maps every token to its job.
A few named GPTBot and left it open, among them wsj.com, theverge.com and theatlantic.com. Two sites, wsj.com and reuters.com, run default-deny: User-agent: * with Disallow: /, then an allowlist of named crawlers. Any client not on that list, including yours, is refused everything by the file.
Here is every site, grouped:
| Site | GPTBot | OAI-SearchBot | ClaudeBot | Google-Extended | Bingbot |
|---|---|---|---|---|---|
| Plaintiffs in the New York case (7) | |||||
| The New York Times (nytimes.com) | blocked | blocked | blocked | blocked | open |
| New York Daily News (nydailynews.com) | blocked | blocked | blocked | blocked | open |
| Chicago Tribune (chicagotribune.com) | blocked | blocked | blocked | blocked | open |
| The Denver Post (denverpost.com) | blocked | blocked | blocked | blocked | open |
| The Mercury News (mercurynews.com) | blocked | blocked | blocked | blocked | open |
| Orange County Register (ocregister.com) | blocked | blocked | blocked | blocked | open |
| Reveal (CIR) (revealnews.org) | blocked | open | blocked | open | open |
| Other US and UK news sites (23) | |||||
| apnews.com | blocked | open | blocked | open | open |
| arstechnica.com | open | open | blocked | blocked | open |
| axios.com | open | open | open | open | open |
| bbc.com | blocked | blocked | blocked | blocked | open |
| bloomberg.com | blocked | blocked | blocked | blocked | open |
| businessinsider.com | open | open | blocked | open | open |
| cbsnews.com | blocked | open | open | open | open |
| cnn.com | blocked | blocked | blocked | blocked | open |
| forbes.com | blocked | open | blocked | open | open |
| foxnews.com | open | open | open | open | open |
| latimes.com | blocked | open | blocked | open | open |
| nbcnews.com | blocked | blocked | blocked | blocked | open |
| npr.org | blocked | blocked | blocked | blocked | open |
| politico.com | blocked | blocked | blocked | blocked | open |
| reuters.com | blocked | open | blocked | blocked | open |
| techcrunch.com | blocked | open | blocked | blocked | open |
| theatlantic.com | open | open | blocked | blocked | open |
| theguardian.com | open | open | blocked | open | open |
| theverge.com | open | open | blocked | blocked | open |
| usatoday.com | blocked | blocked | blocked | blocked | open |
| washingtonpost.com | open | open | blocked | open | open |
| wired.com | open | open | blocked | blocked | open |
| wsj.com | open | open | blocked | blocked | open |
What the case means for your scraper
The lawsuits target the copying of expressive work to build a substitute product. Collecting public facts, such as prices, availability, rankings, headlines and links, sits on different ground, and our pieces on whether AI web scraping is legal and web scraping law in the US cover the doctrine. Four rules come straight from what the filings chose to highlight:
- Never go around a paywall. Nadella's own testimony draws the line at paywalled content: license it. A proxy changes where a request comes from; it does not change what you are entitled to read, and no setting we sell is for that.
- Read robots.txt and the terms before the first request. Eleven of the 30 files open with a written legal notice in comments, six of the seven plaintiffs among them; the New York Times file prohibits automated collection without written permission. Our free robots.txt tester shows the exact rule that allows or blocks a URL.
- Keep purpose and provenance. Log what each crawl was for and what the file said on the day. The filings show a dataset outliving its consent; your logs are how you prove yours did not.
- Keep attribution. Store the source URL and publisher with every record. Stripping notices is one of the allegations.
The setting that works on news sites
For monitoring what publishers allow, or collecting headlines and links rather than article bodies, this is what our measurement supports:
- Network: residential Basic from $0.80/GB. The 30 robots.txt files came to 163,199 bytes of text, about $0.00013 of bandwidth at that rate, counting 1 GB as 10^9 bytes.
- Fetch mode: plain HTTP (
engine: tls). All 30 files answered 200 without a browser; rendering robots.txt buys nothing. Through the web scraping API the 30 fetches cost $0.006 at $0.0002 per page. - Country: a US exit for US publishers. Robots rules are the same worldwide, but paywall and consent layers on the pages themselves vary by country.
- When you need the news itself: the Google News collector returns dated articles with title, source, snippet and link for $0.0005 per delivered article, which keeps you on metadata and links rather than copied text. To check a single site quickly, the free AI crawler checker reads its live robots.txt.
- Free tier: $2 of free API usage per month, and failed requests are never billed.
Sources & further reading
- Microsoft exec called AI scraping the "largest theft of labor in human history", Ars Technica, 17 September 2026
- Microsoft exec called AI scraping "the largest theft of labor in human history," new unredacted filings reveal, TechCrunch, 17 September 2026
- OpenAI, Microsoft Sued by Publishers for Scraping Articles, Bloomberg Law
- Announcing new options for webmasters to control usage of their content in Bing Chat, Bing Webmaster Blog
- nytimes.com/robots.txt (fetched 29 September 2026)
- wsj.com/robots.txt (fetched 29 September 2026)