NYT Filings Reveal Microsoft, OpenAI Knew of 'Theft' Risks in AI Training

Unsealed filings expose internal fears at Microsoft and OpenAI over AI's impact on publishers.

Last Updated: September 19, 2026 Editorial Process
Editorial Process
See more of Inside AI's trusted news by adding us as a preferred source on Google.
AI neural network visualization
Published on: September 19, 2026

September 19, 2026, (Inside AI) — Newly unsealed court documents in The New York Times' copyright lawsuit against OpenAI and Microsoft reveal that executives at both companies privately acknowledged the legal and ethical risks of training AI models on millions of news articles. The filings, made public this week, include internal memos, deposition transcripts, and company data that could reshape the fair use debate at the heart of the case.

The documents show that Microsoft's own director of applied science, Brent Hecht, described large-scale copying of online content for AI training as an "astonishing theft of unprecedented proportions" and perhaps the "largest theft of labor in human history" in a January 2023 memo. OpenAI's head of ChatGPT, Nick Turley, wrote that the company's products were "largely substitutive, period," predicting they would become more so as they improved. Microsoft CEO Satya Nadella acknowledged in a deposition that chatbot answers could replace visits to the underlying website.

The Times sued both companies in December 2023, alleging unauthorized use of its copyrighted articles to train AI models and that products like ChatGPT could reproduce or closely mimic its journalism. Microsoft and OpenAI have argued that training on copyrighted material qualifies as fair use under US law. The case hinges on whether that training is transformative and whether the resulting products harm the market for the original works.

The unsealed filings provide the publishers with new ammunition. Microsoft's own data shows that click-through rates to Times websites from Bing's AI product were between 87% and 93% lower than from conventional Bing Search. Other publisher groups cited in the litigation saw similar declines. These figures do not by themselves prove infringement, but they support the argument that AI products substitute for and damage the market for the journalism used to build them.

Read: ACC Sues The L Suite for Training Chatbot on Copyrighted Legal Materials

"Microsoft's position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers' journalism," a Microsoft spokesman said. A separate company filing described Hecht as a research academic who also worked at Northwestern University, and said he was employed by Microsoft "to present divergent and asymmetric perspectives," not to be a decision-maker.

The filings also detail how the companies acquired training data. Between 2019 and 2022, Microsoft supplied OpenAI with data from its Bing Index, a compilation of billions of webpages, under a codename Project Taxi. The transfer included Times content, which Microsoft sold to OpenAI in an undisclosed transaction. Microsoft later launched Project Mango to help OpenAI "collect as many of the documents as possible for...training the large language models." The Mango web crawler copied web content on OpenAI's behalf, producing a dataset containing at least 160,903 unique works belonging to the news publishers involved in the litigation.

The documents also contain allegations that OpenAI employees found ways to access paywalled material. In one exchange, an OpenAI researcher told company president Greg Brockman about a "hack to get around" the Times paywall; Brockman replied, "ah nice". In a deposition, Nadella said that "anything that is paywalled should be licensed by anyone who wants to use it." He also said that if he "had been made aware that OpenAI had scraped and trained on information that was behind a paywall," he would have invoked Microsoft's right to require OpenAI to retrain its models.

The legal question remains unresolved. An employee's description of copying as "theft" does not determine whether it constitutes copyright infringement. Courts weigh four statutory factors: the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the potential market. The judge will decide whether the use was fair based on the circumstances.

Read: DOJ Will Probe AI-Related Crimes, Attorney General Blanche Says

For users, the issue is straightforward: AI tools can increasingly answer questions without requiring visits to the websites where that information originated. The filings show that OpenAI and Microsoft were aware of what this could mean for the publishers whose work these systems rely on. The case continues to test how US copyright law applies to generative AI, with implications for the entire news industry and the future of AI training data.

More from Inside AI

  • AI Safety

    Google’s Gemini AI Escapes Sandbox, Hacks Three Companies

    September 19, 2026
  • AI Policy & Regulation

    Pakistan Bans Officials From Using Public AI Tools for Classified Data

    September 19, 2026
  • AI Safety

    Ten Days That Changed AI: Labs Admit They Can’t Control Their Models

    September 19, 2026
  • Machine Learning

    Jev: ChatGPT Inventor’s New AI Model 100x Cheaper

    September 19, 2026
  • AI Policy & Regulation

    IIT Bombay Student Dies After ChatGPT Exam Cheating Incident, Protests Erupt

    September 19, 2026
  • AI Tools

    Plaud Note Pro Review: AI Dictaphone Returns with Steep Subscription Costs

    September 19, 2026
  • Artificial Intelligence (AI)

    Meghna Gulzar Says AI Can Never Replace Human Instinct in Cinema

    September 19, 2026
  • Artificial Intelligence (AI)

    AI can map hazards during disasters like Nepal floods: Kamal Bawa

    September 19, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Policy & Regulation
  • AI Hardware & Infrastructure
  • AI Tools
  • AI In Business
  • Robotics
  • Cybersecurity AI
  • AI Safety
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases
  • Features
  • Sponsored Content
  • Newsletter

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital