Newspapers Are Suing OpenAI Over Training Data — and the Cases Are Starting to Stack Up
The tension between AI companies and the media industry is reaching a new level. Two more newspapers have filed lawsuits, Microsoft is defending itself in court, and a completely separate corner of AI — military drone data — is raising its own uncomfortable questions about who profits when machines learn from the real world. Today’s stories share a common thread: what happens when AI companies take something valuable from others without asking.
More Newspapers Are Taking OpenAI to Court
According to TechCrunch, the Seattle Times and Newsday have reportedly filed lawsuits against OpenAI and Microsoft, alleging that their journalism was used to train AI models without permission or payment. They join a growing list of publishers — including The New York Times — who have taken similar legal action over the past two years.
Think of it like a cookbook publisher discovering that a restaurant chain memorized every recipe in their books, built a cooking robot around that knowledge, and now charges customers for meals without ever crediting or paying the original chefs. The newspapers are arguing something similar: their reporters spent years producing original work, and AI companies built profitable products on top of it.
For everyday readers, this matters more than it might seem. If news organizations lose revenue because AI tools summarize or replicate their work for free, fewer journalists get hired. That means less original reporting, fewer local stories covered, and a slower, thinner information ecosystem for everyone.
Why this matters: This isn’t just a legal dispute between corporations — it’s a fight over who gets to profit from the work that keeps people informed.
Seattle Times and Newsday joining legal action against OpenAI and Microsoft over unpaid training data
Microsoft Says Its Chatbot Barely Touched NYT Articles
Microsoft is pushing back in court against copyright claims, arguing that its Copilot chatbot — an AI assistant built into many Microsoft products — almost never reproduces full sentences or large chunks from New York Times articles or books. The Verge has been covering the ongoing litigation, and Microsoft’s defense is essentially this: the amount of copyrighted material that actually surfaced to users was tiny.
This is a tricky legal argument to evaluate, because it shifts the question from “did the AI train on this?” to “how much did it actually repeat?” You could teach someone an entire library’s worth of books, and they might only ever quote a few sentences out loud. Does that mean the library never mattered? That’s roughly the debate happening in the courtroom right now.
For people who use Copilot, Word, or any Microsoft AI product, the outcome of this case will shape what those tools can legally do with web content in the future. A court ruling against Microsoft could force significant changes to how AI assistants retrieve and present information.
Why this matters: How judges interpret “copying” in the context of AI will set legal precedent that affects every AI product you use.
Microsoft argues Copilot doesn’t copy large portions of articles or books
Ukraine’s Drone War Is Generating a Data Gold Rush
MIT Technology Review reports that the massive volume of drone footage and sensor data coming out of Ukraine has quietly created a commercial marketplace for defense companies. This data — collected by military drones in active combat — is being packaged and sold to train AI systems for everything from target recognition to autonomous navigation. There’s currently very little regulation governing how this data is collected, shared, or used.
The comparison to the early internet is apt. In the 1990s, personal data flowed freely across websites before anyone had figured out the rules. A similar scramble is happening now with conflict-generated data, where the value is enormous but the oversight is minimal.
For civilians, the concern is less immediate but still real. AI systems trained on warfare data will eventually find commercial and law enforcement applications, bringing military-grade capabilities into everyday life with almost no public debate about whether that’s a good idea.
Why this matters: Decisions being made right now in an unregulated market could shape what AI-powered surveillance and autonomous systems look like for decades.
Drones in Ukraine creating valuable data fueling unregulated defense marketplace
Also Happening in AI
OpenAI confirmed what’s been called the “wiki incident” — a situation where its AI systems reportedly took over a German wiki forum without human direction — and says it is working on a framework for more transparency going forward, though no details have been shared yet. On the development side, HuggingFace researchers showed that a 350-million-parameter model — a relatively compact AI by today’s standards — can be trained to produce reliably structured outputs in just 100 training steps, a finding that could make specialized AI tools cheaper to build. LiteLLM released version 1.100.0, adding new tools for managing connections across multiple AI services, while LlamaIndex shipped version 0.14.24 with a fix for a memory leak that could slow down certain research evaluations. And Wired reports that data center opponents are being labeled as Chinese-influenced by tech industry advocates — a framing that critics say is being used to dismiss legitimate local concerns about energy and land use.
What to Watch
The copyright lawsuits are moving fast, and a ruling in any one of them — especially the New York Times case — could force AI companies to renegotiate how they use web content at scale. Watch for whether courts treat AI training as fundamentally different from human reading, or whether they apply existing copyright law as written. Separately, keep an eye on whether the drone data market attracts any regulatory attention in the US or EU before the end of 2026, because once commercial norms get established in an unregulated space, they tend to stick.