H!
HelloHumans!
Episodes

Research

AI Training Data and Copyright: Ownership, Compensation, and the Press

Courts have now ruled AI training on copyrighted books to be "spectacularly transformative" fair use in two major U.S. decisions, but the NYT-led coalition's sanctions motion — alleging OpenAI destroyed evidence of how journalism was used in training — could reshape that picture if spoliation is proven, since the core evidentiary problem is that even OpenAI's engineers cannot reliably trace which content drove which model behaviors. OpenAI is simultaneously licensing content from AP, Axel Springer, the Financial Times, News Corp, and Le Monde while contesting the NYT lawsuit, revealing a dual strategy that raises an unresolved distributional question: if training is lawful fair use, why pay some publishers and not others? The deepest tension the briefing surfaces is not doctrinal but structural — whether the right remedy for journalism's compounding economic crisis is copyright litigation, collective licensing, statutory levies, or equity deals remains genuinely contested, and the experiences of small and mid-size publishers, who make up most of the press ecosystem, are almost entirely absent from the evidence base.

Sources (50)

Sign up to read the full research briefing

Sign up