Is training AI on copyrighted books legal? Courts are still figuring it out
US court rulings suggest training AI models on copyrighted works is often legal, but the overall legal landscape remains murky and inconsistent.

Large language models behind chatbots like ChatGPT, Gemini and Claude are trained on vast troves of text, including hundreds of millions of books and articles, much of it used without authors' knowledge or consent. Yet legal experts say that doesn't automatically make the practice illegal.
Last year, Judge William Alsup ordered Anthropic to pay $1.5 billion to a group of writers whose work was used in training its AI models. Despite the size of the settlement, the ruling actually found that training AI on the books themselves was lawful — Anthropic was penalized specifically for obtaining the books through pirated "shadow library" sources, not for using them to train its models. Alsup compared how an LLM absorbs huge volumes of text to a writer studying literature, noting the goal was to create something new rather than reproduce existing work.
A legal framework built for a different era
US copyright law hasn't been substantially updated since 1976, forcing judges to apply decades-old standards to novel AI disputes. Much of the debate centers on "fair use" — a doctrine permitting use of copyrighted material without permission when it is sufficiently "transformative." Courts weigh factors such as the purpose of the use, how much material was used, and the effect on the original market.
A contrasting outcome came in Thomson Reuters v. Ross Intelligence, where a judge ruled that using Reuters' content to build a directly competing AI legal research platform was not fair use. Attorneys note that courts tend to favor AI companies when their products don't directly compete with the source material, and are more skeptical when they do.
Another unresolved question concerns AI-generated content itself: in Thaler v. Perlmutter, a court ruled that works generated entirely by AI are not eligible for copyright protection, raising further questions about how to determine what portion of a work was AI-assisted.
With most major AI companies still facing ongoing litigation, a definitive legal resolution remains far off. In the meantime, early rulings continue to shape how the industry operates, even though later court decisions could still overturn or reshape today's precedents.

