The legal fight over how chatbots learn to write produced its first real precedents in 2026. They don't all point the same way.
For two years, the question of whether an AI company can train its models on copyrighted books sat unanswered in courtrooms across the country. In 2026 the answers started arriving.
On July 20, a federal judge in San Francisco signed off on a $1.5 billion settlement between Anthropic and a class of authors whose books were pulled into the training data behind its Claude chatbot. It is the largest copyright recovery ever reported in the United States. Four months earlier, the Supreme Court quietly declined to review whether a fully machine-generated artwork could be copyrighted. And in June, a federal appeals court in Philadelphia became the first to hear arguments over whether feeding copyrighted text into an AI system qualifies as fair use.
Read together, the rulings send a message that sounds contradictory until you look closely. Training an AI on copyrighted work is often legal. The way companies obtain that work can still cost them billions.
The record payout that dodged the biggest question
Anthropic's settlement is widely misread.
The case, Bartz v. Anthropic, was filed in August 2024 by novelist Andrea Bartz along with nonfiction writers Charles Graeber and Kirk Wallace Johnson. They alleged the company had copied their books without permission to build its models.
In June 2025, Judge William Alsup issued the ruling that mattered most, and it broke in Anthropic's favor on the core issue. He found that using copyrighted books to train a large language model was transformative enough to count as fair use, comparing the process to the way a writer reads widely before producing something original. What Alsup would not excuse was where the books came from. Anthropic had downloaded millions of them from pirate repositories called Library Genesis and Pirate Library Mirror, and the judge described those pirated copies as inherently, irredeemably infringing.
That distinction is the whole story. The $1.5 billion was never a penalty a judge imposed for training on books. It was a negotiated settlement, approved by the court, resolving the piracy claims specifically. Anthropic's deputy general counsel, Aparna Sridhar, said the deal "resolves narrow claims about how certain materials were obtained."
The scale of the response surprised even the lawyers who ran it. The final Works List covered 482,460 books, and rightsholders filed claims on roughly 91 percent of them, an extraordinary rate for a class action, where participation often hovers near 10 percent. Eligible works are expected to draw about $3,000 each, though a single title's payout can be divided among authors, publishers, co-authors or estates. Anthropic agreed to destroy the pirated datasets as part of the deal. Judge Araceli Martínez-Olguín granted final approval after Alsup, who has since retired, shepherded the case through its earlier stages.
One detail undercuts the idea that authors won a sweeping legal victory. The class was certified only over the piracy, not the training. Alsup's fair use finding on training formally binds just the three named plaintiffs, which leaves the larger question technically open for everyone else.
The line courts keep drawing: competition
If there is a through-line in the early rulings, it is market competition.
The clearest example came before the Anthropic decision. In February 2025, Judge Stephanos Bibas ruled that the startup Ross Intelligence had infringed Thomson Reuters' copyrights by using editorial summaries from the Westlaw legal database, known as headnotes, to build a rival legal-research tool. Bibas rejected Ross's fair use defense. He found the use commercial and not transformative, because Ross set out to build a direct market substitute for Westlaw.
That case carried an asterisk worth understanding. Ross's system was not a generative chatbot. It was a search tool trained to recognize relationships in legal language, which makes it an imperfect stand-in for the ChatGPT-style products most people picture. Even so, it produced the first substantive American ruling on fair use in AI training, and the reasoning has echoed through later disputes.
Ross appealed. On June 11, 2026, the Third Circuit Court of Appeals in Philadelphia heard argument in the case, the first time a federal appeals court has taken up whether AI training can be fair use. Lawyers who attended reported that the panel's questions suggested skepticism that a product serving the same function as the original can claim fair use protection. No decision had been issued as of late August.
The competition principle helps explain why authors suing chatbot makers have struggled to win on the training theory itself. To prevail, they would need to show that models trained on their books flood the market with substitute titles that displace their sales. Courts have not accepted that argument.
Who owns what a machine makes
A separate question runs alongside the training fight. Once an AI produces something, who, if anyone, owns it?
The answer, for now, is settled at the top. Stephen Thaler, a computer scientist, spent years trying to register a copyright for an image his AI system generated on its own, listing the machine as the author. The Copyright Office refused. A district court agreed. In March 2025 the D.C. Circuit affirmed that a copyrighted work must be authored by a human being, and on March 2, 2026, the Supreme Court declined to hear the case, leaving that rule in place.
Thaler's situation was clean because he conceded that no human shaped the output. The messier cases are still coming. The Copyright Office has said that typing prompts into a generator does not, by itself, make a person the author of the result, reasoning that prompts function as instructions rather than creative execution. Where the line falls between a human using AI as a tool and a machine doing the creating remains open, and it is the fight likely to define the next several years.
A 50-year-old law doing new work
Every one of these disputes runs through a statute written in 1976, before personal computers and long before software that could digest a library and produce prose on demand.
Judges are left interpreting decades-old fair use rules to resolve cases that will shape a trillion-dollar industry. The doctrine asks them to weigh several factors, chief among them the purpose behind the copying and the effect on the market for the original work.
The result is a patchwork. Ross lost because it built a competitor. Anthropic won on training and paid for piracy. Thaler lost because no human held the pen. Each ruling turns on its own facts, which means the law reads differently depending on which courtroom you are standing in.
What is still unsettled
The biggest cases have not been decided.
Anthropic's settlement resolved one company's exposure over pirated books, but it set no binding precedent on the training question, because settlements do not create law. Parallel suits against other major AI developers over their use of copyrighted material are still moving through the courts, and a split among them could unsettle the tentative view that training is transformative. The Third Circuit's coming decision in the Ross appeal could push the law in either direction, and because it is the first appellate word on the subject, it will carry weight far beyond the two companies fighting it out.
For authors, the settlement delivered money without the ruling they wanted. For AI companies, the lesson is more practical: the training itself may survive in court, while sourcing data from pirate libraries is a liability worth avoiding.
Comments
Join the discussion and share your perspective.