Hearings to examine the AI industry's mass ingestion of copyrighted works for AI training

Immigration Enforcement and Sanctuary PoliciesSenate Judiciary Subcommittee on Crime and Counterterrorism · 2025-07-16 · 119th Congress
The Senate Judiciary Subcommittee on Crime and Counterterrorism held this hearing to examine allegations that major AI companies, particularly Meta, used torrented and pirated copyrighted books and articles — obtained from "shadow libraries" like Library Genesis and Anna's Archive — to train generative AI models. Begins at 0:18:58
Transcript
Highlights

Title

"Too Big to Prosecute?": AI companies' use of pirated works to train models

Purpose

The Senate Judiciary Subcommittee on Crime and Counterterrorism held this hearing to examine allegations that major AI companies, particularly Meta, used torrented and pirated copyrighted books and articles — obtained from "shadow libraries" like Library Genesis and Anna's Archive — to train generative AI models. Witnesses included an author's attorney litigating against Meta, an economist studying digital piracy, a law professor specializing in copyright and fair use, and bestselling novelist David Baldacci, who testified about AI's effect on creators. Begins at0:18:58

Who spoke

Chair Josh Hawley (R-MO)0:19:11: Opened by calling the AI industry's use of pirated material "the largest intellectual property theft in American history," alleging Meta and Anthropic used torrent networks to steal enough copyrighted material to fill 22 Libraries of Congress0:21:11; later walked through internal Meta chat logs showing employees flagged the practice as "beyond our ethical threshold"1:03:43 and that decisions to proceed were escalated to Mark Zuckerberg1:38:11.

Sen. Dick Durbin (D-IL), Ranking Member0:24:16: Framed the hearing around balancing AI innovation with creators' rights, noting copyrighted industries contribute over a trillion dollars annually0:24:16; said Anthropic pirated over 7 million books from shadow libraries despite having legal purchase options0:26:46, and later pressed Edward Lee on whether the burden improperly falls on authors to prove market harm1:12:07.

Maxwell Pritt, attorney representing authors suing Meta0:28:26: Testified Meta pirated over 200 terabytes of books and articles — comparable to the Library of Congress's printed collection 20 times over0:30:39 — and that the decision was approved by Zuckerberg despite employee objections0:31:55; said Meta routed torrenting through Amazon Web Services rather than its own servers to avoid being traced1:33:41.

Prof. Michael Smith, Carnegie Mellon University0:34:10: Cited 25 years of research and a USPTO-commissioned study concluding piracy harms creators' earnings and reduces incentives for creative and economic investment0:35:21; warned AI-enabled piracy could flood markets with machine-generated output that displaces human creators, especially emerging artists0:37:19.

Prof. Bhamati Viswanathan, author and copyright law professor0:39:18: Argued AI companies compound a pre-existing crime by taking material from illegal shadow libraries, comparing it to knowingly buying a stolen car0:41:07; said Meta's conduct meets both prongs of criminal copyright liability — willfulness and commercial gain1:00:36.

David Baldacci, novelist0:44:23: Said his son's ChatGPT query produced an outline of a Baldacci-style novel, "like somebody backed up a truck to my imagination and had stolen everything I had ever written"0:46:09; testified a class-action database showed AI companies used at least 44 of his books to train large language models1:22:46.

Prof. Edward Lee, Santa Clara University School of Law0:50:44: Argued courts in the Anthropic and Meta cases correctly found AI training highly "transformative" under fair-use factor one0:51:33, but urged Congress to let ongoing litigation (44 pending lawsuits) resolve the question rather than legislate now1:17:26; acknowledged under questioning that AI training's direct financial benefit accrues to AI companies, not authors1:14:13.

Sen. Peter Welch (D-VT)0:20:46: Noted his TRAIN Act with Sen. Blackburn would require AI platforms to disclose use of copyrighted material upon reasonable suspicion of infringement1:22:21; had Baldacci confirm at least 44 of his books were used without compensation1:22:46.

Key moments

Hawley presented internal Meta chat logs in which an engineer wrote "I don't think we should use pirated material... it should be beyond our ethical threshold," and another responded they needed to "cut some corners here and there"1:03:431:04:56.

Pritt testified Meta considered licensing (with internal documents suggesting tens to hundreds of millions of dollars contemplated) but abandoned it because piracy was faster and paid copyright holders zero1:02:481:03:13.

Documents shown by Hawley indicated Meta deliberately routed torrenting through Amazon Web Services instead of its own servers specifically to avoid tracing the activity back to the company1:33:411:36:14.

Hawley disclosed that after failing to acquire licenses, Meta's decision to use torrented works as training data was escalated to and approved by Mark Zuckerberg in spring 20231:38:111:39:00.

Sharp exchange between Hawley and Lee: Lee said AI training serves the "national interest" in competing with China; Hawley countered this meant enriching corporations "may impoverish American citizens" like Baldacci, a U.S. citizen1:15:451:16:08.

Viswanathan testified Anna's Archive, a shadow library, explicitly advertises to AI companies "train on us" and solicits donations from them, arguing this helps pirate sites "thrive and proliferate"0:59:291:00:00.

Pritt said Meta's copying totaled "well over 200 terabytes" of pirated material, including 12 books authored by subcommittee members and every 21st-century U.S. president and vice president0:31:031:02:05.

Lee cited "AI czar" David Sacks's statement that without a fair-use pathway for AI training, "we will lose the race with China," prompting Hawley to ask whether an "unelected AI czar" should decide Americans' rights1:16:491:17:18.

Durbin and Lee debated whether the burden improperly falls on authors to prove market harm once fair use is asserted, with Lee responding that AI companies could still lose if plaintiffs show "cognizable market harm"1:12:071:13:16.

Smith testified that AI licensing deals amount to contracts "signed with a gun to your head," since companies can simply steal the content if authors refuse to license it1:24:53.

Metadata

CommitteeSenate Judiciary Subcommittee on Crime and Counterterrorism
Chamber / CongressSenate · 119th Congress
Date2025-07-16
TypeMeeting
Witnesses
(none listed in event metadata)
Videosenate-isvp
Transcript221 caption blocks · 11,261 words · 1:40:52 runtime
EventCongress.gov 337253