Writers, photographers, musicians, illustrators and publishers are being approached by AI developers who want to license their catalogs, and many more are finding their work in datasets they never agreed to. Our page on IP licensing and assignments explains how a license is built. This page applies those basics to training deals, drawing on the Copyright Act and the Copyright Office's three-part report on copyright and artificial intelligence.
How the law treats training, step by step
- Training involves copying. The copyright owner has the exclusive right to reproduce the work and to prepare derivative works (17 U.S.C. 106). The Copyright Office's Part 3 report concludes that "several stages in the development of generative AI involve using copyrighted works in ways that implicate the owners' exclusive rights."
- The question becomes fair use. "The key question, as most commenters agreed, is whether those acts of prima facie infringement can be excused as fair use," judged under the four factors in 17 U.S.C. 107.
- The answer depends on the use. The Office says fairness "will depend on what works were used, from what source, for what purpose, and with what controls on the outputs."
- A license removes the question. With a license, the developer does not need fair use for the licensed uses. The Office reports that licensing agreements, individual and collective, are "fast emerging in certain sectors, although their availability so far is inconsistent."
- Unlicensed use is decided in court. An owner who wants to sue over a U.S. work must first register it (17 U.S.C. 411(a)).
What did the Copyright Office conclude?
The Office has issued its report in three parts: Part 1 on digital replicas (July 31, 2024), Part 2 on the copyrightability of AI outputs (January 29, 2025), and a pre-publication version of Part 3 on generative AI training (May 9, 2025), with a final version expected "without any substantive changes expected in the analysis or conclusions" (Copyright Office, Copyright and Artificial Intelligence).
On fair use, the Part 3 conclusion makes three points. "Various uses of copyrighted works in AI training are likely to be transformative." When a model is used for purposes such as analysis or research, "the outputs are unlikely to substitute for expressive works used in training." But "making commercial use of vast troves of copyrighted works to produce expressive content that competes with them in existing markets, especially where this is accomplished through illegal access, goes beyond established fair use boundaries."
On licensing, the Office found voluntary licensing growing and said government intervention "would be premature at this time," recommending that licensing markets keep developing and that alternatives such as extended collective licensing be considered where gaps are unlikely to be filled. The report is the Office's analysis, not a court ruling, and courts will decide the cases in front of them.
| Situation | Part 3 view | What it means for an owner |
|---|---|---|
| Model used for analysis or research, outputs not substituting for the works | Outputs "unlikely to substitute"; more likely fair | Harder to stop; licensing may still be negotiated |
| Commercial model producing expressive content that competes with the training works | "Goes beyond established fair use boundaries," especially with illegal access | Stronger position to demand a license or sue |
| Works obtained from pirated or unlawful sources | Source is a factor that "can affect the market" analysis | Keep evidence of where your work was obtained |
| Licensed training under an agreement | Licensing "fast emerging in certain sectors" | Terms are set by contract, not by fair use |
What should a training license cover?
- Exactly what is licensed. Which works, in what format, and whether metadata, captions or annotations come with them.
- Which uses. Training only, or also fine-tuning, evaluation, retrieval at the time of a user's request, and display of excerpts in outputs. Each is a different use with a different value.
- Which models. One model, a family of models, or any future model, and whether sublicensing to other developers is allowed.
- Exclusivity. A nonexclusive license lets you license others; an exclusive one is a transfer of ownership that needs a signed writing (17 U.S.C. 204(a)), as our page on whether a license or assignment must be in writing explains.
- Output controls. Commitments about outputs that reproduce or closely imitate your works, which the Office treats as relevant to market harm.
- Term, deletion and audits. What happens to the data and to trained models at the end, and how compliance is checked.
- Money. A flat fee, per-work fee, or ongoing royalty, with reporting, the variables our licensing page discusses.
- Attribution and identity. If voices, faces or names are involved, terms on likeness, which the Office's Part 1 report on digital replicas addresses and our personal brand and NIL page covers.
Can you stop unlicensed training?
Options exist, but none is a switch. Registration comes first, because an owner of a U.S. work cannot sue for infringement until the work is registered, and statutory damages and fees depend on registering in time. Our page on registration cost and processing time compares the filing options. The remedies are set out on our page on copyright infringement damages. Removing copyright management information, such as titles, author names and copyright notices, can be a separate violation when done knowingly and with reason to know it will facilitate infringement (17 U.S.C. 1202(b)), and the Part 3 report notes that some training practices may implicate that provision.
Technical opt-outs are imperfect. The Part 3 report records commenters' views that ignoring opt-outs, such as robots.txt instructions or terms of use, "might inform the fair use analysis," and others' concerns that robots.txt works only if respected, cannot be used by owners who do not control the platform hosting their work, and also blocks search engines. Abroad, the report notes that European Union law conditions a text and data mining exception on respecting owners' opt-outs; our page on whether a U.S. license covers other countries explains why territory matters in licensing.
What changes the answer
- Who owns the work. Only the owner of the reproduction right, or an exclusive licensee of it, can license or sue. Check publishing, label and platform agreements before signing.
- What the model does with the work. The Office treats the purpose and the effect on markets for the works as central.
- Where the data came from. Lawfully accessed copies and pirated copies are treated differently in the Office's analysis.
- AI-generated outputs. Material generated without sufficient human control is not protected; the D.C. Circuit affirmed that human authorship is required in Thaler v. Perlmutter (2025), and the Supreme Court denied review on March 2, 2026. Our post Creating With AI: Where Copyright Protection Stops or Thins covers what that means for creators using these tools.
- The other side's finances. A license with an early-stage developer should consider what happens if it fails; our page on licenses in bankruptcy explains section 365.
A worked example
For example, suppose an Atlanta photographer with 8,000 registered images is offered a flat fee by an AI developer for "all rights to use the images in connection with artificial intelligence."
She counters with a nonexclusive license limited to training one named family of image models, no use for retrieval or display of her images in outputs, a commitment that the developer will filter outputs that reproduce her images, deletion of her files at the end of the term, an annual compliance statement, and a fee per image plus an annual renewal fee. She reserves all other rights, including licensing to other developers.
Separately, she finds her images in a dataset compiled by a different company she never dealt with, with her names and copyright notices stripped. Her registrations mean she can sue if needed, and the stripped notices may support a section 1202(b) claim; the Copyright Office's analysis gives her a basis to argue that a commercial image generator trained on her work competes with her market.
Common mistakes
- Signing "all AI uses" language. It sweeps in uses you may want to price separately or refuse.
- Licensing what you do not own. A publisher or label may hold the rights a developer needs.
- Skipping registration. Without it, a U.S. owner cannot sue and may lose statutory damages.
- Assuming robots.txt solves it. The Office's report describes its limits.
- Granting exclusivity cheaply. An exclusive training license can block every other deal; see our page on what exclusivity gives a licensee.
- No end-of-term terms. Without deletion and audit clauses, the license may never really end.
What to do this week
- Inventory your works and confirm you own the reproduction rights, not a publisher, label or client.
- Register your most valuable works, or group-register where the Copyright Office allows.
- Keep copyright notices and author information attached to your files.
- Decide your position: license, refuse, or license only certain uses.
- If approached, ask for the developer's draft and list of uses before talking price.
- Document any evidence that your work appears in a dataset or output.
Frequently asked questions
Is AI training fair use?
It depends. The Copyright Office says some training uses are likely transformative and fair, while commercial uses that produce competing content, especially from unlawfully accessed works, go beyond fair use. Courts decide each case on its facts.
Is the Copyright Office report binding?
No. It is the Office's analysis and recommendations; courts decide fair use. Part 3 was released in pre-publication form, with no substantive changes expected in the final version.
Can I register work I made with AI tools?
Only the human-authored parts. The Office requires applicants to disclose AI-generated material that is more than de minimis, and human authorship is required.
Does a training license need to be in writing?
An exclusive license does, under section 204(a). A nonexclusive one does not, but a signed writing is the only practical way to fix the scope and gives priority over later transfers (17 U.S.C. 205(e)).
What about my voice or likeness?
Those raise right of publicity and digital replica issues. The Copyright Office's Part 1 report recommended a federal digital replica law; contracts remain the main protection. Our page on legal protection against an AI copy of your voice or face covers digital replicas.
Who should negotiate for a large catalog?
Whoever owns the rights. For catalogs held by publishers or labels, the agreements with them decide who can license training, which is why reviewing those contracts comes first.
Zala IP Law advises creators and companies on licensing works for new technologies and on responding to unlicensed use, and Shreepal J. Zala practices federal intellectual property law nationally. If a developer has approached you, request a consultation or call 404-313-1701 before you sign.
Sources
- Copyright and Artificial Intelligence, Part 3: Generative AI Training, pre-publication version, May 2025 (U.S. Copyright Office)
- Copyright and Artificial Intelligence hub (U.S. Copyright Office)
- Copyright and Artificial Intelligence, Part 2: Copyrightability, January 2025 (U.S. Copyright Office)
- 17 U.S.C. 106: exclusive rights (GovInfo)
- 17 U.S.C. 107: fair use (GovInfo)
- 17 U.S.C. 204: execution of transfers (GovInfo)
- 17 U.S.C. 205: recordation and priority (GovInfo)
- 17 U.S.C. 411: registration before suit (GovInfo)
- 17 U.S.C. 1202: integrity of copyright management information (GovInfo)
- Thaler v. Perlmutter, No. 23-5233 (D.C. Cir. March 18, 2025)
- Supreme Court docket 25-449, petition denied March 2, 2026
- Copyright Registration Guidance: Works Containing Material Generated by Artificial Intelligence, 88 FR 16190 (2023)