FIELD GUIDE · AI, ART & COPYRIGHT

IS AI STEALING FROM ARTISTS?

Training looks like learning to one side and industrial-scale appropriation to the other. Copyright law answers only part of the fight.

UPDATED 2026-09-21

NO PHD REQUIRED

ELI5

Generative AI is trained by adjusting model parameters after exposure to enormous collections of text, images, music, code, and other data. The model usually does not store each work like a folder of files, but training can still involve making copies, models can sometimes reproduce or memorize source material, and new outputs can compete with the people whose work helped train the system. Copyright law asks several different questions depending on what was copied, how it was acquired, what the model does with it, and what market harm can be shown.

TRAINING

A model changes its internal parameters after processing examples so it can predict or generate new outputs.

Training is not the same thing as simply searching a database, but copyrighted works may still be copied during the process.

FAIR USE

A U.S. copyright doctrine that can allow some unlicensed uses after weighing purpose, the work used, amount, and market effect.

It is a case-specific legal test, not a universal AI exemption.

MEMORIZATION

A model can sometimes reproduce unusually close or verbatim fragments of material seen during training.

The existence and frequency of memorization matter because 'learning patterns' and 'reproducing works' are different legal and moral claims.

STYLE

The recognizable aesthetic traits people associate with a creator or genre.

Copyright generally protects particular expression, not an artistic style by itself, which is why the moral and legal arguments do not line up perfectly.

WHY THIS BECOMES A FIGHT

WHY ARE PEOPLE FIGHTING ABOUT THIS?

Creators see companies building valuable systems from a cultural archive they did not create, often without asking first. Developers see learning from existing works as a basic ingredient of useful models and warn that mandatory licensing for every input could entrench the richest companies. The moral question is broader than the legal one: when a new industry gets rich by learning from everybody, what does it owe the people whose work made the learning possible?

GET THESE OFF THE TABLE

THE STRAW MEN

The bad arguments first. Nobody gets to win by beating these.

THE PIRATE-HARD-DRIVE VERSION

“THE MODEL IS JUST A PIRATE HARD DRIVE THAT SPITS BACK STOLEN ART.”

THE INTERNET-WAS-FREE VERSION

“IF IT WAS ON THE INTERNET, IT WAS FREE. CASE CLOSED.”

NOW MAKE THE GOOD ARGUMENT

STEEL MAN THE CASE

Give the people you disagree with the version they would actually defend.

THE CONSENT + COMPENSATION CASE

CREATORS DID NOT VOLUNTEER TO BECOME RAW MATERIAL

Training systems can require large-scale copying, some datasets have included unlawfully acquired material, and generated outputs can compete in the same markets as the originals. Even when a particular training use is legally fair, creators can reasonably argue that an industry extracting commercial value from their work should offer consent, attribution, licensing, or compensation where practical.

Reading every book in the library is one thing. Building a competing publishing machine from a pirated copy of the library is a different moral picture.

THE FAIR-USE + INNOVATION CASE

LEARNING FROM CULTURE IS NOT THE SAME AS REPUBLISHING IT

Models can learn statistical relationships that support new expression, analysis, translation, accessibility, and other socially useful applications. Requiring a negotiated license for every training work could make frontier development available only to incumbents with the money and rights infrastructure to clear enormous catalogs. Courts have already found some model-training uses fair on specific records.

A painter can study thousands of paintings without owing every painter a fee for the ideas learned. The difficult question is when machine training stops looking like study and starts looking like substitution or copying.

FOLLOW THE MONEY

WHO PAYS? WHO WINS?

CREATORS

They bear uncompensated training use and potential market substitution when their work is used without a license and the resulting system competes for similar commissions, audiences, or sales.

AI DEVELOPERS

Licensing, provenance, filtering, and rights management can be expensive, especially at web scale. Broad mandatory licensing could advantage companies with the deepest pockets.

RIGHTSHOLDERS + PLATFORMS

Publishers, labels, stock libraries, and platforms can gain new licensing revenue and bargaining power, but deals made at the catalog level may not distribute value evenly to individual creators.

USERS + THE PUBLIC

They benefit from cheaper creative tools, translation, search, accessibility, and experimentation. They can also lose if a thinner creative economy produces fewer original works or concentrates cultural production in a few model providers.

RECEIPTS, NOT VIBES

WHAT DO WE ACTUALLY KNOW?

U.S. COPYRIGHT LAW DOES NOT HAVE ONE BLANKET ANSWER FOR AI TRAINING

The U.S. Copyright Office's 2025 generative-AI training report says fair use must be analyzed for the particular use and circumstances rather than treating every stage of model development and deployment as identical.

The source of the data, the purpose of copying, the kind of work, and market effects can all matter.

COURTS HAVE ALREADY REACHED DIFFERENT RESULTS ON DIFFERENT COPIES

In Bartz v. Anthropic, the district court held training on the books at issue was fair use while treating Anthropic's acquisition and retention of pirated library copies separately. In Kadrey v. Meta, Meta won summary judgment on the plaintiffs' fair-use record, while the court stressed that the ruling did not establish that all AI training is categorically lawful.

The cases make the same point the slogans miss: acquisition, purpose, evidence, and market harm matter.

MEMORIZATION IS REAL, BUT IT IS NOT THE WHOLE MODEL

Research has shown that language models can memorize particular training examples and that extraction risk rises for some repeated or distinctive sequences. That supports concern about reproducing protected expression without reducing all model behavior to retrieval.

How often meaningful copyrighted content can be elicited, and under what conditions, remains an empirical question.

WHAT WOULD SETTLE SOME OF THIS?

WHAT WOULD CHANGE THE ARGUMENT?

TAKE THIS TO DINNER: Fair use can answer whether a use is legal. It cannot, by itself, answer what a trillion-dollar industry owes the people whose work fed the machine.

RECEIPTS

The guide is the map. These are the sources behind the substantive claims.

  1. Copyright and Artificial Intelligence, Part 3: Generative AI TrainingU.S. Copyright Office
  2. Bartz v. Anthropic PBC, order on summary judgmentU.S. District Court / Justia
  3. Kadrey v. Meta Platforms, Inc., order on fair useU.S. District Court / Justia
  4. Quantifying Memorization Across Neural Language ModelsNeurIPS
  5. Scalable Extraction of Training Data from (Production) Language ModelsarXiv