Copyright & Data Theft

Copyright and AI: two concepts that haven't figured out how to deal with each other yet.

Ethics 8 min Beginner June 8, 2026

AI models learn from billions of images, texts, and songs — most of them copyrighted. Artists discover their styles replicated in AI generators without ever giving consent. Companies generate millions in content but legally own none of it. Three questions define this conflict: Is AI training on copyrighted data legal? Can creators opt out? And who owns what an AI produces? The answers are far less clear than either side claims.

Legal status as of early 2026. The legal landscape for AI copyright is evolving rapidly — court rulings may change the positions described here.

Fair Use — When May AI Use Your Data?

Fair Use

AnalogyDefinition
Imagine reading every cookbook ever published, absorbing techniques and flavor combinations, then writing your own recipes from memory. Did you steal or did you learn? That is the core dispute in AI training — and courts have not yet decided the answer. Breaking point: A human genuinely transforms and reinterprets. An AI model can sometimes reproduce near-exact copies (as the Getty watermark case shows), which considerably weakens the "transformation" argument.

US Fair Use vs. EU TDM Compared

US: Fair Use Doctrine

Case-by-case assessment using four factors (purpose, nature, amount, market impact). AI companies argue: training is "transformative" — the model extracts patterns, not copies. No statutory opt-out right. Pending cases: NYT vs. OpenAI, Getty vs. Stability AI.

EU: TDM Directive (DSM 2019)

Statutory regulation via Text and Data Mining. Art. 3: TDM for scientific research permitted. Art. 4: TDM also commercially permitted — but with an explicit opt-out right for rights holders. Legally binding within the EU.

Example: Getty Images vs. Stability AI

Stability AI trained Stable Diffusion on the LAION-5B dataset — a collection of over five billion images from the internet. Generated images repeatedly contained distorted versions of Getty watermarks. For the plaintiffs, this is evidence that the model memorized specific copyrighted images rather than learning only abstract patterns.

Misconception: "AI training is clearly illegal"

Wrong. The legal status is genuinely unsettled. Both sides present valid legal arguments. In the US, no supreme court ruling exists yet. In the EU, the TDM Directive addresses at least part of the question. It will likely take years before established precedents exist.

Opt-Out — The Right to Say No

Opt-out is the right of creators to prohibit the use of their works for AI training. Since the default on the internet is that anything accessible gets crawled, creators must take active steps. Four mechanisms are available — but all share a fundamental limitation.

robots.txt Website operators can deny access to AI crawlers. Based on voluntary compliance — not legally binding per se.
EU TDM Opt-Out Art. 4.3 of the DSM Directive: rights holders can object in machine-readable form. Legally binding within the EU.
Platform Tags DeviantArt, ArtStation, and others offer "No AI Training" tags. Activatable per click, but only effective on the respective platform.
Glaze Tool from the University of Chicago: Alters pixel structure invisibly to the human eye, but deliberately disrupts the AI's learning process.

Example: The Photographer and His Stolen Style

A photographer discovers that his unique style is perfectly replicated through the prompt "in the style of [Name]." He adds a restrictive robots.txt to his website. This blocks future crawling — but the existing model has already encoded his style in billions of weights. Only training a completely new model from scratch without his data could remove his style — practically infeasible.

Misconception: "Opt-out fully protects me"

No. Opt-out mechanisms only protect against future training runs. Most widely-used AI models today were trained on datasets collected before these protection mechanisms existed. Retroactive opt-out is practically impossible because the extracted patterns are embedded in billions of model weights.

AI-Generated Content — Who Owns the Output?

The US Copyright Office established in 2023: Purely AI-generated works without significant human creative contribution are not copyrightable. The rationale: copyright protects human creativity — a machine is not an author. For mixed works (human arrangement + AI generation), the human portion can be protected.

5 Bn
LAION-5B Images Over five billion images from the internet form the training dataset of Stable Diffusion
0%
Copyright Protection (US) Purely AI-generated images receive no copyright protection under US law — they effectively belong to nobody
2023
Zarya Ruling US Copyright Office: text and layout protected, AI-generated images not

Case Study: Zarya of the Dawn

Kris Kashtanova created a comic in 2023 using her own text and Midjourney-generated images. The US Copyright Office ruled: the text and page layout (human creation) are copyrighted. The individual AI-generated images are not — they lack the required human authorship.

Misconception: "AI images are free of copyright issues"

Dangerous half-truth. The generated image itself may not be copyrightable — but if it closely resembles an existing copyrighted work, the original rights holder can still sue for infringement. Regardless of whether a human or a machine created the copy.

Interactive: Is Your AI Output Original Enough?

Use this checklist to evaluate whether your AI-generated content is on solid legal and ethical ground. Each item carries a different weight — at the end you will see an overall assessment.

Originality Check for AI Content

Check your AI-generated content for copyright risks. Answer the questions honestly — you will receive an assessment at the end.

1.Have you verified that the output does not closely resemble any specific existing work?
Use reverse image search or text comparison. If a recognizable existing work is being reproduced, there is a high risk of infringement.
2.Have you added your own creative contributions (editing, selection, arrangement)?
According to the Zarya decision, only the human-created portions are eligible for copyright protection. The more original creativity, the better.
3.Is the output free of recognizable trademarks, logos, or watermarks?
The Getty case shows: if the AI model reproduces protected watermarks, this is strong evidence of memorization rather than transformation.
4.Are you familiar with the AI tool's terms of service regarding ownership rights?
Some tools reserve rights to the output or restrict commercial use. Check the terms before publishing.
5.Are you using the output transformatively — in a new context, not as a replacement for an original?
Fair use factor 1: Transformative use (new purpose, new meaning) is more likely to be considered permissible than simple copying.
6.Are you confident the output contains no verbatim copyrighted text passages?
AI models can reproduce training texts verbatim. For longer text outputs, spot-check for known passages.
7.Have you documented your creative process (prompts, editing steps)?
In legal disputes, documented human creativity can make the difference. Keep a record of prompts and post-processing steps.
0 / 7 answered

The distorted Getty watermarks in AI-generated images are legally significant: they demonstrate that the model memorized specific copyrighted images — not merely extracted abstract patterns. This memorization considerably weakens the central "transformative use" argument of AI companies. If a model stores a work precisely enough to reproduce watermarks, it becomes difficult to argue that it merely "transformed" the work.

The legal landscape for AI copyright differs fundamentally from country to country. US: Purely AI-generated works receive no copyright protection (Copyright Office 2023). UK: Offers a limited protection framework for "computer-generated works" — even without human authorship. EU: Largely unsettled, national courts issue partly contradictory rulings. China: Early court decisions suggest that AI works guided by precise human prompts may be protectable. There is no global consensus — jurisdiction matters enormously.

Key Takeaways

  1. Legal status globally unsettled: Whether AI training on copyrighted data is legal remains disputed worldwide. Neither "clearly legal" nor "clearly illegal" is accurate.
  2. Opt-out has limits: Protection mechanisms exist (robots.txt, EU TDM, platform tags, Glaze) — but they cannot retroactively remove patterns already learned by existing models.
  3. Pure AI works are not protected: Under US law, purely AI-generated works receive no copyright protection. Only the human-created portions of mixed works are protected.
Question 1 / 5
Not completed

What does the US Fair Use Doctrine evaluate when determining whether the use of copyrighted material is permissible?

Select one answer
Answer Key: 1) B · 2) B · 3) B · 4) B · 5) B

Learning Goals

  • Can you name the four factors of the Fair Use Doctrine and explain why AI companies consider their training "transformative"?
  • Can you name two opt-out mechanisms and explain why they don't help against already-trained models?
  • Can you summarize the Zarya of the Dawn ruling: what is protected and what is not — and why?