Ai Training Data Use Demand Letters
Copyright Infringement, Unauthorized Scraping & Licensing Disputes
| Plaintiff Type | Defendants | Core Claims | Status (2025) |
|---|---|---|---|
| News organizations (NYT, etc.) | OpenAI, Microsoft | Copyright infringement via scraping paywalled articles for training; outputs reproduce works | Pending in SDNY |
| Authors (Silverman, others) | OpenAI, Meta | Books scraped without permission; ChatGPT outputs infringe | Mixed motions to dismiss rulings |
| Visual artists | Stability AI, Midjourney, DeviantArt | Artwork scraped for Stable Diffusion training; outputs are derivative works | Partially survived motions to dismiss |
| Programmers (GitHub Copilot) | Microsoft, GitHub, OpenAI | Code scraped; Copilot outputs infringing code | Ongoing |
| Music publishers | Anthropic, others | Lyrics reproduced in outputs without license | Pending |
- Direct copyright infringement: Copying entire works into training datasets without authorization
- Derivative works: AI outputs are derivative works of training data
- DMCA §1202 violations: Removing or altering copyright management information (CMI) during scraping
- Misappropriation: Unfair competition, unjust enrichment for commercial use of creative works
- Privacy violations: Scraping and using personal data without consent (CCPA, other privacy laws)
- Terms of Service violations: Scraping paywalled or ToS-restricted content
- Fair use: Training is transformative; creates new outputs, doesn’t substitute for originals
- No substantial similarity: Outputs don’t reproduce training data (except in rare hallucination cases)
- Publicly available data: Scraping public internet content is lawful
- No market harm: AI tools complement rather than substitute for original works
- Licensing deals: Growing number of licensing agreements with publishers (e.g., OpenAI-News Corp, Google-publishers)
- Is training on copyrighted works fair use or infringement?
- Does scraping alone constitute infringement, or only if outputs reproduce?
- What level of similarity between output and training data is actionable?
- Can rights-holders opt out of AI training?
- Are DMCA circumvention claims viable (bypassing paywalls, robots.txt)?
| Rights-Holder Type | Strongest Claims | Evidence Needed |
|---|---|---|
| News publishers / journalists | Copyright infringement (articles); DMCA §1202 (CMI removal); ToS breach (paywall circumvention) | Articles in training datasets; paywall breach evidence; outputs reproducing articles |
| Book authors | Copyright infringement; derivative works | Books in datasets (e.g., Books3 corpus); AI outputs containing passages; registration certificates |
| Visual artists / photographers | Copyright infringement; derivative works; right of publicity (if person depicted) | Images in training sets (LAION, etc.); outputs mimicking style; similarity analysis |
| Musicians / composers | Copyright infringement (compositions, sound recordings) | Music in training data; outputs reproducing melodies/lyrics |
| Programmers / software developers | Copyright infringement (code); license violations (GPL, etc.) | Code repositories scraped; Copilot outputs containing copyrighted code |
| Individuals (privacy claims) | CCPA violations; misappropriation of likeness; privacy torts | Personal data/images in training sets; lack of consent; outputs using likeness |
Challenges for plaintiffs:
- Black box problem: Training datasets and model architectures often not publicly disclosed
- Discovery needed: Requires litigation to compel disclosure of training data sources
- Probabilistic outputs: Hard to prove specific output is “copy” vs. coincidental similarity
Evidence plaintiffs can gather:
- Public dataset disclosures: Common Crawl, LAION, Books3 have been documented; check if your works are included
- Prompting for reproduction: Test AI with prompts designed to elicit your copyrighted work (e.g., “Write article about [topic] in style of [Your Name]”)
- Metadata analysis: Some outputs contain artifacts suggesting training on specific sources
- Company statements: Public disclosures about training data sources
- Scraped content logs: Web server logs showing AI company bots scraping your site
| Damage Theory | Calculation | Challenges |
|---|---|---|
| Statutory damages (copyright) | $750–$30k per work ($150k willful) | Requires timely registration; “per work” definition unclear for massive datasets |
| Licensing fees | What you would have charged for AI training license | No established market rates yet; AI companies argue $0 (fair use) |
| Lost market value | AI outputs compete with your work, reducing sales/licensing | Hard to prove causation; substitution effect |
| Unjust enrichment | AI company’s profits attributable to using your work | Difficult to trace profits to specific training data |
AI training cases are natural class actions:
- Common questions: Did defendants scrape/train on works without permission? Is it fair use?
- Large classes: Millions of creators whose works were scraped
- Settlement leverage: Class certification creates existential risk for AI companies
- Opt-out rights: Class members can opt out to pursue individual claims if they have strong damages cases
- Individual vs. collective action: Consider joining existing class actions vs. individual demand
- Publicity: Public demands/lawsuits attract media attention, putting pressure on AI companies
- Licensing opportunity: Frame as “We’re open to licensing our content for AI training at fair rates”
- Discovery needs: Litigation may be necessary to uncover what training data was used
| Section | Content |
|---|---|
| Your works & ownership | Identify copyrighted works, registration status, commercial value, market position |
| Evidence of use in training | How you know your works were scraped/used (dataset disclosures, outputs, server logs) |
| Infringement theories | Copyright (copying for training), derivative works (outputs), DMCA §1202, ToS violations |
| Fair use rebuttal | Why training is NOT transformative; commercial use; market harm; non-substitution is false |
| Damages calculation | Statutory damages potential, licensing fees, lost market value |
| Demand | Cease using works in training; remove from datasets; destroy derivative models; licensing negotiation OR litigation |
| Deadline | 30–60 days (longer than typical IP demands given complexity) |
- Firm but business-oriented: “We recognize AI’s potential but demand fair compensation for our creative works”
- Open to licensing: “We’re willing to negotiate reasonable licensing terms for training use”
- Cite precedent: Reference licensing deals AI companies have made with other publishers
- Collective strength: If representing multiple creators, emphasize scale of infringement
For News Publishers:
- Emphasize paywall circumvention and ToS violations
- Reference NYT and other publisher litigation as precedent
- Highlight existing licensing deals (OpenAI-News Corp, etc.) as proof of market value
- DMCA §1202 claims for copyright management information removal
For Individual Creators:
- Consider joining class actions rather than individual demands (cost-effective)
- If pursuing individually, focus on works with clear commercial value and registration
- Document attempts to opt out (robots.txt, no-scraping notices)
For Software Developers:
- Open-source license violations (GPL requires attribution/sharing; Copilot doesn’t comply)
- Specific code snippets reproduced in outputs
- Loss of attribution and credit
Four-factor analysis:
| Factor | AI Company Argument | Rights-Holder Counter |
|---|---|---|
| 1. Purpose & character | Transformative: training creates new tool; outputs are new works, not copies | Commercial use; outputs compete with originals; no transformation of individual works |
| 2. Nature of work | Many training works are factual (news, code); less protection | Also includes highly creative works (fiction, art, music); core of copyright |
| 3. Amount used | Entire work needed for training; outputs use minimal amounts | Copied entire works; many outputs substantially reproduce training data |
| 4. Market effect | AI tools complement, don’t substitute; new markets created | Direct substitution; users get content without licensing; lost licensing revenue |
- No substantial similarity: Outputs don’t reproduce copyrighted expression; only rare “memorization” cases show copying
- No market harm: Plaintiffs can’t show lost sales/licenses caused by AI training
- Implied license: Publicly posting content online creates implied license for certain uses
- First sale doctrine: AI company lawfully acquired copies (e.g., purchased books) and can use for training
- Statute of limitations: Claims accrued when training occurred (3 years for copyright)
- Evaluate claim strength: Is plaintiff’s work actually in training data? Can they prove it?
- Discovery burden: Plaintiffs need litigation to compel disclosure of training data; expensive for individuals
- Settlement vs. licensing: Consider whether licensing deal is cheaper than litigation
- Collective approach: Industry-wide licensing standards emerging; individual deals may set precedent
- Policy advocacy: Support legislative solutions creating AI training exemptions or compulsory licensing
Growing trend toward negotiated licenses:
- Publisher deals: OpenAI-News Corp, Google-AP, etc. – typically $X million annually for training access
- Opt-in registries: Some platforms allow creators to register works for AI training for compensation
- Collective licensing: CMOs (collective management organizations) for music/publishing could administer AI licenses
- Statutory licensing: Possible future legislation creating compulsory licenses (like music mechanical licenses)
I represent content creators asserting rights against unauthorized AI training and AI companies defending against infringement claims. This emerging area requires understanding both copyright law and AI technology.
- Evaluate whether your works were used in AI training (dataset analysis, output testing)
- Draft demand letters and licensing proposals to AI companies
- Negotiate licensing agreements for authorized AI training use
- Join or initiate class action litigation
- File individual copyright infringement lawsuits when damages justify
- Pursue DMCA §1202 claims for CMI removal
- Assert ToS violations and breach of contract claims
- Assess fair use and other defenses to training-based infringement claims
- Respond to demand letters and evaluate settlement vs. litigation
- Negotiate licensing agreements with content owners
- Defend copyright infringement and class action lawsuits
- Advise on training data sourcing and documentation
- Implement opt-out mechanisms and respect robots.txt / no-scraping signals
- Develop industry-standard licensing frameworks
- News publisher claims against LLM developers
- Book author and artist class actions
- Software developer Copilot disputes
- Licensing negotiations for AI training data
- Fair use defense in AI training cases
- DMCA §1202 CMI removal claims
- Privacy-based claims (CCPA, BIPA) for facial recognition / personal data training
Book a call to discuss your AI training dispute. I’ll assess the strength of infringement or fair use claims, evaluate litigation vs. licensing options, and recommend strategy for resolution or defense.
Email: owner@terms.law