AI, Music, & Money

Paying Music Creators for the Use of Their Works by Generative AI

Author’s Purpose

  • The Challenge: How to ensure ongoing payment for music creators when their existing music trains (or has been used to train) AI models?
  • The Goal:
    1. Explain the technology (GenAI) and current copyright law.
    2. Propose a new right for creators focused on fair compensation.

The Music Creator’s View (CIAM)

  • Support Existing Rights: All current copyright protections should apply to AI uses.
  • Propose NEW Right: An additional right of remuneration (payment) for the ongoing use of human works by AI platforms.
  • Demand Transparency: AI platforms must track and report which human works they use. Attribution is key for fair payment.

Executive Summary: The Problem

  • Generative AI challenges human creativity (music, art, writing).
  • Creators hone their craft over time, often needing to earn a living from it.
  • AI uses creators’ existing work (“datasets”) to generate new “content” that competes with them.
  • Aim: Find a way for creators to get paid when their work is used by AI, allowing them to retain agency.

Executive Summary: Proposed Solution Concept

  • Pay for Use: Creators should be paid when AI systems use the datasets (containing their work) to generate new output.
  • Mechanism: This payment should take the form of a license.
  • Requirement: To license something, there must be a right that can be licensed.

Two key stages with potential copyright implications:

  1. Training (“Input”): AI models are trained on vast datasets. This involves making copies (Text and Data Mining - TDM).
  2. Generation (“Output”): Users prompt the AI, which then produces new content based on its training.

Both stages can potentially infringe copyright.

Copyright owners generally have exclusive rights to:

  • (1) Reproduce the work (make copies)
  • (2) Prepare Derivative Works (adaptations, arrangements, translations)
  • (3) Distribute copies
  • (4) Perform the work publicly (for music, drama, etc.)
  • (5) Display the work publicly
  • (6) Perform sound recordings publicly via digital transmission

See: 17 U.S.C. § 106

  • Exceptions (e.g., Fair Use): In the US, “Fair Use” might allow limited copying for purposes like criticism, commentary, teaching, or research. Its application to AI training is highly debated and uncertain.
  • “Style” or “Sound”: Copyright generally protects the specific expression (melody, lyrics, recorded sounds), not the underlying idea, genre, style, or a general “sound” or voice.

Discussion Question 1

  1. Is it fair for AI companies to train models on vast amounts of copyrighted music without asking for permission or paying creators? Why or why not?
  2. Should an AI be allowed to generate music that perfectly mimics a specific artist’s unique “sound” or vocal style, even if it doesn’t copy their exact songs?

Part I: How Generative AI Works (Simplified)

  • Generative AI (GenAI): AI that creates new content (text, images, music).
    • Examples: ChatGPT (text), DALL-E (image), various music AI tools.
  • Models:
    • LLMs (Large Language Models): Trained on text (e.g., GPT).
    • Diffusion Models: Often used for images/video.
  • Training: Requires massive computing power and huge datasets. Often dominated by large tech companies.

Part I: The Training Process - Copying & Tokenizing

  1. Copying: Data (music files, text, images) is typically copied locally for efficient training.
  2. Tokenizing: The copied data is broken down into smaller pieces called “tokens”.
    • For text: tokens might be words, parts of words, or characters.
    • For music: tokens might represent notes, chords, timings, or audio features.
  3. Relationships: The AI learns patterns and relationships between these tokens.

Part I: Training Misconception

  • Myth: Training destroys the original work, leaving only abstract learning.
  • Reality: The process preserves representations (embeddings/vectors) of the original data, including relationships between elements.
    • This allows the AI to reproduce parts of the training data or generate statistically similar content.
    • The tokenized dataset itself can be seen as another reproduction.

Part I: How AI Generates Output

  • Prediction Machine: Based on a prompt (user request), the AI uses its training (the learned relationships between tokens) to predict the most likely next token (word, note, pixel).
  • It does this repeatedly to generate a sequence (sentence, melody, image).
  • Copying (Reproduction):
    • Initial copying for training.
    • The tokenized dataset itself.
    • Outputs that are substantially similar to training data.
  • Rights Management Information (RMI): Metadata (author, title, etc.) is often stripped during training, potentially violating laws like the DMCA in the US.
  • AI Company Indemnification: Offers from companies (like Google, Microsoft, OpenAI) to protect users from copyright lawsuits often have significant limitations and exclusions.

Part I: Key Takeaways Summary

  1. Training LLMs involves copying data.
  2. Copying copyrighted data infringes rights unless an exception (like fair use) applies (highly debated).
  3. Tokenization also creates a reproduction, preserving aspects of the original works.
  4. AI outputs can infringe if substantially similar or derivative of training data.
  1. Training may involve illegal removal of RMI (metadata).
  2. AI company indemnification offers have significant limits.
  3. Any legal exception for TDM must likely pass the international “three-step test” (balancing creator rights and public interest).

Discussion Question 2

  • Does understanding how LLMs work (copying, tokenizing, preserving relationships, predicting) change your view on whether using copyrighted music for training is fair or constitutes infringement?
  • How is AI “learning” different from human learning in a way that matters for copyright?
  • Copyright law isn’t static; it evolves with technology:
    • Player Pianos -> Mechanical reproduction rights
    • Radio/TV -> Broadcasting rights
    • Cinema -> Film recognized as copyrighted work
    • Internet -> Digital transmission / “Making available” rights (WCT/WPPT 1996)
  • New rights (exclusive or remuneration) were often created to ensure creators were compensated for new major uses.

Part II: Why Adapt Now for AI?

  • AI is a profound technological change.
  • It uses creators’ past work to generate competing content.
  • This threatens creators’ ability to earn a living, potentially shrinking future creativity.
  • Delegating human interpretation/creation to machines has deep cultural implications.
  • Goal: Not to stop AI, but to ensure human creators can coexist and thrive alongside it, using the established system of copyright.

Part II: The Proposal - A Right to Remuneration

  • Create a NEW right specifically for creators.
  • What it covers: Payment (remuneration) for the use of copyrighted material within an AI model when that model is used to generate competing content.
  • Distinction: This focuses on the output phase’s reliance on the training data, not just the initial input copying (which is covered by existing reproduction rights).

Part II: Advantages of This Proposed Right

  • AI training can continue largely unhindered.
  • AI companies pay for a key, valuable input (the creative data).
  • Exempts non-commercial research uses (e.g., universities).
  • Compensates creators when AI uses their life’s work to compete with them.
  • Applies to: Copyright-protected musical works (likely compositions first).
  • Initial Owner: The human creator(s).
  • Transferable: Creators could assign or license this right (e.g., to publishers, CMOs).
  • Normative Basis: Empowers creators, giving them agency when their work fuels competing AI output.
  • If “Sui Generis” (Unique Right): Countries could choose reciprocity (only pay creators from countries with a similar right). This incentivizes other countries to adopt it.
  • If Part of Copyright: Subject to national treatment (must treat foreign creators the same as domestic ones, under treaties like Berne/TRIPS).
  • Remuneration Right: Can be structured similar to existing compulsory/statutory licenses (use allowed if payment made), which is generally permissible under international law.
  • Better than a Levy: A simple levy (like on blank tapes/hard drives) based on input data size doesn’t reflect actual use or market impact of the AI output. The proposed right connects payment more directly to the AI’s productive use.

Part II: Practicalities - The Distribution Challenge

  • Common Argument Against: How can we possibly track which works were used by the AI to generate a specific output and distribute the money fairly? It’s too complicated!

Part II: Practicalities - Potential Solutions

  • Transparency is Key: AI platforms can be designed to identify source material or influences (though they may resist). Requires legal obligation.
    • EU AI Act mandates some training data disclosure.
  • CMOs: Collective Management Organizations already distribute royalties based on complex usage data/proxies. They have experience.
  • Usage Data: Ideally, track how often tokenized works are “pulled” by the AI during generation.
  • Proxies: If direct tracking is impossible, use proxies (like commercial success of source works, genre representation in dataset).
  • Focus: Ensure money reaches individual creators.

Discussion Question 3

  • How could we realistically track AI’s use of specific songs or musical elements to pay creators fairly?
  • What are the pros and cons of relying on Collective Management Organizations (CMOs) versus requiring AI companies themselves to provide detailed usage data for distribution?

see: CMOs - Collective Management Organizations - CLIP

Conclusion

  • Generative AI poses a significant challenge and opportunity for music creators.
  • Existing copyright law offers partial solutions but may not be sufficient for ongoing, fair compensation.
  • A new right of remuneration, focused on payment when AI uses training data to generate competing content, is proposed.
  • This requires transparency from AI platforms and robust distribution mechanisms.
  • The goal is to adapt copyright to balance AI innovation with the sustainable livelihoods of human creators.