ResearchResearch paperImage, Video & 3D Generation · Speech & Audio · Large Language Models1 source · Oct 6, 2026

Towards AI-Generated Music Plagiarism Detection as a Version Identification Problem

The rapid expansion of text-to-music generative models challenges traditional paradigms of music creation and intellectual property.

Key points

  • Plagiarism in this context is rarely an absolute mathematical binary, but an ambiguous threshold negotiated over harmonic structure, melodic contours, or overall perceived stylistic character.
  • In this work, we test the transferability of state-of-the-art music version identification architectures from the human-to-human cover domain to the human-to-AI plagiarism setting.
  • To evaluate this task, we introduce COPYCAT, a benchmark derived from real-world plagiarism cases and extended through generative re-synthesis and digital signal processing obfuscations, yielding 350,654 evaluation pairs.
  • We show that scalar distance thresholding collapses under generative re-synthesis, while a supervised framework leveraging coordinate-wise embedding shifts recovers the dispersed plagiarism signal, raising overall $F{0.5}$ from $0.612$ to $0.803$.

Sources (1)

  • [1]Towards AI-Generated Music Plagiarism Detection as a Version Identification Problem
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:18 PM
    The rapid expansion of text-to-music generative models challenges traditional paradigms of music creation and intellectual property.
    Plagiarism in this context is rarely an absolute mathematical binary, but an ambiguous threshold negotiated over harmonic structure, melodic contours, or overall perceived stylistic character.

Extractive summary: sentences quoted from the sources.

Related