If you've released music on Spotify or YouTube, there's a real chance it's already sitting inside a dataset that was used to train an AI music generator like Suno or Udio — not "AI" in the abstract, but your specific recordings, listed by name. The good news: many of those datasets are public, so you can actually check. This post covers what AI training data is, which datasets your catalog might be in, how to find out, and — honestly — what that does and doesn't mean. Checking your whole catalog against them is one of the things a SongBounty scan does for you.
What "AI training data" actually is
Generative music models don't invent sound from nothing. They learn patterns from enormous datasets of existing music — audio files, metadata, lyrics, even music videos. Many of the big training corpora were assembled from public platforms and are themselves public and searchable: The Atlantic even built an "AI Watchdog" tool that lets you look up whether a name shows up in the major ones.
This isn't hypothetical. In 2024 the major labels (via the RIAA) sued Suno and Udio, alleging they trained on copyrighted recordings without permission — a case that turns on exactly this question of what went into the training data. Whether or not you care about the lawsuit, one fact is useful to every artist: you can look up whether your music is in the pile.
The datasets your music might be in
A quick tour of the kinds of corpora that sweep up independent music. (These are illustrative — the full list is longer, and it keeps growing.)
Sleeping-Disco 9M
Millions of tracks with audio and metadata — the kind of large audio-text corpus that sits right next to Suno/Udio-style music-generation training.
LAION-Disco 12M
An open dataset of roughly 12 million music tracks assembled by LAION, the group behind several widely used open training sets.
YouTube-8M
Google's massive video dataset — millions of clips, including music videos, tagged and machine-readable for model training.
Books3 & Library Genesis
Text corpora that vacuum up lyrics, sheet music, and songbooks — the words and notation of your songs, not the audio.
How to check if your catalog is in them
You can do a version of this by hand. Public tools like The Atlantic's AI Watchdog let you search one dataset and one name at a time — but they search by artist name, cap how many results they show, and leave you to cross-reference your own catalog track by track. For anyone with more than a handful of releases, that's slow and easy to get wrong.
This is the part SongBounty automates. A scan pulls your full footprint in each dataset, matches each of your tracks to it by its Spotify or YouTube ID (or, failing that, its title), and shows you the exact matched entry — the title and the creators listed on it — as on-screen evidence. Instead of guessing, you get a per-song "here's where your music appears" view. (It's the same scan that audits your royalties across every pot — the AI check rides along for free.)
What it means — and what it doesn't
Here's the honest part, and we hold to it in the product:
- Appearing in a dataset is evidence, not proof. A work showing up in a training corpus doesn't prove any specific model was trained on it, or trained on it from that dataset. That's The Atlantic's own disclaimer, and it's ours too. We surface it as exposure, never as infringement.
- It is not a royalty you can invoice today. There's no button that turns "my song is in LAION-Disco" into a check. Anyone telling you otherwise is selling something.
So why bother? Because it's documentation, and documentation is leverage. AI licensing deals are already being negotiated, the Suno/Udio litigation is live, and opt-out ("do not train") mechanisms are appearing. If and when money or settlements flow from any of that, the artists who can show their work was in the data are the ones positioned to act. Knowing where you stand beats finding out later.
What you can actually do about it now
- Document your exposure. Save the evidence — which datasets, which of your tracks, the creators listed. A record today is leverage tomorrow.
- Use opt-outs where they exist. More platforms and datasets are adding "do not train" flags; take them where offered (with no guarantee of removal from copies already made).
- Keep your metadata clean. Your music can only be identified — and defended — if your ISRCs, titles, and credits are correct. That's the same metadata hygiene that keeps your royalties routed correctly.
- Watch the licensing and the lawsuits. The framework for paying artists for training use doesn't fully exist yet. When it arrives, be ready to claim your place in it.
See your AI exposure
A free SongBounty scan checks your catalog against these training datasets and shows you which of your tracks appear, with the evidence attached — as part of the same scan that audits what your music has earned across streaming, mechanical, performance, and radio.
To be clear about what we're handing you: this is exposure, not infringement; evidence, not proof; and leverage, not revenue. We surface where your music shows up and let you decide what to do with it — we never take a percentage, and we don't act on your behalf.
The AI was trained on somebody's music. Every artist deserves to know whether it was theirs.
FAQ
No public tool can prove that. What you can establish is whether your recordings appear in the public datasets associated with music-AI training. That's evidence of exposure, not proof a specific model used your work — an important distinction we keep in the product and you should keep too.
You can search public indexes like The Atlantic's AI Watchdog one dataset and one name at a time. SongBounty does it across your whole catalog at once, matching each track by its platform ID and showing the matched dataset entry as evidence — the free scan includes it.
Not today. There's no established mechanism that turns dataset inclusion into a payout, and we won't pretend otherwise. It matters as documentation and leverage for the licensing deals and litigation now taking shape — not as current revenue.
Sometimes partially. A growing number of platforms and datasets offer "do not train" opt-outs going forward, but no one can guarantee removal from copies that already exist. Documenting where you appear is the realistic first step.
Yes — the AI-exposure check is part of the standard artist scan, alongside the royalty audit. You get a per-song view of which datasets your catalog turns up in, with evidence, and we never take a cut of anything you pursue from it.