The Atlantic Built a Searchable Database of Music Used to Train AI Models

Atlantic reporter Alex Reisner has uncovered four datasets of music being used to train AI models and made them fully searchable for the public, raising fresh questions about how AI developers source their training audio.

Two of the datasets are particularly large, containing 12 million and 9 million tracks respectively. The other two are smaller but still substantial, each containing over 100,000 songs. Reisner reports the datasets have been downloaded thousands of times, and both Google and Stability have confirmed using them in published research papers.

Some of the music included comes from sources such as the Free Music Archive, which permits personal streaming but requires licensing for commercial use. That distinction matters because using the music as AI training data likely constitutes a commercial application.

Accessing the data is also not straightforward. According to Reisner, three of the four datasets are distributed as lists of links pointing to songs hosted on YouTube or Spotify, rather than as direct audio files. Developers then use automated tools to download the actual audio — tools that, in some cases, are designed to bypass logins, advertisements, and other platform mechanisms that would otherwise generate revenue or subscribers for creators. Reisner notes that using such tools violates the terms of service of those platforms.

The database was made public in June 2026. By making the datasets searchable, Reisner’s work may allow artists and rights holders to check whether their music appears in AI training data — a question that has become increasingly significant as legal and regulatory scrutiny of AI training practices continues to grow.

Source: The Verge

This article was generated by AI and cites original sources.
Scroll to Top