Activist Pirate Group Claims it Scraped Spotify’s Entire Music Library

news
Activist Pirate Group Claims it Scraped Spotify’s Entire Music Library

An activist pirate collective claims it scraped Spotify’s catalogue on a huge scale, raising fresh questions about copyright enforcement, scraping defenses, and AI training data. Spotify confirmed unauthorized scraping activity and says it disabled the accounts involved.

Anna’s Archive, a group known for running large shadow libraries, describes the project as a “preservation archive” for music. Rights holders and security researchers say the effort still involves copying and redistributing copyrighted works without permission.

What the group claims it collected

Anna’s Archive says it scraped metadata for about 256 million tracks, which it frames as near-complete coverage of Spotify’s public catalogue. Reports also cite Spotify’s own “100 million+ tracks” figure, which makes the group’s metadata claim difficult to reconcile without broader definitions of track entries.

The group also claims it downloaded roughly 86 million audio files. It argues that the subset represents about 37% of tracks but roughly 99.6% of total listening activity based on stream concentration in popular music.

Multiple reports put the total dataset size at about 300TB. Several outlets also note the archive may stop around July 2025 for the audio-file portion, which could leave out newer releases if the claim holds.

What has been released so far

Anna’s Archive has made the metadata searchable, while it has not publicly released the full set of audio files for people to download the Spotify songs. The group says it plans to distribute audio in stages through torrent-style sharing.

Spotify’s response

Spotify acknowledged unauthorized scraping and says it disabled the user accounts involved. The company says it added safeguards to reduce the chance of similar scraping at scale.

Spotify also says it has not seen evidence that private user data was compromised, since the scraping targeted publicly accessible parts of the service. Some reporting suggests the activity also involved tactics that bypassed technical protections on audio access, though Spotify has not shared full technical details.

Why this matters beyond piracy

If the audio archive circulates widely, it could fuel large-scale infringement across mirrors, torrents, and rehosting services. It could also become attractive as training material for AI systems, which would add another layer of legal and ethical disputes around copyrighted works.

The incident also puts pressure on streaming platforms to tighten anti-scraping controls, rate limits, and abuse detection. Those moves often collide with legitimate needs like research access, accessibility tooling, and third-party integrations.

What to watch next

Watch for signs of a broader audio release, coordinated takedowns, or court action tied to distribution infrastructure. Also watch for platform-side changes that restrict catalogue visibility, metadata access, and automated requests as Spotify and peers try to reduce scraping risk.

Discover: News

Discussion (0)

Be the first to comment.