You are currently viewing Amazon’s AI Training Data Scandal: Why CSAM Findings Are Raising Global Alarm

Amazon’s AI Training Data Scandal: Why CSAM Findings Are Raising Global Alarm

Amazon is facing intense scrutiny after reporting a high volume of child sexual abuse material (CSAM) found in Amazon AI training data, a revelation that has stunned child safety experts and raised fresh concerns about how tech giants source data in the race to build powerful AI models.

According to disclosures made to the National Center for Missing and Exploited Children (NCMEC), Amazon accounted for the vast majority of AI-related CSAM reports in 2025, submitting hundreds of thousands of flagged items. While the company says the material was removed before any AI model training occurred, critics argue that Amazon’s lack of detail about the content’s origin may be undermining law enforcement efforts to protect victims.

A Massive Spike in AI-Related CSAM Reports

NCMEC confirmed that AI-related CSAM reports jumped more than 15x year over year, with Amazon responsible for most of the increase. This marks a dramatic rise from just 67,000 AI-related reports across the tech industry in 2024 — and only 4,700 in 2023.

“This is really an outlier,” said Fallon McNulty, executive director of NCMEC’s CyberTipline. “Having such a high volume come in begs serious questions about where the data is coming from and what safeguards exist.”

How Amazon Says It Happened

Amazon says the CSAM was discovered while scanning externally sourced, non-proprietary AI training data, much of it scraped from the public web. The company relies on automated “hashing” tools that compare datasets against databases of known abuse material involving real victims.

An Amazon spokesperson emphasized that 99.97% of flagged content came from non-Amazon sources and insisted the company intentionally over-reported potential matches to avoid missing anything.

“We use an over-inclusive threshold for scanning, which results in a high rate of false positives,” the spokesperson said.

The Transparency Problem

While Amazon complied with legal reporting requirements, child safety officials say its reports often lacked actionable details, such as where the content originated or whether it remains accessible online.

“That makes the reports essentially inactionable,” McNulty said. “Without source data, law enforcement can’t identify perpetrators or rescue victims.”

Other major AI developers — including Google, OpenAI, Meta, and Anthropic — have also scanned training data for CSAM. But NCMEC says those companies submitted far fewer reports and typically provided more detailed information.

Why Amazon AI Training Data CSAM Is a Bigger Issue

Experts warn that training AI on illegal content, even unintentionally, poses serious risks:

  • Reinforcing harmful patterns in AI behavior
  • Enabling more realistic AI-generated abuse imagery
  • Re-circulating exploitative content and re-victimizing survivors

David Thiel, former chief technologist at Stanford Internet Observatory, says the rush to scale AI models has outpaced safety practices.

“When speed matters more than safeguards, errors are inevitable,” Thiel said. “Companies must be transparent about how data is sourced and cleaned.”

Industry Wake-Up Call

The Amazon AI training data CSAM findings highlight a growing problem across the AI industry: mass data ingestion without sufficient provenance controls. As companies scramble to outbuild competitors, child safety groups fear that safeguards are becoming an afterthought.

“There’s been a clear shift this past year,” said Thorn data scientist David Rust-Smith. “Companies are finally realizing that scraping the internet at scale means you will encounter CSAM.”

What Comes Next?

Amazon says it is committed to responsible AI and claims it is unaware of any instance where its models generated CSAM. Still, regulators and safety advocates argue that voluntary reporting isn’t enough — and that clearer standards for AI training data transparency are urgently needed.

As governments worldwide debate AI regulation, the Amazon AI training data CSAM controversy may become a defining case in how seriously the industry treats child protection in the age of generative AI.

Google Preferred Source

Don’t miss out on our latest news—follow us for the latest AI newsbreakthroughs, and insights that matter.

Leave a Reply