You are currently viewing Wikipedia AI Policy Calls on AI Companies to Stop Scraping and Use Its Paid API

Wikipedia AI Policy Calls on AI Companies to Stop Scraping and Use Its Paid API

Wikipedia is drawing a line in the digital sand. In a new statement, the Wikimedia Foundation — the nonprofit behind the world’s largest online encyclopedia — unveiled its Wikipedia AI policy, urging artificial intelligence companies to stop scraping its site and instead access its data through the Wikimedia Enterprise API, a paid product designed for large-scale use.

In its blog post, the foundation emphasized that AI developers should “use Wikipedia content responsibly” — meaning pay for structured access, attribute sources clearly, and respect the work of human contributors.

“For people to trust information shared on the internet, platforms should make it clear where the information is sourced from,” the organization wrote.

The new Wikipedia AI policy aims to protect the site’s servers, maintain content integrity, and ensure fair compensation for the volunteers and editors who have helped build one of the most trusted information repositories on the web.

Why Wikipedia’s AI Policy Matters

Generative AI systems — from ChatGPT and Gemini to Perplexity and Claude — rely heavily on Wikipedia’s open data for training and reference. But as AI companies scale up, Wikipedia’s infrastructure has taken a hit.

The Wikimedia Foundation revealed that its servers saw massive surges in fake “human” traffic earlier this year, which turned out to be AI bots scraping content while evading detection.

This artificial spike in traffic masked a worrying trend: human page views dropped 8% year-over-year, a decline that could threaten the site’s sustainability in the long run.

By enforcing its new Wikipedia AI policy, the foundation hopes to ensure that AI companies support the very ecosystem they rely on.

Enter the Wikimedia Enterprise API

The cornerstone of the new Wikipedia AI policy is the Wikimedia Enterprise platform — a paid API that gives AI companies structured access to Wikipedia’s vast database without straining its servers.

This product was created to make partnerships between AI developers and Wikipedia more sustainable. The organization notes that this opt-in, paid approach supports its nonprofit mission while still allowing access to high-quality data.

It’s worth noting that several major tech players, including Google and OpenAI, already have Enterprise partnerships with Wikimedia. The foundation hopes others will follow suit instead of relying on unsanctioned scraping.

Responsible AI Use and Attribution

Beyond the infrastructure issue, the Wikipedia AI policy also addresses a growing ethical concern: attribution.

Generative AI often produces content that draws heavily from Wikipedia articles — but without acknowledging its human contributors.

The Wikimedia Foundation insists that AI models must give credit where it’s due, saying:

“With fewer visits to Wikipedia, fewer volunteers may grow and enrich the content, and fewer donors may support this work.”

In other words, when AI companies strip out attribution, they’re not just reusing data — they’re eroding the community model that makes Wikipedia possible.

Wikipedia’s Own Use of AI

Interestingly, the Wikipedia AI policy doesn’t reject AI outright. Earlier this year, Wikimedia announced that it’s exploring AI to assist editors — but in a supportive, not replacement-based, role.

AI tools are being developed to automate translations, flag vandalism, and simplify tedious editing tasks, allowing human contributors to focus on improving content rather than repetitive maintenance.

This aligns with Wikimedia’s broader philosophy: AI should empower people, not replace them.

The Bigger Picture

Wikipedia’s move reflects a broader trend of content platforms demanding fair use from AI companies. As AI models grow more powerful — and more data-hungry — creators and organizations are pushing for clear rules, attribution standards, and compensation models.

The Wikipedia AI policy sends a clear message: open knowledge doesn’t mean free-for-all exploitation.

If AI systems continue to rely on Wikipedia’s vast human-generated content, they’ll now need to play by Wikimedia’s rules — or risk losing access altogether.

Goodle Preferred Source

Don’t miss out on our latest news—follow us for the latest AI newsbreakthroughs, and insights that matter.

Leave a Reply