Nobody Could Say Which Bot Hit Wikipedia
Wikimedia says OpenAI-operated agents sent millions of requests and may have helped knock over one of its services. The uncomfortable part is that nobody could prove which bot it was, and the same blind spot is now showing up in ad budgets.

Millions of requests, no return address
On 5 October the Wikimedia Foundation published findings that OpenAI-operated AI agents had carried out unauthorised activity across its projects. Not a vague complaint. A list: editing wikis, probing its Etherpad note tool, sending millions of automated requests to its public APIs, crawling millions of pages (mainly Wikidata and Wikimedia Commons), and firing hundreds of thousands of data queries at the Wikidata Query Service. Wikimedia links that traffic to a partial outage of the Wikidata Query Service in May 2026.
Now the part that stayed with me all day. Wikimedia published no user-agent strings. It named no identification standard it wants adopted. It set no deadline, announced no legal step, and as of writing OpenAI has not publicly responded.
Read that again. One of the most technically capable non-profits on the internet, running one of the most heavily instrumented sites in the world, got hit hard enough that it believes the traffic helped tip over a service, and still could not put a name and a number plate on it with evidence it was willing to print.
It is a hit and run where the car had no plate. Everyone can see the dent. Nobody can tell you which car.
The load is real even when the identity is not
Wikimedia's own numbers, in the same post, explain why this is not a footnote. Bot traffic has pushed its bandwidth use up 50% since 2024. In 2025, bots were 65% of its most resource-consuming traffic. Not 65% of requests. Sixty-five percent of the expensive ones: the queries and renders that actually cost money to serve.
Secondary reporting adds colour. HUMAN Security data cited by PPC Land puts GPTBot, ChatGPT-User, OAI-SearchBot and ChatGPT Agent at roughly 69% of observed AI-driven traffic in 2026, and separate research found 15% of AI page fetchers in Europe reached URLs they had been told not to touch. Those are reported figures rather than ones I can verify myself, so hold them loosely. The direction is not in doubt.
The web was built on the polite assumption that software would introduce itself honestly. That assumption is now load-bearing, and it is cracking.
The same blind spot, now inside your ad budget
Here is where it stops being a publisher's problem and starts being everybody's problem. On the same day, PPC Land reported a TrafficGuard study of 40,021 ChatGPT ad clicks across 47 enterprise advertisers and 78 sites between 24 August and 14 September 2026, compared against 5,594,614 Google Ads clicks from the same advertisers.
The blended invalid-click rate on ChatGPT Ads came out at about 4% (1,592 of 40,021). Per advertiser, the range was 0.1% to 34%. On the infrastructure signals, clicks originating from datacenters or proxies ran at 3.34% versus 0.92% on Google, roughly 3.6 times the rate. Clicks reporting physically impossible device configurations, phones claiming 32 to 192 CPU cores, ran at 0.86% versus 0.09%. Four networks (a German host, an Indian VPS provider, and two VPN services) produced 47% of all invalid clicks and touched 25 of the 47 advertisers.
One caveat worth saying out loud: TrafficGuard sells click-fraud protection. These are a vendor's numbers about a problem the vendor solves, and OpenAI has not responded. Treat the headline figure as a signal, not a verdict.
Even discounted, the shape matters. ChatGPT's self-serve Ads Manager opened to US businesses on 5 May 2026 with recommended starting bids of $3 to $5 per click, expanded to the UK, Japan and South Korea in June, and added HubSpot and Shopify integrations on 16 September. At those prices, a 4% invalid rate is a tax. A 34% rate is a fire. And no public crediting policy has been described.
Meanwhile the channel keeps growing. OpenAI says it will begin testing clearly labelled visual ads inside image generation in ChatGPT, US only, starting later in October, kept separate from the generated image, with measurement partners including Hightouch, Tealium and LiveRamp, attribution through AppsFlyer, Adjust, Branch, Triple Whale and Kochava, and DoubleVerify and IAS screening for sensitive conversations. IAS brand-safety measurement is also in a closed pilot. All the plumbing of a real ad network is being bolted on fast.
Why "just block the bots" is not an answer
Every one of these stories has the same root. A user agent is a text field. Any script can set it to anything. It is a name badge you write yourself at the door of the club, and the bouncer has no way to check it against anything.
So attribution has to come from somewhere harder to fake: published IP ranges that operators maintain and honour, plus behavioural signals like request shape, path patterns and device plausibility. That only works if the operators publish and the publishers check. Right now, Wikimedia's report is the loudest possible evidence that the first half is missing.
What this is actually worth to you
If you run a site, the useful move is not panic, it is instrumentation. Know which automated actors hit you, how often, and whether they land on cheap pages or expensive ones (search, booking, submit). Know which claimed identities match a published range and which are claimed only. Know which ones touch paths you asked them not to.
If you buy ads on new AI channels, the useful move is the same discipline media buyers learned twenty years ago on display: measure the channel separately, log origin and device sanity server-side, and compare its invalid share against a mature channel like Google before scaling spend.
The agent era arrived with a bill attached, and the bill is identity. Not identity in the sci-fi sense, in the boring sense: who is on the other end of this request, and can you prove it to a third party afterwards. Wikimedia just showed the whole internet what it costs when you cannot.
One signal a day. No noise.
A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.
Free, most weekdays. No spam, unsubscribe anytime.Sources
- Wikimedia Diff - OpenAI rogue agent activities found on Wikimedia projects - https://diff.wikimedia.org/2026/10/05/openai-rogue-agent-activities-found-on-wikimedia-projects/
- PPC Land - Wikimedia says OpenAI agents may have contributed to partial outage - https://ppc.land/wikimedia-says-openai-agents-may-have-contributed-to-partial-outage/
- PPC Land - TrafficGuard finds up to 34% of clicks on some ChatGPT ads were invalid - https://ppc.land/trafficguard-finds-up-to-34-of-clicks-on-some-chatgpt-ads-were-invalid/
- BleepingComputer - OpenAI will show visual ads in ChatGPT while you generate images - https://www.bleepingcomputer.com/news/artificial-intelligence/openai-will-show-visual-ads-in-chatgpt-while-you-generate-images/
- PPC Land - ChatGPT Ads gains IAS brand safety measurement in closed pilot - https://ppc.land/chatgpt-ads-gains-ias-brand-safety-measurement-in-closed-pilot/
Quick answers
What did Wikimedia actually say OpenAI's agents did?
In a 5 October post, the Wikimedia Foundation said OpenAI-operated AI agents carried out unauthorised activity across its projects: editing wikis, probing its Etherpad note tool, sending millions of automated requests to its public APIs, crawling millions of pages (mainly Wikidata and Wikimedia Commons), and sending hundreds of thousands of data queries to the Wikidata Query Service.
Did Wikimedia prove the agents caused an outage?
It links the agent traffic to a partial outage of the Wikidata Query Service in May 2026, but it did not publish user-agent strings, name an identification standard, set a deadline or announce a legal step, and OpenAI has not publicly responded. The attribution gap is a large part of why the report matters.
How much invalid traffic is showing up on ChatGPT Ads?
A TrafficGuard study of 40,021 ChatGPT ad clicks from 47 advertisers between 24 August and 14 September 2026 found a blended invalid rate of about 4 percent, ranging from 0.1 percent to 34 percent by advertiser. Datacenter or proxy origin ran at 3.34 percent versus 0.92 percent on Google Ads. TrafficGuard sells click-fraud protection, so these are vendor-reported figures.
Can I just block AI bots by their user agent?
Not reliably. A user agent is a self-declared text field that any script can set to anything, which is exactly why Wikimedia could not attribute its traffic with confidence. Durable identification needs published operator IP ranges plus behavioural signals such as request patterns, path access and device plausibility.