Twenty websites, two hundred decoys: a study shows who really reads your content
A research team planted marked content and then asked 22 AI systems whether they knew it. The result refutes two assumptions that appear in almost every bot discussion: that blocking after the fact helps, and that the bot doing the reading works for the same company that later answers.
We have written in this magazine several times that a user agent is a claim, not an identity document. Until now that was a reasoned inference from our own request path. There is now a measurement, and it is less comfortable than the inference.
Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu and Emily Wenger published Identifying AI Web Scrapers Using Canary Tokens in May 2026, revised in September 2026. The setup is simple enough that any shop could rebuild it, and that is exactly what makes it interesting.
The setup
The team put 20 websites online, each carrying 10 canary tokens. A canary token is a unique and otherwise meaningless piece of content: an invented product name, a numeric code, a sentence that exists nowhere else. The decisive part is that every visitor is served a different variant. Whoever fetches the page unknowingly carries away a serial number.
Then comes the second half, and that is the actual trick: afterwards you ask the AI systems about the topic of the test page. If one of those serial numbers turns up in the answer, it is proven which fetch is behind that answer. Not suspected. Proven.
Over the run time the test pages saw an average of 592 distinct combinations of user agent and network operator per page. For twenty pages with no audience, no advertising and no inbound links, that is a number worth pausing on. The machine share of traffic is no longer a footnote.
Finding 1: the reader is often not the provider
For 18 of the 22 systems tested, the team could determine the user agent. The expected mappings were there: ChatGPT via OAI-SearchBot, Copilot via Bingbot, Gemini via Googlebot, Perplexity via PerplexityBot.
The others are the remarkable part. 10 of the 18 systems relied on scrapers belonging to other search engines, without that relationship being publicly documented anywhere. The paper names Qwen and Perplexity among others, both returning content that had been collected via Googlebot.
And an entire group did not arrive as bots at all: ERNIE, Grok, Solar, Qwen and Kimi returned content that had been collected under ordinary browser identifiers, meaning Chrome, Safari, Firefox, Edge.
That has an immediate practical consequence. If you maintain your robots.txt or your firewall against a list of known AI bots, you are working with a list that does not describe the process. The fetch that later produces the answer looks in your log file like a person using Chrome.
Finding 2: blocking after the fact does not work
The team then blocked and measured what changed. Result: 12 systems kept serving the content under both blocking variants. Exactly one system stopped, Duck.ai.
The last point is sharper still: many chatbots were still returning the content of the test pages a week after those pages had been taken down. The pages no longer existed. The answers did.
This is where the mental model has to be corrected. A block is not a recall button. At best it governs what goes in from today onward. What was read has been read, sits in caches and in model weights, and comes back without your server ever being asked again.
What this means for a shop or a mid-sized company
Three things, and none of them is a block list.
First: access control is the wrong question. If ten of eighteen systems source their content through third parties, and five of those travel under normal browser identifiers, the idea that you can govern the inflow does not hold. The question that carries weight is not who may read, but what they find when they do. An incomplete product record is passed along exactly like a complete one, only with a worse outcome for you.
Second: you can measure this yourself, and cheaply. The study's setup is not laboratory equipment. One unambiguous expression that appears nowhere else, on a page of your own, a few weeks of patience, then the question put to the assistants. If the expression comes back, you know the page is being read and used, and you know it independently of any bot list. We consider this the most honest method available for checking your own visibility, because it does not rest on self-declaration.
Third: the window is narrower than you think. If content keeps being served for weeks after collection, every correction to your data takes effect with a delay, and so does every mistake. A wrong price, a stale availability, an incomplete product description: that then sits in answers you no longer have any access to. Anyone who tends their structured data only once they see a problem has already been shipping it for several weeks.
What the study does not say
Two limits, so the numbers are not made larger than they are.
The test pages were purpose-built and had no real audience. The 592 combinations per page therefore describe the baseline machine traffic of a new site, not the traffic mix of an established shop. On a page with actual visitors the ratio shifts.
And the attribution proves which fetch is behind an answer, not whether the content went into training or was retrieved at the moment of the question. For the practical conclusion that makes little difference. For the legal assessment it makes a great deal.
Where this leads
The debate about AI crawlers is usually run as a defence topic: how do I keep them out, how much does the traffic cost me, which list do I maintain. This study shows that the defensive half is the weaker one. Blocks barely work, identifiers are unreliable, caches outlive the shutdown.
What remains is the other half, and that one is yours to shape: making sure that what does get collected works in your favour. Complete product data, clean structured markup, answers to the questions customers actually ask. That is not a security measure. It is the only lever that demonstrably still belongs to you.
Source: Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger: Identifying AI Web Scrapers Using Canary Tokens, arXiv:2605.13706, submitted 13 May 2026, revised 3 September 2026. arxiv.org/abs/2605.13706
Related articles in this magazine: "A user agent is a claim: the bot compendium", "Somebody is using the names of the AI crawlers", "When bots work around the rules: why AI crawlers become a governance question"
Ask your question about Klariton.
Grounded in Klariton’s own knowledge, cited rather than invented.
Klariton is an Answer Engine Optimization platform. It adds to your existing shop or website without replacing anything or migrating data. Your own knowledge becomes verified, source-based answers (BIQs) that AI assistants and your visitors use directly.
Product copy, FAQs, spec sheets, your CMS or shop catalog. Klariton processes your material read-only and turns it into cited answers. Your data is never used for model training.
All AI calls run EU-hosted. Personal data is pseudonymized before every call. Klariton reads your sources read-only and stores no plain-text customer data.
Every answer is source-backed and passes a review gate. Nothing is published unchecked. Safe Guard continuously watches published answers against your material and flags outdated or weak claims.
No. Klariton integrates with your existing site, with no platform switch or migration. One embed is enough, and your CMS stays the source of truth.
Plans scale with touchpoints and answer volume. The AI visibility check is free and needs no sign-up. For everything else, just talk to us.
How visible is your brand to AI?
The free AI visibility check shows you in under a minute how AI assistants see your shop today.