Method, disclosed

Which bots we detect, and how we verify them

A reference table of every bot we detect: which provider, which purpose, which verification method. Publicly accessible, no login. The table is evidence for our method and at the same time a page that can be cited itself.

  • 18providers listed
  • 46individual bot identifiers
  • 7verified on every request
  • 4route published, not yet wired up
  • 7without a route, reason given

Detected is not verified

  1. Stage 1

    Detected, by user agent

    Every bot names itself in the user-agent header. That is how we detect it, and a bot showing up for the first time tomorrow is in the data immediately. But anyone can write a user agent into a header. A detected fetch is a claim first of all.

  2. Stage 2

    Verified, via the sender address

    Where the provider publishes a route, we check every fetch against it: against their address list, or via reverse DNS with forward confirmation. Only then does the fetch count as confirmed rather than merely claimed.

"Not verifiable" therefore does not mean "invisible". We see the bot, we count it, we attribute it to the provider its user agent names. What is missing is the proof, and that is owed by the provider, not by the bot. Such fetches are carried as detected and unverified, kept apart from the confirmed ones.

That also settles what happens with new bots: they land in stage 1 the moment they first knock, and move to stage 2 as soon as their provider publishes a route. The catalogue grows by itself. The verification does not.

ProviderBot and purposeVerification methodState
OpenAI
  • oai-searchbotansweringsearch indexMozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
  • oai-adsbotopenad review
  • chatgpt-useransweringlive fetchMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
  • gptbottrainingtrainingMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.4; +https://openai.com/gptbot)

Abgleich gegen die vom Anbieter veröffentlichte Adressliste, bei jedem Abruf.

Checked against: openai.com/gptbot.json, openai.com/searchbot.json, openai.com/chatgpt-user.json, openai.com/adsbot.json

active
Anthropic
  • claude-searchbotansweringsearch indexMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Claude-SearchBot/1.0; +mailto:support@anthropic.com
  • claude-useransweringlive fetchMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +claude-user@anthropic.com)
  • claude-webansweringlive fetchMozilla/5.0 (compatible; Claude-Web/1.0; +https://www.anthropic.com)
  • anthropic-aitrainingtraininganthropic-ai
  • claudebottrainingtrainingMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)

Abgleich gegen die vom Anbieter veröffentlichte Adressliste, bei jedem Abruf.

Checked against: claude.com/crawling/bots.json

active
Google
  • google-cloudvertexbottrainingtrainingMozilla/5.0 (compatible; Google-CloudVertexBot; +https://cloud.google.com/vertex-ai-bot)
  • google-gemininotebookansweringlive fetch
  • google-notebooklmansweringlive fetchGoogle-NotebookLM
  • google-agentansweringagent
  • google-extendedopentrainingMozilla/5.0 (compatible; Google-Extended/1.0; +http://www.google.com/bot.html)
  • googleothertrainingtrainingGoogleOther
  • googlebotopensearch indexMozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/150.0.7871.186 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

Googlebot zählt bei uns als Suchmaschinen-Crawler und nicht als KI-Abruf. Der Zweck steht auf offen, weil derselbe Crawl die klassische Suche und die KI-Übersichten bedient und am User-Agent nicht zu erkennen ist, wofür er diesmal lief.

Rückwärts-DNS mit Vorwärtsbestätigung: Der Name zur Adresse muss dem Anbieter gehören und wieder auf dieselbe Adresse zeigen.

Permitted hostnames: *.googlebot.com, *.google.com, *.googleusercontent.com

active
Perplexity
  • perplexitybottrainingsearch indexMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
  • perplexity-useransweringlive fetchMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user

Abgleich gegen die vom Anbieter veröffentlichte Adressliste, bei jedem Abruf.

Checked against: www.perplexity.ai/perplexitybot.json, www.perplexity.ai/perplexity-user.json

active
Microsoft
  • microsoft-aiansweringlive fetch
  • bingbottrainingsearch indexMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/116.0.1938.76 Safari/537.36

Copilot-Verweise weisen wir über den Referrer nach, nicht über einen eigenen Bot.

Rückwärts-DNS mit Vorwärtsbestätigung: Der Name zur Adresse muss dem Anbieter gehören und wieder auf dieselbe Adresse zeigen.

Permitted hostnames: *.search.msn.com

active
Meta
  • meta-externalfetcheransweringlive fetch
  • meta-webindexeransweringsearch index
  • facebookexternalhitopenlive fetch
  • meta-externalagenttrainingtrainingMozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36 (compatible; meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler))
  • facebookbottrainingtraining

Kein prüfbarer Weg vorhanden.

Meta veröffentlicht keine Datei, sondern verweist auf die Netzkennung AS32934, abzufragen per whois über TCP-Port 43. Ein Cloudflare-Worker spricht nur HTTP. Über einen Dritt-Dienst zu gehen hieße, die Verifikation von jemandem abhängig zu machen, den weder wir noch Meta kontrollieren.

not verifiable
Apple
  • applebot-extendedopentrainingMozilla/5.0 (compatible; Applebot-Extended/0.1; +http://www.apple.com/go/applebot)
  • applebotopensearch indexMozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)

Rückwärts-DNS mit Vorwärtsbestätigung: Der Name zur Adresse muss dem Anbieter gehören und wieder auf dieselbe Adresse zeigen.

Permitted hostnames: *.applebot.apple.com

active
Amazon
  • amazonbottrainingsearch indexMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) Chrome/119.0.6045.214 Safari/537.36

Kein prüfbarer Weg vorhanden.

Es gibt keinen maschinenlesbaren Endpunkt. Die Liste steht nur als Codeblock innerhalb einer HTML-Seite, ohne CIDR-Angaben und mit uneinheitlichen Feldnamen. Sie aus dem Seitenlayout zu kratzen hieße, dass die Prüfung beim nächsten Redesign still ausfällt.

not verifiable
Mistral
  • mistralai-useransweringlive fetchMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; MistralAI-User/1.0
  • mistralai-indexansweringsearch index
  • mistralai-trainingtrainingtraining

Der Anbieter veroeffentlicht einen Pruefweg, bei uns ist er noch nicht angeschlossen. Abrufe dieses Anbieters fuehren wir deshalb als erkannt und unverifiziert, nicht als bestaetigt.

Checked against: mistral.ai/mistralai-user-ips.json, mistral.ai/mistralai-index-ips.json

not verifiable
DuckDuckGo
  • duckassistbotansweringlive fetchDuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)

Der Anbieter veroeffentlicht einen Pruefweg, bei uns ist er noch nicht angeschlossen. Abrufe dieses Anbieters fuehren wir deshalb als erkannt und unverifiziert, nicht als bestaetigt.

Checked against: duckduckgo.com/duckassistbot.json

not verifiable
You.com
  • youbottrainingtrainingMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; YouBot/1.0; +https://docs.you.com/youbot; env:prod) Chrome/142.0.0.0 Safari/537.36

Der Anbieter veroeffentlicht einen Pruefweg, bei uns ist er noch nicht angeschlossen. Abrufe dieses Anbieters fuehren wir deshalb als erkannt und unverifiziert, nicht als bestaetigt.

Permitted hostnames: *.search.you.com

not verifiable
Common Crawl
  • ccbottrainingarchiveCCBot/2.0 (https://commoncrawl.org/faq/)

Kein Assistent, sondern ein Archiv, aus dem viele Modelle ihre Trainingsdaten beziehen.

Abgleich gegen die vom Anbieter veröffentlichte Adressliste, bei jedem Abruf.

Checked against: index.commoncrawl.org/ccbot.json

active
ByteDance
  • bytespidertrainingtrainingMozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; https://zhanzhang.toutiao.com/)

Kein prüfbarer Weg vorhanden.

Keine belastbare Primärquelle auffindbar.

not verifiable
Huawei
  • pangubottrainingtraining

Der Anbieter veroeffentlicht einen Pruefweg, bei uns ist er noch nicht angeschlossen. Abrufe dieses Anbieters fuehren wir deshalb als erkannt und unverifiziert, nicht als bestaetigt.

Permitted hostnames: *.petalsearch.com

not verifiable
xAI
  • grokansweringlive fetch

Kein prüfbarer Weg vorhanden.

Verifikationsweg noch nicht untersucht.

not verifiable
Cohere
  • cohere-training-data-crawlertrainingtraining
  • cohere-aitrainingtrainingMozilla/5.0 (compatible; cohere-ai/1.0; +https://cohere.com)

Kein prüfbarer Weg vorhanden.

Keine belastbare Primärquelle auffindbar. Cohere gibt an, derzeit keine Trainings-Crawler zu betreiben, und rät von IP-Sperren ab.

not verifiable
Baidu
  • baiduspider-renderopensearch indexMozilla/5.0 (Linux; u; Android 4.2.2; zh-cn;) AppleWebKit/534.46 (KHTML, like Gecko) Version/5.1 Mobile Safari/10600.6.3 (compatible; Baiduspider-render/2.0; +http://www.baidu.com/search/spider.html)
  • baiduspideropensearch indexMozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)

Kein prüfbarer Weg vorhanden.

Suchmaschinen-Crawler ohne KI-Zweck. Wir führen ihn, damit er nicht unter „unbekannt" verschwindet und damit er in keiner KI-Zahl mitzählt. Einen von Baidu veröffentlichten Prüfweg haben wir nicht gefunden.

not verifiable
Yandex
  • yandexbotopensearch indexMozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)

Kein prüfbarer Weg vorhanden.

Suchmaschinen-Crawler ohne KI-Zweck. Erkannt wird ausdrücklich „yandexbot" und nicht „yandex": der Yandex-BROWSER trägt den Namen ebenfalls im User-Agent, und ein zu weites Muster würde echte Besucher zu Bots erklären.

not verifiable
SEO- und Analyse-Werkzeuge
  • ahrefsbotopen
  • semrushbotopen
  • dataforseobotopen
  • serankingbacklinksbotopen

Kein prüfbarer Weg vorhanden.

Diese Werkzeuge vermitteln keine KI-Antworten. Wir erkennen sie als Bots, ordnen sie aber bewusst keinem Anbieter zu und zählen sie in keiner Reichweiten-Zahl mit.

not verifiable

"Active" means every fetch is checked; the result per address is cached for at most seven days. "Not verifiable" means there is no route we can take, and the reason is stated next to it. The full identifiers below are the ones these bots used when they reached us, observed between 7 July and 11 August 2026. Where nothing is shown, that identifier has never visited us, and we do not copy one in from elsewhere.

Why this table is public

Anyone publishing numbers about AI visibility should say how they count. Detection that only reads a user agent can be faked by anyone who writes the matching text into a header. So this page states not only which bots we know, but which fetch we actually believe and which one only until proven otherwise.

The gaps are listed too, and that is deliberate. A provider without a verification route is not missing here, it is listed with its reason. A table that only shows what works is an advertisement.

Why Google is not checked against an address list

Google publishes an address list, but only for its user-triggered agents, not for Googlebot and not for Google-Extended. We still do not use it, and the reason belongs on this page: as soon as any list exists for a provider, our check runs against that list before reverse DNS. A regular Googlebot fetch would then be checked against a list it has no business being in, and would come out as a forgery.

A wrong "failed" costs more than a missing "passed". One accuses a genuine request of deception, the other only says we do not know. So Google stays on reverse DNS, which covers every Google crawler. The address list becomes usable once we resolve per bot identifier instead of per provider, and that is a rebuild, not an entry.

What we do with the addresses

The check needs the sender address, otherwise there would be nothing to check. It is not stored. What sits in our cache for at most seven days is the result of the check, keyed by a hash (SHA-256) of the address, with no link to a person, a session or a customer.

And a caveat you rarely see stated: a hash of an IPv4 address without a secret is not anonymisation. There are only about four billion possible inputs, and they can be tried one by one. Our calculation therefore includes a secret that makes the value unresolvable for anyone who does not hold it. It stays a pseudonym nonetheless, because we hold that secret ourselves. So we call it pseudonymised rather than anonymous, even though that sounds weaker.