Microsoft Copilot SEO: how Copilot picks brands and sources
How Microsoft Copilot grounds answers in Bing, how it differs from ChatGPT search, and how Bing Webmaster Tools, IndexNow and AI Performance get you cited.
How to find GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot in your logs, verify their IPs, and read training versus live-fetch patterns.
AI crawlers show up in your server logs under their own user-agent tokens: GPTBot, OAI-SearchBot and ChatGPT-User for OpenAI, ClaudeBot, Claude-SearchBot and Claude-User for Anthropic, PerplexityBot and Perplexity-User for Perplexity. Each does one of three jobs: collecting training data, building an AI search index, or fetching a page live because a person asked an assistant a question. To read them, search your access or CDN logs for those tokens, verify the IPs against the ranges each vendor publishes, and look at the mix. User-triggered fetches are the closest thing you have to live evidence that assistants are reading your pages to answer real questions.
This post is about reading the logs. Whether to let these bots in at all is a separate decision, covered in should you block AI crawlers? - short version: for most brands, no.
The big three assistant vendors each run separate crawlers for separate jobs and document them. This is what their own pages say, as of October 2026:
| User agent | Vendor | Job | robots.txt |
|---|---|---|---|
GPTBot | OpenAI | Crawls content "that may be used in training" | Honored |
OAI-SearchBot | OpenAI | Surfaces sites in ChatGPT's search features | Honored; changes take ~24 hours |
ChatGPT-User | OpenAI | User actions in ChatGPT and Custom GPTs | "robots.txt rules may not apply" |
ClaudeBot | Anthropic | Collects web content for model training | Honored |
Claude-SearchBot | Anthropic | Improves Claude's search results | Honored |
Claude-User | Anthropic | Fetches pages when people ask Claude questions | Honored |
PerplexityBot | Perplexity | Surfaces and links sites in search results; "not used to crawl content for AI foundation models" | Honored; up to 24 hours |
Perplexity-User | Perplexity | Visits a page to answer a user's question | "Generally ignores robots.txt" |
OpenAI also runs OAI-AdsBot, which checks pages submitted as ChatGPT ads; you'll only see it if you advertise there. Beyond the assistant vendors, expect Meta's meta-externalagent, Amazon's Amazonbot, Common Crawl's CCBot and ByteDance's Bytespider. And remember the dual-use crawlers: Googlebot feeds both classic Google Search and Google's AI features, and Bingbot feeds Bing, which Microsoft Copilot draws on.
Notice the robots.txt column. The two "user" fetchers from OpenAI and Perplexity are explicitly treated as acting on behalf of a person, so their vendors say robots.txt may not stop them. Anthropic says all three of its bots honor it. That difference matters when you try to explain a hit you thought you had blocked.
The three jobs operate on different timescales, and they tell you different things when they show up.
By volume, training dominates. Cloudflare, which sees a large share of web traffic, reported that over the 12 months to August 2025, 80% of AI crawling was for training, 18% for search and 2% for user actions. The smallest slice is also growing fastest: Cloudflare's 2025 Year in Review found that AI "user action" crawling increased by more than 15x over the year.
The same Cloudflare analysis put a number on the imbalance between crawling and visits. In July 2025, for every visitor it referred to a site, Anthropic's crawlers made about 38,066 requests, OpenAI's about 1,091 and Perplexity's about 195. Those are network-wide ratios, not a forecast for your site, but they explain why a busy crawler line in your logs doesn't mean much traffic is coming back. To see the clicks, you need your analytics - see how to track AI traffic in GA4.
Two names on every list of AI crawlers will never show up in an access log, because they don't crawl. Google-Extended is, in Google's words, a standalone product token that controls whether content Google has already crawled may be used to train future Gemini models and to ground Gemini answers. The fetching itself is done with Google's existing user agents, and Google says the token doesn't affect inclusion or ranking in Google Search.
Apple's equivalent works the same way. Apple says Applebot-Extended "does not crawl webpages"; disallowing it opts your content out of training Apple's foundation models, while the regular Applebot keeps crawling for Siri, Spotlight and Safari. So if you search your logs for "Google-Extended" and find nothing, nothing is broken. Those tokens are switches you set in robots.txt, not visitors.
A related point about Google: AI Overviews and AI Mode are part of Search and draw on pages indexed for Google Search. There's no separate "AI Overviews bot" to look for - Googlebot is the one that matters.
Where the logs live depends on your stack. On your own server (nginx, Apache), the access log has every request, including the user agent. On a CDN or platform, raw logs usually need to be exported: Cloudflare offers Logpush, Vercel offers log drains, and most managed hosts have an equivalent. Your analytics tool won't help here - crawlers don't run your tracking script, and GA4 excludes known bot traffic anyway.
On a standard access log, one grep pulls out the AI lines and a second counts them per bot:
# every AI crawler request grep -Ei "gptbot|oai-searchbot|chatgpt-user|claudebot|claude-searchbot|claude-user|perplexitybot|perplexity-user" access.log # hits per bot, most active first grep -oEi "gptbot|oai-searchbot|chatgpt-user|claudebot|claude-searchbot|claude-user|perplexitybot|perplexity-user" access.log \ | tr 'A-Z' 'a-z' | sort | uniq -c | sort -rn
Then group by path and status code for the bots you care about. The most useful single view is user-triggered fetches by page: grep -i chatgpt-user access.log, pull the request path, count. Not sure what an unfamiliar user agent is? Paste it here:
Paste a user-agent string from your access log. It's matched in your browser against the same crawler list Zene's log ingest uses - nothing is sent anywhere.
ChatGPT-User
ChatGPT-User - the vendor says robots.txt may not apply to it, since a person triggered the fetch.A user agent is just a string, and anyone can send one. Google says it plainly: the HTTP user agent string can be spoofed. Scrapers routinely pretend to be GPTBot or Googlebot because some sites wave those through. Before you draw conclusions - or write firewall rules - check where the request came from.
The vendors make this straightforward by publishing their IP ranges as JSON files:
openai.com/gptbot.json, openai.com/searchbot.json and openai.com/chatgpt-user.json, linked from its crawler docs.claude.com/crawling/bots.json; it says a crawler whose source IP is on that list is coming from Anthropic.googlebot.com, google.com or googleusercontent.com, then run a forward lookup and confirm it returns the original IP.applebot.apple.com and publishes a CIDR list too.Matching an IP against a CIDR list is a one-line job in most languages, and it's worth automating. Download the files on a schedule rather than hard-coding them, since the ranges change. Some vendors publish no machine-readable list at all, so their hits stay unverifiable - count them, but treat them with suspicion.
Once the lines are verified, the patterns become useful. These are the ones we'd look at first:
/pricing means an assistant was answering a question that needed your pricing. The pages that collect these fetches are the pages that participate in answers - usually pricing, comparisons, docs and integration pages. They deserve the most care./llms.txt, that tells you more than any opinion - see does llms.txt work?Most fixes are small. Unblock the search and user-triggered bots if a CDN rule is catching them. Add redirects for the URLs assistants keep requesting and getting 404s on. Make the pages that attract user-triggered fetches easy to quote - plain-text prices, a clear one-sentence description, current facts. And after any robots.txt change, wait before you judge it: OpenAI says OAI-SearchBot can take about 24 hours to adjust, and Perplexity says up to 24 hours.
Then connect the layers. Crawler logs show what was read, analytics shows who clicked, and answer tracking shows whether you were named. A page that gets plenty of ChatGPT-User fetches but never appears in ChatGPT's recommendations is being read and passed over - a content problem, not an access problem.
Zene can take your crawler logs and do the classification and verification for you. You point a log drain at it - Vercel log drains, a Cloudflare Logpush HTTP destination, or any script that posts JSON lines - using a per-brand token. The ingest keeps only the AI crawler lines and checks each request IP against the vendors' published ranges in memory. It stores daily counts per bot, page and status class, plus how many hits passed verification - never the IPs themselves. The AI traffic page then shows the crawlers by purpose, next to the AI referral clicks and your visibility in the answers. Bots whose vendors publish no IP list are marked unverifiable rather than guessed.
Crawler access is the first step; being recommended is the goal. Run a free audit to see what ChatGPT and Gemini say about you today, or check whether your pages are ready to be read with the AI readiness checker.
They do three different jobs for OpenAI. GPTBot crawls content that may be used to train future models. OAI-SearchBot crawls to surface sites in ChatGPT's search features, so blocking it keeps you out of ChatGPT search citations. ChatGPT-User fetches a page live when a ChatGPT user's request needs it, and OpenAI says robots.txt rules may not apply to it because a person initiated the action.
Search your access log or exported CDN logs for the user-agent tokens - GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User - for example with grep -Ei. Then group the matches by bot, path and status code. Analytics tools like GA4 won't show them, because crawlers don't run your tracking script and GA4 excludes known bots.
Check the request IP against the ranges OpenAI publishes for each bot (gptbot.json, searchbot.json and chatgpt-user.json on openai.com). Anthropic and Perplexity publish similar JSON lists, and Google and Apple also support reverse-DNS checks. A user agent alone proves nothing - scrapers routinely impersonate well-known crawlers.
Because they don't crawl. Google-Extended is a robots.txt product token that controls whether content Google already crawled may be used for Gemini training and grounding; the fetching is done by Google's normal user agents. Apple says Applebot-Extended does not crawl webpages and only opts content out of training Apple's foundation models.
It depends on the vendor. OpenAI says robots.txt rules may not apply to ChatGPT-User and Perplexity says Perplexity-User generally ignores robots.txt, because both act on a person's request. Anthropic says its bots, including Claude-User, honor robots.txt. Training and search crawlers from all three vendors honor it.
Put this guide into practice - get your free visibility score in minutes.
