CCBot user agent
Operated by Common Crawl · Model training · Learning from you
CCBotCCBot/2.0 (https://commoncrawl.org/faq/)CCBot identifies itself with the token CCBot. That is the case-insensitive substring to grep for in access logs, and the exact name to put after User-agent: in robots.txt. It collects pages as training data for a model.
Allow or block CCBot in robots.txt
Robots.txt applies the most specific matching group, so a group naming CCBot overrides your User-agent: * rules for this bot alone.
User-agent: CCBot Allow: /
User-agent: CCBot Disallow: /
What blocking costs you: Excluded from the Common Crawl corpus that feeds most open-source LLMs.
To see which AI crawlers your robots.txt allows right now, run your domain through the AI crawler access checker. To check whether AI answers actually cite you, use the AI citation checker.
Common Crawl's own documentation: https://commoncrawl.org/ccbot
FAQ
What is the CCBot user agent?
Who operates CCBot and what does it collect?
How do I block CCBot?
Should I block CCBot?
Can CCBot be spoofed?
Paste a full User-Agent line and identify any crawler with the AI bot user agent list and lookup, or see every token side by side.