CCBot user agent

Operated by Common Crawl · Model training · Learning from you

robots.txt token
CCBot
Full User-Agent string
CCBot/2.0 (https://commoncrawl.org/faq/)

CCBot identifies itself with the token CCBot. That is the case-insensitive substring to grep for in access logs, and the exact name to put after User-agent: in robots.txt. It collects pages as training data for a model.

Allow or block CCBot in robots.txt

Robots.txt applies the most specific matching group, so a group naming CCBot overrides your User-agent: * rules for this bot alone.

Allow
User-agent: CCBot
Allow: /
Block
User-agent: CCBot
Disallow: /

What blocking costs you: Excluded from the Common Crawl corpus that feeds most open-source LLMs.

To see which AI crawlers your robots.txt allows right now, run your domain through the AI crawler access checker. To check whether AI answers actually cite you, use the AI citation checker.

Common Crawl's own documentation: https://commoncrawl.org/ccbot

FAQ

What is the CCBot user agent?
CCBot announces itself with the User-Agent token "CCBot". That token is the case-insensitive substring to match in server logs and the exact name to put after "User-agent:" in robots.txt. The full line Common Crawl publishes is: CCBot/2.0 (https://commoncrawl.org/faq/)
Who operates CCBot and what does it collect?
CCBot is run by Common Crawl and collects pages as training data for a model. In the AI picture that puts it under "Learning from you": Feeding the models that power future answers.
How do I block CCBot?
Add a robots.txt group naming the token exactly: User-agent: CCBot followed by Disallow: /. Robots.txt matches the most specific group, so a CCBot group overrides whatever your User-agent: * group says. Excluded from the Common Crawl corpus that feeds most open-source LLMs.
Should I block CCBot?
That depends on what you lose. Excluded from the Common Crawl corpus that feeds most open-source LLMs. Sites chasing AI citations usually keep the live-answer and AI-search bots open and make a separate decision about the training crawlers. Run your domain through the AI crawler access checker to see which ones your robots.txt lets in today.
Can CCBot be spoofed?
Yes. A User-Agent header is self-reported, so any scraper can claim to be CCBot. The token is fine for reporting and for robots.txt, but if you are rate-limiting or firewalling on it, verify the request IP against Common Crawl's published ranges before you trust the name.

Paste a full User-Agent line and identify any crawler with the AI bot user agent list and lookup, or see every token side by side.