GPTBot user agent

Operated by OpenAI · Model training · Learning from you

robots.txt token
GPTBot
Full User-Agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot

GPTBot identifies itself with the token GPTBot. That is the case-insensitive substring to grep for in access logs, and the exact name to put after User-agent: in robots.txt. It collects pages as training data for a model.

Allow or block GPTBot in robots.txt

Robots.txt applies the most specific matching group, so a group naming GPTBot overrides your User-agent: * rules for this bot alone.

Allow
User-agent: GPTBot
Allow: /
Block
User-agent: GPTBot
Disallow: /

What blocking costs you: Your content won't train future GPT models. ChatGPT browsing may still reach you via ChatGPT-User.

To see which AI crawlers your robots.txt allows right now, run your domain through the AI crawler access checker. To check whether AI answers actually cite you, use the AI citation checker.

OpenAI's other crawlers

Blocking one OpenAI bot does not block the others — each token is a separate group.

OpenAI's own documentation: https://platform.openai.com/docs/gptbot

FAQ

What is the GPTBot user agent?
GPTBot announces itself with the User-Agent token "GPTBot". That token is the case-insensitive substring to match in server logs and the exact name to put after "User-agent:" in robots.txt. The full line OpenAI publishes is: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot
Who operates GPTBot and what does it collect?
GPTBot is run by OpenAI and collects pages as training data for a model. In the AI picture that puts it under "Learning from you": Feeding the models that power future answers.
How do I block GPTBot?
Add a robots.txt group naming the token exactly: User-agent: GPTBot followed by Disallow: /. Robots.txt matches the most specific group, so a GPTBot group overrides whatever your User-agent: * group says. Your content won't train future GPT models. ChatGPT browsing may still reach you via ChatGPT-User.
Should I block GPTBot?
That depends on what you lose. Your content won't train future GPT models. ChatGPT browsing may still reach you via ChatGPT-User. Sites chasing AI citations usually keep the live-answer and AI-search bots open and make a separate decision about the training crawlers. Run your domain through the AI crawler access checker to see which ones your robots.txt lets in today.
Can GPTBot be spoofed?
Yes. A User-Agent header is self-reported, so any scraper can claim to be GPTBot. The token is fine for reporting and for robots.txt, but if you are rate-limiting or firewalling on it, verify the request IP against OpenAI's published ranges before you trust the name.

Paste a full User-Agent line and identify any crawler with the AI bot user agent list and lookup, or see every token side by side.