Google-Extended user agent

Operated by Google · Model training · Learning from you

robots.txt token
Google-Extended

Google-Extended is a robots.txt control token, not a crawler. No request ever arrives with Google-Extended in its User-Agent header, so grepping your access logs for it will always return nothing. Google reads the token from your robots.txt and applies it to content its ordinary crawler has already fetched.

Allow or block Google-Extended in robots.txt

Robots.txt applies the most specific matching group, so a group naming Google-Extended overrides your User-agent: * rules for this bot alone.

Allow
User-agent: Google-Extended
Allow: /
Block
User-agent: Google-Extended
Disallow: /

What blocking costs you: Won't train Gemini or feed AI Overviews. Does NOT remove you from Search — Googlebot is separate.

To see which AI crawlers your robots.txt allows right now, run your domain through the AI crawler access checker. To check whether AI answers actually cite you, use the AI citation checker.

Google's other crawlers

Blocking one Google bot does not block the others — each token is a separate group.

Google's own documentation: https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers#google-extended

FAQ

What is the Google-Extended user agent?
Google-Extended is not a crawler and sends no requests, so no User-Agent header ever contains it. Google reads the token "Google-Extended" from your robots.txt as an opt-out switch, and honours it using the pages its ordinary crawler already fetched. Searching your access logs for it will always come up empty.
Who operates Google-Extended and what does it collect?
Google-Extended is run by Google and collects pages as training data for a model. In the AI picture that puts it under "Learning from you": Feeding the models that power future answers.
How do I block Google-Extended?
Add a robots.txt group naming the token exactly: User-agent: Google-Extended followed by Disallow: /. Robots.txt matches the most specific group, so a Google-Extended group overrides whatever your User-agent: * group says. Won't train Gemini or feed AI Overviews. Does NOT remove you from Search — Googlebot is separate.
Should I block Google-Extended?
That depends on what you lose. Won't train Gemini or feed AI Overviews. Does NOT remove you from Search — Googlebot is separate. Sites chasing AI citations usually keep the live-answer and AI-search bots open and make a separate decision about the training crawlers. Run your domain through the AI crawler access checker to see which ones your robots.txt lets in today.
Can Google-Extended be spoofed?
Yes. A User-Agent header is self-reported, so any scraper can claim to be Google-Extended. The token is fine for reporting and for robots.txt, but if you are rate-limiting or firewalling on it, verify the request IP against Google's published ranges before you trust the name.

Paste a full User-Agent line and identify any crawler with the AI bot user agent list and lookup, or see every token side by side.