Bytespider user agent

Operated by ByteDance · Model training · Learning from you

robots.txt token
Bytespider
Full User-Agent string
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)

Bytespider identifies itself with the token Bytespider. That is the case-insensitive substring to grep for in access logs, and the exact name to put after User-agent: in robots.txt. It collects pages as training data for a model.

Allow or block Bytespider in robots.txt

Robots.txt applies the most specific matching group, so a group naming Bytespider overrides your User-agent: * rules for this bot alone.

Allow
User-agent: Bytespider
Allow: /
Block
User-agent: Bytespider
Disallow: /

What blocking costs you: ByteDance (TikTok) can't crawl you for training data; historically aggressive — most sites block it.

To see which AI crawlers your robots.txt allows right now, run your domain through the AI crawler access checker. To check whether AI answers actually cite you, use the AI citation checker.

FAQ

What is the Bytespider user agent?
Bytespider announces itself with the User-Agent token "Bytespider". That token is the case-insensitive substring to match in server logs and the exact name to put after "User-agent:" in robots.txt. The full line ByteDance publishes is: Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)
Who operates Bytespider and what does it collect?
Bytespider is run by ByteDance and collects pages as training data for a model. In the AI picture that puts it under "Learning from you": Feeding the models that power future answers.
How do I block Bytespider?
Add a robots.txt group naming the token exactly: User-agent: Bytespider followed by Disallow: /. Robots.txt matches the most specific group, so a Bytespider group overrides whatever your User-agent: * group says. ByteDance (TikTok) can't crawl you for training data; historically aggressive — most sites block it.
Should I block Bytespider?
That depends on what you lose. ByteDance (TikTok) can't crawl you for training data; historically aggressive — most sites block it. Sites chasing AI citations usually keep the live-answer and AI-search bots open and make a separate decision about the training crawlers. Run your domain through the AI crawler access checker to see which ones your robots.txt lets in today.
Can Bytespider be spoofed?
Yes. A User-Agent header is self-reported, so any scraper can claim to be Bytespider. The token is fine for reporting and for robots.txt, but if you are rate-limiting or firewalling on it, verify the request IP against ByteDance's published ranges before you trust the name.

Paste a full User-Agent line and identify any crawler with the AI bot user agent list and lookup, or see every token side by side.