Feed your content into an AI model's training data. Blocking these doesn't affect your search rankings.
OpenAI — crawls to train ChatGPT and other OpenAI models on your content.
Anthropic — crawls to train Claude models on your content.
Controls use of your content for training Google's generative AI (Gemini), separate from normal Google Search indexing.
Controls use of your content to train Apple Intelligence models, separate from normal Siri/Spotlight indexing.
Meta — crawls to train Meta AI models on your content.
Common Crawl — a public dataset widely reused to train many different AI models, not just one company's.
ByteDance (TikTok) — crawls for ByteDance's AI and search products.
Fetch a page live only when a person asks an AI assistant to browse or search it — a different use than training.
OpenAI — fetches a page live only when a person asks ChatGPT to browse or read it.
OpenAI — crawls to power ChatGPT's search results (separate from GPTBot training).
Anthropic — fetches a page live only when a person asks Claude to browse or read it.
Anthropic — crawls to power Claude's search features (separate from ClaudeBot training).
Perplexity — crawls to power its AI answer engine.
Amazon — crawls for product discovery and Alexa answers.
Blocking a major search engine here removes your site from that engine's results.
Google Search's main crawler — blocking this removes your site from Google Search entirely.
Microsoft Bing's main crawler.
DuckDuckGo's crawler.
Yahoo's crawler.
Baidu Search's crawler — the dominant search engine in China.
Yandex Search's crawler — widely used in Russia.
Generate the link preview card shown when your pages are shared on social platforms.
Generates link preview cards when your pages are shared on Facebook or Instagram.
Generates link preview cards when your pages are shared on X (Twitter).
Generates link preview cards when your pages are shared on LinkedIn.
Applies to every crawler not given its own rules above. Leave empty to allow everything.
# robots.txt generated with CrawlOut — 2026-09-24 User-agent: * Allow: /
A robots.txt file tells search engines and AI crawlers which parts of your site they may access — get it wrong and you can accidentally block the pages you most want indexed.
Read our guide