Try CrawlOut
All tools
Advance Robots.txt Generator

Control what crawls your site

Choose what search engines, AI crawlers and social bots can access — individually or in bulk — then export a ready-to-use robots.txt. Optionally import your site's existing file first and adjust it from there.

AI Crawlers
Named toggles for GPTBot, ClaudeBot & more.
Search Engines
Keep search bots on by default.
User-agent: GPTBot
Disallow: /

User-agent: *
Allow: /

Sitemap: yoursite.com/sitemap.xml
Granular Control
Allow or block each crawler individually.
AI Crawler Aware
Covers GPTBot, ClaudeBot and more by name.
Import Existing
Start from your current robots.txt, not scratch.
Instant Export
Copy or download a ready-to-use file.
Start from your site (optional)
AI training crawlers

Feed your content into an AI model's training data. Blocking these doesn't affect your search rankings.

GPTBot

OpenAI — crawls to train ChatGPT and other OpenAI models on your content.

ClaudeBot

Anthropic — crawls to train Claude models on your content.

Google-Extended

Controls use of your content for training Google's generative AI (Gemini), separate from normal Google Search indexing.

Applebot-Extended

Controls use of your content to train Apple Intelligence models, separate from normal Siri/Spotlight indexing.

Meta-ExternalAgent

Meta — crawls to train Meta AI models on your content.

CCBot

Common Crawl — a public dataset widely reused to train many different AI models, not just one company's.

Bytespider

ByteDance (TikTok) — crawls for ByteDance's AI and search products.

AI assistant crawlers

Fetch a page live only when a person asks an AI assistant to browse or search it — a different use than training.

ChatGPT-User

OpenAI — fetches a page live only when a person asks ChatGPT to browse or read it.

OAI-SearchBot

OpenAI — crawls to power ChatGPT's search results (separate from GPTBot training).

Claude-User

Anthropic — fetches a page live only when a person asks Claude to browse or read it.

Claude-SearchBot

Anthropic — crawls to power Claude's search features (separate from ClaudeBot training).

PerplexityBot

Perplexity — crawls to power its AI answer engine.

Amazonbot

Amazon — crawls for product discovery and Alexa answers.

Search engine crawlers

Blocking a major search engine here removes your site from that engine's results.

Googlebot

Google Search's main crawler — blocking this removes your site from Google Search entirely.

Bingbot

Microsoft Bing's main crawler.

DuckDuckBot

DuckDuckGo's crawler.

Slurp

Yahoo's crawler.

Baiduspider

Baidu Search's crawler — the dominant search engine in China.

YandexBot

Yandex Search's crawler — widely used in Russia.

Social preview crawlers

Generate the link preview card shown when your pages are shared on social platforms.

facebookexternalhit

Generates link preview cards when your pages are shared on Facebook or Instagram.

Twitterbot

Generates link preview cards when your pages are shared on X (Twitter).

LinkedInBot

Generates link preview cards when your pages are shared on LinkedIn.

Default rules — User-agent: *

Applies to every crawler not given its own rules above. Leave empty to allow everything.

Crawl-delay (seconds, optional)Most major bots ignore this now — Bing and some others still honor it.
Sitemaps
Live preview
# robots.txt generated with CrawlOut — 2026-09-24

User-agent: *
Allow: /
Learn more

Understanding robots.txt

A robots.txt file tells search engines and AI crawlers which parts of your site they may access — get it wrong and you can accidentally block the pages you most want indexed.

Read our guide