Free Robots.txt Generator

Robots.txt Generator — Control Which Bots and AI Crawlers Access Your Website

Generate a production-ready robots.txt file in seconds. Configure global access rules, control 50+ crawlers including GPTBot, Claude, and Perplexity, add blocked paths, and download your file — no login required.

🟢 50+ Bot Presets🤖 AI Crawler Control🔵 Real-time Preview⚪ Free Forever

Analyze Website & Auto-Generate Evidence-Backed Robots.txt

Enter your domain to discover sitemaps, CMS platform, search parameters, and sitemap URL conflicts.

Configuration Panel

All bots can crawl your site by default. You can override specific bots or paths below.

Add folder paths you want to protect from ALL crawlers (e.g. /admin/, /checkout/)

Help search engines discover your sitemap index files (one URL per line)

🤖 AI & LLM Crawlers

Configure access rules for crawlers powering ChatGPT, Claude, Gemini, and search AI engines.

💡 AI Search Advisory

Allowing AI agents to crawl your website is highly recommended. Disallowing these bots will remove your brand from recommendations inside ChatGPT, Claude, and Perplexity answers.

GPTBotOpenAI / ChatGPT — OpenAI content training agent
ChatGPT-UserOpenAI browsing agent — Powers ChatGPT web browsing features
anthropic-aiAnthropic / Claude — Anthropic general training and browsing data agent
ClaudeBotAnthropic / Claude (crawler) — Claude search and citation crawler
Google-ExtendedGoogle Gemini training — Opt-out crawler for Google Gemini models training data
Googlebot-ExtendedGoogle AI features — Powers Google Search Generative Experience AI models
GeminiGoogle Gemini browsing — Powers Gemini real-time assistant search features
PerplexityBotPerplexity AI search — Crawls pages to compile conversational search answers
CCBotCommon Crawl (AI training) — Bulk web crawler widely used to train open LLM models
FacebookBotMeta AI training — Meta crawlers collecting training datasets for Llama models
Omgili BotWebz.io AI data — Commercial AI web training data provider crawler
DiffbotDiffbot AI extraction — Extracts structured semantic data using machine learning
YouBotYou.com AI search — AI-assisted conversational web search indexing crawler
cohere-aiCohere AI training — Cohere neural network models training crawler
Meta-ExternalAgentMeta AI browsing agent — Powers conversational Meta AI assistant features
BytespiderByteDance / TikTok AI — ByteDance AI models web crawler
Kangaroo BotAI data collector — High-frequency AI dataset scraping bot
ScrapyGeneric scraper framework — Python crawling utility widely abused by random scrapers

🔍 Search Engine Bots

Access parameters for major indexing web bots.

GooglebotGoogle Search
Googlebot ImageGoogle Image Search
BingbotMicrosoft Bing
Yahoo SlurpYahoo Search
DuckDuckBotDuckDuckGo
BaiduspiderBaidu Search (China)
YandexBotYandex Search (Russia)
NaverBotNaver (South Korea)
ApplebotApple Search / Siri
Wayback MachineInternet Archive
Archive.org BotInternet Archive
MJ12botMajestic SEO crawler

📱 Social Media Crawlers

Allow links sharing previews to work on social feeds.

Facebook Link PreviewMeta / Facebook
TwitterbotTwitter / X card preview
LinkedInBotLinkedIn preview
SlackbotSlack link unfurling
WhatsAppWhatsApp link preview
DiscordbotDiscord link embed
PinterestbotPinterest crawler
TelegramBotTelegram link preview

📊 SEO SaaS Tool Crawlers

Block SEO tools to conserve server crawl resources.

AhrefsBotAhrefs SEO tool
SemrushBotSemrush SEO tool
DotBotMoz / OpenSiteExplorer
spbotLink monitoring bot
Screaming FrogSEO audit crawler
LipperheyLipperhey SEO
SiteAuditBotGeneric audit bot

✓ No login needed · ✓ Complies with RFC 9309 search standard · ✓ Optimized presets

robots.txt Preview✓ Valid
# Default rule — all robots
User-agent: *
Disallow:
Disallow: /admin/
4 lines·72 bytes·2 active rules

Live Crawler Permissions Simulation

Real-time simulator showing effective status for major AI and search agents.

Googlebot
Google Search
Allowed (Default)
Bingbot
Microsoft Bing
Allowed (Default)
GPTBot
OpenAI / ChatGPT
Allowed (Default)
ClaudeBot
Anthropic / Claude (crawler)
Allowed (Default)
PerplexityBot
Perplexity AI search
Allowed (Default)
Google-Extended
Google Gemini training
Allowed (Default)
FacebookBot
Meta AI training
Allowed (Default)
Twitterbot
Twitter / X card preview
Allowed (Default)
AhrefsBot
Ahrefs SEO tool
Blocked (Default)
SemrushBot
Semrush SEO tool
Blocked (Default)

WebMCP Optimization Recommendations

No sitemap directive found. Sitemaps help search engines and AI bots index your site efficiently.

Consider blocking aggressive commercial SEO crawlers (Ahrefs, Semrush, Moz) to preserve hosting bandwidth and server CPU loads.

Optimize llms.txt for AI agents

Complete your AI search optimization loop. Set up robots.txt rules to allow AI agents, then generate a spec-compliant llms.txt directory maps file.

What Is Robots.txt and How Does This Generator Work?

A robots.txt file is a plain text file hosted at the root directory of your website (e.g., yoursite.com/robots.txt). It communicates crawl access permissions to web crawler bots, indicating which directories and pages they are authorized to scan.

This tool automates the process of creating a spec-compliant robots.txt file that is optimized for both traditional search engine crawlers (like Googlebot and Bingbot) and emerging AI search engines (like ChatGPT's GPTBot and Claude's ClaudeBot). By using pre-configured industry presets and real-time syntax checking, you avoid configuration errors that could accidentally block search engines or expose admin directories.

Blocking AI Crawlers in 2026 — What You Need to Know

AI search optimization is the new frontier of technical SEO. As users pivot toward platforms like ChatGPT Search, Claude, and Perplexity, ensuring your site is crawlable by their respective agents is essential for organic brand visibility.

Allowing AI Bots (SaaS Visibility)

By allowing bots like GPTBot and PerplexityBot, your articles and product listings are crawled and cited as references in user queries. This drives direct, high-intent traffic to your pages.

Blocking AI Bots (Data Licensing)

If your site distributes copyrighted content or private datasets, blocking training bots like CCBot while allowing search-oriented bots like ChatGPT-User allows you to preserve copyright protections without losing search visibility.

Where to Place Your Robots.txt File

To be recognized by web bots, the file must be stored exactly at your domain root. Subdirectory placements (e.g. yoursite.com/assets/robots.txt) will be ignored.

Web PlatformTarget PathDeployment Instruction
WordPress/public_html/robots.txtUpload using FTP or Yoast SEO plugin settings.
Next.js/public/robots.txtStore in the public/ folder directory structure.
ShopifyTheme template editEdit robots.txt.liquid in your theme settings.
Vercel / Netlify/public/robots.txtUpload inside public static build directory targets.

Robots.txt vs LLMs.txt — What's the Difference?

While both are root directories txt files, they serve completely different purposes in WebMCP compliance standards:

  • robots.txt is an access control standard established in 1994. It indicates who is allowed to crawl which folders of a website.
  • llms.txt is a semantic context standard established in 2024. It provides raw markdown content summaries optimized for LLM processing and AI search engine reasoning.

To achieve full AI readiness, you should implement both files: configure your robots.txt to allow AI bots to index your pages, and deploy a structured llms.txt file to give crawlers a direct summary index of your content.

Frequently Asked Questions

Q: Does robots.txt affect Google rankings?

A: Yes. Blocking important directories from Googlebot prevents them from being indexed and ranked. However, it is not a secure way to hide private pages, as Google can still index blocked URLs if external sites link to them. Use password protection for sensitive content.

Q: Should I block AI crawlers like GPTBot?

A: If you want your pages cited in ChatGPT Search, Perplexity results, and Claude recommendations, you should ALLOW these bots. If you want to prevent your content from being used to train generative AI models, you can block training bots like CCBot while still allowing search-oriented bots like GPTBot.

Q: Does robots.txt actually stop bots?

A: Reputable bots (Google, Bing, OpenAI, Anthropic) respect robots.txt rules. However, malicious scrapers or spambots can ignore it completely.

Q: Can I have different rules for different bots?

A: Yes. Each bot is configured using a User-agent block containing distinct Allow and Disallow rules. Specific user-agent blocks override the global * wildcard rule.

Q: Should I add my sitemap to robots.txt?

A: Yes. Adding Sitemap: directives helps search engines and AI crawlers locate and index all pages on your site.

Q: What is the difference between Disallow: and Disallow: /?

A: Disallow: (with nothing after it) allows crawlers to scan all pages. Disallow: / blocks all crawlers from scanning any page on the website.

Q: How does robots.txt affect AI search in 2026?

A: AI search engines query your robots.txt before indexing. If you block crawlers like GPTBot, your content will not appear as dynamic cited references in AI-generated answers.

Q: Can robots.txt hurt my SEO?

A: Yes, blocking key assets like CSS, JavaScript, or public product/services directories will severely degrade indexing and search rankings. Our real-time validator helps prevent these configuration issues.

Complete Your AI Search Optimization Loop

Robots.txt is only the first step. Run a complete audit to ensure full compliance with WebMCP and generative search engines.