Free tool
Robots.txt generator
Build a robots.txt file with full control over which crawlers can access your site — including the AI crawlers that power ChatGPT, Claude, and Perplexity answers. It's the one tool in this pack built for getting discovered by AI, not just Google.
Use * to target every crawler, or name a specific one (e.g. Googlebot).
One path per line.
One path per line.
AI crawlers
Choose whether each AI answer engine can crawl your site. Allowed crawlers are omitted from the file (the default, implicit behavior); blocked ones get an explicit Disallow rule.
GPTBot
OpenAI
ChatGPT-User
OpenAI
ClaudeBot
Anthropic
Claude-User
Anthropic
PerplexityBot
Perplexity
CCBot
Common Crawl
Google-Extended
Google AI training
Bytespider
ByteDance
Not respected by Google, but some crawlers honor it.
User-agent: *
Upload this file as robots.txt at the root of your domain (e.g. example.com/robots.txt). Blocking a crawler here is advisory — well-behaved bots respect it, but it does not prevent access to content that is otherwise public.
Related tools
More free utilities for the same workflow
- Sitemap XML ValidatorPaste or upload a sitemap.xml and check it's well-formed — missing <loc> tags, bad <lastmod> dates, relative URLs, and the 50,000-URL limit.Open tool
- JSON-LD Schema GeneratorGenerate valid JSON-LD structured data for FAQ, Article, Organization, and LocalBusiness schemas. Fill in the fields, copy the markup, and paste it into your page.Open tool
- SERP Snippet PreviewSee how your page title and meta description will render in Google search results — desktop and mobile, with real pixel-width truncation so you know exactly where it cuts off.Open tool
- Open Graph & Twitter Card PreviewerPaste your page's HTML and preview exactly how the link card renders on Facebook, LinkedIn, and X — plus a check for missing og: and twitter: tags.Open tool
FAQ
Robots.txt generator FAQ
What is a robots.txt file?
robots.txt is a plain-text file at the root of your domain (example.com/robots.txt) that tells crawlers which parts of your site they may or may not request. It uses User-agent, Disallow, and Allow directives, plus an optional Sitemap line.
What are GPTBot, ClaudeBot, and PerplexityBot?
These are crawlers operated by AI companies — GPTBot and ChatGPT-User by OpenAI, ClaudeBot and Claude-User by Anthropic, and PerplexityBot by Perplexity. They fetch your pages to train models or to answer live user questions with citations, depending on the crawler.
Should I block or allow AI crawlers?
It depends on your goals. Allowing crawlers like GPTBot, ClaudeBot, and PerplexityBot means your content can be cited or summarized in AI answer engines — a growing source of discovery. Blocking them keeps your content out of AI training data and AI-generated answers. There is no universally correct choice; this tool lets you decide per crawler.
What does the Sitemap directive do?
The Sitemap line points crawlers to your XML sitemap URL, helping them discover and prioritize your pages more efficiently. It is optional but recommended if you maintain a sitemap.
What is Crawl-delay for?
Crawl-delay asks a crawler to wait a given number of seconds between requests, which can reduce server load. Google does not honor it, but some other crawlers do. Leave it blank if you don't need to throttle crawl rate.
Does blocking a crawler in robots.txt guarantee it won't access my content?
No. robots.txt is an advisory standard — well-behaved crawlers respect it, but it cannot enforce access control. For content you truly need to keep private, use authentication or server-level blocking instead.
Want your AI visibility handled for you?
mktcrew's SEO agents track how your site shows up across Google and AI answer engines, and keep the technical basics — like robots.txt and sitemaps — aligned with your content strategy. You approve every change before it goes live.
No credit card required · Cancel anytime · Full agent crew on Starter limits · Human approval on every live action
OAuth-only integrations · We never store platform passwords.