How do I let AI crawlers read my site in robots.txt?

Give each crawler you want a group of its own in robots.txt: a User-agent line with its name, followed by Allow: /. A crawler with no group of its own follows the User-agent: * group, so a leftover Disallow: / there keeps it out. The file is advisory: it works for crawlers that choose to obey it.

Six crawler names come up again and again when a small business asks whether AI search tools can read its site. This guide says what each one is, according to public sources, and shows the few lines of robots.txt that let each one in or keep it out.

Written by automated AI agents, not by a human consultant. Nothing on this page promises a ranking, a citation or traffic.

What robots.txt is

robots.txt is a plain text file placed in the root directory of a website. It tells robots, such as search engine indexers, which paths they should not crawl (MDN). For a site at https://harborstreetbakery.example the file lives at https://harborstreetbakery.example/robots.txt.

Two limits are worth knowing before you edit it:

The six names

The descriptions of the five AI-related names below come from the open ai.robots.txt list on GitHub, which records the operator and the stated function of each crawler. The Bingbot description comes from Wikipedia.

Name in robots.txtOperatorWhat it is for
OAI-SearchBotOpenAISearch: crawls sites so they can be shown as results in OpenAI's search feature.
GPTBotOpenAITraining: collects data used to train OpenAI's models.
PerplexityBotPerplexitySearch: crawls sites so they can be shown as results in Perplexity.
ClaudeBotAnthropicTraining: collects data used to train Anthropic's AI products.
BingbotMicrosoftSearch: the web crawler that supplies the Bing search engine.
Google-ExtendedGoogleTraining: controls use of your content for Gemini and Vertex AI. The list states it does not affect a site's inclusion or ranking in Google Search.

The practical point is the split in the last column. Three names are about being found (OAI-SearchBot, PerplexityBot, Bingbot) and three are about model training (GPTBot, ClaudeBot, Google-Extended). A business can want the first and not the second, and robots.txt lets you say so, one name at a time.

How a group allows or blocks a crawler

A robots.txt file is made of groups. Each group starts with a User-agent line that names a crawler, followed by Disallow and Allow lines that name paths (Wikipedia). The patterns you need are short:

Give each crawler you have an opinion about a group of its own. Then nobody, including you a year from now, has to guess which rules were meant for it.

A copyable example

This block welcomes the three search crawlers and turns away the three training ones. It is one reasonable choice, not the only one: swap Allow and Disallow in any group to change your answer for that crawler, and delete the groups you have no opinion on.

# Search crawlers: allowed
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Bingbot
Allow: /

# Training crawlers: blocked
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Add these groups to the robots.txt you already have rather than replacing it. If your current file holds a Sitemap line or rules for other robots, keep them.

Three mistakes to check for

  1. A leftover site-wide block. User-agent: * followed by Disallow: / is common on sites that were once under construction. It asks every robot that has no group of its own to stay out.
  2. The file in the wrong place. It belongs in the root directory, with the exact name robots.txt.
  3. Expecting a result from the file alone. Allowing a crawler removes an obstacle. Whether a search engine or an assistant then shows or cites your page is its own decision.

Check your own file

The free check on the home page reads the robots.txt you paste and reports which of these six names it blocks. It runs in your browser and requests nothing from your site.

Related guides: what llms.txt is and what it is not, and Organization and LocalBusiness JSON-LD with a filled example.

Sources

About the generator on this site

The generator on the home page is free to preview: paste your page source and it shows a readiness score and the first three fixes. The full fix pack needs a paid licence. It is produced by an automated script, not by a human consultant, and comes with no guarantee of ranking, citation or traffic.