AEO guide 06

AI Crawlers and Website Access for Local Businesses

Let the crawlers that feed AI search reach your site, and decide separately whether to allow crawlers that only train AI models. Then ask your host or content delivery network to confirm its bot settings match your robots.txt, because a firewall rule can block AI search even when robots.txt allows it.

← All answer engine optimization guidance
The operating sequence
  1. Know the crawler types

  2. Allow search crawlers

  3. Decide on training

  4. Check host and CDN

  5. Review Google settings

Why this matters

AI assistants can only cite pages they can reach.

ChatGPT, Claude, Perplexity, Microsoft Copilot, and Google's AI Overviews and AI Mode all depend on crawlers: automated programs that fetch web pages. If your site blocks the crawler behind an assistant's search, your own pages can drop out of its answers, and the assistant leans on whatever directories and review sites say about you instead. OpenAI, for example, says sites that opt out of its search crawler, OAI-SearchBot, will not be shown in ChatGPT search answers.

Blocking often happens by accident. robots.txt, a small text file at the root of your site, tells crawlers what they may fetch. But your host, content delivery network (CDN), or security plugin can turn bots away before robots.txt is ever read. Cloudflare changed its AI-crawler options in July 2026, and Cloudflare says customers who block AI training crawlers also block multi-purpose crawlers such as Googlebot and Bingbot. Nothing on your site looks broken when this goes wrong, so it's worth checking on purpose.

Responsible practices

Allow search, decide on training, and check your host.

01

Know the three kinds of AI crawler

Search crawlers build the index an assistant searches when it answers, such as OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, and Applebot. Training crawlers and tokens collect or permit pages for training future AI models, such as GPTBot, ClaudeBot, Google-Extended, and Amazonbot. User-triggered fetchers, such as ChatGPT-User, Claude-User, and Perplexity-User, read a page only when a person asks the assistant to, and OpenAI and Perplexity say robots.txt may not apply to them.

02

Allow search crawlers, and decide on training separately

Each company lists its search and training crawlers as separate controls, so you can allow one and block the other. Blocking a training crawler does not remove you from that company's search. Whether to allow training is a business choice. For being found, what matters is the search crawlers.

03

Know what Google's controls actually do

Google says Google-Extended "does not impact a site's inclusion in Google Search" and isn't a ranking signal; it covers Gemini model training and some other Google AI products. The control for AI Overviews and AI Mode is in Google Search Console under Settings > Search generative AI, available worldwide since August 31, 2026. Google says sites that choose Exclude get no traffic or impressions from those features, so most local businesses should leave it on Include.

04

Think twice before blocking Applebot-Extended

Apple says pages that block Applebot-Extended can still appear in its search results. But Apple also says Applebot then won't use those pages as context when Apple's AI models generate answers. That is a stricter trade-off than Google-Extended, so leave it allowed if you want Siri and Apple's AI answers to draw on your site.

05

Ask your host or CDN what it blocks

Cloudflare now groups AI traffic as Search, Agent, and Training, and says new domains from September 15, 2026 default to blocking Training and Agent on pages that show ads. Reports differ on whether existing free accounts are affected, so don't assume either way. Other hosts, website builders, and security plugins have their own AI-bot switches. Ask whoever manages your site to check and tell you in writing what is on.

06

Build pages that people and agents can use

Google's web.dev guidance for AI agents recommends semantic HTML: real buttons and links, labeled form fields, stable layouts, and no hover-only menus. OpenAI says its Atlas browser agent reads ARIA roles and labels to understand pages. These are also accessibility basics, so the same work helps your customers.

Worked example

Sample robots.txt rules for AI crawlers

Hypothetical example: a two-person HVAC company wants its site to appear in AI search answers but prefers that its pages not be used to train AI models. Its web provider drafts the rules below. Each rule is two lines in the real file; they are shown here on one line, separated by dots.

Example — adapt with your provider
User-agent: OAI-SearchBot / Allow: / · User-agent: Claude-SearchBot / Allow: / · User-agent: PerplexityBot / Allow: / · User-agent: GPTBot / Disallow: / · User-agent: ClaudeBot / Disallow: / · User-agent: Google-Extended / Disallow: / · User-agent: * / Allow: / · Sitemap: https://www.example.com/sitemap.xml
What the Allow rules do
They keep ChatGPT, Claude, and Perplexity search able to read the site. Googlebot, Bingbot, and Applebot are covered by the final rule for everyone else.
What the Disallow rules do
They opt out of training for OpenAI, Anthropic, and Google's Gemini models. Delete them if you're happy to allow training. None of them affects Google Search.
What the file leaves alone
Applebot-Extended stays allowed, so Apple can still use the pages for Siri answers. Existing rules for private or admin pages must be copied into each named group, because named groups don't inherit from the everyone-else rule.
What robots.txt can't do
It states a preference. A CDN or firewall setting can still block these crawlers, so the provider also checks the host's bot settings.

robots.txt says what you want. Your host and CDN settings decide what actually gets through, so check both.

Access check

Check your site's AI access this month.

What you can do yourself

  1. 01

    Open your site's /robots.txt in a browser and note any lines naming GPTBot, OAI-SearchBot, or Google-Extended, or disallowing all bots.

  2. 02

    Decide in writing whether to allow AI training crawlers, and keep search crawlers allowed either way.

  3. 03

    In Google Search Console, open Settings > Search generative AI and confirm it says Include, unless you chose otherwise.

  4. 04

    Ask whoever runs your hosting or CDN which AI-bot settings are on, and get the answer in writing.

  5. 05

    Set a yearly reminder to recheck robots.txt and host bot settings, since platform defaults change.

What to ask your provider or staff to handle

  1. 01

    Check the CDN, firewall, and any security plugin for AI-bot blocks, and confirm Googlebot, Bingbot, Applebot, and OAI-SearchBot aren't being refused.

  2. 02

    Draft robots.txt rules that match your written decision, keep existing private-page rules, and show you before publishing.

  3. 03

    Review key pages for real buttons, labeled form fields, and contact details in text that people and AI agents can read.

Avoid these mistakes

Don't block the crawlers that send you customers.

  • Turning on a one-click "block AI bots" setting without checking what it blocks
  • Blocking Google-Extended and assuming it removes you from AI Overviews
  • Blocking Applebot-Extended without knowing it limits Apple's AI answers
  • Paying for an llms.txt file as a visibility fix
  • Adding noindex or nosnippet to key pages and losing regular search listings

AI assistants decide what to show, change often, and can give different answers to different people, so mentions, citations, and recommendations are not guaranteed.

Questions business owners ask

Common questions about AI crawlers.

01Should a local business block AI crawlers?

Not the search crawlers. Blocking OAI-SearchBot, Claude-SearchBot, PerplexityBot, Bingbot, or Applebot can keep your own pages out of the answers customers see. Training crawlers are a separate choice; blocking them doesn't remove you from search, and allowing them is fine if you're comfortable with it.

02Do I need an llms.txt file?

No. llms.txt is a proposed file that lists a site's pages for AI tools. Google says its Search features don't use it and that adding one will neither help nor harm. The crawler documentation from OpenAI, Anthropic, Perplexity, and Microsoft doesn't list it as something their answers rely on. If your provider already generates one at no cost, it does no harm.

03Can I keep some text out of AI answers without leaving Google?

Partly. Google says nosnippet, data-nosnippet, max-snippet, and noindex limit what it shows from your pages in Search, including in AI features. data-nosnippet can cover just one section of a page. These also limit your regular search snippets, so use them only on text you truly don't want quoted. The Search Console setting removes your whole site from AI features instead.

04How can I tell whether my host is blocking AI search?

robots.txt alone won't tell you, because a firewall can refuse crawlers the file allows. OpenAI tells site owners to confirm that their host or CDN allows traffic from its published search crawler addresses. Ask your provider to check the firewall or security logs for refused requests from the crawlers you want to allow.

Primary guidance

Crawler documentation and settings

Check your site's access

Make sure AI search can read your website.

Start the conversation