AI Crawler Access Analysis Skill

AI crawler access analysis.

How to use it

  1. Hit Copy SKILL.md — or use the Claude Code line below to get every file.
  2. Claude: ⋯ → Download .md, then Customize → Skills → Add → Upload skill.
    ChatGPT: make a Project and paste it into Instructions.
    Neither? Paste it at the top of a new chat — it works for that chat.
  3. Describe your job in plain words. The AI follows the skill from there.
Claude Code — installs the whole folder, not just SKILL.md
npx degit zubair-trabzada/geo-seo-claude/skills/geo-crawlers#main ~/.claude/skills/geo-crawlers

For one project only, change the path to .claude/skills/geo-crawlers. This skill also uses robots.txt, ads.txt, GEO-CRAWLER-ACCESS.md, llms.txt — copying SKILL.md alone won't be enough. See the folder on GitHub.

Not working?
  • Check which app you pasted it into — the steps above name the right one.
  • Some skills need the paid tier of Claude or ChatGPT.
Step-by-step guide with screenshots · Ask in the forum

Paste into Claude, ChatGPT or Cursor.

Show the full text378 lines
geo-crawlers/SKILL.md378 lines17.5 KBpushed 145d agoRawView on GitHub

AI Crawler Access Analysis Skill

Purpose

This skill analyzes a website's accessibility to AI crawlers -- the bots that AI companies use to discover, index, and train on web content. If AI crawlers are blocked, the site's content cannot appear in AI-generated responses regardless of its quality. Crawler access is the foundational technical requirement for GEO.

Key Insight

As of early 2026, many websites inadvertently block AI crawlers through overly aggressive robots.txt rules, inherited from legacy SEO configurations. An Originality.ai 2025 study found that over 35% of the top 1,000 websites block at least one major AI crawler, and 5-10% block all AI crawlers. Blocking AI crawlers is the single fastest way to become invisible in AI-generated search results.


Complete AI Crawler Reference

Tier 1: Critical for AI Search Visibility (RECOMMEND: ALLOW)

These crawlers power the AI search products where users actively look for answers. Blocking them directly reduces your visibility in AI-generated responses.

GPTBot

  • Operator: OpenAI
  • User-Agent: GPTBot
  • Full User-Agent String: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot)
  • Purpose: Fetches content for ChatGPT's web browsing, plugins, and search features. Content accessed by GPTBot may be used to improve OpenAI models.
  • Impact of Blocking: Content will NOT appear in ChatGPT Search results or be accessible when users ask ChatGPT to browse the web. This is the highest-impact AI crawler to allow.
  • Recommendation: ALLOW -- ChatGPT has 300M+ weekly active users as of 2025. Blocking GPTBot removes your content from one of the largest AI search surfaces.

OAI-SearchBot

  • Operator: OpenAI
  • User-Agent: OAI-SearchBot
  • Full User-Agent String: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; OAI-SearchBot/1.0; +https://docs.openai.com/bots/overview)
  • Purpose: Specifically powers ChatGPT's search feature. Unlike GPTBot, content accessed by OAI-SearchBot is NOT used for model training -- only for live search results.
  • Impact of Blocking: Content will not appear in ChatGPT's search results even if GPTBot is allowed.
  • Recommendation: ALLOW -- This is a search-only crawler with no training implications. There is no strategic reason to block it.

ChatGPT-User

  • Operator: OpenAI
  • User-Agent: ChatGPT-User
  • Full User-Agent String: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ChatGPT-User/1.0; +https://openai.com/bot)
  • Purpose: Used when a ChatGPT user explicitly asks the model to visit a specific URL. Acts like a browser agent on behalf of the user.
  • Impact of Blocking: ChatGPT cannot visit your pages when users ask it to read or summarize them. This prevents direct user-initiated traffic.
  • Recommendation: ALLOW -- Blocking this bot prevents users who are actively trying to engage with your content from accessing it through ChatGPT.

ClaudeBot

  • Operator: Anthropic
  • User-Agent: ClaudeBot
  • Full User-Agent String: ClaudeBot/1.0; +https://www.anthropic.com/claude-bot
  • Purpose: Fetches web content for Claude's features including web search, citations, and analysis tools.
  • Impact of Blocking: Content will not be accessible to Claude for web search or when users ask Claude to analyze specific URLs.
  • Recommendation: ALLOW -- Claude is a major AI assistant with growing market share. Blocking ClaudeBot reduces your AI search footprint.

PerplexityBot

  • Operator: Perplexity AI
  • User-Agent: PerplexityBot
  • Full User-Agent String: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
  • Purpose: Powers Perplexity's AI search engine, which provides sourced answers with direct citations and links back to source pages.
  • Impact of Blocking: Content will not appear in Perplexity search results. Perplexity is one of the best referral traffic sources among AI search products because it always displays source links.
  • Recommendation: ALLOW -- Perplexity drives actual referral traffic and always attributes sources. High-value AI crawler for publishers and businesses.

Tier 2: Important for Broader AI Ecosystem (RECOMMEND: ALLOW)

These crawlers serve large AI platforms or search ecosystems. Allowing them increases your content's reach.

Google-Extended

  • Operator: Google
  • User-Agent: Google-Extended
  • Purpose: Controls whether Google uses your content for Gemini model training and AI Overviews improvement. CRITICAL NOTE: Blocking Google-Extended does NOT affect your Google Search rankings or your appearance in Google Search results. That is controlled by the standard Googlebot.
  • Impact of Blocking: Content may not be used for Gemini training or to improve AI Overviews. However, your content can still appear in AI Overviews based on standard search indexing.
  • Recommendation: ALLOW -- Blocking provides minimal content protection upside while reducing your presence in Google's AI features. Since it does not affect standard search ranking, the only reason to block is philosophical objection to training data usage.

GoogleOther

  • Operator: Google
  • User-Agent: GoogleOther
  • Purpose: Used by Google for various non-search-ranking purposes including research, one-off crawls, and AI-related data collection.
  • Impact of Blocking: Minimal impact on search rankings. May reduce presence in Google's AI research and experimental features.
  • Recommendation: ALLOW -- Low risk, moderate potential benefit for AI feature inclusion.

Applebot-Extended

  • Operator: Apple
  • User-Agent: Applebot-Extended
  • Purpose: Used by Apple to train and improve Apple Intelligence features, Siri, and Apple's AI products. Separate from standard Applebot (which powers Siri search and Spotlight Suggestions).
  • Impact of Blocking: Content may not be used in Apple Intelligence features. Standard Siri and Spotlight functionality is unaffected (controlled by Applebot).
  • Recommendation: ALLOW -- Apple Intelligence is integrated into all Apple devices (2B+ active devices). Presence in Apple's AI features has growing strategic value.

Amazonbot

  • Operator: Amazon
  • User-Agent: Amazonbot
  • Full User-Agent String: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/600.2.5 (KHTML, like Gecko) Version/8.0.2 Safari/600.2.5 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)
  • Purpose: Indexes content for Alexa answers and Amazon's AI features.
  • Impact of Blocking: Content will not appear in Alexa voice responses or Amazon's AI-powered search features.
  • Recommendation: ALLOW -- Relevant for voice search optimization. Lower priority than Tier 1 crawlers but no downside to allowing.

FacebookBot

  • Operator: Meta
  • User-Agent: FacebookBot
  • Purpose: Used by Meta for AI features across Facebook, Instagram, WhatsApp, and Meta AI assistant.
  • Impact of Blocking: Content may not be accessible to Meta AI. Link previews on Facebook/Instagram are handled by a different crawler and are unaffected.
  • Recommendation: ALLOW -- Meta AI is embedded in apps with 3B+ combined users. Growing importance for AI visibility.

Tier 3: Training-Only Crawlers (ALLOW or BLOCK Based on Strategy)

These crawlers are primarily used for AI model training rather than live search features. Blocking them does not affect AI search visibility.

CCBot

  • Operator: Common Crawl (nonprofit)
  • User-Agent: CCBot
  • Full User-Agent String: CCBot/2.0 (https://commoncrawl.org/faq/)
  • Purpose: Builds the Common Crawl dataset, which is used as training data by many AI companies (Google, Meta, Stability AI, and others).
  • Impact of Blocking: Content will not appear in future Common Crawl datasets. Does NOT affect any live AI search product.
  • Recommendation: CONTEXT-DEPENDENT -- Allow if you want maximum long-term AI training presence. Block if you want to control training data usage. No impact on search visibility.

anthropic-ai

  • Operator: Anthropic
  • User-Agent: anthropic-ai
  • Purpose: Used by Anthropic for AI safety research and Claude model training. Separate from ClaudeBot (which powers live features).
  • Impact of Blocking: Content will not be used for Claude training. Does NOT affect Claude's live search or web browsing features (controlled by ClaudeBot).
  • Recommendation: CONTEXT-DEPENDENT -- Similar to CCBot. Allow for training presence, block for training data control. No impact on live AI search.

Bytespider

  • Operator: ByteDance
  • User-Agent: Bytespider
  • Purpose: Used by ByteDance for various AI products including TikTok's AI features and Doubao (their ChatGPT competitor in China).
  • Impact of Blocking: Content will not be used for ByteDance AI products. Minimal impact for Western-market businesses.
  • Recommendation: BLOCK for most Western businesses (aggressive crawling behavior reported, minimal search visibility benefit). ALLOW if targeting Chinese/Asian markets.

cohere-ai

  • Operator: Cohere
  • User-Agent: cohere-ai
  • Purpose: Used by Cohere for model training. Cohere powers enterprise AI solutions and the Coral chat product.
  • Impact of Blocking: Content will not be used for Cohere model training. Minimal direct consumer-facing impact.
  • Recommendation: CONTEXT-DEPENDENT -- Low priority. Allow or block based on general training data stance.

Recommendation Matrix Summary

Crawler Tier Recommendation Reason
GPTBot 1 ALLOW Powers ChatGPT Search (300M+ users)
OAI-SearchBot 1 ALLOW Search-only, no training use
ChatGPT-User 1 ALLOW User-initiated browsing
ClaudeBot 1 ALLOW Claude web search and analysis
PerplexityBot 1 ALLOW Best referral traffic AI search
Google-Extended 2 ALLOW Gemini features; no search rank impact
GoogleOther 2 ALLOW Google AI research
Applebot-Extended 2 ALLOW Apple Intelligence (2B+ devices)
Amazonbot 2 ALLOW Alexa and Amazon AI
FacebookBot 2 ALLOW Meta AI (3B+ app users)
CCBot 3 Context Training data only
anthropic-ai 3 Context Training data only
Bytespider 3 BLOCK Aggressive crawler, low benefit
cohere-ai 3 Context Training data only

Maximum AI Visibility Configuration (robots.txt)

For sites wanting maximum AI search visibility:

# AI Crawlers - ALLOWED for AI search visibility
User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: anthropic-ai
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: GoogleOther
Allow: /

User-agent: Applebot-Extended
Allow: /

User-agent: Amazonbot
Allow: /

User-agent: FacebookBot
Allow: /

# AI Crawlers - BLOCKED (aggressive/low value)
User-agent: Bytespider
Disallow: /

User-agent: CCBot
Disallow: /

Analysis Procedure

Step 1: Fetch and Parse robots.txt

  1. Use WebFetch to retrieve [domain]/robots.txt.
  2. Parse all User-agent directives and their associated Allow/Disallow rules.
  3. For each AI crawler in the reference list above:
    • Check if there is a specific User-agent block for that crawler
    • Check if there is a wildcard (User-agent: *) block that would apply
    • Determine effective access: Allowed, Blocked, or Not Mentioned (inherits wildcard rules)
  4. Note any Crawl-delay directives that may slow AI crawler access.
  5. Check for Sitemap directives (AI crawlers use these for discovery).

Step 2: Check Meta Robots Tags

  1. For a sample of 5-10 key pages, fetch the HTML and check for:
    • <meta name="robots" content="noindex"> -- blocks all bots
    • <meta name="robots" content="nofollow"> -- prevents link following
    • <meta name="robots" content="noai"> -- emerging tag to block AI use
    • <meta name="robots" content="noimageai"> -- blocks AI image training
    • Bot-specific meta tags: <meta name="GPTBot" content="noindex">
  2. Record any page-level overrides of the robots.txt directives.

Step 3: Check HTTP Headers

  1. For the same sample pages, check response headers for:
    • X-Robots-Tag: noindex -- HTTP header equivalent of meta noindex
    • X-Robots-Tag: noai -- HTTP header to block AI use
    • X-Robots-Tag: noimageai -- blocks AI image training
    • Bot-specific headers: X-Robots-Tag: GPTBot: noindex
  2. Note that HTTP headers override meta tags and apply to non-HTML resources too.

Step 4: Check for AI-Specific Files

  1. Check for /llms.txt (emerging standard for AI crawler guidance).
  2. Check for /.well-known/ai-plugin.json (OpenAI plugin manifest).
  3. Check for /ai.txt (proposed standard, similar to ads.txt for AI).
  4. Record presence/absence and quality of each file.

Step 5: Assess JavaScript Rendering Requirements

  1. Check if the site is a Single Page Application (SPA) or heavily JavaScript-rendered.
  2. AI crawlers vary in their JavaScript rendering capabilities:
    • GPTBot: Limited JS rendering
    • ClaudeBot: Limited JS rendering
    • PerplexityBot: Limited JS rendering
    • Googlebot: Full JS rendering (but Google-Extended inherits this)
  3. If critical content requires JS rendering, flag this as a potential issue.
  4. Check for Server-Side Rendering (SSR) or Static Site Generation (SSG) as mitigations.

Step 6: Parse Content Signals

Using the already-fetched robots.txt from Step 1, scan for Content-Signal: directives (IETF draft draft-romm-aipref-contentsignals).

  1. Scan every line for a line starting with Content-Signal: (case-insensitive).
  2. If found:
    • Parse all key=value pairs (split on , then on =).
    • Validate keys against the known set: ai-train, search, ai-personalization, ai-retrieval.
    • Validate values: only yes and no are valid.
    • Flag any unknown keys or invalid values as a warning — the spec is still an IETF draft.
    • Record the result as Pass and surface parsed values with plain-English meaning.
  3. If absent: record as Recommendation — the site has not declared AI usage preferences.

No additional HTTP request is needed. robots.txt is already fetched in Step 1.


Output Format

Generate a file called GEO-CRAWLER-ACCESS.md:

# AI Crawler Access Report: [Domain]

**Analysis Date:** [Date]
**Domain:** [Domain]
**robots.txt Status:** [Found/Not Found/Error]

---

## Crawler Access Summary

| Crawler | Operator | Tier | Status | Impact |
|---|---|---|---|---|
| GPTBot | OpenAI | 1 | [Allowed/Blocked/Not Mentioned] | [Impact description] |
| OAI-SearchBot | OpenAI | 1 | [Status] | [Impact] |
| ChatGPT-User | OpenAI | 1 | [Status] | [Impact] |
| ClaudeBot | Anthropic | 1 | [Status] | [Impact] |
| PerplexityBot | Perplexity | 1 | [Status] | [Impact] |
| Google-Extended | Google | 2 | [Status] | [Impact] |
| GoogleOther | Google | 2 | [Status] | [Impact] |
| Applebot-Extended | Apple | 2 | [Status] | [Impact] |
| Amazonbot | Amazon | 2 | [Status] | [Impact] |
| FacebookBot | Meta | 2 | [Status] | [Impact] |
| CCBot | Common Crawl | 3 | [Status] | [Impact] |
| anthropic-ai | Anthropic | 3 | [Status] | [Impact] |
| Bytespider | ByteDance | 3 | [Status] | [Impact] |
| cohere-ai | Cohere | 3 | [Status] | [Impact] |

## AI Visibility Score: [X]/100

**Tier 1 Access:** [X/5 crawlers allowed]
**Tier 2 Access:** [X/5 crawlers allowed]
**Tier 3 Access:** [X/4 crawlers allowed]

---

## Critical Issues

[List any Tier 1 crawlers that are blocked]

## Recommendations

### Immediate Actions
[Specific robots.txt changes needed]

### robots.txt Recommendation

[Complete recommended robots.txt content for AI crawlers]


### Additional Technical Findings
- **Meta Robots Tags:** [Findings]
- **X-Robots-Tag Headers:** [Findings]
- **JavaScript Rendering:** [Assessment]
- **llms.txt:** [Present/Absent]
- **Sitemap Accessibility:** [Assessment]

### Content Signals (IETF Draft)

**Status:** Present / Absent

<!-- If present: -->
| Signal Key | Value | Meaning |
|---|---|---|
| ai-train | no | Opted out of AI model training |
| search | yes | Permits use in AI-powered search results |

<!-- If absent: -->
**Recommendation:** Add a `Content-Signal:` directive to robots.txt to declare AI usage preferences explicitly. Example:

`Content-Signal: ai-train=no, search=yes, ai-retrieval=yes`

See https://contentsignals.org/ for the full specification.

Scoring for Crawler Access

The AI Crawler Access Score is calculated as:

Component Weight Scoring
Tier 1 Crawlers Allowed 50% 20 points per Tier 1 crawler allowed (5 crawlers = 100 points max, scaled to 50)
Tier 2 Crawlers Allowed 25% 20 points per Tier 2 crawler allowed (5 crawlers = 100 points max, scaled to 25)
No Blanket AI Blocks 15% Full points if no User-agent: * Disallow: / and no noai meta tags
AI-Specific Files Present 10% 5 points for llms.txt, 5 points for sitemap accessible to AI crawlers

Final score = sum of all weighted components, capped at 100.

1---
2name: geo-crawlers
3description: AI crawler access analysis. Checks robots.txt, meta tags, and HTTP headers to determine which AI crawlers can access the site. Provides a complete access map and recommendations for maximizing AI visibility while maintaining appropriate control.
4allowed-tools:
5 - Read
6 - Grep
7 - Glob
8 - Bash
9 - WebFetch
10 - Write
11---
12 
13# AI Crawler Access Analysis Skill
14 
15## Purpose
16 
17This skill analyzes a website's accessibility to AI crawlers -- the bots that AI companies use to discover, index, and train on web content. If AI crawlers are blocked, the site's content cannot appear in AI-generated responses regardless of its quality. Crawler access is the foundational technical requirement for GEO.
18 
19## Key Insight
20 
21As of early 2026, many websites inadvertently block AI crawlers through overly aggressive robots.txt rules, inherited from legacy SEO configurations. An Originality.ai 2025 study found that over 35% of the top 1,000 websites block at least one major AI crawler, and 5-10% block all AI crawlers. Blocking AI crawlers is the single fastest way to become invisible in AI-generated search results.
22 
23---
24 
25## Complete AI Crawler Reference
26 
27### Tier 1: Critical for AI Search Visibility (RECOMMEND: ALLOW)
28 
29These crawlers power the AI search products where users actively look for answers. Blocking them directly reduces your visibility in AI-generated responses.
30 
31#### GPTBot
32- **Operator:** OpenAI
33- **User-Agent:** `GPTBot`
34- **Full User-Agent String:** `Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot)`
35- **Purpose:** Fetches content for ChatGPT's web browsing, plugins, and search features. Content accessed by GPTBot may be used to improve OpenAI models.
36- **Impact of Blocking:** Content will NOT appear in ChatGPT Search results or be accessible when users ask ChatGPT to browse the web. This is the highest-impact AI crawler to allow.
37- **Recommendation:** **ALLOW** -- ChatGPT has 300M+ weekly active users as of 2025. Blocking GPTBot removes your content from one of the largest AI search surfaces.
38 
39#### OAI-SearchBot
40- **Operator:** OpenAI
41- **User-Agent:** `OAI-SearchBot`
42- **Full User-Agent String:** `Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; OAI-SearchBot/1.0; +https://docs.openai.com/bots/overview)`
43- **Purpose:** Specifically powers ChatGPT's search feature. Unlike GPTBot, content accessed by OAI-SearchBot is NOT used for model training -- only for live search results.
44- **Impact of Blocking:** Content will not appear in ChatGPT's search results even if GPTBot is allowed.
45- **Recommendation:** **ALLOW** -- This is a search-only crawler with no training implications. There is no strategic reason to block it.
46 
47#### ChatGPT-User
48- **Operator:** OpenAI
49- **User-Agent:** `ChatGPT-User`
50- **Full User-Agent String:** `Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ChatGPT-User/1.0; +https://openai.com/bot)`
51- **Purpose:** Used when a ChatGPT user explicitly asks the model to visit a specific URL. Acts like a browser agent on behalf of the user.
52- **Impact of Blocking:** ChatGPT cannot visit your pages when users ask it to read or summarize them. This prevents direct user-initiated traffic.
53- **Recommendation:** **ALLOW** -- Blocking this bot prevents users who are actively trying to engage with your content from accessing it through ChatGPT.
54 
55#### ClaudeBot
56- **Operator:** Anthropic
57- **User-Agent:** `ClaudeBot`
58- **Full User-Agent String:** `ClaudeBot/1.0; +https://www.anthropic.com/claude-bot`
59- **Purpose:** Fetches web content for Claude's features including web search, citations, and analysis tools.
60- **Impact of Blocking:** Content will not be accessible to Claude for web search or when users ask Claude to analyze specific URLs.
61- **Recommendation:** **ALLOW** -- Claude is a major AI assistant with growing market share. Blocking ClaudeBot reduces your AI search footprint.
62 
63#### PerplexityBot
64- **Operator:** Perplexity AI
65- **User-Agent:** `PerplexityBot`
66- **Full User-Agent String:** `Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)`
67- **Purpose:** Powers Perplexity's AI search engine, which provides sourced answers with direct citations and links back to source pages.
68- **Impact of Blocking:** Content will not appear in Perplexity search results. Perplexity is one of the best referral traffic sources among AI search products because it always displays source links.
69- **Recommendation:** **ALLOW** -- Perplexity drives actual referral traffic and always attributes sources. High-value AI crawler for publishers and businesses.
70 
71---
72 
73### Tier 2: Important for Broader AI Ecosystem (RECOMMEND: ALLOW)
74 
75These crawlers serve large AI platforms or search ecosystems. Allowing them increases your content's reach.
76 
77#### Google-Extended
78- **Operator:** Google
79- **User-Agent:** `Google-Extended`
80- **Purpose:** Controls whether Google uses your content for Gemini model training and AI Overviews improvement. **CRITICAL NOTE:** Blocking Google-Extended does NOT affect your Google Search rankings or your appearance in Google Search results. That is controlled by the standard Googlebot.
81- **Impact of Blocking:** Content may not be used for Gemini training or to improve AI Overviews. However, your content can still appear in AI Overviews based on standard search indexing.
82- **Recommendation:** **ALLOW** -- Blocking provides minimal content protection upside while reducing your presence in Google's AI features. Since it does not affect standard search ranking, the only reason to block is philosophical objection to training data usage.
83 
84#### GoogleOther
85- **Operator:** Google
86- **User-Agent:** `GoogleOther`
87- **Purpose:** Used by Google for various non-search-ranking purposes including research, one-off crawls, and AI-related data collection.
88- **Impact of Blocking:** Minimal impact on search rankings. May reduce presence in Google's AI research and experimental features.
89- **Recommendation:** **ALLOW** -- Low risk, moderate potential benefit for AI feature inclusion.
90 
91#### Applebot-Extended
92- **Operator:** Apple
93- **User-Agent:** `Applebot-Extended`
94- **Purpose:** Used by Apple to train and improve Apple Intelligence features, Siri, and Apple's AI products. Separate from standard Applebot (which powers Siri search and Spotlight Suggestions).
95- **Impact of Blocking:** Content may not be used in Apple Intelligence features. Standard Siri and Spotlight functionality is unaffected (controlled by Applebot).
96- **Recommendation:** **ALLOW** -- Apple Intelligence is integrated into all Apple devices (2B+ active devices). Presence in Apple's AI features has growing strategic value.
97 
98#### Amazonbot
99- **Operator:** Amazon
100- **User-Agent:** `Amazonbot`
101- **Full User-Agent String:** `Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/600.2.5 (KHTML, like Gecko) Version/8.0.2 Safari/600.2.5 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)`
102- **Purpose:** Indexes content for Alexa answers and Amazon's AI features.
103- **Impact of Blocking:** Content will not appear in Alexa voice responses or Amazon's AI-powered search features.
104- **Recommendation:** **ALLOW** -- Relevant for voice search optimization. Lower priority than Tier 1 crawlers but no downside to allowing.
105 
106#### FacebookBot
107- **Operator:** Meta
108- **User-Agent:** `FacebookBot`
109- **Purpose:** Used by Meta for AI features across Facebook, Instagram, WhatsApp, and Meta AI assistant.
110- **Impact of Blocking:** Content may not be accessible to Meta AI. Link previews on Facebook/Instagram are handled by a different crawler and are unaffected.
111- **Recommendation:** **ALLOW** -- Meta AI is embedded in apps with 3B+ combined users. Growing importance for AI visibility.
112 
113---
114 
115### Tier 3: Training-Only Crawlers (ALLOW or BLOCK Based on Strategy)
116 
117These crawlers are primarily used for AI model training rather than live search features. Blocking them does not affect AI search visibility.
118 
119#### CCBot
120- **Operator:** Common Crawl (nonprofit)
121- **User-Agent:** `CCBot`
122- **Full User-Agent String:** `CCBot/2.0 (https://commoncrawl.org/faq/)`
123- **Purpose:** Builds the Common Crawl dataset, which is used as training data by many AI companies (Google, Meta, Stability AI, and others).
124- **Impact of Blocking:** Content will not appear in future Common Crawl datasets. Does NOT affect any live AI search product.
125- **Recommendation:** **CONTEXT-DEPENDENT** -- Allow if you want maximum long-term AI training presence. Block if you want to control training data usage. No impact on search visibility.
126 
127#### anthropic-ai
128- **Operator:** Anthropic
129- **User-Agent:** `anthropic-ai`
130- **Purpose:** Used by Anthropic for AI safety research and Claude model training. Separate from ClaudeBot (which powers live features).
131- **Impact of Blocking:** Content will not be used for Claude training. Does NOT affect Claude's live search or web browsing features (controlled by ClaudeBot).
132- **Recommendation:** **CONTEXT-DEPENDENT** -- Similar to CCBot. Allow for training presence, block for training data control. No impact on live AI search.
133 
134#### Bytespider
135- **Operator:** ByteDance
136- **User-Agent:** `Bytespider`
137- **Purpose:** Used by ByteDance for various AI products including TikTok's AI features and Doubao (their ChatGPT competitor in China).
138- **Impact of Blocking:** Content will not be used for ByteDance AI products. Minimal impact for Western-market businesses.
139- **Recommendation:** **BLOCK** for most Western businesses (aggressive crawling behavior reported, minimal search visibility benefit). **ALLOW** if targeting Chinese/Asian markets.
140 
141#### cohere-ai
142- **Operator:** Cohere
143- **User-Agent:** `cohere-ai`
144- **Purpose:** Used by Cohere for model training. Cohere powers enterprise AI solutions and the Coral chat product.
145- **Impact of Blocking:** Content will not be used for Cohere model training. Minimal direct consumer-facing impact.
146- **Recommendation:** **CONTEXT-DEPENDENT** -- Low priority. Allow or block based on general training data stance.
147 
148---
149 
150## Recommendation Matrix Summary
151 
152| Crawler | Tier | Recommendation | Reason |
153|---|---|---|---|
154| GPTBot | 1 | **ALLOW** | Powers ChatGPT Search (300M+ users) |
155| OAI-SearchBot | 1 | **ALLOW** | Search-only, no training use |
156| ChatGPT-User | 1 | **ALLOW** | User-initiated browsing |
157| ClaudeBot | 1 | **ALLOW** | Claude web search and analysis |
158| PerplexityBot | 1 | **ALLOW** | Best referral traffic AI search |
159| Google-Extended | 2 | **ALLOW** | Gemini features; no search rank impact |
160| GoogleOther | 2 | **ALLOW** | Google AI research |
161| Applebot-Extended | 2 | **ALLOW** | Apple Intelligence (2B+ devices) |
162| Amazonbot | 2 | **ALLOW** | Alexa and Amazon AI |
163| FacebookBot | 2 | **ALLOW** | Meta AI (3B+ app users) |
164| CCBot | 3 | Context | Training data only |
165| anthropic-ai | 3 | Context | Training data only |
166| Bytespider | 3 | **BLOCK** | Aggressive crawler, low benefit |
167| cohere-ai | 3 | Context | Training data only |
168 
169### Maximum AI Visibility Configuration (robots.txt)
170 
171For sites wanting maximum AI search visibility:
172 
173```
174# AI Crawlers - ALLOWED for AI search visibility
175User-agent: GPTBot
176Allow: /
177 
178User-agent: OAI-SearchBot
179Allow: /
180 
181User-agent: ChatGPT-User
182Allow: /
183 
184User-agent: ClaudeBot
185Allow: /
186 
187User-agent: anthropic-ai
188Allow: /
189 
190User-agent: PerplexityBot
191Allow: /
192 
193User-agent: Google-Extended
194Allow: /
195 
196User-agent: GoogleOther
197Allow: /
198 
199User-agent: Applebot-Extended
200Allow: /
201 
202User-agent: Amazonbot
203Allow: /
204 
205User-agent: FacebookBot
206Allow: /
207 
208# AI Crawlers - BLOCKED (aggressive/low value)
209User-agent: Bytespider
210Disallow: /
211 
212User-agent: CCBot
213Disallow: /
214```
215 
216---
217 
218## Analysis Procedure
219 
220### Step 1: Fetch and Parse robots.txt
221 
2221. Use WebFetch to retrieve `[domain]/robots.txt`.
2232. Parse all User-agent directives and their associated Allow/Disallow rules.
2243. For each AI crawler in the reference list above:
225 - Check if there is a specific User-agent block for that crawler
226 - Check if there is a wildcard (`User-agent: *`) block that would apply
227 - Determine effective access: **Allowed**, **Blocked**, or **Not Mentioned** (inherits wildcard rules)
2284. Note any `Crawl-delay` directives that may slow AI crawler access.
2295. Check for `Sitemap` directives (AI crawlers use these for discovery).
230 
231### Step 2: Check Meta Robots Tags
232 
2331. For a sample of 5-10 key pages, fetch the HTML and check for:
234 - `<meta name="robots" content="noindex">` -- blocks all bots
235 - `<meta name="robots" content="nofollow">` -- prevents link following
236 - `<meta name="robots" content="noai">` -- emerging tag to block AI use
237 - `<meta name="robots" content="noimageai">` -- blocks AI image training
238 - Bot-specific meta tags: `<meta name="GPTBot" content="noindex">`
2392. Record any page-level overrides of the robots.txt directives.
240 
241### Step 3: Check HTTP Headers
242 
2431. For the same sample pages, check response headers for:
244 - `X-Robots-Tag: noindex` -- HTTP header equivalent of meta noindex
245 - `X-Robots-Tag: noai` -- HTTP header to block AI use
246 - `X-Robots-Tag: noimageai` -- blocks AI image training
247 - Bot-specific headers: `X-Robots-Tag: GPTBot: noindex`
2482. Note that HTTP headers override meta tags and apply to non-HTML resources too.
249 
250### Step 4: Check for AI-Specific Files
251 
2521. Check for `/llms.txt` (emerging standard for AI crawler guidance).
2532. Check for `/.well-known/ai-plugin.json` (OpenAI plugin manifest).
2543. Check for `/ai.txt` (proposed standard, similar to ads.txt for AI).
2554. Record presence/absence and quality of each file.
256 
257### Step 5: Assess JavaScript Rendering Requirements
258 
2591. Check if the site is a Single Page Application (SPA) or heavily JavaScript-rendered.
2602. AI crawlers vary in their JavaScript rendering capabilities:
261 - GPTBot: Limited JS rendering
262 - ClaudeBot: Limited JS rendering
263 - PerplexityBot: Limited JS rendering
264 - Googlebot: Full JS rendering (but Google-Extended inherits this)
2653. If critical content requires JS rendering, flag this as a potential issue.
2664. Check for Server-Side Rendering (SSR) or Static Site Generation (SSG) as mitigations.
267 
268### Step 6: Parse Content Signals
269 
270Using the already-fetched robots.txt from Step 1, scan for `Content-Signal:` directives (IETF draft `draft-romm-aipref-contentsignals`).
271 
2721. Scan every line for a line starting with `Content-Signal:` (case-insensitive).
2732. If found:
274 - Parse all key=value pairs (split on `,` then on `=`).
275 - Validate keys against the known set: `ai-train`, `search`, `ai-personalization`, `ai-retrieval`.
276 - Validate values: only `yes` and `no` are valid.
277 - Flag any unknown keys or invalid values as a warning — the spec is still an IETF draft.
278 - Record the result as **Pass** and surface parsed values with plain-English meaning.
2793. If absent: record as **Recommendation** — the site has not declared AI usage preferences.
280 
281No additional HTTP request is needed. robots.txt is already fetched in Step 1.
282 
283---
284 
285## Output Format
286 
287Generate a file called `GEO-CRAWLER-ACCESS.md`:
288 
289```markdown
290# AI Crawler Access Report: [Domain]
291 
292**Analysis Date:** [Date]
293**Domain:** [Domain]
294**robots.txt Status:** [Found/Not Found/Error]
295 
296---
297 
298## Crawler Access Summary
299 
300| Crawler | Operator | Tier | Status | Impact |
301|---|---|---|---|---|
302| GPTBot | OpenAI | 1 | [Allowed/Blocked/Not Mentioned] | [Impact description] |
303| OAI-SearchBot | OpenAI | 1 | [Status] | [Impact] |
304| ChatGPT-User | OpenAI | 1 | [Status] | [Impact] |
305| ClaudeBot | Anthropic | 1 | [Status] | [Impact] |
306| PerplexityBot | Perplexity | 1 | [Status] | [Impact] |
307| Google-Extended | Google | 2 | [Status] | [Impact] |
308| GoogleOther | Google | 2 | [Status] | [Impact] |
309| Applebot-Extended | Apple | 2 | [Status] | [Impact] |
310| Amazonbot | Amazon | 2 | [Status] | [Impact] |
311| FacebookBot | Meta | 2 | [Status] | [Impact] |
312| CCBot | Common Crawl | 3 | [Status] | [Impact] |
313| anthropic-ai | Anthropic | 3 | [Status] | [Impact] |
314| Bytespider | ByteDance | 3 | [Status] | [Impact] |
315| cohere-ai | Cohere | 3 | [Status] | [Impact] |
316 
317## AI Visibility Score: [X]/100
318 
319**Tier 1 Access:** [X/5 crawlers allowed]
320**Tier 2 Access:** [X/5 crawlers allowed]
321**Tier 3 Access:** [X/4 crawlers allowed]
322 
323---
324 
325## Critical Issues
326 
327[List any Tier 1 crawlers that are blocked]
328 
329## Recommendations
330 
331### Immediate Actions
332[Specific robots.txt changes needed]
333 
334### robots.txt Recommendation
335```
336[Complete recommended robots.txt content for AI crawlers]
337```
338 
339### Additional Technical Findings
340- **Meta Robots Tags:** [Findings]
341- **X-Robots-Tag Headers:** [Findings]
342- **JavaScript Rendering:** [Assessment]
343- **llms.txt:** [Present/Absent]
344- **Sitemap Accessibility:** [Assessment]
345 
346### Content Signals (IETF Draft)
347 
348**Status:** Present / Absent
349 
350<!-- If present: -->
351| Signal Key | Value | Meaning |
352|---|---|---|
353| ai-train | no | Opted out of AI model training |
354| search | yes | Permits use in AI-powered search results |
355 
356<!-- If absent: -->
357**Recommendation:** Add a `Content-Signal:` directive to robots.txt to declare AI usage preferences explicitly. Example:
358 
359`Content-Signal: ai-train=no, search=yes, ai-retrieval=yes`
360 
361See https://contentsignals.org/ for the full specification.
362```
363 
364---
365 
366## Scoring for Crawler Access
367 
368The AI Crawler Access Score is calculated as:
369 
370| Component | Weight | Scoring |
371|---|---|---|
372| Tier 1 Crawlers Allowed | 50% | 20 points per Tier 1 crawler allowed (5 crawlers = 100 points max, scaled to 50) |
373| Tier 2 Crawlers Allowed | 25% | 20 points per Tier 2 crawler allowed (5 crawlers = 100 points max, scaled to 25) |
374| No Blanket AI Blocks | 15% | Full points if no `User-agent: *` Disallow: / and no noai meta tags |
375| AI-Specific Files Present | 10% | 5 points for llms.txt, 5 points for sitemap accessible to AI crawlers |
376 
377Final score = sum of all weighted components, capped at 100.
378 

Discussion

Alternatives

Also in Scraping & extraction