Technical AI Visibility: Crawler Access & Rendering skill
- robots.txt Template for AI Crawlers
by AgriciDaniel·MIT license·★ 2,219 Stars on the repo·GitHub ↗
Files of Technical AI Visibility: Crawler Access & Rendering
AgriciDaniel/
Show the full text469 lines
Technical AI Visibility: Crawler Access & Rendering
Contents
- robots.txt Template for AI Crawlers
- Cloudflare AI Crawl Control: CRITICAL
- Google Gen-AI Guidance
- llms.txt Implementation
- Server-Side Rendering Requirements
- Passage-Level Extractability
- Performance Requirements
- Testing AI Crawler Visibility
- AI Crawler Traffic Growth
- AI Crawler Checklist
robots.txt Template for AI Crawlers
Allow documented AI crawlers explicitly when you want access. For compliant
crawlers, an absent Disallow usually means allowed; explicit Allow rules are
optional documentation and help teams audit intent.
# ===========================================
# AI Search & LLM Crawlers: Explicitly Allow
# ===========================================
# OpenAI
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
# Anthropic documented crawler families
User-agent: ClaudeBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
# Deprecated Anthropic strings (kept for legacy compatibility):
# User-agent: Claude-Web
# User-agent: anthropic-ai
# Google AI product token (Gemini/Vertex training and non-Search grounding controls)
# Google Search AI features use Googlebot plus preview controls:
# https://developers.google.com/search/docs/appearance/ai-features
User-agent: Google-Extended
Allow: /
# Perplexity
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
# Meta
User-agent: Meta-ExternalAgent
Allow: /
# ByteDance
User-agent: Bytespider
Allow: /
# Google AI agents (Project Mariner)
User-agent: Google-Agent
Allow: /
# DuckDuckGo AI
User-agent: DuckAssistBot
Allow: /
# Apple (Siri, Apple Intelligence)
User-agent: Applebot-Extended
Allow: /
# Amazon (Alexa, product search)
User-agent: Amazonbot
Allow: /
# You.com
User-agent: YouBot
Allow: /
# Phind (developer search)
User-agent: PhindBot
Allow: /
# Exa (AI-native search engine)
User-agent: ExaBot
Allow: /
# Common Crawl (used by many AI models)
User-agent: CCBot
Allow: /
# ===========================================
# Traditional Search Engines
# ===========================================
User-agent: Googlebot
Allow: /
User-agent: Bingbot
Allow: /
User-agent: *
Allow: /
# ===========================================
# Sitemap
# ===========================================
Sitemap: https://example.com/sitemap.xml
Crawler Identification Reference
Providers expose different crawler classes. Some split training, search indexing, and user-triggered retrieval; others publish only one bot or a product token. Blocking a documented search/indexing bot can reduce visibility in that platform's answers. User-triggered retrieval may not fully respect robots.txt. OpenAI bot details: https://platform.openai.com/docs/bots.
| Crawler | Operator | Type | Respects robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Training | Yes |
| OAI-SearchBot | OpenAI | Search indexing | Yes |
| ChatGPT-User | OpenAI | User-triggered retrieval | Not guaranteed |
| ClaudeBot | Anthropic | Training | Yes |
| Claude-SearchBot | Anthropic | Search indexing | Yes |
| Claude-User | Anthropic | User retrieval | Yes |
| Anthropic | Deprecated | - | |
| Anthropic | Deprecated | - | |
| Google-Extended | Gemini/Vertex training and some non-Search grounding controls; not Search AI inclusion | Yes | |
| Google-Agent | Project Mariner agentic (2026) | Yes | |
| PerplexityBot | Perplexity | Search indexing | Yes |
| Perplexity-User | Perplexity | User retrieval | Partial |
| Applebot-Extended | Apple | Apple Intelligence training | Yes |
| Meta-ExternalAgent | Meta | High-volume data collection | Yes |
| Bytespider | ByteDance | Training/indexing | Partial (documented issues) |
| Amazonbot | Amazon | Alexa / product search | Yes |
| DuckAssistBot | DuckDuckGo | DuckAssist AI answers | Yes |
| YouBot | You.com | AI search engine | Yes |
| PhindBot | Phind | Developer-focused AI search | Yes |
| ExaBot | Exa | Neural search engine | Yes |
| CCBot | Common Crawl | Open dataset (used by many LLMs) | Yes |
robots.txt Strategy by Bot Type
Treat each bot category differently based on your goals:
- Training/product tokens (GPTBot, ClaudeBot, CCBot, Google-Extended): Your choice. Blocking affects training or non-Search product use as documented by each provider, but Google-Extended does not control Google Search AI inclusion.
- Search/indexing bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot): Allow these. Blocking means your content won't appear in ChatGPT, Claude, or Perplexity answers.
- Retrieval bots (ChatGPT-User, Perplexity-User): May not fully respect robots.txt. These are triggered by live user queries and may fetch content regardless of directives.
Cloudflare AI Crawl Control: CRITICAL
Since July 2025, Cloudflare blocks AI crawlers by default on new domains. This is the single most common reason blogs are invisible to AI systems despite having correct robots.txt configuration.
How to Fix
- Log in to Cloudflare dashboard
- Navigate to Security > Bots > AI Crawlers
- Review the list of AI crawlers
- Toggle "Allow" for each AI crawler you want to permit
- Save changes
What Cloudflare Blocks by Default
| Crawler or token | Default Status (New Domains) |
|---|---|
| GPTBot | Blocked |
| ClaudeBot | Blocked |
| PerplexityBot | Blocked |
| CCBot | Blocked |
| Google-Extended | Blocked |
| Applebot-Extended | Allowed |
| Googlebot | Allowed (not an AI crawler) |
Verification
After updating Cloudflare settings, verify access:
# Simulate GPTBot user-agent
curl -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" https://yourdomain.com/blog/test-post | head -50
# Check for Cloudflare block page (403 or challenge page)
curl -s -o /dev/null -w "%{http_code}" -A "Mozilla/5.0 (compatible; ClaudeBot/1.0)" https://yourdomain.com/
If you get a 403 or an HTML page with "Cloudflare" in it, the crawler is blocked.
Google Gen-AI Guidance
Google Search Central's AI features guidance says optimization for AI Overviews and AI Mode is normal SEO: https://developers.google.com/search/docs/appearance/ai-features. Google does not require special AI schema or llms.txt for Search AI features. Use crawlable HTML, standard Article schema with author Person and publisher Organization, clear source attribution, helpful trustworthy content, and fast server responses.
llms.txt Implementation
The llms.txt standard (proposed by llmstxt.org, Sep 2024) provides a machine-readable
summary of your site for LLMs. Place at site root: https://example.com/llms.txt.
Important caveat: Google's current stance is no llms.txt needed for AI Overviews or AI Mode per Google Search Central's AI features guidance. No major AI platform has confirmed relying on it. Treat it as an optional site inventory for non-Google tools, not a ranking, indexing, or citation requirement.
Specification
- Plain text file, UTF-8
- Under 10KB total
- Structured list of important URLs with brief descriptions
- Helps LLMs understand site structure and find authoritative content
Template
# Example Blog
> A blog about modern web development, SEO, and content strategy.
## Main Pages
- [Home](https://example.com/): Main landing page with latest articles
- [About](https://example.com/about): Company information and mission
- [Blog](https://example.com/blog): All published articles
## Popular Articles
- [Complete Guide to Technical SEO in 2026](https://example.com/blog/technical-seo-guide): Comprehensive technical SEO guide covering Core Web Vitals, crawlability, and schema markup.
- [How AI Overviews Changed Search](https://example.com/blog/ai-overviews-impact): Data-driven analysis of AI Overview impact on organic traffic with case studies.
- [Content Strategy for B2B SaaS](https://example.com/blog/b2b-saas-content-strategy): Framework for building a content program that drives pipeline.
## Topic Clusters
- [SEO](https://example.com/topics/seo): All articles about search engine optimization
- [Content Strategy](https://example.com/topics/content-strategy): Content planning and execution
- [Web Development](https://example.com/topics/web-development): Frontend and backend development guides
## Authors
- [Sarah Chen](https://example.com/author/sarah-chen): Content strategist, B2B SaaS specialist
- [Marcus Rivera](https://example.com/author/marcus-rivera): Senior frontend engineer, React expert
Key Rules
- Do not exceed 10KB (LLMs may truncate or ignore larger files)
- Use markdown-style links:
[Title](URL): Description - Include only your most important and highest-quality pages
- Update when you publish significant new content
- This is NOT a sitemap replacement: it supplements sitemap.xml
- Do not treat a missing llms.txt file as an AI visibility blocker
Server-Side Rendering Requirements
Standard non-Google AI crawlers generally should be assumed not to execute JavaScript unless their documentation says otherwise. Content rendered only via client-side JavaScript is risky for AI visibility; render important blog content into initial HTML and verify per crawler.
Rendering Strategy Ranking
| Strategy | AI Visibility | Performance | Recommendation |
|---|---|---|---|
| SSG (Static Site Generation) | Best | Best | Preferred for blogs |
| SSR (Server-Side Rendering) | Excellent | Good | Good for dynamic content |
| ISR (Incremental Static Regeneration) | Excellent | Good | Good for large sites |
| CSR (Client-Side Rendering) | None | Poor for crawlers | Never use for content |
JavaScript Execution by Crawler
| Crawler | Executes JavaScript | Renders Pages |
|---|---|---|
| GPTBot | No | No |
| OAI-SearchBot | No | No |
| ChatGPT-User | No | No |
| ClaudeBot | No | No |
| Claude-SearchBot | No | No |
| Claude-User | No | No |
| PerplexityBot | No | No |
| Perplexity-User | No | No |
| Meta-ExternalAgent | No | No |
| Bytespider | No | No |
| Amazonbot | No | No |
| CCBot | No | No |
| Googlebot | Yes | Yes |
| AppleBot | Yes | Yes |
| OpenAI agentic browsing surfaces | Yes | Yes |
| Google-Agent (agentic) | Yes | Yes |
Vercel Findings
Vercel analyzed 500M+ GPTBot fetches and found zero evidence of JavaScript execution. GPTBot reads raw HTML only. Content loaded via React hydration, Vue mounting, or any client-side framework is completely invisible.
Exception: Agentic Tools
Standard AI crawlers generally do not execute JavaScript. However, agentic tools are different:
- OpenAI agentic browsing surfaces: Full JS rendering may be available depending on product mode.
- Google-Agent / Project Mariner (Google, 2026): Operates through Chrome with full rendering.
These are user-directed agents, not automated crawlers. They can see JS-rendered content, but they do not replace the need for SSR - standard crawlers still dominate citation indexing.
Passage-Level Extractability
Crawler access gets a page into the candidate set. Citation selection depends on whether the page contains self-contained answer passages AI systems can extract. Target 120-180 word passages that answer one question without relying on the surrounding article.
Under each H2, start with an approximately 50-word direct-answer sentence that gives the answer, the year, the named entity, and the source attribution. Follow with specific entities, dates, original examples, and first-hand Experience markers. A clean passage can earn an AI Overview citation even when the full page is not cited.
AI Overviews also began highlighting links from a user's subscribed publications in 2026, so publisher trust and subscriptions can affect which citations users notice (Nieman Lab, 2026-05).
Performance Requirements
AI retrieval systems have practical latency budgets. Slow sites may reduce crawl, fetch, and extraction reliability before content quality is evaluated.
Note: The thresholds below are industry best practices and observations from SEO tooling (Discovered Labs, Prerender.io, Kevin Indig). They are NOT officially published specifications from OpenAI, Anthropic, or Perplexity. Treat as directional targets, not guaranteed cutoffs.
Thresholds
| Metric | Target | Risk threshold | Consequence |
|---|---|---|---|
| TTFB (Time to First Byte) | < 200ms | > 600ms | May reduce crawl or extraction reliability |
| Full page load (HTML) | < 500ms | > 1,000ms | May reduce crawl frequency |
| Response size (HTML) | < 200KB | > 500KB | May cause partial content extraction |
Optimization Priorities
- Use a CDN: Content must be served from edge locations
- Enable compression: gzip or Brotli for all text responses
- Minimize HTML bloat: Remove unused CSS/JS from HTML response
- Cache aggressively: Static pages should have long cache headers
- Pre-render: Use SSG or SSR, never CSR for content pages
Testing AI Crawler Visibility
Quick Test: See What AI Crawlers See
# Basic: view raw HTML (what all AI crawlers receive)
curl -s https://yourdomain.com/blog/your-post | head -200
# Check if main content is in HTML source
curl -s https://yourdomain.com/blog/your-post | grep -c "<article"
# Check for JS-only rendering indicators
curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"__next\""
curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"root\""
curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"app\""
# If the above returns content in a <noscript> tag or empty divs,
# your content is behind JS and invisible to AI crawlers.
Full Crawler Simulation
# Simulate GPTBot
curl -s -H "User-Agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" \
https://yourdomain.com/blog/your-post > /tmp/gptbot-view.html
# Simulate ClaudeBot
curl -s -H "User-Agent: Mozilla/5.0 (compatible; ClaudeBot/1.0; +https://claudebot.ai)" \
https://yourdomain.com/blog/your-post > /tmp/claudebot-view.html
# Check if content exists
wc -l /tmp/gptbot-view.html
grep -c "your-expected-heading-text" /tmp/gptbot-view.html
Red Flags (Content Invisible to AI)
| Indicator | What It Means |
|---|---|
Empty <div id="root"></div> |
React CSR: content loads via JS only |
Empty <div id="__next"></div> without SSR/RSC/static output |
Next.js App Router or Pages Router shipping content client-side only |
<noscript> contains the content |
Content explicitly hidden from non-JS clients |
<script> tags contain all content as JSON |
Data fetched client-side, not in HTML |
| HTML under 5KB for a full blog post | Content not rendered server-side |
Next.js App Router Guidance
- Prefer static rendering for blog routes. Use Server Components for article
content and
generateStaticParams()for known slugs. - Use ISR for large blogs when content changes after build. Keep the article body in server-rendered HTML.
- Use dynamic rendering only when the page genuinely depends on request-time data. Do not move the article body behind client-only data fetching.
generateMetadata()should emit canonical, Open Graph, and Article metadata server-side.
AI Crawler Traffic Growth
Traffic from AI crawlers is growing exponentially. Sites that block or fail to serve these crawlers are losing compounding visibility.
| Metric | Value | Source |
|---|---|---|
| GPTBot traffic growth | +305% YoY | Cloudflare Radar, 2025 |
| PerplexityBot traffic growth | +157,490% YoY | Cloudflare Radar, 2025 |
| AI crawling volume overall | +32% YoY | Cloudflare, 2025 |
| Top 10 domains' citation share | 46% of all ChatGPT citations per topic | Growth Memo, Mar 2026 |
| AI referral traffic share | Small but fastest-growing; no standardized total-web share | Similarweb, 2026-05-28 |
| AI referral traffic growth | 3x+ YoY from September 2024 to September 2025 | Similarweb, 2026-05-28 |
| Gemini referral trend | About 18% share, +237% YoY | Similarweb, 2026-05-28 |
| ChatGPT referral trend | Share slid from about 87% to the high-60s | Similarweb, 2026-05-28 |
AI Crawler Checklist
| Check | Pass | Fail |
|---|---|---|
| robots.txt allows AI crawlers | All major bots listed with Allow: / |
Missing entries or Disallow: / |
| Cloudflare AI settings reviewed | AI crawlers explicitly allowed in dashboard | Default block left in place |
| llms.txt treated as optional | Not required for Google AI visibility | Treating a missing file as a blocker |
| Content in HTML source | curl returns full content |
Empty divs, JS-only rendering |
| TTFB under 200ms | Measured from CDN edge | Over 600ms increases crawl or extraction risk |
| Schema in HTML source | Standard Article, Person, Organization JSON-LD in source HTML | Special AI-only schema or JS-injected schema |
| Sitemap.xml accessible | Valid XML, all blog URLs included | Missing or returns 404 |
| No Cloudflare challenge on bot UA | 200 status code | 403 or challenge page |
| 1 | # Technical AI Visibility: Crawler Access & Rendering |
| 2 | |
| 3 | ## Contents |
| 4 | |
| 5 | [robots.txt Template for AI Crawlers] |
| 6 | [Cloudflare AI Crawl Control: CRITICAL] |
| 7 | [Google Gen-AI Guidance] |
| 8 | [llms.txt Implementation] |
| 9 | [Server-Side Rendering Requirements] |
| 10 | [Passage-Level Extractability] |
| 11 | [Performance Requirements] |
| 12 | [Testing AI Crawler Visibility] |
| 13 | [AI Crawler Traffic Growth] |
| 14 | [AI Crawler Checklist] |
| 15 | |
| 16 | ## robots.txt Template for AI Crawlers |
| 17 | |
| 18 | Allow documented AI crawlers explicitly when you want access. For compliant |
| 19 | crawlers, an absent `Disallow` usually means allowed; explicit `Allow` rules are |
| 20 | optional documentation and help teams audit intent. |
| 21 | |
| 22 | |
| 23 | # =========================================== |
| 24 | # AI Search & LLM Crawlers: Explicitly Allow |
| 25 | # =========================================== |
| 26 | |
| 27 | # OpenAI |
| 28 | User-agent: GPTBot |
| 29 | Allow: / |
| 30 | |
| 31 | User-agent: OAI-SearchBot |
| 32 | Allow: / |
| 33 | |
| 34 | User-agent: ChatGPT-User |
| 35 | Allow: / |
| 36 | |
| 37 | # Anthropic documented crawler families |
| 38 | User-agent: ClaudeBot |
| 39 | Allow: / |
| 40 | |
| 41 | User-agent: Claude-SearchBot |
| 42 | Allow: / |
| 43 | |
| 44 | User-agent: Claude-User |
| 45 | Allow: / |
| 46 | |
| 47 | # Deprecated Anthropic strings (kept for legacy compatibility): |
| 48 | # User-agent: Claude-Web |
| 49 | # User-agent: anthropic-ai |
| 50 | |
| 51 | # Google AI product token (Gemini/Vertex training and non-Search grounding controls) |
| 52 | # Google Search AI features use Googlebot plus preview controls: |
| 53 | # https://developers.google.com/search/docs/appearance/ai-features |
| 54 | User-agent: Google-Extended |
| 55 | Allow: / |
| 56 | |
| 57 | # Perplexity |
| 58 | User-agent: PerplexityBot |
| 59 | Allow: / |
| 60 | |
| 61 | User-agent: Perplexity-User |
| 62 | Allow: / |
| 63 | |
| 64 | # Meta |
| 65 | User-agent: Meta-ExternalAgent |
| 66 | Allow: / |
| 67 | |
| 68 | # ByteDance |
| 69 | User-agent: Bytespider |
| 70 | Allow: / |
| 71 | |
| 72 | # Google AI agents (Project Mariner) |
| 73 | User-agent: Google-Agent |
| 74 | Allow: / |
| 75 | |
| 76 | # DuckDuckGo AI |
| 77 | User-agent: DuckAssistBot |
| 78 | Allow: / |
| 79 | |
| 80 | # Apple (Siri, Apple Intelligence) |
| 81 | User-agent: Applebot-Extended |
| 82 | Allow: / |
| 83 | |
| 84 | # Amazon (Alexa, product search) |
| 85 | User-agent: Amazonbot |
| 86 | Allow: / |
| 87 | |
| 88 | # You.com |
| 89 | User-agent: YouBot |
| 90 | Allow: / |
| 91 | |
| 92 | # Phind (developer search) |
| 93 | User-agent: PhindBot |
| 94 | Allow: / |
| 95 | |
| 96 | # Exa (AI-native search engine) |
| 97 | User-agent: ExaBot |
| 98 | Allow: / |
| 99 | |
| 100 | # Common Crawl (used by many AI models) |
| 101 | User-agent: CCBot |
| 102 | Allow: / |
| 103 | |
| 104 | # =========================================== |
| 105 | # Traditional Search Engines |
| 106 | # =========================================== |
| 107 | |
| 108 | User-agent: Googlebot |
| 109 | Allow: / |
| 110 | |
| 111 | User-agent: Bingbot |
| 112 | Allow: / |
| 113 | |
| 114 | User-agent: * |
| 115 | Allow: / |
| 116 | |
| 117 | # =========================================== |
| 118 | # Sitemap |
| 119 | # =========================================== |
| 120 | Sitemap: https://example.com/sitemap.xml |
| 121 | |
| 122 | |
| 123 | ### Crawler Identification Reference |
| 124 | |
| 125 | Providers expose different crawler classes. Some split training, search indexing, |
| 126 | and user-triggered retrieval; others publish only one bot or a product token. |
| 127 | Blocking a documented search/indexing bot can reduce visibility in that platform's |
| 128 | answers. User-triggered retrieval may not fully respect robots.txt. |
| 129 | OpenAI bot details: https://platform.openai.com/docs/bots. |
| 130 | |
| 131 | | Crawler | Operator | Type | Respects robots.txt | |
| 132 | |---------|----------|------|---------------------| |
| 133 | | GPTBot | OpenAI | Training | Yes | |
| 134 | | OAI-SearchBot | OpenAI | Search indexing | Yes | |
| 135 | | ChatGPT-User | OpenAI | User-triggered retrieval | Not guaranteed | |
| 136 | | ClaudeBot | Anthropic | Training | Yes | |
| 137 | | Claude-SearchBot | Anthropic | Search indexing | Yes | |
| 138 | | Claude-User | Anthropic | User retrieval | Yes | |
| 139 | | ~~Claude-Web~~ | Anthropic | Deprecated | - | |
| 140 | | ~~anthropic-ai~~ | Anthropic | Deprecated | - | |
| 141 | | Google-Extended | Google | Gemini/Vertex training and some non-Search grounding controls; not Search AI inclusion | Yes | |
| 142 | | Google-Agent | Google | Project Mariner agentic (2026) | Yes | |
| 143 | | PerplexityBot | Perplexity | Search indexing | Yes | |
| 144 | | Perplexity-User | Perplexity | User retrieval | Partial | |
| 145 | | Applebot-Extended | Apple | Apple Intelligence training | Yes | |
| 146 | | Meta-ExternalAgent | Meta | High-volume data collection | Yes | |
| 147 | | Bytespider | ByteDance | Training/indexing | Partial (documented issues) | |
| 148 | | Amazonbot | Amazon | Alexa / product search | Yes | |
| 149 | | DuckAssistBot | DuckDuckGo | DuckAssist AI answers | Yes | |
| 150 | | YouBot | You.com | AI search engine | Yes | |
| 151 | | PhindBot | Phind | Developer-focused AI search | Yes | |
| 152 | | ExaBot | Exa | Neural search engine | Yes | |
| 153 | | CCBot | Common Crawl | Open dataset (used by many LLMs) | Yes | |
| 154 | |
| 155 | ### robots.txt Strategy by Bot Type |
| 156 | |
| 157 | Treat each bot category differently based on your goals: |
| 158 | **Training/product tokens** (GPTBot, ClaudeBot, CCBot, Google-Extended): Your |
| 159 | choice. Blocking affects training or non-Search product use as documented by |
| 160 | each provider, but Google-Extended does not control Google Search AI inclusion. |
| 161 | **Search/indexing bots** (OAI-SearchBot, Claude-SearchBot, PerplexityBot): **Allow these.** |
| 162 | Blocking means your content won't appear in ChatGPT, Claude, or Perplexity answers. |
| 163 | **Retrieval bots** (ChatGPT-User, Perplexity-User): May not fully respect robots.txt. These |
| 164 | are triggered by live user queries and may fetch content regardless of directives. |
| 165 | |
| 166 | |
| 167 | |
| 168 | ## Cloudflare AI Crawl Control: CRITICAL |
| 169 | |
| 170 | **Since July 2025, Cloudflare blocks AI crawlers by default on new domains.** |
| 171 | This is the single most common reason blogs are invisible to AI systems despite |
| 172 | having correct robots.txt configuration. |
| 173 | |
| 174 | ### How to Fix |
| 175 | |
| 176 | Log in to Cloudflare dashboard |
| 177 | Navigate to **Security > Bots > AI Crawlers** |
| 178 | Review the list of AI crawlers |
| 179 | **Toggle "Allow" for each AI crawler you want to permit** |
| 180 | Save changes |
| 181 | |
| 182 | ### What Cloudflare Blocks by Default |
| 183 | |
| 184 | | Crawler or token | Default Status (New Domains) | |
| 185 | |------------------|------------------------------| |
| 186 | | GPTBot | Blocked | |
| 187 | | ClaudeBot | Blocked | |
| 188 | | PerplexityBot | Blocked | |
| 189 | | CCBot | Blocked | |
| 190 | | Google-Extended | Blocked | |
| 191 | | Applebot-Extended | Allowed | |
| 192 | | Googlebot | Allowed (not an AI crawler) | |
| 193 | |
| 194 | ### Verification |
| 195 | |
| 196 | After updating Cloudflare settings, verify access: |
| 197 | |
| 198 | |
| 199 | # Simulate GPTBot user-agent |
| 200 | curl -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" https://yourdomain.com/blog/test-post | head -50 |
| 201 | |
| 202 | # Check for Cloudflare block page (403 or challenge page) |
| 203 | curl -s -o /dev/null -w "%{http_code}" -A "Mozilla/5.0 (compatible; ClaudeBot/1.0)" https://yourdomain.com/ |
| 204 | |
| 205 | |
| 206 | If you get a 403 or an HTML page with "Cloudflare" in it, the crawler is blocked. |
| 207 | |
| 208 | |
| 209 | |
| 210 | ## Google Gen-AI Guidance |
| 211 | |
| 212 | Google Search Central's AI features guidance says optimization for AI Overviews |
| 213 | and AI Mode is normal SEO: https://developers.google.com/search/docs/appearance/ai-features. |
| 214 | Google does not require special AI schema or llms.txt for Search AI features. |
| 215 | Use crawlable HTML, standard Article schema with author Person and publisher |
| 216 | Organization, clear source attribution, helpful trustworthy content, and fast |
| 217 | server responses. |
| 218 | |
| 219 | |
| 220 | |
| 221 | ## llms.txt Implementation |
| 222 | |
| 223 | The `llms.txt` standard (proposed by llmstxt.org, Sep 2024) provides a machine-readable |
| 224 | summary of your site for LLMs. Place at site root: `https://example.com/llms.txt`. |
| 225 | |
| 226 | **Important caveat:** Google's current stance is no llms.txt needed for AI |
| 227 | Overviews or AI Mode per Google Search Central's AI features guidance. No major |
| 228 | AI platform has confirmed relying on it. Treat it as an optional site inventory |
| 229 | for non-Google tools, not a ranking, indexing, or citation requirement. |
| 230 | |
| 231 | ### Specification |
| 232 | |
| 233 | Plain text file, UTF-8 |
| 234 | Under 10KB total |
| 235 | Structured list of important URLs with brief descriptions |
| 236 | Helps LLMs understand site structure and find authoritative content |
| 237 | |
| 238 | ### Template |
| 239 | |
| 240 | |
| 241 | # Example Blog |
| 242 | |
| 243 | > A blog about modern web development, SEO, and content strategy. |
| 244 | |
| 245 | ## Main Pages |
| 246 | |
| 247 | - [Home](https://example.com/): Main landing page with latest articles |
| 248 | - [About](https://example.com/about): Company information and mission |
| 249 | - [Blog](https://example.com/blog): All published articles |
| 250 | |
| 251 | ## Popular Articles |
| 252 | |
| 253 | - [Complete Guide to Technical SEO in 2026](https://example.com/blog/technical-seo-guide): Comprehensive technical SEO guide covering Core Web Vitals, crawlability, and schema markup. |
| 254 | - [How AI Overviews Changed Search](https://example.com/blog/ai-overviews-impact): Data-driven analysis of AI Overview impact on organic traffic with case studies. |
| 255 | - [Content Strategy for B2B SaaS](https://example.com/blog/b2b-saas-content-strategy): Framework for building a content program that drives pipeline. |
| 256 | |
| 257 | ## Topic Clusters |
| 258 | |
| 259 | - [SEO](https://example.com/topics/seo): All articles about search engine optimization |
| 260 | - [Content Strategy](https://example.com/topics/content-strategy): Content planning and execution |
| 261 | - [Web Development](https://example.com/topics/web-development): Frontend and backend development guides |
| 262 | |
| 263 | ## Authors |
| 264 | |
| 265 | - [Sarah Chen](https://example.com/author/sarah-chen): Content strategist, B2B SaaS specialist |
| 266 | - [Marcus Rivera](https://example.com/author/marcus-rivera): Senior frontend engineer, React expert |
| 267 | |
| 268 | |
| 269 | ### Key Rules |
| 270 | |
| 271 | Do not exceed 10KB (LLMs may truncate or ignore larger files) |
| 272 | Use markdown-style links: `[Title]: Description` |
| 273 | Include only your most important and highest-quality pages |
| 274 | Update when you publish significant new content |
| 275 | This is NOT a sitemap replacement: it supplements sitemap.xml |
| 276 | Do not treat a missing llms.txt file as an AI visibility blocker |
| 277 | |
| 278 | |
| 279 | |
| 280 | ## Server-Side Rendering Requirements |
| 281 | |
| 282 | Standard non-Google AI crawlers generally should be assumed not to execute |
| 283 | JavaScript unless their documentation says otherwise. Content rendered only via |
| 284 | client-side JavaScript is risky for AI visibility; render important blog content |
| 285 | into initial HTML and verify per crawler. |
| 286 | |
| 287 | ### Rendering Strategy Ranking |
| 288 | |
| 289 | | Strategy | AI Visibility | Performance | Recommendation | |
| 290 | |----------|--------------|-------------|----------------| |
| 291 | | **SSG** (Static Site Generation) | Best | Best | Preferred for blogs | |
| 292 | | **SSR** (Server-Side Rendering) | Excellent | Good | Good for dynamic content | |
| 293 | | **ISR** (Incremental Static Regeneration) | Excellent | Good | Good for large sites | |
| 294 | | **CSR** (Client-Side Rendering) | None | Poor for crawlers | Never use for content | |
| 295 | |
| 296 | ### JavaScript Execution by Crawler |
| 297 | |
| 298 | | Crawler | Executes JavaScript | Renders Pages | |
| 299 | |---------|-------------------|---------------| |
| 300 | | GPTBot | No | No | |
| 301 | | OAI-SearchBot | No | No | |
| 302 | | ChatGPT-User | No | No | |
| 303 | | ClaudeBot | No | No | |
| 304 | | Claude-SearchBot | No | No | |
| 305 | | Claude-User | No | No | |
| 306 | | PerplexityBot | No | No | |
| 307 | | Perplexity-User | No | No | |
| 308 | | Meta-ExternalAgent | No | No | |
| 309 | | Bytespider | No | No | |
| 310 | | Amazonbot | No | No | |
| 311 | | CCBot | No | No | |
| 312 | | **Googlebot** | **Yes** | **Yes** | |
| 313 | | **AppleBot** | **Yes** | **Yes** | |
| 314 | | **OpenAI agentic browsing surfaces** | **Yes** | **Yes** | |
| 315 | | **Google-Agent** (agentic) | **Yes** | **Yes** | |
| 316 | |
| 317 | ### Vercel Findings |
| 318 | |
| 319 | Vercel analyzed 500M+ GPTBot fetches and found **zero evidence of JavaScript |
| 320 | execution**. GPTBot reads raw HTML only. Content loaded via React hydration, |
| 321 | Vue mounting, or any client-side framework is completely invisible. |
| 322 | |
| 323 | ### Exception: Agentic Tools |
| 324 | |
| 325 | Standard AI crawlers generally do not execute JavaScript. However, **agentic tools** are different: |
| 326 | **OpenAI agentic browsing surfaces**: Full JS rendering may be available depending on product mode. |
| 327 | **Google-Agent / Project Mariner** (Google, 2026): Operates through Chrome with full rendering. |
| 328 | |
| 329 | These are user-directed agents, not automated crawlers. They can see JS-rendered content, |
| 330 | but they do not replace the need for SSR - standard crawlers still dominate citation indexing. |
| 331 | |
| 332 | |
| 333 | |
| 334 | ## Passage-Level Extractability |
| 335 | |
| 336 | Crawler access gets a page into the candidate set. Citation selection depends on |
| 337 | whether the page contains self-contained answer passages AI systems can extract. |
| 338 | Target 120-180 word passages that answer one question without relying on the |
| 339 | surrounding article. |
| 340 | |
| 341 | Under each H2, start with an approximately 50-word direct-answer sentence that |
| 342 | gives the answer, the year, the named entity, and the source attribution. Follow |
| 343 | with specific entities, dates, original examples, and first-hand Experience |
| 344 | markers. A clean passage can earn an AI Overview citation even when the full |
| 345 | page is not cited. |
| 346 | |
| 347 | AI Overviews also began highlighting links from a user's subscribed |
| 348 | publications in 2026, so publisher trust and subscriptions can affect which |
| 349 | citations users notice (Nieman Lab, 2026-05). |
| 350 | |
| 351 | |
| 352 | |
| 353 | ## Performance Requirements |
| 354 | |
| 355 | AI retrieval systems have practical latency budgets. Slow sites may reduce crawl, |
| 356 | fetch, and extraction reliability before content quality is evaluated. |
| 357 | |
| 358 | **Note:** The thresholds below are industry best practices and observations from SEO tooling |
| 359 | (Discovered Labs, Prerender.io, Kevin Indig). They are NOT officially published specifications |
| 360 | from OpenAI, Anthropic, or Perplexity. Treat as directional targets, not guaranteed cutoffs. |
| 361 | |
| 362 | ### Thresholds |
| 363 | |
| 364 | | Metric | Target | Risk threshold | Consequence | |
| 365 | |--------|--------|----------------|-------------| |
| 366 | | TTFB (Time to First Byte) | < 200ms | > 600ms | May reduce crawl or extraction reliability | |
| 367 | | Full page load (HTML) | < 500ms | > 1,000ms | May reduce crawl frequency | |
| 368 | | Response size (HTML) | < 200KB | > 500KB | May cause partial content extraction | |
| 369 | |
| 370 | ### Optimization Priorities |
| 371 | |
| 372 | **Use a CDN**: Content must be served from edge locations |
| 373 | **Enable compression**: gzip or Brotli for all text responses |
| 374 | **Minimize HTML bloat**: Remove unused CSS/JS from HTML response |
| 375 | **Cache aggressively**: Static pages should have long cache headers |
| 376 | **Pre-render**: Use SSG or SSR, never CSR for content pages |
| 377 | |
| 378 | |
| 379 | |
| 380 | ## Testing AI Crawler Visibility |
| 381 | |
| 382 | ### Quick Test: See What AI Crawlers See |
| 383 | |
| 384 | |
| 385 | # Basic: view raw HTML (what all AI crawlers receive) |
| 386 | curl -s https://yourdomain.com/blog/your-post | head -200 |
| 387 | |
| 388 | # Check if main content is in HTML source |
| 389 | curl -s https://yourdomain.com/blog/your-post | grep -c "<article" |
| 390 | |
| 391 | # Check for JS-only rendering indicators |
| 392 | curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"__next\"" |
| 393 | curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"root\"" |
| 394 | curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"app\"" |
| 395 | |
| 396 | # If the above returns content in a <noscript> tag or empty divs, |
| 397 | # your content is behind JS and invisible to AI crawlers. |
| 398 | |
| 399 | |
| 400 | ### Full Crawler Simulation |
| 401 | |
| 402 | |
| 403 | # Simulate GPTBot |
| 404 | curl -s -H "User-Agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" \ |
| 405 | https://yourdomain.com/blog/your-post > /tmp/gptbot-view.html |
| 406 | |
| 407 | # Simulate ClaudeBot |
| 408 | curl -s -H "User-Agent: Mozilla/5.0 (compatible; ClaudeBot/1.0; +https://claudebot.ai)" \ |
| 409 | https://yourdomain.com/blog/your-post > /tmp/claudebot-view.html |
| 410 | |
| 411 | # Check if content exists |
| 412 | wc -l /tmp/gptbot-view.html |
| 413 | grep -c "your-expected-heading-text" /tmp/gptbot-view.html |
| 414 | |
| 415 | |
| 416 | ### Red Flags (Content Invisible to AI) |
| 417 | |
| 418 | | Indicator | What It Means | |
| 419 | |-----------|---------------| |
| 420 | | Empty `<div id="root"></div>` | React CSR: content loads via JS only | |
| 421 | | Empty `<div id="__next"></div>` without SSR/RSC/static output | Next.js App Router or Pages Router shipping content client-side only | |
| 422 | | `<noscript>` contains the content | Content explicitly hidden from non-JS clients | |
| 423 | | `<script>` tags contain all content as JSON | Data fetched client-side, not in HTML | |
| 424 | | HTML under 5KB for a full blog post | Content not rendered server-side | |
| 425 | |
| 426 | ### Next.js App Router Guidance |
| 427 | |
| 428 | Prefer static rendering for blog routes. Use Server Components for article |
| 429 | content and `generateStaticParams()` for known slugs. |
| 430 | Use ISR for large blogs when content changes after build. Keep the article body |
| 431 | in server-rendered HTML. |
| 432 | Use dynamic rendering only when the page genuinely depends on request-time data. |
| 433 | Do not move the article body behind client-only data fetching. |
| 434 | `generateMetadata()` should emit canonical, Open Graph, and Article metadata |
| 435 | server-side. |
| 436 | |
| 437 | |
| 438 | |
| 439 | ## AI Crawler Traffic Growth |
| 440 | |
| 441 | Traffic from AI crawlers is growing exponentially. Sites that block or fail |
| 442 | to serve these crawlers are losing compounding visibility. |
| 443 | |
| 444 | | Metric | Value | Source | |
| 445 | |--------|-------|--------| |
| 446 | | GPTBot traffic growth | +305% YoY | Cloudflare Radar, 2025 | |
| 447 | | PerplexityBot traffic growth | +157,490% YoY | Cloudflare Radar, 2025 | |
| 448 | | AI crawling volume overall | +32% YoY | Cloudflare, 2025 | |
| 449 | | Top 10 domains' citation share | 46% of all ChatGPT citations per topic | Growth Memo, Mar 2026 | |
| 450 | | AI referral traffic share | Small but fastest-growing; no standardized total-web share | Similarweb, 2026-05-28 | |
| 451 | | AI referral traffic growth | 3x+ YoY from September 2024 to September 2025 | Similarweb, 2026-05-28 | |
| 452 | | Gemini referral trend | About 18% share, +237% YoY | Similarweb, 2026-05-28 | |
| 453 | | ChatGPT referral trend | Share slid from about 87% to the high-60s | Similarweb, 2026-05-28 | |
| 454 | |
| 455 | |
| 456 | |
| 457 | ## AI Crawler Checklist |
| 458 | |
| 459 | | Check | Pass | Fail | |
| 460 | |-------|------|------| |
| 461 | | robots.txt allows AI crawlers | All major bots listed with `Allow: /` | Missing entries or `Disallow: /` | |
| 462 | | Cloudflare AI settings reviewed | AI crawlers explicitly allowed in dashboard | Default block left in place | |
| 463 | | llms.txt treated as optional | Not required for Google AI visibility | Treating a missing file as a blocker | |
| 464 | | Content in HTML source | `curl` returns full content | Empty divs, JS-only rendering | |
| 465 | | TTFB under 200ms | Measured from CDN edge | Over 600ms increases crawl or extraction risk | |
| 466 | | Schema in HTML source | Standard Article, Person, Organization JSON-LD in source HTML | Special AI-only schema or JS-injected schema | |
| 467 | | Sitemap.xml accessible | Valid XML, all blog URLs included | Missing or returns 404 | |
| 468 | | No Cloudflare challenge on bot UA | 200 status code | 403 or challenge page | |
| 469 |
Discussion
Alternatives
Browse more free Claude skills or everything in Marketing.