Technical AI Visibility: Crawler Access & Rendering skill

- robots.txt Template for AI Crawlers

by AgriciDaniel·MIT license·★ 2,219 Stars on the repo·GitHub ↗

Use now

Files of Technical AI Visibility: Crawler Access & Rendering

AgriciDaniel/main1 file
ai-crawler-guide.md
Show the full text469 lines

Technical AI Visibility: Crawler Access & Rendering

Contents

robots.txt Template for AI Crawlers

Allow documented AI crawlers explicitly when you want access. For compliant crawlers, an absent Disallow usually means allowed; explicit Allow rules are optional documentation and help teams audit intent.

# ===========================================
# AI Search & LLM Crawlers: Explicitly Allow
# ===========================================

# OpenAI
User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

# Anthropic documented crawler families
User-agent: ClaudeBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

# Deprecated Anthropic strings (kept for legacy compatibility):
# User-agent: Claude-Web
# User-agent: anthropic-ai

# Google AI product token (Gemini/Vertex training and non-Search grounding controls)
# Google Search AI features use Googlebot plus preview controls:
# https://developers.google.com/search/docs/appearance/ai-features
User-agent: Google-Extended
Allow: /

# Perplexity
User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

# Meta
User-agent: Meta-ExternalAgent
Allow: /

# ByteDance
User-agent: Bytespider
Allow: /

# Google AI agents (Project Mariner)
User-agent: Google-Agent
Allow: /

# DuckDuckGo AI
User-agent: DuckAssistBot
Allow: /

# Apple (Siri, Apple Intelligence)
User-agent: Applebot-Extended
Allow: /

# Amazon (Alexa, product search)
User-agent: Amazonbot
Allow: /

# You.com
User-agent: YouBot
Allow: /

# Phind (developer search)
User-agent: PhindBot
Allow: /

# Exa (AI-native search engine)
User-agent: ExaBot
Allow: /

# Common Crawl (used by many AI models)
User-agent: CCBot
Allow: /

# ===========================================
# Traditional Search Engines
# ===========================================

User-agent: Googlebot
Allow: /

User-agent: Bingbot
Allow: /

User-agent: *
Allow: /

# ===========================================
# Sitemap
# ===========================================
Sitemap: https://example.com/sitemap.xml
Crawler Identification Reference

Providers expose different crawler classes. Some split training, search indexing, and user-triggered retrieval; others publish only one bot or a product token. Blocking a documented search/indexing bot can reduce visibility in that platform's answers. User-triggered retrieval may not fully respect robots.txt. OpenAI bot details: https://platform.openai.com/docs/bots.

Crawler Operator Type Respects robots.txt
GPTBot OpenAI Training Yes
OAI-SearchBot OpenAI Search indexing Yes
ChatGPT-User OpenAI User-triggered retrieval Not guaranteed
ClaudeBot Anthropic Training Yes
Claude-SearchBot Anthropic Search indexing Yes
Claude-User Anthropic User retrieval Yes
Claude-Web Anthropic Deprecated -
anthropic-ai Anthropic Deprecated -
Google-Extended Google Gemini/Vertex training and some non-Search grounding controls; not Search AI inclusion Yes
Google-Agent Google Project Mariner agentic (2026) Yes
PerplexityBot Perplexity Search indexing Yes
Perplexity-User Perplexity User retrieval Partial
Applebot-Extended Apple Apple Intelligence training Yes
Meta-ExternalAgent Meta High-volume data collection Yes
Bytespider ByteDance Training/indexing Partial (documented issues)
Amazonbot Amazon Alexa / product search Yes
DuckAssistBot DuckDuckGo DuckAssist AI answers Yes
YouBot You.com AI search engine Yes
PhindBot Phind Developer-focused AI search Yes
ExaBot Exa Neural search engine Yes
CCBot Common Crawl Open dataset (used by many LLMs) Yes
robots.txt Strategy by Bot Type

Treat each bot category differently based on your goals:

  • Training/product tokens (GPTBot, ClaudeBot, CCBot, Google-Extended): Your choice. Blocking affects training or non-Search product use as documented by each provider, but Google-Extended does not control Google Search AI inclusion.
  • Search/indexing bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot): Allow these. Blocking means your content won't appear in ChatGPT, Claude, or Perplexity answers.
  • Retrieval bots (ChatGPT-User, Perplexity-User): May not fully respect robots.txt. These are triggered by live user queries and may fetch content regardless of directives.

Cloudflare AI Crawl Control: CRITICAL

Since July 2025, Cloudflare blocks AI crawlers by default on new domains. This is the single most common reason blogs are invisible to AI systems despite having correct robots.txt configuration.

How to Fix
  1. Log in to Cloudflare dashboard
  2. Navigate to Security > Bots > AI Crawlers
  3. Review the list of AI crawlers
  4. Toggle "Allow" for each AI crawler you want to permit
  5. Save changes
What Cloudflare Blocks by Default
Crawler or token Default Status (New Domains)
GPTBot Blocked
ClaudeBot Blocked
PerplexityBot Blocked
CCBot Blocked
Google-Extended Blocked
Applebot-Extended Allowed
Googlebot Allowed (not an AI crawler)
Verification

After updating Cloudflare settings, verify access:

# Simulate GPTBot user-agent
curl -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" https://yourdomain.com/blog/test-post | head -50

# Check for Cloudflare block page (403 or challenge page)
curl -s -o /dev/null -w "%{http_code}" -A "Mozilla/5.0 (compatible; ClaudeBot/1.0)" https://yourdomain.com/

If you get a 403 or an HTML page with "Cloudflare" in it, the crawler is blocked.


Google Gen-AI Guidance

Google Search Central's AI features guidance says optimization for AI Overviews and AI Mode is normal SEO: https://developers.google.com/search/docs/appearance/ai-features. Google does not require special AI schema or llms.txt for Search AI features. Use crawlable HTML, standard Article schema with author Person and publisher Organization, clear source attribution, helpful trustworthy content, and fast server responses.


llms.txt Implementation

The llms.txt standard (proposed by llmstxt.org, Sep 2024) provides a machine-readable summary of your site for LLMs. Place at site root: https://example.com/llms.txt.

Important caveat: Google's current stance is no llms.txt needed for AI Overviews or AI Mode per Google Search Central's AI features guidance. No major AI platform has confirmed relying on it. Treat it as an optional site inventory for non-Google tools, not a ranking, indexing, or citation requirement.

Specification
  • Plain text file, UTF-8
  • Under 10KB total
  • Structured list of important URLs with brief descriptions
  • Helps LLMs understand site structure and find authoritative content
Template
# Example Blog

> A blog about modern web development, SEO, and content strategy.

## Main Pages

- [Home](https://example.com/): Main landing page with latest articles
- [About](https://example.com/about): Company information and mission
- [Blog](https://example.com/blog): All published articles

## Popular Articles

- [Complete Guide to Technical SEO in 2026](https://example.com/blog/technical-seo-guide): Comprehensive technical SEO guide covering Core Web Vitals, crawlability, and schema markup.
- [How AI Overviews Changed Search](https://example.com/blog/ai-overviews-impact): Data-driven analysis of AI Overview impact on organic traffic with case studies.
- [Content Strategy for B2B SaaS](https://example.com/blog/b2b-saas-content-strategy): Framework for building a content program that drives pipeline.

## Topic Clusters

- [SEO](https://example.com/topics/seo): All articles about search engine optimization
- [Content Strategy](https://example.com/topics/content-strategy): Content planning and execution
- [Web Development](https://example.com/topics/web-development): Frontend and backend development guides

## Authors

- [Sarah Chen](https://example.com/author/sarah-chen): Content strategist, B2B SaaS specialist
- [Marcus Rivera](https://example.com/author/marcus-rivera): Senior frontend engineer, React expert
Key Rules
  • Do not exceed 10KB (LLMs may truncate or ignore larger files)
  • Use markdown-style links: [Title](URL): Description
  • Include only your most important and highest-quality pages
  • Update when you publish significant new content
  • This is NOT a sitemap replacement: it supplements sitemap.xml
  • Do not treat a missing llms.txt file as an AI visibility blocker

Server-Side Rendering Requirements

Standard non-Google AI crawlers generally should be assumed not to execute JavaScript unless their documentation says otherwise. Content rendered only via client-side JavaScript is risky for AI visibility; render important blog content into initial HTML and verify per crawler.

Rendering Strategy Ranking
Strategy AI Visibility Performance Recommendation
SSG (Static Site Generation) Best Best Preferred for blogs
SSR (Server-Side Rendering) Excellent Good Good for dynamic content
ISR (Incremental Static Regeneration) Excellent Good Good for large sites
CSR (Client-Side Rendering) None Poor for crawlers Never use for content
JavaScript Execution by Crawler
Crawler Executes JavaScript Renders Pages
GPTBot No No
OAI-SearchBot No No
ChatGPT-User No No
ClaudeBot No No
Claude-SearchBot No No
Claude-User No No
PerplexityBot No No
Perplexity-User No No
Meta-ExternalAgent No No
Bytespider No No
Amazonbot No No
CCBot No No
Googlebot Yes Yes
AppleBot Yes Yes
OpenAI agentic browsing surfaces Yes Yes
Google-Agent (agentic) Yes Yes
Vercel Findings

Vercel analyzed 500M+ GPTBot fetches and found zero evidence of JavaScript execution. GPTBot reads raw HTML only. Content loaded via React hydration, Vue mounting, or any client-side framework is completely invisible.

Exception: Agentic Tools

Standard AI crawlers generally do not execute JavaScript. However, agentic tools are different:

  • OpenAI agentic browsing surfaces: Full JS rendering may be available depending on product mode.
  • Google-Agent / Project Mariner (Google, 2026): Operates through Chrome with full rendering.

These are user-directed agents, not automated crawlers. They can see JS-rendered content, but they do not replace the need for SSR - standard crawlers still dominate citation indexing.


Passage-Level Extractability

Crawler access gets a page into the candidate set. Citation selection depends on whether the page contains self-contained answer passages AI systems can extract. Target 120-180 word passages that answer one question without relying on the surrounding article.

Under each H2, start with an approximately 50-word direct-answer sentence that gives the answer, the year, the named entity, and the source attribution. Follow with specific entities, dates, original examples, and first-hand Experience markers. A clean passage can earn an AI Overview citation even when the full page is not cited.

AI Overviews also began highlighting links from a user's subscribed publications in 2026, so publisher trust and subscriptions can affect which citations users notice (Nieman Lab, 2026-05).


Performance Requirements

AI retrieval systems have practical latency budgets. Slow sites may reduce crawl, fetch, and extraction reliability before content quality is evaluated.

Note: The thresholds below are industry best practices and observations from SEO tooling (Discovered Labs, Prerender.io, Kevin Indig). They are NOT officially published specifications from OpenAI, Anthropic, or Perplexity. Treat as directional targets, not guaranteed cutoffs.

Thresholds
Metric Target Risk threshold Consequence
TTFB (Time to First Byte) < 200ms > 600ms May reduce crawl or extraction reliability
Full page load (HTML) < 500ms > 1,000ms May reduce crawl frequency
Response size (HTML) < 200KB > 500KB May cause partial content extraction
Optimization Priorities
  1. Use a CDN: Content must be served from edge locations
  2. Enable compression: gzip or Brotli for all text responses
  3. Minimize HTML bloat: Remove unused CSS/JS from HTML response
  4. Cache aggressively: Static pages should have long cache headers
  5. Pre-render: Use SSG or SSR, never CSR for content pages

Testing AI Crawler Visibility

Quick Test: See What AI Crawlers See
# Basic: view raw HTML (what all AI crawlers receive)
curl -s https://yourdomain.com/blog/your-post | head -200

# Check if main content is in HTML source
curl -s https://yourdomain.com/blog/your-post | grep -c "<article"

# Check for JS-only rendering indicators
curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"__next\""
curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"root\""
curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"app\""

# If the above returns content in a <noscript> tag or empty divs,
# your content is behind JS and invisible to AI crawlers.
Full Crawler Simulation
# Simulate GPTBot
curl -s -H "User-Agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" \
  https://yourdomain.com/blog/your-post > /tmp/gptbot-view.html

# Simulate ClaudeBot
curl -s -H "User-Agent: Mozilla/5.0 (compatible; ClaudeBot/1.0; +https://claudebot.ai)" \
  https://yourdomain.com/blog/your-post > /tmp/claudebot-view.html

# Check if content exists
wc -l /tmp/gptbot-view.html
grep -c "your-expected-heading-text" /tmp/gptbot-view.html
Red Flags (Content Invisible to AI)
Indicator What It Means
Empty <div id="root"></div> React CSR: content loads via JS only
Empty <div id="__next"></div> without SSR/RSC/static output Next.js App Router or Pages Router shipping content client-side only
<noscript> contains the content Content explicitly hidden from non-JS clients
<script> tags contain all content as JSON Data fetched client-side, not in HTML
HTML under 5KB for a full blog post Content not rendered server-side
Next.js App Router Guidance
  • Prefer static rendering for blog routes. Use Server Components for article content and generateStaticParams() for known slugs.
  • Use ISR for large blogs when content changes after build. Keep the article body in server-rendered HTML.
  • Use dynamic rendering only when the page genuinely depends on request-time data. Do not move the article body behind client-only data fetching.
  • generateMetadata() should emit canonical, Open Graph, and Article metadata server-side.

AI Crawler Traffic Growth

Traffic from AI crawlers is growing exponentially. Sites that block or fail to serve these crawlers are losing compounding visibility.

Metric Value Source
GPTBot traffic growth +305% YoY Cloudflare Radar, 2025
PerplexityBot traffic growth +157,490% YoY Cloudflare Radar, 2025
AI crawling volume overall +32% YoY Cloudflare, 2025
Top 10 domains' citation share 46% of all ChatGPT citations per topic Growth Memo, Mar 2026
AI referral traffic share Small but fastest-growing; no standardized total-web share Similarweb, 2026-05-28
AI referral traffic growth 3x+ YoY from September 2024 to September 2025 Similarweb, 2026-05-28
Gemini referral trend About 18% share, +237% YoY Similarweb, 2026-05-28
ChatGPT referral trend Share slid from about 87% to the high-60s Similarweb, 2026-05-28

AI Crawler Checklist

Check Pass Fail
robots.txt allows AI crawlers All major bots listed with Allow: / Missing entries or Disallow: /
Cloudflare AI settings reviewed AI crawlers explicitly allowed in dashboard Default block left in place
llms.txt treated as optional Not required for Google AI visibility Treating a missing file as a blocker
Content in HTML source curl returns full content Empty divs, JS-only rendering
TTFB under 200ms Measured from CDN edge Over 600ms increases crawl or extraction risk
Schema in HTML source Standard Article, Person, Organization JSON-LD in source HTML Special AI-only schema or JS-injected schema
Sitemap.xml accessible Valid XML, all blog URLs included Missing or returns 404
No Cloudflare challenge on bot UA 200 status code 403 or challenge page
1# Technical AI Visibility: Crawler Access & Rendering
2 
3## Contents
4 
5- [robots.txt Template for AI Crawlers](#robotstxt-template-for-ai-crawlers)
6- [Cloudflare AI Crawl Control: CRITICAL](#cloudflare-ai-crawl-control-critical)
7- [Google Gen-AI Guidance](#google-gen-ai-guidance)
8- [llms.txt Implementation](#llmstxt-implementation)
9- [Server-Side Rendering Requirements](#server-side-rendering-requirements)
10- [Passage-Level Extractability](#passage-level-extractability)
11- [Performance Requirements](#performance-requirements)
12- [Testing AI Crawler Visibility](#testing-ai-crawler-visibility)
13- [AI Crawler Traffic Growth](#ai-crawler-traffic-growth)
14- [AI Crawler Checklist](#ai-crawler-checklist)
15 
16## robots.txt Template for AI Crawlers
17 
18Allow documented AI crawlers explicitly when you want access. For compliant
19crawlers, an absent `Disallow` usually means allowed; explicit `Allow` rules are
20optional documentation and help teams audit intent.
21 
22```
23# ===========================================
24# AI Search & LLM Crawlers: Explicitly Allow
25# ===========================================
26 
27# OpenAI
28User-agent: GPTBot
29Allow: /
30 
31User-agent: OAI-SearchBot
32Allow: /
33 
34User-agent: ChatGPT-User
35Allow: /
36 
37# Anthropic documented crawler families
38User-agent: ClaudeBot
39Allow: /
40 
41User-agent: Claude-SearchBot
42Allow: /
43 
44User-agent: Claude-User
45Allow: /
46 
47# Deprecated Anthropic strings (kept for legacy compatibility):
48# User-agent: Claude-Web
49# User-agent: anthropic-ai
50 
51# Google AI product token (Gemini/Vertex training and non-Search grounding controls)
52# Google Search AI features use Googlebot plus preview controls:
53# https://developers.google.com/search/docs/appearance/ai-features
54User-agent: Google-Extended
55Allow: /
56 
57# Perplexity
58User-agent: PerplexityBot
59Allow: /
60 
61User-agent: Perplexity-User
62Allow: /
63 
64# Meta
65User-agent: Meta-ExternalAgent
66Allow: /
67 
68# ByteDance
69User-agent: Bytespider
70Allow: /
71 
72# Google AI agents (Project Mariner)
73User-agent: Google-Agent
74Allow: /
75 
76# DuckDuckGo AI
77User-agent: DuckAssistBot
78Allow: /
79 
80# Apple (Siri, Apple Intelligence)
81User-agent: Applebot-Extended
82Allow: /
83 
84# Amazon (Alexa, product search)
85User-agent: Amazonbot
86Allow: /
87 
88# You.com
89User-agent: YouBot
90Allow: /
91 
92# Phind (developer search)
93User-agent: PhindBot
94Allow: /
95 
96# Exa (AI-native search engine)
97User-agent: ExaBot
98Allow: /
99 
100# Common Crawl (used by many AI models)
101User-agent: CCBot
102Allow: /
103 
104# ===========================================
105# Traditional Search Engines
106# ===========================================
107 
108User-agent: Googlebot
109Allow: /
110 
111User-agent: Bingbot
112Allow: /
113 
114User-agent: *
115Allow: /
116 
117# ===========================================
118# Sitemap
119# ===========================================
120Sitemap: https://example.com/sitemap.xml
121```
122 
123### Crawler Identification Reference
124 
125Providers expose different crawler classes. Some split training, search indexing,
126and user-triggered retrieval; others publish only one bot or a product token.
127Blocking a documented search/indexing bot can reduce visibility in that platform's
128answers. User-triggered retrieval may not fully respect robots.txt.
129OpenAI bot details: https://platform.openai.com/docs/bots.
130 
131| Crawler | Operator | Type | Respects robots.txt |
132|---------|----------|------|---------------------|
133| GPTBot | OpenAI | Training | Yes |
134| OAI-SearchBot | OpenAI | Search indexing | Yes |
135| ChatGPT-User | OpenAI | User-triggered retrieval | Not guaranteed |
136| ClaudeBot | Anthropic | Training | Yes |
137| Claude-SearchBot | Anthropic | Search indexing | Yes |
138| Claude-User | Anthropic | User retrieval | Yes |
139| ~~Claude-Web~~ | Anthropic | Deprecated | - |
140| ~~anthropic-ai~~ | Anthropic | Deprecated | - |
141| Google-Extended | Google | Gemini/Vertex training and some non-Search grounding controls; not Search AI inclusion | Yes |
142| Google-Agent | Google | Project Mariner agentic (2026) | Yes |
143| PerplexityBot | Perplexity | Search indexing | Yes |
144| Perplexity-User | Perplexity | User retrieval | Partial |
145| Applebot-Extended | Apple | Apple Intelligence training | Yes |
146| Meta-ExternalAgent | Meta | High-volume data collection | Yes |
147| Bytespider | ByteDance | Training/indexing | Partial (documented issues) |
148| Amazonbot | Amazon | Alexa / product search | Yes |
149| DuckAssistBot | DuckDuckGo | DuckAssist AI answers | Yes |
150| YouBot | You.com | AI search engine | Yes |
151| PhindBot | Phind | Developer-focused AI search | Yes |
152| ExaBot | Exa | Neural search engine | Yes |
153| CCBot | Common Crawl | Open dataset (used by many LLMs) | Yes |
154 
155### robots.txt Strategy by Bot Type
156 
157Treat each bot category differently based on your goals:
158- **Training/product tokens** (GPTBot, ClaudeBot, CCBot, Google-Extended): Your
159 choice. Blocking affects training or non-Search product use as documented by
160 each provider, but Google-Extended does not control Google Search AI inclusion.
161- **Search/indexing bots** (OAI-SearchBot, Claude-SearchBot, PerplexityBot): **Allow these.**
162 Blocking means your content won't appear in ChatGPT, Claude, or Perplexity answers.
163- **Retrieval bots** (ChatGPT-User, Perplexity-User): May not fully respect robots.txt. These
164 are triggered by live user queries and may fetch content regardless of directives.
165 
166---
167 
168## Cloudflare AI Crawl Control: CRITICAL
169 
170**Since July 2025, Cloudflare blocks AI crawlers by default on new domains.**
171This is the single most common reason blogs are invisible to AI systems despite
172having correct robots.txt configuration.
173 
174### How to Fix
175 
1761. Log in to Cloudflare dashboard
1772. Navigate to **Security > Bots > AI Crawlers**
1783. Review the list of AI crawlers
1794. **Toggle "Allow" for each AI crawler you want to permit**
1805. Save changes
181 
182### What Cloudflare Blocks by Default
183 
184| Crawler or token | Default Status (New Domains) |
185|------------------|------------------------------|
186| GPTBot | Blocked |
187| ClaudeBot | Blocked |
188| PerplexityBot | Blocked |
189| CCBot | Blocked |
190| Google-Extended | Blocked |
191| Applebot-Extended | Allowed |
192| Googlebot | Allowed (not an AI crawler) |
193 
194### Verification
195 
196After updating Cloudflare settings, verify access:
197 
198```bash
199# Simulate GPTBot user-agent
200curl -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" https://yourdomain.com/blog/test-post | head -50
201 
202# Check for Cloudflare block page (403 or challenge page)
203curl -s -o /dev/null -w "%{http_code}" -A "Mozilla/5.0 (compatible; ClaudeBot/1.0)" https://yourdomain.com/
204```
205 
206If you get a 403 or an HTML page with "Cloudflare" in it, the crawler is blocked.
207 
208---
209 
210## Google Gen-AI Guidance
211 
212Google Search Central's AI features guidance says optimization for AI Overviews
213and AI Mode is normal SEO: https://developers.google.com/search/docs/appearance/ai-features.
214Google does not require special AI schema or llms.txt for Search AI features.
215Use crawlable HTML, standard Article schema with author Person and publisher
216Organization, clear source attribution, helpful trustworthy content, and fast
217server responses.
218 
219---
220 
221## llms.txt Implementation
222 
223The `llms.txt` standard (proposed by llmstxt.org, Sep 2024) provides a machine-readable
224summary of your site for LLMs. Place at site root: `https://example.com/llms.txt`.
225 
226**Important caveat:** Google's current stance is no llms.txt needed for AI
227Overviews or AI Mode per Google Search Central's AI features guidance. No major
228AI platform has confirmed relying on it. Treat it as an optional site inventory
229for non-Google tools, not a ranking, indexing, or citation requirement.
230 
231### Specification
232 
233- Plain text file, UTF-8
234- Under 10KB total
235- Structured list of important URLs with brief descriptions
236- Helps LLMs understand site structure and find authoritative content
237 
238### Template
239 
240```
241# Example Blog
242 
243> A blog about modern web development, SEO, and content strategy.
244 
245## Main Pages
246 
247- [Home](https://example.com/): Main landing page with latest articles
248- [About](https://example.com/about): Company information and mission
249- [Blog](https://example.com/blog): All published articles
250 
251## Popular Articles
252 
253- [Complete Guide to Technical SEO in 2026](https://example.com/blog/technical-seo-guide): Comprehensive technical SEO guide covering Core Web Vitals, crawlability, and schema markup.
254- [How AI Overviews Changed Search](https://example.com/blog/ai-overviews-impact): Data-driven analysis of AI Overview impact on organic traffic with case studies.
255- [Content Strategy for B2B SaaS](https://example.com/blog/b2b-saas-content-strategy): Framework for building a content program that drives pipeline.
256 
257## Topic Clusters
258 
259- [SEO](https://example.com/topics/seo): All articles about search engine optimization
260- [Content Strategy](https://example.com/topics/content-strategy): Content planning and execution
261- [Web Development](https://example.com/topics/web-development): Frontend and backend development guides
262 
263## Authors
264 
265- [Sarah Chen](https://example.com/author/sarah-chen): Content strategist, B2B SaaS specialist
266- [Marcus Rivera](https://example.com/author/marcus-rivera): Senior frontend engineer, React expert
267```
268 
269### Key Rules
270 
271- Do not exceed 10KB (LLMs may truncate or ignore larger files)
272- Use markdown-style links: `[Title](URL): Description`
273- Include only your most important and highest-quality pages
274- Update when you publish significant new content
275- This is NOT a sitemap replacement: it supplements sitemap.xml
276- Do not treat a missing llms.txt file as an AI visibility blocker
277 
278---
279 
280## Server-Side Rendering Requirements
281 
282Standard non-Google AI crawlers generally should be assumed not to execute
283JavaScript unless their documentation says otherwise. Content rendered only via
284client-side JavaScript is risky for AI visibility; render important blog content
285into initial HTML and verify per crawler.
286 
287### Rendering Strategy Ranking
288 
289| Strategy | AI Visibility | Performance | Recommendation |
290|----------|--------------|-------------|----------------|
291| **SSG** (Static Site Generation) | Best | Best | Preferred for blogs |
292| **SSR** (Server-Side Rendering) | Excellent | Good | Good for dynamic content |
293| **ISR** (Incremental Static Regeneration) | Excellent | Good | Good for large sites |
294| **CSR** (Client-Side Rendering) | None | Poor for crawlers | Never use for content |
295 
296### JavaScript Execution by Crawler
297 
298| Crawler | Executes JavaScript | Renders Pages |
299|---------|-------------------|---------------|
300| GPTBot | No | No |
301| OAI-SearchBot | No | No |
302| ChatGPT-User | No | No |
303| ClaudeBot | No | No |
304| Claude-SearchBot | No | No |
305| Claude-User | No | No |
306| PerplexityBot | No | No |
307| Perplexity-User | No | No |
308| Meta-ExternalAgent | No | No |
309| Bytespider | No | No |
310| Amazonbot | No | No |
311| CCBot | No | No |
312| **Googlebot** | **Yes** | **Yes** |
313| **AppleBot** | **Yes** | **Yes** |
314| **OpenAI agentic browsing surfaces** | **Yes** | **Yes** |
315| **Google-Agent** (agentic) | **Yes** | **Yes** |
316 
317### Vercel Findings
318 
319Vercel analyzed 500M+ GPTBot fetches and found **zero evidence of JavaScript
320execution**. GPTBot reads raw HTML only. Content loaded via React hydration,
321Vue mounting, or any client-side framework is completely invisible.
322 
323### Exception: Agentic Tools
324 
325Standard AI crawlers generally do not execute JavaScript. However, **agentic tools** are different:
326- **OpenAI agentic browsing surfaces**: Full JS rendering may be available depending on product mode.
327- **Google-Agent / Project Mariner** (Google, 2026): Operates through Chrome with full rendering.
328 
329These are user-directed agents, not automated crawlers. They can see JS-rendered content,
330but they do not replace the need for SSR - standard crawlers still dominate citation indexing.
331 
332---
333 
334## Passage-Level Extractability
335 
336Crawler access gets a page into the candidate set. Citation selection depends on
337whether the page contains self-contained answer passages AI systems can extract.
338Target 120-180 word passages that answer one question without relying on the
339surrounding article.
340 
341Under each H2, start with an approximately 50-word direct-answer sentence that
342gives the answer, the year, the named entity, and the source attribution. Follow
343with specific entities, dates, original examples, and first-hand Experience
344markers. A clean passage can earn an AI Overview citation even when the full
345page is not cited.
346 
347AI Overviews also began highlighting links from a user's subscribed
348publications in 2026, so publisher trust and subscriptions can affect which
349citations users notice (Nieman Lab, 2026-05).
350 
351---
352 
353## Performance Requirements
354 
355AI retrieval systems have practical latency budgets. Slow sites may reduce crawl,
356fetch, and extraction reliability before content quality is evaluated.
357 
358**Note:** The thresholds below are industry best practices and observations from SEO tooling
359(Discovered Labs, Prerender.io, Kevin Indig). They are NOT officially published specifications
360from OpenAI, Anthropic, or Perplexity. Treat as directional targets, not guaranteed cutoffs.
361 
362### Thresholds
363 
364| Metric | Target | Risk threshold | Consequence |
365|--------|--------|----------------|-------------|
366| TTFB (Time to First Byte) | < 200ms | > 600ms | May reduce crawl or extraction reliability |
367| Full page load (HTML) | < 500ms | > 1,000ms | May reduce crawl frequency |
368| Response size (HTML) | < 200KB | > 500KB | May cause partial content extraction |
369 
370### Optimization Priorities
371 
3721. **Use a CDN**: Content must be served from edge locations
3732. **Enable compression**: gzip or Brotli for all text responses
3743. **Minimize HTML bloat**: Remove unused CSS/JS from HTML response
3754. **Cache aggressively**: Static pages should have long cache headers
3765. **Pre-render**: Use SSG or SSR, never CSR for content pages
377 
378---
379 
380## Testing AI Crawler Visibility
381 
382### Quick Test: See What AI Crawlers See
383 
384```bash
385# Basic: view raw HTML (what all AI crawlers receive)
386curl -s https://yourdomain.com/blog/your-post | head -200
387 
388# Check if main content is in HTML source
389curl -s https://yourdomain.com/blog/your-post | grep -c "<article"
390 
391# Check for JS-only rendering indicators
392curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"__next\""
393curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"root\""
394curl -s https://yourdomain.com/blog/your-post | grep -c "id=\"app\""
395 
396# If the above returns content in a <noscript> tag or empty divs,
397# your content is behind JS and invisible to AI crawlers.
398```
399 
400### Full Crawler Simulation
401 
402```bash
403# Simulate GPTBot
404curl -s -H "User-Agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" \
405 https://yourdomain.com/blog/your-post > /tmp/gptbot-view.html
406 
407# Simulate ClaudeBot
408curl -s -H "User-Agent: Mozilla/5.0 (compatible; ClaudeBot/1.0; +https://claudebot.ai)" \
409 https://yourdomain.com/blog/your-post > /tmp/claudebot-view.html
410 
411# Check if content exists
412wc -l /tmp/gptbot-view.html
413grep -c "your-expected-heading-text" /tmp/gptbot-view.html
414```
415 
416### Red Flags (Content Invisible to AI)
417 
418| Indicator | What It Means |
419|-----------|---------------|
420| Empty `<div id="root"></div>` | React CSR: content loads via JS only |
421| Empty `<div id="__next"></div>` without SSR/RSC/static output | Next.js App Router or Pages Router shipping content client-side only |
422| `<noscript>` contains the content | Content explicitly hidden from non-JS clients |
423| `<script>` tags contain all content as JSON | Data fetched client-side, not in HTML |
424| HTML under 5KB for a full blog post | Content not rendered server-side |
425 
426### Next.js App Router Guidance
427 
428- Prefer static rendering for blog routes. Use Server Components for article
429 content and `generateStaticParams()` for known slugs.
430- Use ISR for large blogs when content changes after build. Keep the article body
431 in server-rendered HTML.
432- Use dynamic rendering only when the page genuinely depends on request-time data.
433 Do not move the article body behind client-only data fetching.
434- `generateMetadata()` should emit canonical, Open Graph, and Article metadata
435 server-side.
436 
437---
438 
439## AI Crawler Traffic Growth
440 
441Traffic from AI crawlers is growing exponentially. Sites that block or fail
442to serve these crawlers are losing compounding visibility.
443 
444| Metric | Value | Source |
445|--------|-------|--------|
446| GPTBot traffic growth | +305% YoY | Cloudflare Radar, 2025 |
447| PerplexityBot traffic growth | +157,490% YoY | Cloudflare Radar, 2025 |
448| AI crawling volume overall | +32% YoY | Cloudflare, 2025 |
449| Top 10 domains' citation share | 46% of all ChatGPT citations per topic | Growth Memo, Mar 2026 |
450| AI referral traffic share | Small but fastest-growing; no standardized total-web share | Similarweb, 2026-05-28 |
451| AI referral traffic growth | 3x+ YoY from September 2024 to September 2025 | Similarweb, 2026-05-28 |
452| Gemini referral trend | About 18% share, +237% YoY | Similarweb, 2026-05-28 |
453| ChatGPT referral trend | Share slid from about 87% to the high-60s | Similarweb, 2026-05-28 |
454 
455---
456 
457## AI Crawler Checklist
458 
459| Check | Pass | Fail |
460|-------|------|------|
461| robots.txt allows AI crawlers | All major bots listed with `Allow: /` | Missing entries or `Disallow: /` |
462| Cloudflare AI settings reviewed | AI crawlers explicitly allowed in dashboard | Default block left in place |
463| llms.txt treated as optional | Not required for Google AI visibility | Treating a missing file as a blocker |
464| Content in HTML source | `curl` returns full content | Empty divs, JS-only rendering |
465| TTFB under 200ms | Measured from CDN edge | Over 600ms increases crawl or extraction risk |
466| Schema in HTML source | Standard Article, Person, Organization JSON-LD in source HTML | Special AI-only schema or JS-injected schema |
467| Sitemap.xml accessible | Valid XML, all blog URLs included | Missing or returns 404 |
468| No Cloudflare challenge on bot UA | 200 status code | 403 or challenge page |
469 

Discussion

Alternatives