Technical SEO Audit

Technical SEO audit across 9 categories: crawlability, indexability, security, URL structure, mobile, Core Web Vitals, structured data, JavaScript rendering, and IndexNow protocol.

How to use it

  1. Hit Copy SKILL.md — or use the Claude Code line below to get every file.
  2. Claude: ⋯ → Download .md, then Customize → Skills → Add → Upload skill.
    ChatGPT: make a Project and paste it into Instructions.
    Neither? Paste it at the top of a new chat — it works for that chat.
  3. Describe your job in plain words. The AI follows the skill from there.
Claude Code — installs the whole folder, not just SKILL.md
npx degit AgriciDaniel/claude-seo/skills/seo-technical#main ~/.claude/skills/seo-technical

For one project only, change the path to .claude/skills/seo-technical. This skill also uses robots.txt, sitemap_discovery.py, googlebot.json, common-crawlers.json, Next.js, llms.txt — copying SKILL.md alone won't be enough. See the folder on GitHub.

Not working?
  • Check which app you pasted it into — the steps above name the right one.
  • Some skills need the paid tier of Claude or ChatGPT.
Step-by-step guide with screenshots · Ask in the forum

Paste into Claude, ChatGPT or Cursor.

Show the full text287 lines
seo-technical/SKILL.md287 lines16.9 KBpushed 11d agoRawView on GitHub

Technical SEO Audit

Categories

1. Crawlability

  • robots.txt: exists, valid, not blocking important resources
  • XML sitemap: run "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run sitemap_discovery.py <url> --json; require a valid entry in found, and report stale or unsafe robots.txt declarations separately from working fallback locations
  • Noindex tags: intentional vs accidental
  • Crawl depth: important pages within 3 clicks of homepage
  • JavaScript rendering: check if critical content requires JS execution
  • Crawl budget: for large sites (>10k pages), efficiency matters
  • Googlebot fetch limits: Googlebot fetches the first 2MB of HTML and first 64MB of a PDF (uncompressed; 15MB is the broader crawler-infra default). Long-standing, not a 2026 change, but inline base64 images, oversized inline CSS/JS, or bloated nav can push critical content/JSON-LD past the cap and out of the index. Keep key content + structured data within the first 2MB.
  • Crawl rate auto-adjusts (backs off on 5xx/slow responses); there is no manual crawl-rate control (the legacy Search Console setting was removed Jan 2024). Influence crawling via sitemaps, server responsiveness, and robots controls.
  • Google's canonical crawling/robots reference moved to developers.google.com/crawling (migrated 2025-11-20); IP-range files relocated to /crawling/ipranges/ and googlebot.json was renamed common-crawlers.json.
  • AMP has no separate ranking advantage. Since 2026-07-01, Google Search sends users directly to publisher-hosted AMP URLs, so do not recommend AMP Cache, AMP Viewer, or signed exchange maintenance. Audit AMP against the same content, action-parity, and quality requirements as other pages.

AI Crawler Management

As of 2025-2026, AI companies actively crawl the web to train models and power AI search. Managing these crawlers via robots.txt is a critical technical SEO consideration.

Known AI crawlers:

Crawler Company robots.txt token Purpose
GPTBot OpenAI GPTBot Model training (NOT ChatGPT Search)
OAI-SearchBot OpenAI OAI-SearchBot ChatGPT Search citability
ChatGPT-User OpenAI ChatGPT-User Real-time browsing (user-triggered)
ClaudeBot Anthropic ClaudeBot Model training (NOT Claude search citability)
Claude-SearchBot Anthropic Claude-SearchBot Claude search-result citability
PerplexityBot Perplexity PerplexityBot Search index + training
Bytespider ByteDance Bytespider Model training
Google-Extended Google Google-Extended Gemini training (NOT search)
Applebot-Extended Apple Applebot-Extended Apple Intelligence training opt-out (NOT Siri/Spotlight/Safari)
CCBot Common Crawl CCBot Open dataset

Key distinctions:

  • Blocking Google-Extended prevents Gemini training use but does NOT affect Google Search indexing or AI Overviews (those use Googlebot)
  • Blocking GPTBot prevents OpenAI training but does NOT affect ChatGPT Search citability, which is governed by OAI-SearchBot, nor user-triggered browsing (ChatGPT-User). Check OAI-SearchBot for any citability claim; GPTBot status is evidence about training use only
  • Blocking ClaudeBot prevents Anthropic model training but does NOT affect citability in Claude's own search features, which is governed by Claude-SearchBot (per Anthropic's crawler support article). Check Claude-SearchBot for any Claude-search citability claim; ClaudeBot status is evidence about training use only
  • Blocking Applebot-Extended opts out of Apple Intelligence / generative-model training use but does NOT affect discoverability via Siri, Spotlight, or Safari, which follows Applebot (per Apple's support article); Applebot-Extended does not itself crawl
  • ~3-5% of websites now use AI-specific robots.txt rules

Example, selective AI crawler blocking:

# Allow search indexing, block AI training crawlers
User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Bytespider
Disallow: /

# Allow all other crawlers (including Googlebot for search)
User-agent: *
Allow: /

Recommendation: Consider your AI visibility strategy before blocking. Being cited by AI systems drives brand awareness and referral traffic. Cross-reference the seo-geo skill for the full AI crawler/fetcher taxonomy.

User-triggered fetchers ignore robots.txt by design. Google now documents Google-Agent (Project Mariner, agentic browsing) plus Google-NotebookLM and Google Messages as user-triggered fetchers that cannot be blocked via robots.txt. Use server-side access controls instead. By contrast, Google-Extended and Google-CloudVertexBot obey robots.txt. Emerging: Web Bot Auth (RFC 9421) lets bots authenticate cryptographically via a Signature-Agent header + key directory at agent.bot.goog (used by Google-Agent); reverse-DNS verification remains the fallback.

2. Indexability

  • Canonical tags: self-referencing, no conflicts with noindex
  • Duplicate content: near-duplicates, parameter URLs, www vs non-www
  • Canonicalization fixes can take time: Google may retain corrected pages in a duplicate cluster for up to two weeks while re-evaluating them. Do not interpret an unchanged canonical immediately after a fix as proof that the fix failed.
  • Thin content: pages below minimum word counts per type
  • Pagination: rel=next/prev or load-more pattern
  • Hreflang: correct for multi-language/multi-region sites
  • Index bloat: unnecessary pages consuming crawl budget

3. Security

  • HTTPS: enforced, valid SSL certificate, no mixed content
  • Security headers:
    • Content-Security-Policy (CSP)
    • Strict-Transport-Security (HSTS)
    • X-Frame-Options
    • X-Content-Type-Options
    • Referrer-Policy
  • HSTS preload: check preload list inclusion for high-security sites
  • Back-button hijacking (spam-policy violation, malicious practices): flag pages that defeat the Back button via history.pushState/replaceState (including scripts injected by third-party ad/library platforms). Added to Google's spam policies 2026-04-13; enforcement live since 2026-06-15 (manual actions + automated demotions): treat as Critical.

4. URL Structure

  • Clean URLs: descriptive, hyphenated, no query parameters for content
  • Hierarchy: logical folder structure reflecting site architecture
  • Redirects: no chains (max 1 hop), 301 for permanent moves
  • URL length: flag >100 characters
  • Trailing slashes: consistent usage

5. Mobile Optimization & Page Experience

  • Responsive design: viewport meta tag, responsive CSS
  • Touch targets: minimum 48x48px with 8px spacing
  • Font size: minimum 16px base
  • No horizontal scroll
  • Mobile-first indexing: Googlebot Smartphone is the primary crawler (rollout completed 2024). A mobile version is not strictly required (Google says "very strongly recommended"), sites that don't work on mobile can still be indexed, but the real risk is content/parity loss, not hard exclusion.
  • Mobile/desktop content parity (highest-value mobile check): equivalent primary content, matching robots meta tags, matching titles/descriptions, equivalent structured data, crawlable resources; avoid lazy-loading primary content that requires user interaction.
  • Intrusive interstitials / ad density: flag full-page interstitials, standalone consent-redirect pages, persistent blocking dialogs, and excessive/distracting ad density (a named page-experience aspect). Acceptable: small banners, standard CMS/legal dialogs.
  • "Read more" deep links: keep key content immediately visible on load (not behind tabs/accordions), don't hijack scroll on load, and preserve URL hash fragments, content hidden behind expandable sections is less likely to qualify.

Page experience is guidance, not a single ranking system. Only Core Web Vitals feeds ranking directly; HTTPS is a confirmed but lightweight signal (affects <~1% of queries). Relevance can still win even when page experience is sub-par, so don't over-weight security headers. Note: the standalone Page Experience report was removed from Search Console (monitor via the Core Web Vitals + HTTPS reports).

6. Core Web Vitals

  • LCP (Largest Contentful Paint): target <=2.5s
  • INP (Interaction to Next Paint): target <=200ms
    • INP replaced FID on March 12, 2024. FID was removed from Chrome's field-data tools (CrUX API, PageSpeed Insights) on September 9, 2024 (Lighthouse is a lab tool that never reported FID). Do NOT reference FID anywhere.
  • CLS (Cumulative Layout Shift): target <=0.1
  • Evaluation uses 75th percentile of real user data
  • Use PageSpeed Insights API or CrUX data if MCP available

7. Structured Data

  • Detection: JSON-LD (preferred), Microdata, RDFa
  • Validation against Google's supported types
  • See seo-schema skill for full analysis

8. JavaScript Rendering

  • Check if content visible in initial HTML vs requires JS
  • Identify client-side rendered (CSR) vs server-side rendered (SSR)
  • Flag SPA frameworks (React, Vue, Angular) that may cause indexing issues
  • If dynamic rendering is detected, flag it as technical debt rather than a valid setup. Google documents it as "a workaround and not a recommended solution" because of the added complexity and resource cost. See https://developers.google.com/search/docs/crawling-indexing/javascript/dynamic-rendering

Recommended rendering strategy:

Strategy Use Case
SSR Public SEO content, dynamic pages
SSG Static content, blogs, docs
CSR Authenticated / behind-login content only

Preferred frameworks: Next.js, Astro, React Router v7 (Remix), SvelteKit

JavaScript SEO: Canonical & Indexing Guidance (December 2025)

Google updated its JavaScript SEO documentation in December 2025 with critical clarifications:

  1. Canonical conflicts: If a canonical tag in raw HTML differs from one injected by JavaScript, Google may use EITHER one. Ensure canonical tags are identical between server-rendered HTML and JS-rendered output.
  2. noindex with JavaScript: If raw HTML contains <meta name="robots" content="noindex"> but JavaScript removes it, Google MAY still honor the noindex from raw HTML. Serve correct robots directives in the initial HTML response.
  3. Non-200 status codes: Google does NOT render JavaScript on pages returning non-200 HTTP status codes. Any content or meta tags injected via JS on error pages will be invisible to Googlebot.
  4. Structured data in JavaScript: Product, Article, and other structured data injected via JS may face delayed processing. For time-sensitive structured data (especially e-commerce Product markup), include it in the initial server-rendered HTML.

Best practice: Serve critical SEO elements (canonical, meta robots, structured data, title, meta description) in the initial server-rendered HTML rather than relying on JavaScript injection.

9. IndexNow Protocol

  • Check if site supports IndexNow for Bing, Yandex, Naver
  • Supported by search engines other than Google
  • Recommend implementation for faster indexing on non-Google engines

Agent-Friendly Pages & Agentic Browsing

AI agents (not just AI summarizers) increasingly read sites through three channels: vision models on screenshots, raw HTML/DOM, and the accessibility tree (the cleanest signal). Audit criteria: semantic HTML (real <button> and , not ), label associations, interactive target sizing, layout stability across templates, cursor: pointer correctness, live in references/agent-friendly-pages.md.

Google now ships a Lighthouse Agentic Browsing category (default-on since Lighthouse 13.3.0, Chrome 150+; buckets: agent-centric accessibility, CLS + llms.txt, three WebMCP audits). It reports a fractional pass-ratio (X of N), not a 0-100 score, keep that distinct from this skill's own Agent-UX 0-100 heuristic below. Lighthouse 13.4.1 re-enabled the category through the PSI API. It is also available through Lighthouse CLI with --only-categories=agentic-browsing, DevTools, and the PSI web UI. See references/agent-friendly-pages.md.

Audit command

# Render with Playwright + capture accessibility tree, then score
"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run agent_ux_check.py https://example.com --json

The scanner outputs an Agent-UX score (0-100) plus itemized issues:

  • HTML findings: real buttons / anchors, `` widgets, semantic landmarks, inputs without <label for>, inputs without ARIA labels
  • Accessibility tree findings: total nodes, interactive nodes, unnamed interactive elements, role="generic" ratio

The accessibility-tree snapshot uses Chromium's Accessibility.getFullAXTree CDP command through Playwright. To capture the tree without scoring, use "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run render_page.py <url> --a11y-tree --json.

Surface findings as opportunities, not failures; don't gate audits on a sub-100 Agent-UX score. WebMCP origin-trial/sign-up status needs verification, and absence of WebMCP support is still an opportunity, not a defect.

Output

Technical Score: XX/100

Category Breakdown

Category Status Score
Crawlability pass/warn/fail XX/100
Indexability pass/warn/fail XX/100
Security pass/warn/fail XX/100
URL Structure pass/warn/fail XX/100
Mobile pass/warn/fail XX/100
Core Web Vitals pass/warn/fail XX/100
Structured Data pass/warn/fail XX/100
JS Rendering pass/warn/fail XX/100
IndexNow pass/warn/fail XX/100

Critical Issues (fix immediately)

High Priority (fix within 1 week)

Medium Priority (fix within 1 month)

Low Priority (backlog)

DataForSEO Integration (Optional)

If DataForSEO MCP tools are available, use on_page_instant_pages for real page analysis (status codes, page timing, broken links, on-page checks), on_page_lighthouse for Lighthouse audits (performance, accessibility, SEO scores), and domain_analytics_technologies_domain_technologies for technology stack detection.

Google API Integration (Optional)

If Google API credentials are configured, use "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run pagespeed_check.py <url> --json for real PSI + CrUX field data (replaces lab-only CWV estimates), "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run crux_history.py <url> --json for 25-week CWV trends, and "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run gsc_inspect.py <url> --json for real indexation status per URL.

Auditing a Local or Private Host

url_safety refuses loopback and private addresses by default, so http://localhost:3000 and a staging host on Tailscale fail with "Blocked hostname" or "Blocked IP literal". That default is deliberate: these scripts follow URLs found on the pages they crawl.

To audit a pre-deployment host, the operator names it in CLAUDE_SEO_LOCAL_TARGETS, a comma-separated list of host or host:port entries:

CLAUDE_SEO_LOCAL_TARGETS="localhost:3000,127.0.0.1:8080,100.101.102.103" \
  "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run fetch_page.py http://localhost:3000/

What it does and does not cover:

Behaviour Allowlisted host
First, top-level URL over raw HTTP Allowed
Redirect target reached from that URL Refused
Subresource fetched by a rendered page Refused
Playwright renders (--render, screenshots) Refused; use the raw-HTTP path
A host not named in the variable Refused
Cloud metadata endpoints, even when listed Refused

host:port matches that port only; a bare host matches any port. With the variable unset the policy is unchanged. Never suggest setting it for a host the user does not control. See SECURITY.md.

Error Handling

Scenario Action
URL unreachable Report connection error with status code. Suggest verifying URL, checking DNS resolution, and confirming the site is publicly accessible.
robots.txt not found Note that no robots.txt was detected at the root domain. Recommend creating one with appropriate directives. Continue audit on remaining categories.
HTTPS not configured Flag as a critical issue. Report whether HTTP is served without redirect, mixed content exists, or SSL certificate is missing/expired.
Core Web Vitals data unavailable Note that CrUX data is not available (common for low-traffic sites). Suggest using Lighthouse lab data as a proxy and recommend increasing traffic before re-testing.
1---
2name: seo-technical
3description: >
4 Technical SEO audit across 9 categories: crawlability, indexability, security,
5 URL structure, mobile, Core Web Vitals, structured data, JavaScript rendering,
6 and IndexNow protocol. Use when user says "technical SEO", "crawl issues",
7 "robots.txt", "Core Web Vitals", "site speed", or "security headers".
8user-invocable: true
9argument-hint: "[url]"
10license: MIT
11metadata:
12 author: AgriciDaniel
13 version: "2.3.1"
14 category: seo
15---
16 
17# Technical SEO Audit
18 
19## Categories
20 
21### 1. Crawlability
22- robots.txt: exists, valid, not blocking important resources
23- XML sitemap: run `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run sitemap_discovery.py <url> --json`; require a
24 valid entry in `found`, and report stale or unsafe robots.txt declarations
25 separately from working fallback locations
26- Noindex tags: intentional vs accidental
27- Crawl depth: important pages within 3 clicks of homepage
28- JavaScript rendering: check if critical content requires JS execution
29- Crawl budget: for large sites (>10k pages), efficiency matters
30- Googlebot **fetch limits**: Googlebot fetches the first **2MB of HTML** and first **64MB of a PDF** (uncompressed; 15MB is the broader crawler-infra default). Long-standing, not a 2026 change, but inline base64 images, oversized inline CSS/JS, or bloated nav can push critical content/JSON-LD past the cap and out of the index. Keep key content + structured data within the first 2MB.
31- Crawl rate **auto-adjusts** (backs off on 5xx/slow responses); there is **no manual crawl-rate control** (the legacy Search Console setting was removed Jan 2024). Influence crawling via sitemaps, server responsiveness, and robots controls.
32- Google's canonical crawling/robots reference moved to **developers.google.com/crawling** (migrated 2025-11-20); IP-range files relocated to `/crawling/ipranges/` and `googlebot.json` was renamed `common-crawlers.json`.
33- AMP has no separate ranking advantage. Since 2026-07-01, Google Search sends
34 users directly to publisher-hosted AMP URLs, so do not recommend AMP Cache,
35 AMP Viewer, or signed exchange maintenance. Audit AMP against the same content,
36 action-parity, and quality requirements as other pages.
37 
38#### AI Crawler Management
39 
40As of 2025-2026, AI companies actively crawl the web to train models and power AI search. Managing these crawlers via robots.txt is a critical technical SEO consideration.
41 
42**Known AI crawlers:**
43 
44| Crawler | Company | robots.txt token | Purpose |
45|---------|---------|-----------------|---------|
46| GPTBot | OpenAI | `GPTBot` | Model training (NOT ChatGPT Search) |
47| OAI-SearchBot | OpenAI | `OAI-SearchBot` | ChatGPT Search citability |
48| ChatGPT-User | OpenAI | `ChatGPT-User` | Real-time browsing (user-triggered) |
49| ClaudeBot | Anthropic | `ClaudeBot` | Model training (NOT Claude search citability) |
50| Claude-SearchBot | Anthropic | `Claude-SearchBot` | Claude search-result citability |
51| PerplexityBot | Perplexity | `PerplexityBot` | Search index + training |
52| Bytespider | ByteDance | `Bytespider` | Model training |
53| Google-Extended | Google | `Google-Extended` | Gemini training (NOT search) |
54| Applebot-Extended | Apple | `Applebot-Extended` | Apple Intelligence training opt-out (NOT Siri/Spotlight/Safari) |
55| CCBot | Common Crawl | `CCBot` | Open dataset |
56 
57**Key distinctions:**
58- Blocking `Google-Extended` prevents Gemini training use but does NOT affect Google Search indexing or AI Overviews (those use `Googlebot`)
59- Blocking `GPTBot` prevents OpenAI training but does NOT affect ChatGPT Search
60 citability, which is governed by `OAI-SearchBot`, nor user-triggered browsing
61 (`ChatGPT-User`). Check `OAI-SearchBot` for any citability claim; `GPTBot`
62 status is evidence about training use only
63- Blocking `ClaudeBot` prevents Anthropic model training but does NOT affect
64 citability in Claude's own search features, which is governed by
65 `Claude-SearchBot` (per Anthropic's crawler support article). Check
66 `Claude-SearchBot` for any Claude-search citability claim; `ClaudeBot` status
67 is evidence about training use only
68- Blocking `Applebot-Extended` opts out of Apple Intelligence / generative-model
69 training use but does NOT affect discoverability via Siri, Spotlight, or Safari,
70 which follows `Applebot` (per Apple's support article); `Applebot-Extended` does
71 not itself crawl
72- ~3-5% of websites now use AI-specific robots.txt rules
73 
74**Example, selective AI crawler blocking:**
75```
76# Allow search indexing, block AI training crawlers
77User-agent: GPTBot
78Disallow: /
79 
80User-agent: Google-Extended
81Disallow: /
82 
83User-agent: Bytespider
84Disallow: /
85 
86# Allow all other crawlers (including Googlebot for search)
87User-agent: *
88Allow: /
89```
90 
91**Recommendation:** Consider your AI visibility strategy before blocking. Being cited by AI systems drives brand awareness and referral traffic. Cross-reference the `seo-geo` skill for the full AI crawler/fetcher taxonomy.
92 
93> **User-triggered fetchers ignore robots.txt by design.** Google now documents **Google-Agent** (Project Mariner, agentic browsing) plus **Google-NotebookLM** and **Google Messages** as *user-triggered* fetchers that **cannot be blocked via robots.txt**. Use server-side access controls instead. By contrast, `Google-Extended` and `Google-CloudVertexBot` obey robots.txt. Emerging: **Web Bot Auth** (RFC 9421) lets bots authenticate cryptographically via a `Signature-Agent` header + key directory at `agent.bot.goog` (used by Google-Agent); reverse-DNS verification remains the fallback.
94 
95### 2. Indexability
96- Canonical tags: self-referencing, no conflicts with noindex
97- Duplicate content: near-duplicates, parameter URLs, www vs non-www
98- Canonicalization fixes can take time: Google may retain corrected pages in a
99 duplicate cluster for **up to two weeks** while re-evaluating them. Do not
100 interpret an unchanged canonical immediately after a fix as proof that the
101 fix failed.
102- Thin content: pages below minimum word counts per type
103- Pagination: rel=next/prev or load-more pattern
104- Hreflang: correct for multi-language/multi-region sites
105- Index bloat: unnecessary pages consuming crawl budget
106 
107### 3. Security
108- HTTPS: enforced, valid SSL certificate, no mixed content
109- Security headers:
110 - Content-Security-Policy (CSP)
111 - Strict-Transport-Security (HSTS)
112 - X-Frame-Options
113 - X-Content-Type-Options
114 - Referrer-Policy
115- HSTS preload: check preload list inclusion for high-security sites
116- **Back-button hijacking** (spam-policy violation, malicious practices): flag pages that defeat the Back button via `history.pushState`/`replaceState` (including scripts injected by third-party ad/library platforms). Added to Google's spam policies 2026-04-13; **enforcement live since 2026-06-15** (manual actions + automated demotions): treat as Critical.
117 
118### 4. URL Structure
119- Clean URLs: descriptive, hyphenated, no query parameters for content
120- Hierarchy: logical folder structure reflecting site architecture
121- Redirects: no chains (max 1 hop), 301 for permanent moves
122- URL length: flag >100 characters
123- Trailing slashes: consistent usage
124 
125### 5. Mobile Optimization & Page Experience
126- Responsive design: viewport meta tag, responsive CSS
127- Touch targets: minimum 48x48px with 8px spacing
128- Font size: minimum 16px base
129- No horizontal scroll
130- Mobile-first indexing: Googlebot Smartphone is the primary crawler (rollout completed 2024). A mobile version is **not strictly required** (Google says "very strongly recommended"), sites that don't work on mobile can still be indexed, but the real risk is **content/parity loss**, not hard exclusion.
131- **Mobile/desktop content parity** (highest-value mobile check): equivalent primary content, matching robots meta tags, matching titles/descriptions, equivalent structured data, crawlable resources; avoid lazy-loading primary content that requires user interaction.
132- **Intrusive interstitials / ad density**: flag full-page interstitials, standalone consent-redirect pages, persistent blocking dialogs, and excessive/distracting ad density (a named page-experience aspect). Acceptable: small banners, standard CMS/legal dialogs.
133- **"Read more" deep links**: keep key content **immediately visible on load** (not behind tabs/accordions), don't hijack scroll on load, and preserve URL hash fragments, content hidden behind expandable sections is less likely to qualify.
134 
135> **Page experience is guidance, not a single ranking system.** Only **Core Web Vitals** feeds ranking directly; **HTTPS** is a confirmed but lightweight signal (affects <~1% of queries). Relevance can still win even when page experience is sub-par, so don't over-weight security headers. Note: the standalone **Page Experience report was removed** from Search Console (monitor via the Core Web Vitals + HTTPS reports).
136 
137### 6. Core Web Vitals
138- **LCP** (Largest Contentful Paint): target <=2.5s
139- **INP** (Interaction to Next Paint): target <=200ms
140 - INP replaced FID on March 12, 2024. FID was removed from Chrome's field-data tools (CrUX API, PageSpeed Insights) on September 9, 2024 (Lighthouse is a lab tool that never reported FID). Do NOT reference FID anywhere.
141- **CLS** (Cumulative Layout Shift): target <=0.1
142- Evaluation uses 75th percentile of real user data
143- Use PageSpeed Insights API or CrUX data if MCP available
144 
145### 7. Structured Data
146- Detection: JSON-LD (preferred), Microdata, RDFa
147- Validation against Google's supported types
148- See seo-schema skill for full analysis
149 
150### 8. JavaScript Rendering
151- Check if content visible in initial HTML vs requires JS
152- Identify client-side rendered (CSR) vs server-side rendered (SSR)
153- Flag SPA frameworks (React, Vue, Angular) that may cause indexing issues
154- If dynamic rendering is detected, flag it as technical debt rather than a valid setup.
155 Google documents it as "a workaround and not a recommended solution" because of the added
156 complexity and resource cost.
157 See https://developers.google.com/search/docs/crawling-indexing/javascript/dynamic-rendering
158 
159**Recommended rendering strategy:**
160 
161| Strategy | Use Case |
162|----------|----------|
163| **SSR** | Public SEO content, dynamic pages |
164| **SSG** | Static content, blogs, docs |
165| **CSR** | Authenticated / behind-login content only |
166 
167**Preferred frameworks:** Next.js, Astro, React Router v7 (Remix), SvelteKit
168 
169#### JavaScript SEO: Canonical & Indexing Guidance (December 2025)
170 
171Google updated its JavaScript SEO documentation in December 2025 with critical clarifications:
172 
1731. **Canonical conflicts:** If a canonical tag in raw HTML differs from one injected by JavaScript, Google may use EITHER one. Ensure canonical tags are identical between server-rendered HTML and JS-rendered output.
1742. **noindex with JavaScript:** If raw HTML contains `<meta name="robots" content="noindex">` but JavaScript removes it, Google MAY still honor the noindex from raw HTML. Serve correct robots directives in the initial HTML response.
1753. **Non-200 status codes:** Google does NOT render JavaScript on pages returning non-200 HTTP status codes. Any content or meta tags injected via JS on error pages will be invisible to Googlebot.
1764. **Structured data in JavaScript:** Product, Article, and other structured data injected via JS may face delayed processing. For time-sensitive structured data (especially e-commerce Product markup), include it in the initial server-rendered HTML.
177 
178**Best practice:** Serve critical SEO elements (canonical, meta robots, structured data, title, meta description) in the initial server-rendered HTML rather than relying on JavaScript injection.
179 
180### 9. IndexNow Protocol
181- Check if site supports IndexNow for Bing, Yandex, Naver
182- Supported by search engines other than Google
183- Recommend implementation for faster indexing on non-Google engines
184 
185## Agent-Friendly Pages & Agentic Browsing
186 
187AI agents (not just AI summarizers) increasingly read sites through three
188channels: vision models on screenshots, raw HTML/DOM, and the **accessibility
189tree** (the cleanest signal). Audit criteria: semantic HTML (real `<button>`
190and `<a>`, not `<div onclick>`), label associations, interactive target sizing,
191layout stability across templates, `cursor: pointer` correctness, live in
192`references/agent-friendly-pages.md`.
193 
194Google now ships a Lighthouse **Agentic Browsing** category (default-on since
195Lighthouse 13.3.0, Chrome 150+; buckets: agent-centric accessibility, CLS +
196llms.txt, three WebMCP audits). It reports a **fractional pass-ratio (X of N),
197not a 0-100 score**, keep that distinct from this skill's own Agent-UX 0-100
198heuristic below. Lighthouse 13.4.1 re-enabled the category through the PSI API.
199It is also available through Lighthouse CLI with
200`--only-categories=agentic-browsing`, DevTools, and the PSI web UI. See
201`references/agent-friendly-pages.md`.
202 
203### Audit command
204 
205```bash
206# Render with Playwright + capture accessibility tree, then score
207"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run agent_ux_check.py https://example.com --json
208```
209 
210The scanner outputs an Agent-UX score (0-100) plus itemized issues:
211- HTML findings: real buttons / anchors, `<div onclick>` widgets, semantic
212 landmarks, inputs without `<label for>`, inputs without ARIA labels
213- Accessibility tree findings: total nodes, interactive nodes, unnamed
214 interactive elements, `role="generic"` ratio
215 
216The accessibility-tree snapshot uses Chromium's
217`Accessibility.getFullAXTree` CDP command through Playwright. To capture the
218tree without scoring, use
219`"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run render_page.py <url> --a11y-tree --json`.
220 
221Surface findings as **opportunities**, not failures; don't gate audits on a
222sub-100 Agent-UX score. WebMCP origin-trial/sign-up status needs verification,
223and absence of WebMCP support is still an opportunity, not a defect.
224 
225## Output
226 
227### Technical Score: XX/100
228 
229### Category Breakdown
230| Category | Status | Score |
231|----------|--------|-------|
232| Crawlability | pass/warn/fail | XX/100 |
233| Indexability | pass/warn/fail | XX/100 |
234| Security | pass/warn/fail | XX/100 |
235| URL Structure | pass/warn/fail | XX/100 |
236| Mobile | pass/warn/fail | XX/100 |
237| Core Web Vitals | pass/warn/fail | XX/100 |
238| Structured Data | pass/warn/fail | XX/100 |
239| JS Rendering | pass/warn/fail | XX/100 |
240| IndexNow | pass/warn/fail | XX/100 |
241 
242### Critical Issues (fix immediately)
243### High Priority (fix within 1 week)
244### Medium Priority (fix within 1 month)
245### Low Priority (backlog)
246 
247## DataForSEO Integration (Optional)
248 
249If DataForSEO MCP tools are available, use `on_page_instant_pages` for real page analysis (status codes, page timing, broken links, on-page checks), `on_page_lighthouse` for Lighthouse audits (performance, accessibility, SEO scores), and `domain_analytics_technologies_domain_technologies` for technology stack detection.
250 
251## Google API Integration (Optional)
252 
253If Google API credentials are configured, use `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run pagespeed_check.py <url> --json` for real PSI + CrUX field data (replaces lab-only CWV estimates), `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run crux_history.py <url> --json` for 25-week CWV trends, and `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run gsc_inspect.py <url> --json` for real indexation status per URL.
254 
255## Auditing a Local or Private Host
256 
257`url_safety` refuses loopback and private addresses by default, so `http://localhost:3000` and a staging host on Tailscale fail with "Blocked hostname" or "Blocked IP literal". That default is deliberate: these scripts follow URLs found on the pages they crawl.
258 
259To audit a pre-deployment host, the operator names it in `CLAUDE_SEO_LOCAL_TARGETS`, a comma-separated list of `host` or `host:port` entries:
260 
261```bash
262CLAUDE_SEO_LOCAL_TARGETS="localhost:3000,127.0.0.1:8080,100.101.102.103" \
263 "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run fetch_page.py http://localhost:3000/
264```
265 
266What it does and does not cover:
267 
268| Behaviour | Allowlisted host |
269|-----------|------------------|
270| First, top-level URL over raw HTTP | Allowed |
271| Redirect target reached from that URL | Refused |
272| Subresource fetched by a rendered page | Refused |
273| Playwright renders (`--render`, screenshots) | Refused; use the raw-HTTP path |
274| A host not named in the variable | Refused |
275| Cloud metadata endpoints, even when listed | Refused |
276 
277`host:port` matches that port only; a bare `host` matches any port. With the variable unset the policy is unchanged. Never suggest setting it for a host the user does not control. See SECURITY.md.
278 
279## Error Handling
280 
281| Scenario | Action |
282|----------|--------|
283| URL unreachable | Report connection error with status code. Suggest verifying URL, checking DNS resolution, and confirming the site is publicly accessible. |
284| robots.txt not found | Note that no robots.txt was detected at the root domain. Recommend creating one with appropriate directives. Continue audit on remaining categories. |
285| HTTPS not configured | Flag as a critical issue. Report whether HTTP is served without redirect, mixed content exists, or SSL certificate is missing/expired. |
286| Core Web Vitals data unavailable | Note that CrUX data is not available (common for low-traffic sites). Suggest using Lighthouse lab data as a proxy and recommend increasing traffic before re-testing. |
287 

Discussion

Alternatives

Also in SEO & keywords