Technical SEO Audit
Technical SEO audit across 9 categories: crawlability, indexability, security, URL structure, mobile, Core Web Vitals, structured data, JavaScript rendering, and IndexNow protocol.
How to use it
- Hit Copy SKILL.md — or use the Claude Code line below to get every file.
- Claude: ⋯ → Download .md, then Customize → Skills → Add → Upload skill.
ChatGPT: make a Project and paste it into Instructions.
Neither? Paste it at the top of a new chat — it works for that chat. - Describe your job in plain words. The AI follows the skill from there.
npx degit AgriciDaniel/claude-seo/skills/seo-technical#main ~/.claude/skills/seo-technicalFor one project only, change the path to .claude/skills/seo-technical. This skill also uses robots.txt, sitemap_discovery.py, googlebot.json, common-crawlers.json, Next.js, llms.txt — copying SKILL.md alone won't be enough. See the folder on GitHub.
Not working?
- Check which app you pasted it into — the steps above name the right one.
- Some skills need the paid tier of Claude or ChatGPT.
Paste into Claude, ChatGPT or Cursor.
Show the full text287 lines
Technical SEO Audit
Categories
1. Crawlability
- robots.txt: exists, valid, not blocking important resources
- XML sitemap: run
"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run sitemap_discovery.py <url> --json; require a valid entry infound, and report stale or unsafe robots.txt declarations separately from working fallback locations - Noindex tags: intentional vs accidental
- Crawl depth: important pages within 3 clicks of homepage
- JavaScript rendering: check if critical content requires JS execution
- Crawl budget: for large sites (>10k pages), efficiency matters
- Googlebot fetch limits: Googlebot fetches the first 2MB of HTML and first 64MB of a PDF (uncompressed; 15MB is the broader crawler-infra default). Long-standing, not a 2026 change, but inline base64 images, oversized inline CSS/JS, or bloated nav can push critical content/JSON-LD past the cap and out of the index. Keep key content + structured data within the first 2MB.
- Crawl rate auto-adjusts (backs off on 5xx/slow responses); there is no manual crawl-rate control (the legacy Search Console setting was removed Jan 2024). Influence crawling via sitemaps, server responsiveness, and robots controls.
- Google's canonical crawling/robots reference moved to developers.google.com/crawling (migrated 2025-11-20); IP-range files relocated to
/crawling/ipranges/andgooglebot.jsonwas renamedcommon-crawlers.json. - AMP has no separate ranking advantage. Since 2026-07-01, Google Search sends users directly to publisher-hosted AMP URLs, so do not recommend AMP Cache, AMP Viewer, or signed exchange maintenance. Audit AMP against the same content, action-parity, and quality requirements as other pages.
AI Crawler Management
As of 2025-2026, AI companies actively crawl the web to train models and power AI search. Managing these crawlers via robots.txt is a critical technical SEO consideration.
Known AI crawlers:
| Crawler | Company | robots.txt token | Purpose |
|---|---|---|---|
| GPTBot | OpenAI | GPTBot |
Model training (NOT ChatGPT Search) |
| OAI-SearchBot | OpenAI | OAI-SearchBot |
ChatGPT Search citability |
| ChatGPT-User | OpenAI | ChatGPT-User |
Real-time browsing (user-triggered) |
| ClaudeBot | Anthropic | ClaudeBot |
Model training (NOT Claude search citability) |
| Claude-SearchBot | Anthropic | Claude-SearchBot |
Claude search-result citability |
| PerplexityBot | Perplexity | PerplexityBot |
Search index + training |
| Bytespider | ByteDance | Bytespider |
Model training |
| Google-Extended | Google-Extended |
Gemini training (NOT search) | |
| Applebot-Extended | Apple | Applebot-Extended |
Apple Intelligence training opt-out (NOT Siri/Spotlight/Safari) |
| CCBot | Common Crawl | CCBot |
Open dataset |
Key distinctions:
- Blocking
Google-Extendedprevents Gemini training use but does NOT affect Google Search indexing or AI Overviews (those useGooglebot) - Blocking
GPTBotprevents OpenAI training but does NOT affect ChatGPT Search citability, which is governed byOAI-SearchBot, nor user-triggered browsing (ChatGPT-User). CheckOAI-SearchBotfor any citability claim;GPTBotstatus is evidence about training use only - Blocking
ClaudeBotprevents Anthropic model training but does NOT affect citability in Claude's own search features, which is governed byClaude-SearchBot(per Anthropic's crawler support article). CheckClaude-SearchBotfor any Claude-search citability claim;ClaudeBotstatus is evidence about training use only - Blocking
Applebot-Extendedopts out of Apple Intelligence / generative-model training use but does NOT affect discoverability via Siri, Spotlight, or Safari, which followsApplebot(per Apple's support article);Applebot-Extendeddoes not itself crawl - ~3-5% of websites now use AI-specific robots.txt rules
Example, selective AI crawler blocking:
# Allow search indexing, block AI training crawlers
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Bytespider
Disallow: /
# Allow all other crawlers (including Googlebot for search)
User-agent: *
Allow: /
Recommendation: Consider your AI visibility strategy before blocking. Being cited by AI systems drives brand awareness and referral traffic. Cross-reference the seo-geo skill for the full AI crawler/fetcher taxonomy.
User-triggered fetchers ignore robots.txt by design. Google now documents Google-Agent (Project Mariner, agentic browsing) plus Google-NotebookLM and Google Messages as user-triggered fetchers that cannot be blocked via robots.txt. Use server-side access controls instead. By contrast,
Google-ExtendedandGoogle-CloudVertexBotobey robots.txt. Emerging: Web Bot Auth (RFC 9421) lets bots authenticate cryptographically via aSignature-Agentheader + key directory atagent.bot.goog(used by Google-Agent); reverse-DNS verification remains the fallback.
2. Indexability
- Canonical tags: self-referencing, no conflicts with noindex
- Duplicate content: near-duplicates, parameter URLs, www vs non-www
- Canonicalization fixes can take time: Google may retain corrected pages in a duplicate cluster for up to two weeks while re-evaluating them. Do not interpret an unchanged canonical immediately after a fix as proof that the fix failed.
- Thin content: pages below minimum word counts per type
- Pagination: rel=next/prev or load-more pattern
- Hreflang: correct for multi-language/multi-region sites
- Index bloat: unnecessary pages consuming crawl budget
3. Security
- HTTPS: enforced, valid SSL certificate, no mixed content
- Security headers:
- Content-Security-Policy (CSP)
- Strict-Transport-Security (HSTS)
- X-Frame-Options
- X-Content-Type-Options
- Referrer-Policy
- HSTS preload: check preload list inclusion for high-security sites
- Back-button hijacking (spam-policy violation, malicious practices): flag pages that defeat the Back button via
history.pushState/replaceState(including scripts injected by third-party ad/library platforms). Added to Google's spam policies 2026-04-13; enforcement live since 2026-06-15 (manual actions + automated demotions): treat as Critical.
4. URL Structure
- Clean URLs: descriptive, hyphenated, no query parameters for content
- Hierarchy: logical folder structure reflecting site architecture
- Redirects: no chains (max 1 hop), 301 for permanent moves
- URL length: flag >100 characters
- Trailing slashes: consistent usage
5. Mobile Optimization & Page Experience
- Responsive design: viewport meta tag, responsive CSS
- Touch targets: minimum 48x48px with 8px spacing
- Font size: minimum 16px base
- No horizontal scroll
- Mobile-first indexing: Googlebot Smartphone is the primary crawler (rollout completed 2024). A mobile version is not strictly required (Google says "very strongly recommended"), sites that don't work on mobile can still be indexed, but the real risk is content/parity loss, not hard exclusion.
- Mobile/desktop content parity (highest-value mobile check): equivalent primary content, matching robots meta tags, matching titles/descriptions, equivalent structured data, crawlable resources; avoid lazy-loading primary content that requires user interaction.
- Intrusive interstitials / ad density: flag full-page interstitials, standalone consent-redirect pages, persistent blocking dialogs, and excessive/distracting ad density (a named page-experience aspect). Acceptable: small banners, standard CMS/legal dialogs.
- "Read more" deep links: keep key content immediately visible on load (not behind tabs/accordions), don't hijack scroll on load, and preserve URL hash fragments, content hidden behind expandable sections is less likely to qualify.
Page experience is guidance, not a single ranking system. Only Core Web Vitals feeds ranking directly; HTTPS is a confirmed but lightweight signal (affects <~1% of queries). Relevance can still win even when page experience is sub-par, so don't over-weight security headers. Note: the standalone Page Experience report was removed from Search Console (monitor via the Core Web Vitals + HTTPS reports).
6. Core Web Vitals
- LCP (Largest Contentful Paint): target <=2.5s
- INP (Interaction to Next Paint): target <=200ms
- INP replaced FID on March 12, 2024. FID was removed from Chrome's field-data tools (CrUX API, PageSpeed Insights) on September 9, 2024 (Lighthouse is a lab tool that never reported FID). Do NOT reference FID anywhere.
- CLS (Cumulative Layout Shift): target <=0.1
- Evaluation uses 75th percentile of real user data
- Use PageSpeed Insights API or CrUX data if MCP available
7. Structured Data
- Detection: JSON-LD (preferred), Microdata, RDFa
- Validation against Google's supported types
- See seo-schema skill for full analysis
8. JavaScript Rendering
- Check if content visible in initial HTML vs requires JS
- Identify client-side rendered (CSR) vs server-side rendered (SSR)
- Flag SPA frameworks (React, Vue, Angular) that may cause indexing issues
- If dynamic rendering is detected, flag it as technical debt rather than a valid setup. Google documents it as "a workaround and not a recommended solution" because of the added complexity and resource cost. See https://developers.google.com/search/docs/crawling-indexing/javascript/dynamic-rendering
Recommended rendering strategy:
| Strategy | Use Case |
|---|---|
| SSR | Public SEO content, dynamic pages |
| SSG | Static content, blogs, docs |
| CSR | Authenticated / behind-login content only |
Preferred frameworks: Next.js, Astro, React Router v7 (Remix), SvelteKit
JavaScript SEO: Canonical & Indexing Guidance (December 2025)
Google updated its JavaScript SEO documentation in December 2025 with critical clarifications:
- Canonical conflicts: If a canonical tag in raw HTML differs from one injected by JavaScript, Google may use EITHER one. Ensure canonical tags are identical between server-rendered HTML and JS-rendered output.
- noindex with JavaScript: If raw HTML contains
<meta name="robots" content="noindex">but JavaScript removes it, Google MAY still honor the noindex from raw HTML. Serve correct robots directives in the initial HTML response. - Non-200 status codes: Google does NOT render JavaScript on pages returning non-200 HTTP status codes. Any content or meta tags injected via JS on error pages will be invisible to Googlebot.
- Structured data in JavaScript: Product, Article, and other structured data injected via JS may face delayed processing. For time-sensitive structured data (especially e-commerce Product markup), include it in the initial server-rendered HTML.
Best practice: Serve critical SEO elements (canonical, meta robots, structured data, title, meta description) in the initial server-rendered HTML rather than relying on JavaScript injection.
9. IndexNow Protocol
- Check if site supports IndexNow for Bing, Yandex, Naver
- Supported by search engines other than Google
- Recommend implementation for faster indexing on non-Google engines
Agent-Friendly Pages & Agentic Browsing
AI agents (not just AI summarizers) increasingly read sites through three
channels: vision models on screenshots, raw HTML/DOM, and the accessibility
tree (the cleanest signal). Audit criteria: semantic HTML (real <button>
and , not ), label associations, interactive target sizing,
layout stability across templates, cursor: pointer correctness, live in
references/agent-friendly-pages.md.
Google now ships a Lighthouse Agentic Browsing category (default-on since
Lighthouse 13.3.0, Chrome 150+; buckets: agent-centric accessibility, CLS +
llms.txt, three WebMCP audits). It reports a fractional pass-ratio (X of N),
not a 0-100 score, keep that distinct from this skill's own Agent-UX 0-100
heuristic below. Lighthouse 13.4.1 re-enabled the category through the PSI API.
It is also available through Lighthouse CLI with
--only-categories=agentic-browsing, DevTools, and the PSI web UI. See
references/agent-friendly-pages.md.
Audit command
# Render with Playwright + capture accessibility tree, then score
"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run agent_ux_check.py https://example.com --json
The scanner outputs an Agent-UX score (0-100) plus itemized issues:
- HTML findings: real buttons / anchors, `` widgets, semantic
landmarks, inputs without
<label for>, inputs without ARIA labels - Accessibility tree findings: total nodes, interactive nodes, unnamed
interactive elements,
role="generic"ratio
The accessibility-tree snapshot uses Chromium's
Accessibility.getFullAXTree CDP command through Playwright. To capture the
tree without scoring, use
"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run render_page.py <url> --a11y-tree --json.
Surface findings as opportunities, not failures; don't gate audits on a sub-100 Agent-UX score. WebMCP origin-trial/sign-up status needs verification, and absence of WebMCP support is still an opportunity, not a defect.
Output
Technical Score: XX/100
Category Breakdown
| Category | Status | Score |
|---|---|---|
| Crawlability | pass/warn/fail | XX/100 |
| Indexability | pass/warn/fail | XX/100 |
| Security | pass/warn/fail | XX/100 |
| URL Structure | pass/warn/fail | XX/100 |
| Mobile | pass/warn/fail | XX/100 |
| Core Web Vitals | pass/warn/fail | XX/100 |
| Structured Data | pass/warn/fail | XX/100 |
| JS Rendering | pass/warn/fail | XX/100 |
| IndexNow | pass/warn/fail | XX/100 |
Critical Issues (fix immediately)
High Priority (fix within 1 week)
Medium Priority (fix within 1 month)
Low Priority (backlog)
DataForSEO Integration (Optional)
If DataForSEO MCP tools are available, use on_page_instant_pages for real page analysis (status codes, page timing, broken links, on-page checks), on_page_lighthouse for Lighthouse audits (performance, accessibility, SEO scores), and domain_analytics_technologies_domain_technologies for technology stack detection.
Google API Integration (Optional)
If Google API credentials are configured, use "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run pagespeed_check.py <url> --json for real PSI + CrUX field data (replaces lab-only CWV estimates), "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run crux_history.py <url> --json for 25-week CWV trends, and "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run gsc_inspect.py <url> --json for real indexation status per URL.
Auditing a Local or Private Host
url_safety refuses loopback and private addresses by default, so http://localhost:3000 and a staging host on Tailscale fail with "Blocked hostname" or "Blocked IP literal". That default is deliberate: these scripts follow URLs found on the pages they crawl.
To audit a pre-deployment host, the operator names it in CLAUDE_SEO_LOCAL_TARGETS, a comma-separated list of host or host:port entries:
CLAUDE_SEO_LOCAL_TARGETS="localhost:3000,127.0.0.1:8080,100.101.102.103" \
"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run fetch_page.py http://localhost:3000/
What it does and does not cover:
| Behaviour | Allowlisted host |
|---|---|
| First, top-level URL over raw HTTP | Allowed |
| Redirect target reached from that URL | Refused |
| Subresource fetched by a rendered page | Refused |
Playwright renders (--render, screenshots) |
Refused; use the raw-HTTP path |
| A host not named in the variable | Refused |
| Cloud metadata endpoints, even when listed | Refused |
host:port matches that port only; a bare host matches any port. With the variable unset the policy is unchanged. Never suggest setting it for a host the user does not control. See SECURITY.md.
Error Handling
| Scenario | Action |
|---|---|
| URL unreachable | Report connection error with status code. Suggest verifying URL, checking DNS resolution, and confirming the site is publicly accessible. |
| robots.txt not found | Note that no robots.txt was detected at the root domain. Recommend creating one with appropriate directives. Continue audit on remaining categories. |
| HTTPS not configured | Flag as a critical issue. Report whether HTTP is served without redirect, mixed content exists, or SSL certificate is missing/expired. |
| Core Web Vitals data unavailable | Note that CrUX data is not available (common for low-traffic sites). Suggest using Lighthouse lab data as a proxy and recommend increasing traffic before re-testing. |
| 1 | |
| 2 | name seo-technical |
| 3 | description > |
| 4 | Technical SEO audit across 9 categories: crawlability, indexability, security, |
| 5 | URL structure, mobile, Core Web Vitals, structured data, JavaScript rendering, |
| 6 | and IndexNow protocol. Use when user says "technical SEO", "crawl issues", |
| 7 | "robots.txt", "Core Web Vitals", "site speed", or "security headers". |
| 8 | user-invocable true |
| 9 | argument-hint "[url]" |
| 10 | license MIT |
| 11 | metadata |
| 12 | author AgriciDaniel |
| 13 | version "2.3.1" |
| 14 | category seo |
| 15 | |
| 16 | |
| 17 | # Technical SEO Audit |
| 18 | |
| 19 | ## Categories |
| 20 | |
| 21 | ### 1. Crawlability |
| 22 | robots.txt: exists, valid, not blocking important resources |
| 23 | XML sitemap: run `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run sitemap_discovery.py <url> --json`; require a |
| 24 | valid entry in `found`, and report stale or unsafe robots.txt declarations |
| 25 | separately from working fallback locations |
| 26 | Noindex tags: intentional vs accidental |
| 27 | Crawl depth: important pages within 3 clicks of homepage |
| 28 | JavaScript rendering: check if critical content requires JS execution |
| 29 | Crawl budget: for large sites (>10k pages), efficiency matters |
| 30 | Googlebot **fetch limits**: Googlebot fetches the first **2MB of HTML** and first **64MB of a PDF** (uncompressed; 15MB is the broader crawler-infra default). Long-standing, not a 2026 change, but inline base64 images, oversized inline CSS/JS, or bloated nav can push critical content/JSON-LD past the cap and out of the index. Keep key content + structured data within the first 2MB. |
| 31 | Crawl rate **auto-adjusts** (backs off on 5xx/slow responses); there is **no manual crawl-rate control** (the legacy Search Console setting was removed Jan 2024). Influence crawling via sitemaps, server responsiveness, and robots controls. |
| 32 | Google's canonical crawling/robots reference moved to **developers.google.com/crawling** (migrated 2025-11-20); IP-range files relocated to `/crawling/ipranges/` and `googlebot.json` was renamed `common-crawlers.json`. |
| 33 | AMP has no separate ranking advantage. Since 2026-07-01, Google Search sends |
| 34 | users directly to publisher-hosted AMP URLs, so do not recommend AMP Cache, |
| 35 | AMP Viewer, or signed exchange maintenance. Audit AMP against the same content, |
| 36 | action-parity, and quality requirements as other pages. |
| 37 | |
| 38 | #### AI Crawler Management |
| 39 | |
| 40 | As of 2025-2026, AI companies actively crawl the web to train models and power AI search. Managing these crawlers via robots.txt is a critical technical SEO consideration. |
| 41 | |
| 42 | **Known AI crawlers:** |
| 43 | |
| 44 | | Crawler | Company | robots.txt token | Purpose | |
| 45 | |---------|---------|-----------------|---------| |
| 46 | | GPTBot | OpenAI | `GPTBot` | Model training (NOT ChatGPT Search) | |
| 47 | | OAI-SearchBot | OpenAI | `OAI-SearchBot` | ChatGPT Search citability | |
| 48 | | ChatGPT-User | OpenAI | `ChatGPT-User` | Real-time browsing (user-triggered) | |
| 49 | | ClaudeBot | Anthropic | `ClaudeBot` | Model training (NOT Claude search citability) | |
| 50 | | Claude-SearchBot | Anthropic | `Claude-SearchBot` | Claude search-result citability | |
| 51 | | PerplexityBot | Perplexity | `PerplexityBot` | Search index + training | |
| 52 | | Bytespider | ByteDance | `Bytespider` | Model training | |
| 53 | | Google-Extended | Google | `Google-Extended` | Gemini training (NOT search) | |
| 54 | | Applebot-Extended | Apple | `Applebot-Extended` | Apple Intelligence training opt-out (NOT Siri/Spotlight/Safari) | |
| 55 | | CCBot | Common Crawl | `CCBot` | Open dataset | |
| 56 | |
| 57 | **Key distinctions:** |
| 58 | Blocking `Google-Extended` prevents Gemini training use but does NOT affect Google Search indexing or AI Overviews (those use `Googlebot`) |
| 59 | Blocking `GPTBot` prevents OpenAI training but does NOT affect ChatGPT Search |
| 60 | citability, which is governed by `OAI-SearchBot`, nor user-triggered browsing |
| 61 | (`ChatGPT-User`). Check `OAI-SearchBot` for any citability claim; `GPTBot` |
| 62 | status is evidence about training use only |
| 63 | Blocking `ClaudeBot` prevents Anthropic model training but does NOT affect |
| 64 | citability in Claude's own search features, which is governed by |
| 65 | `Claude-SearchBot` (per Anthropic's crawler support article). Check |
| 66 | `Claude-SearchBot` for any Claude-search citability claim; `ClaudeBot` status |
| 67 | is evidence about training use only |
| 68 | Blocking `Applebot-Extended` opts out of Apple Intelligence / generative-model |
| 69 | training use but does NOT affect discoverability via Siri, Spotlight, or Safari, |
| 70 | which follows `Applebot` (per Apple's support article); `Applebot-Extended` does |
| 71 | not itself crawl |
| 72 | ~3-5% of websites now use AI-specific robots.txt rules |
| 73 | |
| 74 | **Example, selective AI crawler blocking:** |
| 75 | |
| 76 | # Allow search indexing, block AI training crawlers |
| 77 | User-agent: GPTBot |
| 78 | Disallow: / |
| 79 | |
| 80 | User-agent: Google-Extended |
| 81 | Disallow: / |
| 82 | |
| 83 | User-agent: Bytespider |
| 84 | Disallow: / |
| 85 | |
| 86 | # Allow all other crawlers (including Googlebot for search) |
| 87 | User-agent: * |
| 88 | Allow: / |
| 89 | |
| 90 | |
| 91 | **Recommendation:** Consider your AI visibility strategy before blocking. Being cited by AI systems drives brand awareness and referral traffic. Cross-reference the `seo-geo` skill for the full AI crawler/fetcher taxonomy. |
| 92 | |
| 93 | > **User-triggered fetchers ignore robots.txt by design.** Google now documents **Google-Agent** (Project Mariner, agentic browsing) plus **Google-NotebookLM** and **Google Messages** as *user-triggered* fetchers that **cannot be blocked via robots.txt**. Use server-side access controls instead. By contrast, `Google-Extended` and `Google-CloudVertexBot` obey robots.txt. Emerging: **Web Bot Auth** (RFC 9421) lets bots authenticate cryptographically via a `Signature-Agent` header + key directory at `agent.bot.goog` (used by Google-Agent); reverse-DNS verification remains the fallback. |
| 94 | |
| 95 | ### 2. Indexability |
| 96 | Canonical tags: self-referencing, no conflicts with noindex |
| 97 | Duplicate content: near-duplicates, parameter URLs, www vs non-www |
| 98 | Canonicalization fixes can take time: Google may retain corrected pages in a |
| 99 | duplicate cluster for **up to two weeks** while re-evaluating them. Do not |
| 100 | interpret an unchanged canonical immediately after a fix as proof that the |
| 101 | fix failed. |
| 102 | Thin content: pages below minimum word counts per type |
| 103 | Pagination: rel=next/prev or load-more pattern |
| 104 | Hreflang: correct for multi-language/multi-region sites |
| 105 | Index bloat: unnecessary pages consuming crawl budget |
| 106 | |
| 107 | ### 3. Security |
| 108 | HTTPS: enforced, valid SSL certificate, no mixed content |
| 109 | Security headers: |
| 110 | Content-Security-Policy (CSP) |
| 111 | Strict-Transport-Security (HSTS) |
| 112 | X-Frame-Options |
| 113 | X-Content-Type-Options |
| 114 | Referrer-Policy |
| 115 | HSTS preload: check preload list inclusion for high-security sites |
| 116 | **Back-button hijacking** (spam-policy violation, malicious practices): flag pages that defeat the Back button via `history.pushState`/`replaceState` (including scripts injected by third-party ad/library platforms). Added to Google's spam policies 2026-04-13; **enforcement live since 2026-06-15** (manual actions + automated demotions): treat as Critical. |
| 117 | |
| 118 | ### 4. URL Structure |
| 119 | Clean URLs: descriptive, hyphenated, no query parameters for content |
| 120 | Hierarchy: logical folder structure reflecting site architecture |
| 121 | Redirects: no chains (max 1 hop), 301 for permanent moves |
| 122 | URL length: flag >100 characters |
| 123 | Trailing slashes: consistent usage |
| 124 | |
| 125 | ### 5. Mobile Optimization & Page Experience |
| 126 | Responsive design: viewport meta tag, responsive CSS |
| 127 | Touch targets: minimum 48x48px with 8px spacing |
| 128 | Font size: minimum 16px base |
| 129 | No horizontal scroll |
| 130 | Mobile-first indexing: Googlebot Smartphone is the primary crawler (rollout completed 2024). A mobile version is **not strictly required** (Google says "very strongly recommended"), sites that don't work on mobile can still be indexed, but the real risk is **content/parity loss**, not hard exclusion. |
| 131 | **Mobile/desktop content parity** (highest-value mobile check): equivalent primary content, matching robots meta tags, matching titles/descriptions, equivalent structured data, crawlable resources; avoid lazy-loading primary content that requires user interaction. |
| 132 | **Intrusive interstitials / ad density**: flag full-page interstitials, standalone consent-redirect pages, persistent blocking dialogs, and excessive/distracting ad density (a named page-experience aspect). Acceptable: small banners, standard CMS/legal dialogs. |
| 133 | **"Read more" deep links**: keep key content **immediately visible on load** (not behind tabs/accordions), don't hijack scroll on load, and preserve URL hash fragments, content hidden behind expandable sections is less likely to qualify. |
| 134 | |
| 135 | > **Page experience is guidance, not a single ranking system.** Only **Core Web Vitals** feeds ranking directly; **HTTPS** is a confirmed but lightweight signal (affects <~1% of queries). Relevance can still win even when page experience is sub-par, so don't over-weight security headers. Note: the standalone **Page Experience report was removed** from Search Console (monitor via the Core Web Vitals + HTTPS reports). |
| 136 | |
| 137 | ### 6. Core Web Vitals |
| 138 | **LCP** (Largest Contentful Paint): target <=2.5s |
| 139 | **INP** (Interaction to Next Paint): target <=200ms |
| 140 | INP replaced FID on March 12, 2024. FID was removed from Chrome's field-data tools (CrUX API, PageSpeed Insights) on September 9, 2024 (Lighthouse is a lab tool that never reported FID). Do NOT reference FID anywhere. |
| 141 | **CLS** (Cumulative Layout Shift): target <=0.1 |
| 142 | Evaluation uses 75th percentile of real user data |
| 143 | Use PageSpeed Insights API or CrUX data if MCP available |
| 144 | |
| 145 | ### 7. Structured Data |
| 146 | Detection: JSON-LD (preferred), Microdata, RDFa |
| 147 | Validation against Google's supported types |
| 148 | See seo-schema skill for full analysis |
| 149 | |
| 150 | ### 8. JavaScript Rendering |
| 151 | Check if content visible in initial HTML vs requires JS |
| 152 | Identify client-side rendered (CSR) vs server-side rendered (SSR) |
| 153 | Flag SPA frameworks (React, Vue, Angular) that may cause indexing issues |
| 154 | If dynamic rendering is detected, flag it as technical debt rather than a valid setup. |
| 155 | Google documents it as "a workaround and not a recommended solution" because of the added |
| 156 | complexity and resource cost. |
| 157 | See https://developers.google.com/search/docs/crawling-indexing/javascript/dynamic-rendering |
| 158 | |
| 159 | **Recommended rendering strategy:** |
| 160 | |
| 161 | | Strategy | Use Case | |
| 162 | |----------|----------| |
| 163 | | **SSR** | Public SEO content, dynamic pages | |
| 164 | | **SSG** | Static content, blogs, docs | |
| 165 | | **CSR** | Authenticated / behind-login content only | |
| 166 | |
| 167 | **Preferred frameworks:** Next.js, Astro, React Router v7 (Remix), SvelteKit |
| 168 | |
| 169 | #### JavaScript SEO: Canonical & Indexing Guidance (December 2025) |
| 170 | |
| 171 | Google updated its JavaScript SEO documentation in December 2025 with critical clarifications: |
| 172 | |
| 173 | **Canonical conflicts:** If a canonical tag in raw HTML differs from one injected by JavaScript, Google may use EITHER one. Ensure canonical tags are identical between server-rendered HTML and JS-rendered output. |
| 174 | **noindex with JavaScript:** If raw HTML contains `<meta name="robots" content="noindex">` but JavaScript removes it, Google MAY still honor the noindex from raw HTML. Serve correct robots directives in the initial HTML response. |
| 175 | **Non-200 status codes:** Google does NOT render JavaScript on pages returning non-200 HTTP status codes. Any content or meta tags injected via JS on error pages will be invisible to Googlebot. |
| 176 | **Structured data in JavaScript:** Product, Article, and other structured data injected via JS may face delayed processing. For time-sensitive structured data (especially e-commerce Product markup), include it in the initial server-rendered HTML. |
| 177 | |
| 178 | **Best practice:** Serve critical SEO elements (canonical, meta robots, structured data, title, meta description) in the initial server-rendered HTML rather than relying on JavaScript injection. |
| 179 | |
| 180 | ### 9. IndexNow Protocol |
| 181 | Check if site supports IndexNow for Bing, Yandex, Naver |
| 182 | Supported by search engines other than Google |
| 183 | Recommend implementation for faster indexing on non-Google engines |
| 184 | |
| 185 | ## Agent-Friendly Pages & Agentic Browsing |
| 186 | |
| 187 | AI agents (not just AI summarizers) increasingly read sites through three |
| 188 | channels: vision models on screenshots, raw HTML/DOM, and the **accessibility |
| 189 | tree** (the cleanest signal). Audit criteria: semantic HTML (real `<button>` |
| 190 | and `<a>`, not `<div onclick>`), label associations, interactive target sizing, |
| 191 | layout stability across templates, `cursor: pointer` correctness, live in |
| 192 | `references/agent-friendly-pages.md`. |
| 193 | |
| 194 | Google now ships a Lighthouse **Agentic Browsing** category (default-on since |
| 195 | Lighthouse 13.3.0, Chrome 150+; buckets: agent-centric accessibility, CLS + |
| 196 | llms.txt, three WebMCP audits). It reports a **fractional pass-ratio (X of N), |
| 197 | not a 0-100 score**, keep that distinct from this skill's own Agent-UX 0-100 |
| 198 | heuristic below. Lighthouse 13.4.1 re-enabled the category through the PSI API. |
| 199 | It is also available through Lighthouse CLI with |
| 200 | `--only-categories=agentic-browsing`, DevTools, and the PSI web UI. See |
| 201 | `references/agent-friendly-pages.md`. |
| 202 | |
| 203 | ### Audit command |
| 204 | |
| 205 | |
| 206 | # Render with Playwright + capture accessibility tree, then score |
| 207 | "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run agent_ux_check.py https://example.com --json |
| 208 | |
| 209 | |
| 210 | The scanner outputs an Agent-UX score (0-100) plus itemized issues: |
| 211 | HTML findings: real buttons / anchors, `<div onclick>` widgets, semantic |
| 212 | landmarks, inputs without `<label for>`, inputs without ARIA labels |
| 213 | Accessibility tree findings: total nodes, interactive nodes, unnamed |
| 214 | interactive elements, `role="generic"` ratio |
| 215 | |
| 216 | The accessibility-tree snapshot uses Chromium's |
| 217 | `Accessibility.getFullAXTree` CDP command through Playwright. To capture the |
| 218 | tree without scoring, use |
| 219 | `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run render_page.py <url> --a11y-tree --json`. |
| 220 | |
| 221 | Surface findings as **opportunities**, not failures; don't gate audits on a |
| 222 | sub-100 Agent-UX score. WebMCP origin-trial/sign-up status needs verification, |
| 223 | and absence of WebMCP support is still an opportunity, not a defect. |
| 224 | |
| 225 | ## Output |
| 226 | |
| 227 | ### Technical Score: XX/100 |
| 228 | |
| 229 | ### Category Breakdown |
| 230 | | Category | Status | Score | |
| 231 | |----------|--------|-------| |
| 232 | | Crawlability | pass/warn/fail | XX/100 | |
| 233 | | Indexability | pass/warn/fail | XX/100 | |
| 234 | | Security | pass/warn/fail | XX/100 | |
| 235 | | URL Structure | pass/warn/fail | XX/100 | |
| 236 | | Mobile | pass/warn/fail | XX/100 | |
| 237 | | Core Web Vitals | pass/warn/fail | XX/100 | |
| 238 | | Structured Data | pass/warn/fail | XX/100 | |
| 239 | | JS Rendering | pass/warn/fail | XX/100 | |
| 240 | | IndexNow | pass/warn/fail | XX/100 | |
| 241 | |
| 242 | ### Critical Issues (fix immediately) |
| 243 | ### High Priority (fix within 1 week) |
| 244 | ### Medium Priority (fix within 1 month) |
| 245 | ### Low Priority (backlog) |
| 246 | |
| 247 | ## DataForSEO Integration (Optional) |
| 248 | |
| 249 | If DataForSEO MCP tools are available, use `on_page_instant_pages` for real page analysis (status codes, page timing, broken links, on-page checks), `on_page_lighthouse` for Lighthouse audits (performance, accessibility, SEO scores), and `domain_analytics_technologies_domain_technologies` for technology stack detection. |
| 250 | |
| 251 | ## Google API Integration (Optional) |
| 252 | |
| 253 | If Google API credentials are configured, use `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run pagespeed_check.py <url> --json` for real PSI + CrUX field data (replaces lab-only CWV estimates), `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run crux_history.py <url> --json` for 25-week CWV trends, and `"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run gsc_inspect.py <url> --json` for real indexation status per URL. |
| 254 | |
| 255 | ## Auditing a Local or Private Host |
| 256 | |
| 257 | `url_safety` refuses loopback and private addresses by default, so `http://localhost:3000` and a staging host on Tailscale fail with "Blocked hostname" or "Blocked IP literal". That default is deliberate: these scripts follow URLs found on the pages they crawl. |
| 258 | |
| 259 | To audit a pre-deployment host, the operator names it in `CLAUDE_SEO_LOCAL_TARGETS`, a comma-separated list of `host` or `host:port` entries: |
| 260 | |
| 261 | |
| 262 | CLAUDE_SEO_LOCAL_TARGETS="localhost:3000,127.0.0.1:8080,100.101.102.103" \ |
| 263 | "${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run fetch_page.py http://localhost:3000/ |
| 264 | |
| 265 | |
| 266 | What it does and does not cover: |
| 267 | |
| 268 | | Behaviour | Allowlisted host | |
| 269 | |-----------|------------------| |
| 270 | | First, top-level URL over raw HTTP | Allowed | |
| 271 | | Redirect target reached from that URL | Refused | |
| 272 | | Subresource fetched by a rendered page | Refused | |
| 273 | | Playwright renders (`--render`, screenshots) | Refused; use the raw-HTTP path | |
| 274 | | A host not named in the variable | Refused | |
| 275 | | Cloud metadata endpoints, even when listed | Refused | |
| 276 | |
| 277 | `host:port` matches that port only; a bare `host` matches any port. With the variable unset the policy is unchanged. Never suggest setting it for a host the user does not control. See SECURITY.md. |
| 278 | |
| 279 | ## Error Handling |
| 280 | |
| 281 | | Scenario | Action | |
| 282 | |----------|--------| |
| 283 | | URL unreachable | Report connection error with status code. Suggest verifying URL, checking DNS resolution, and confirming the site is publicly accessible. | |
| 284 | | robots.txt not found | Note that no robots.txt was detected at the root domain. Recommend creating one with appropriate directives. Continue audit on remaining categories. | |
| 285 | | HTTPS not configured | Flag as a critical issue. Report whether HTTP is served without redirect, mixed content exists, or SSL certificate is missing/expired. | |
| 286 | | Core Web Vitals data unavailable | Note that CrUX data is not available (common for low-traffic sites). Suggest using Lighthouse lab data as a proxy and recommend increasing traffic before re-testing. | |
| 287 |