GEO query finder skill

Find which ChatGPT search queries mention a given brand.

by OpenClaudia·MIT license·★ 705 Stars on the repo·GitHub ↗

Use now

Files of GEO query finder

OpenClaudia/main1 file shown
SKILL.md
Show the full text157 lines

GEO Query Finder

Find which ChatGPT search queries mention a given brand. Tests long-tail queries against ChatGPT's web-search-enabled model and reports which ones surface the brand.

Trigger

Use when the user asks to "find queries for [brand]", "check GEO visibility", "which queries mention [brand]", "geo query finder", "find AI mentions", or "test ChatGPT queries for [brand]".

Usage

/geo-query-finder <brand_name> [--industry <industry>] [--features <feature1,feature2,...>] [--queries <custom_query1;custom_query2;...>]

Examples:

  • /geo-query-finder "Acme Corp" — auto-researches the brand and generates queries
  • /geo-query-finder "Acme Corp" --industry "smart TV OS" --features "white-label,voice-control,OEM licensing"
  • /geo-query-finder "Acme Corp" --queries "best regulatory AI;eCTD validation tool;pharma compliance software"

How It Works

Step 0: Pull pre-indexed LLM mentions (DataForSEO) — do this FIRST

Before generating speculative queries, check if DataForSEO already has indexed mentions for the brand's domain. If it does, you get ground-truth queries with search volume in one call instead of burning OpenAI dollars guessing.

Auth via DATAFORSEO_LOGIN / DATAFORSEO_PASSWORD environment variables.

AUTH=$(printf '%s' "$DATAFORSEO_LOGIN:$DATAFORSEO_PASSWORD" | base64)
# Google AI Overview citations
curl -s -X POST "https://api.dataforseo.com/v3/ai_optimization/llm_mentions/search/live" \
  -H "Authorization: Basic $AUTH" -H "Content-Type: application/json" \
  -d '[{"target":[{"domain":"<DOMAIN>","search_filter":"include","include_subdomains":true}],"platform":"google","limit":700}]'
# ChatGPT citations (substitute "platform":"chat_gpt")

Critical flags:

  • "include_subdomains": true — without it, apex domains return 0 results (www.X treated as a different domain).
  • Omit location_code to get global results; add "location_code": 2840 only to scope to US.
  • platform options: "google" (AI Overview), "chat_gpt". Perplexity is NOT supported via this dataset.

Extract from each items[]:

  • question — the real search query where the brand was cited
  • ai_search_volume — monthly AI search volume (use to prioritize)
  • sources[] — entries with domain matching the brand have the exact cited URL
  • location_code, language_code, model_name — for geo/locale breakdown
  • answer — the LLM answer text (for context)

Decision rule:

  • If ≥20 queries returned → skip Steps 1–4 entirely; report these as ground-truth mentions and focus Step 5 on gap analysis (sort by volume, find URL-section winners like /guides/ vs /tools/).
  • If <20 queries → use them as seed input for Step 2 (generate variations of the query themes DataForSEO already confirmed), then run Steps 3–4 only on the gaps.
  • If 0 queries → the domain has no AI citations; proceed with the original Steps 1–5 (speculative testing) as fallback.
Step 1: Research the Brand

If no --industry or --features provided, use web search to understand:

  • What the brand does / what industry it's in
  • Key differentiators vs competitors
  • Unique features that competitors DON'T have
Step 2: Generate Long-Tail Queries

Generate 15-20 long-tail queries across these categories:

  1. Feature-specific (unique capabilities only this brand has)
  2. B2B/decision-maker (queries from buyers, not consumers)
  3. Problem-solving ("how to X without Y")
  4. Comparison/alternative ("alternative to [dominant player]")
  5. Use-case specific (niche scenarios where the brand excels)

Avoid generic queries where dominant players will always win.

Step 3: Query ChatGPT via OpenAI Search API

Use OpenAI's gpt-4o-search-preview model with web search enabled:

OPENAI_API_KEY from environment variable
import json, os, urllib.request, ssl

OPENAI_API_KEY = os.environ["OPENAI_API_KEY"]

data = json.dumps({
    "model": "gpt-4o-search-preview",
    "web_search_options": {"search_context_size": "medium"},
    "messages": [{"role": "user", "content": "<query>"}],
    "max_tokens": 1000
}).encode()

req = urllib.request.Request(
    "https://api.openai.com/v1/chat/completions",
    data=data,
    headers={
        "Authorization": f"Bearer {OPENAI_API_KEY}",
        "Content-Type": "application/json"
    }
)

resp = urllib.request.urlopen(req, context=ssl.create_default_context(), timeout=45)
result = json.loads(resp.read())
answer = result["choices"][0]["message"]["content"]
Step 4: Check Mentions

For each query, check if the brand name (or known aliases) appears in ChatGPT's response:

  • Check case-insensitive match
  • Check variations (with/without spaces, dots, hyphens)
  • If mentioned, extract the surrounding context (200 chars around the mention)
  • Note the position (is it #1 recommended? listed among many? mentioned in passing?)
Step 5: Report Results

Output a summary table:

## GEO Query Finder Results: [Brand Name]

### Mentioned (X/N queries)
| Query | Position | Context |
|-------|----------|---------|
| ... | #1 | "Brand is the leading..." |

### Not Mentioned (Y/N queries)
| Query | What ChatGPT Recommended Instead |
|-------|----------------------------------|
| ... | Competitor A, Competitor B |

### Recommendations
- Queries where brand is ALREADY mentioned: create more authoritative content to maintain/improve position
- Queries where brand is NOT mentioned but SHOULD be: these are content gaps — create targeted pages
- Queries to AVOID: too generic, dominated by big players, not worth the effort

Rate Limiting

  • Run queries sequentially with 1-2 second delays to avoid rate limits
  • Each query costs ~$0.01 via OpenAI API
  • Default: 15-20 queries per run (~$0.15-0.20 per run)

Notes

  • Results reflect ChatGPT with web search enabled (grounded in real-time web results)
  • Results may vary slightly between runs due to search freshness
  • This tests ChatGPT specifically — Gemini and Copilot may give different results
  • For ongoing monitoring, consider scheduling periodic runs to track visibility changes over time
1---
2name: geo-query-finder
3description: >
4 Find which ChatGPT search queries mention a given brand. Tests long-tail
5 queries against ChatGPT's web-search-enabled model and reports which ones
6 surface the brand. Use when the user asks to "find queries for [brand]",
7 "check GEO visibility", "which queries mention [brand]", "geo query finder",
8 "find AI mentions", or "test ChatGPT queries for [brand]".
9---
10 
11# GEO Query Finder
12 
13Find which ChatGPT search queries mention a given brand. Tests long-tail queries against ChatGPT's web-search-enabled model and reports which ones surface the brand.
14 
15## Trigger
16 
17Use when the user asks to "find queries for [brand]", "check GEO visibility", "which queries mention [brand]", "geo query finder", "find AI mentions", or "test ChatGPT queries for [brand]".
18 
19## Usage
20 
21```
22/geo-query-finder <brand_name> [--industry <industry>] [--features <feature1,feature2,...>] [--queries <custom_query1;custom_query2;...>]
23```
24 
25**Examples:**
26- `/geo-query-finder "Acme Corp"` — auto-researches the brand and generates queries
27- `/geo-query-finder "Acme Corp" --industry "smart TV OS" --features "white-label,voice-control,OEM licensing"`
28- `/geo-query-finder "Acme Corp" --queries "best regulatory AI;eCTD validation tool;pharma compliance software"`
29 
30## How It Works
31 
32### Step 0: Pull pre-indexed LLM mentions (DataForSEO) — do this FIRST
33 
34Before generating speculative queries, check if DataForSEO already has indexed mentions for the brand's domain. If it does, you get ground-truth queries with search volume in one call instead of burning OpenAI dollars guessing.
35 
36Auth via `DATAFORSEO_LOGIN` / `DATAFORSEO_PASSWORD` environment variables.
37 
38```bash
39AUTH=$(printf '%s' "$DATAFORSEO_LOGIN:$DATAFORSEO_PASSWORD" | base64)
40# Google AI Overview citations
41curl -s -X POST "https://api.dataforseo.com/v3/ai_optimization/llm_mentions/search/live" \
42 -H "Authorization: Basic $AUTH" -H "Content-Type: application/json" \
43 -d '[{"target":[{"domain":"<DOMAIN>","search_filter":"include","include_subdomains":true}],"platform":"google","limit":700}]'
44# ChatGPT citations (substitute "platform":"chat_gpt")
45```
46 
47**Critical flags:**
48- `"include_subdomains": true` — without it, apex domains return 0 results (www.X treated as a different domain).
49- Omit `location_code` to get global results; add `"location_code": 2840` only to scope to US.
50- `platform` options: `"google"` (AI Overview), `"chat_gpt"`. Perplexity is NOT supported via this dataset.
51 
52**Extract from each `items[]`:**
53- `question` — the real search query where the brand was cited
54- `ai_search_volume` — monthly AI search volume (use to prioritize)
55- `sources[]` — entries with `domain` matching the brand have the exact cited URL
56- `location_code`, `language_code`, `model_name` — for geo/locale breakdown
57- `answer` — the LLM answer text (for context)
58 
59**Decision rule:**
60- If ≥20 queries returned → skip Steps 1–4 entirely; report these as ground-truth mentions and focus Step 5 on gap analysis (sort by volume, find URL-section winners like `/guides/` vs `/tools/`).
61- If <20 queries → use them as seed input for Step 2 (generate variations of the query themes DataForSEO already confirmed), then run Steps 3–4 only on the gaps.
62- If 0 queries → the domain has no AI citations; proceed with the original Steps 1–5 (speculative testing) as fallback.
63 
64### Step 1: Research the Brand
65If no `--industry` or `--features` provided, use web search to understand:
66- What the brand does / what industry it's in
67- Key differentiators vs competitors
68- Unique features that competitors DON'T have
69 
70### Step 2: Generate Long-Tail Queries
71Generate 15-20 long-tail queries across these categories:
721. **Feature-specific** (unique capabilities only this brand has)
732. **B2B/decision-maker** (queries from buyers, not consumers)
743. **Problem-solving** ("how to X without Y")
754. **Comparison/alternative** ("alternative to [dominant player]")
765. **Use-case specific** (niche scenarios where the brand excels)
77 
78Avoid generic queries where dominant players will always win.
79 
80### Step 3: Query ChatGPT via OpenAI Search API
81 
82Use OpenAI's `gpt-4o-search-preview` model with web search enabled:
83 
84```bash
85OPENAI_API_KEY from environment variable
86```
87 
88```python
89import json, os, urllib.request, ssl
90 
91OPENAI_API_KEY = os.environ["OPENAI_API_KEY"]
92 
93data = json.dumps({
94 "model": "gpt-4o-search-preview",
95 "web_search_options": {"search_context_size": "medium"},
96 "messages": [{"role": "user", "content": "<query>"}],
97 "max_tokens": 1000
98}).encode()
99 
100req = urllib.request.Request(
101 "https://api.openai.com/v1/chat/completions",
102 data=data,
103 headers={
104 "Authorization": f"Bearer {OPENAI_API_KEY}",
105 "Content-Type": "application/json"
106 }
107)
108 
109resp = urllib.request.urlopen(req, context=ssl.create_default_context(), timeout=45)
110result = json.loads(resp.read())
111answer = result["choices"][0]["message"]["content"]
112```
113 
114### Step 4: Check Mentions
115 
116For each query, check if the brand name (or known aliases) appears in ChatGPT's response:
117- Check case-insensitive match
118- Check variations (with/without spaces, dots, hyphens)
119- If mentioned, extract the surrounding context (200 chars around the mention)
120- Note the position (is it #1 recommended? listed among many? mentioned in passing?)
121 
122### Step 5: Report Results
123 
124Output a summary table:
125 
126```
127## GEO Query Finder Results: [Brand Name]
128 
129### Mentioned (X/N queries)
130| Query | Position | Context |
131|-------|----------|---------|
132| ... | #1 | "Brand is the leading..." |
133 
134### Not Mentioned (Y/N queries)
135| Query | What ChatGPT Recommended Instead |
136|-------|----------------------------------|
137| ... | Competitor A, Competitor B |
138 
139### Recommendations
140- Queries where brand is ALREADY mentioned: create more authoritative content to maintain/improve position
141- Queries where brand is NOT mentioned but SHOULD be: these are content gaps — create targeted pages
142- Queries to AVOID: too generic, dominated by big players, not worth the effort
143```
144 
145## Rate Limiting
146 
147- Run queries sequentially with 1-2 second delays to avoid rate limits
148- Each query costs ~$0.01 via OpenAI API
149- Default: 15-20 queries per run (~$0.15-0.20 per run)
150 
151## Notes
152 
153- Results reflect ChatGPT with web search enabled (grounded in real-time web results)
154- Results may vary slightly between runs due to search freshness
155- This tests ChatGPT specifically — Gemini and Copilot may give different results
156- For ongoing monitoring, consider scheduling periodic runs to track visibility changes over time
157 

Discussion

Alternatives