Kol discovery

Find Key Opinion Leaders (KOLs) in a given domain by combining web research with LinkedIn post search.

How to use it

  1. Hit Copy SKILL.md — or use the Claude Code line below to get every file.
  2. Claude: ⋯ → Download .md, then Customize → Skills → Add → Upload skill.
    ChatGPT: make a Project and paste it into Instructions.
    Neither? Paste it at the top of a new chat — it works for that chat.
  3. Describe your job in plain words. The AI follows the skill from there.
Claude Code — installs the whole folder, not just SKILL.md
npx degit gooseworks-ai/goose-skills/skills/social/capabilities/kol-discovery#main ~/.claude/skills/kol-discovery

For one project only, change the path to .claude/skills/kol-discovery. This skill also uses kol-discovery.json, kol-web-kols.json — copying SKILL.md alone won't be enough. See the folder on GitHub.

Not working?
  • Check which app you pasted it into — the steps above name the right one.
  • Some skills need the paid tier of Claude or ChatGPT.
Step-by-step guide with screenshots · Ask in the forum

Paste into Claude, ChatGPT or Cursor.

Show the full text167 lines
kol-discovery/SKILL.md167 lines6.1 KBpushed 96d agoRawView on GitHub

KOL Discovery

Find Key Opinion Leaders in any domain by searching LinkedIn posts for prolific, high-engagement authors and merging with web-researched influencers.

Core principle: Search for authority/thought-leadership keywords, not pain-language. We want people who shape conversation in the space — conference speakers, newsletter writers, podcast hosts, and prolific LinkedIn posters.

Phase 0: Intake

Ask the user these questions:

Domain & Audience

  1. What does your company/product do? What space are you in?
  2. What specific domain or topic are the KOLs you want to find expert in?
  3. Who is your target audience? (The people the KOLs influence)
  4. Any KOLs you already know about? (LinkedIn URLs — these become the baseline)
  5. Anyone to EXCLUDE? (Competitors, your own team, irrelevant voices)

Phase 1: Generate Domain Keywords

Based on intake, generate 15-25 topic/authority keywords. These are NOT pain-language — they're the terms thought leaders use when sharing expertise:

  • Industry terms — "freight tech", "supply chain innovation"
  • Thought leadership signals — "lessons learned in logistics", "future of dispatch"
  • Conference/event terms — "supply chain summit keynote"
  • Content creator signals — "newsletter freight", "podcast logistics"

Also generate:

  • KOL title keywords — titles that signal thought leadership (vp, founder, analyst, editor, host)
  • Vendor exclusion keywords — titles to filter out (software engineer, recruiter, saas)
  • Domain relevance keywords — core industry terms for relevance scoring

Present keywords to user for approval before running.

Save config in the current working directory or wherever the user prefers:

Config JSON structure:

{
  "client_name": "example",
  "domain_keywords": ["\"freight tech\" thought leadership", "supply chain innovation"],
  "exclusion_patterns": ["hiring.*position", "we.re recruiting"],
  "kol_title_keywords": ["vp", "founder", "analyst", "editor", "host"],
  "vendor_exclude_keywords": ["software engineer", "saas", "recruiter"],
  "domain_relevance_keywords": ["freight", "logistics", "supply chain"],
  "country_filter": "",
  "max_posts_per_keyword": 50,
  "min_posts": 2,
  "min_total_engagement": 50,
  "top_n_kols": 50
}

Phase 2: Run KOL Discovery Pipeline

python3 skills/kol-discovery/scripts/kol_discovery.py \
  --config kol-discovery.json \
  --output-dir . \
  [--test] [--web-kols kol-web-kols.json] [--yes]

Flags:

  • --config (required) — path to client config JSON
  • --output-dir — directory for output CSV (default: current working directory)
  • --test — limit to 5 keywords (validation run)
  • --web-kols — path to web-researched KOL JSON (agent generates this)
  • --yes — skip cost confirmation prompts
  • --max-runs — override Apify run limit

What the script does:

  1. Keyword searchapimaestro/linkedin-posts-search-scraper-no-cookies for each domain keyword
  2. Author aggregation — Group posts by author, compute engagement metrics
  3. Scoring — Composite KOL score: engagement volume (log-scaled) + consistency (post count) + quality (avg engagement) + relevance (keyword breadth) + web research bonus
  4. Merge — Combine post-data KOLs with web-researched KOLs, flag overlaps
  5. Export — Ranked CSV

Cost estimate: ~$0.10 per keyword. Full run with 20 keywords: ~$2-3.

Always run with --test first.

Phase 2b: Web Research (Agent-Driven)

Before or alongside the script, do web research to find known KOLs:

  • Search for "top [industry] influencers on LinkedIn"
  • Find conference speakers, newsletter authors, podcast hosts
  • Check industry publications for frequent contributors

Save as JSON in the current working directory:

[
  {
    "name": "Jane Doe",
    "linkedin_url": "https://www.linkedin.com/in/janedoe/",
    "source": "FreightWaves conference speaker 2025",
    "notes": "Hosts weekly logistics podcast"
  }
]

Pass to script via --web-kols.

Phase 3: Review & Refine

Present results:

  • Top 20 KOLs — rank, name, headline, KOL score, total engagement, top post
  • Source breakdown — how many from post-data vs web-research vs both
  • Keyword performance — which keywords surfaced the most KOLs

Common adjustments:

  • Too many irrelevant authors — refine domain keywords, add exclusion patterns
  • Missing known KOLs — add more keyword variants, expand web research
  • Too few results — lower min_posts or min_total_engagement thresholds

Phase 4: Output

CSV exported to the current working directory:

Column Description
Rank Overall rank by KOL Score
Name Full name
LinkedIn URL Profile link
Headline From LinkedIn
KOL Score Composite score
Total Posts Posts found in search
Total Reactions Sum of reactions across posts
Total Comments Sum of comments across posts
Avg Engagement Average reactions+comments per post
Top Post URL Highest engagement post
Top Post Preview First 100 chars of top post
Source post-data / web-research / both

Tools Required

  • Apify API token — set as APIFY_API_TOKEN in .env
  • Apify actors used:
    • apimaestro/linkedin-posts-search-scraper-no-cookies (keyword search)

Example Usage

Trigger phrases:

  • "Find KOLs in the freight/logistics space"
  • "Who are the influencers in [industry]?"
  • "Discover thought leaders for [domain]"
  • "Run KOL discovery for [client]"

With existing config:

python3 skills/kol-discovery/scripts/kol_discovery.py \
  --config clients/example/configs/kol-discovery.json \
  --output-dir clients/example/leads --yes
1---
2name: kol-discovery
3description: >
4 Find Key Opinion Leaders (KOLs) in a given domain by combining web research
5 with LinkedIn post search. Given a company/idea and target domain, generates
6 authority keywords, searches LinkedIn posts to find prolific authors with
7 high engagement, and merges with web-researched influencers. Use when someone
8 wants to "find influencers in X space" or "who are the KOLs for Y industry."
9tags: [outreach]
10---
11 
12# KOL Discovery
13 
14Find Key Opinion Leaders in any domain by searching LinkedIn posts for prolific, high-engagement authors and merging with web-researched influencers.
15 
16**Core principle:** Search for **authority/thought-leadership keywords**, not pain-language. We want people who shape conversation in the space — conference speakers, newsletter writers, podcast hosts, and prolific LinkedIn posters.
17 
18## Phase 0: Intake
19 
20Ask the user these questions:
21 
22### Domain & Audience
23 
241. What does your company/product do? What space are you in?
252. What specific domain or topic are the KOLs you want to find expert in?
263. Who is your target audience? (The people the KOLs influence)
274. Any KOLs you already know about? (LinkedIn URLs — these become the baseline)
285. Anyone to EXCLUDE? (Competitors, your own team, irrelevant voices)
29 
30## Phase 1: Generate Domain Keywords
31 
32Based on intake, generate 15-25 topic/authority keywords. These are NOT pain-language — they're the terms thought leaders use when sharing expertise:
33 
34- **Industry terms** — "freight tech", "supply chain innovation"
35- **Thought leadership signals** — "lessons learned in logistics", "future of dispatch"
36- **Conference/event terms** — "supply chain summit keynote"
37- **Content creator signals** — "newsletter freight", "podcast logistics"
38 
39Also generate:
40- **KOL title keywords** — titles that signal thought leadership (vp, founder, analyst, editor, host)
41- **Vendor exclusion keywords** — titles to filter out (software engineer, recruiter, saas)
42- **Domain relevance keywords** — core industry terms for relevance scoring
43 
44**Present keywords to user for approval before running.**
45 
46Save config in the current working directory or wherever the user prefers:
47 
48Config JSON structure:
49```json
50{
51 "client_name": "example",
52 "domain_keywords": ["\"freight tech\" thought leadership", "supply chain innovation"],
53 "exclusion_patterns": ["hiring.*position", "we.re recruiting"],
54 "kol_title_keywords": ["vp", "founder", "analyst", "editor", "host"],
55 "vendor_exclude_keywords": ["software engineer", "saas", "recruiter"],
56 "domain_relevance_keywords": ["freight", "logistics", "supply chain"],
57 "country_filter": "",
58 "max_posts_per_keyword": 50,
59 "min_posts": 2,
60 "min_total_engagement": 50,
61 "top_n_kols": 50
62}
63```
64 
65## Phase 2: Run KOL Discovery Pipeline
66 
67```bash
68python3 skills/kol-discovery/scripts/kol_discovery.py \
69 --config kol-discovery.json \
70 --output-dir . \
71 [--test] [--web-kols kol-web-kols.json] [--yes]
72```
73 
74**Flags:**
75- `--config` (required) — path to client config JSON
76- `--output-dir` — directory for output CSV (default: current working directory)
77- `--test` — limit to 5 keywords (validation run)
78- `--web-kols` — path to web-researched KOL JSON (agent generates this)
79- `--yes` — skip cost confirmation prompts
80- `--max-runs` — override Apify run limit
81 
82**What the script does:**
83 
841. **Keyword search**`apimaestro/linkedin-posts-search-scraper-no-cookies` for each domain keyword
852. **Author aggregation** — Group posts by author, compute engagement metrics
863. **Scoring** — Composite KOL score: engagement volume (log-scaled) + consistency (post count) + quality (avg engagement) + relevance (keyword breadth) + web research bonus
874. **Merge** — Combine post-data KOLs with web-researched KOLs, flag overlaps
885. **Export** — Ranked CSV
89 
90**Cost estimate:** ~$0.10 per keyword. Full run with 20 keywords: ~$2-3.
91 
92**Always run with `--test` first.**
93 
94## Phase 2b: Web Research (Agent-Driven)
95 
96Before or alongside the script, do web research to find known KOLs:
97- Search for "top [industry] influencers on LinkedIn"
98- Find conference speakers, newsletter authors, podcast hosts
99- Check industry publications for frequent contributors
100 
101Save as JSON in the current working directory:
102 
103```json
104[
105 {
106 "name": "Jane Doe",
107 "linkedin_url": "https://www.linkedin.com/in/janedoe/",
108 "source": "FreightWaves conference speaker 2025",
109 "notes": "Hosts weekly logistics podcast"
110 }
111]
112```
113 
114Pass to script via `--web-kols`.
115 
116## Phase 3: Review & Refine
117 
118Present results:
119- **Top 20 KOLs** — rank, name, headline, KOL score, total engagement, top post
120- **Source breakdown** — how many from post-data vs web-research vs both
121- **Keyword performance** — which keywords surfaced the most KOLs
122 
123Common adjustments:
124- **Too many irrelevant authors** — refine domain keywords, add exclusion patterns
125- **Missing known KOLs** — add more keyword variants, expand web research
126- **Too few results** — lower `min_posts` or `min_total_engagement` thresholds
127 
128## Phase 4: Output
129 
130CSV exported to the current working directory:
131 
132| Column | Description |
133|--------|-------------|
134| Rank | Overall rank by KOL Score |
135| Name | Full name |
136| LinkedIn URL | Profile link |
137| Headline | From LinkedIn |
138| KOL Score | Composite score |
139| Total Posts | Posts found in search |
140| Total Reactions | Sum of reactions across posts |
141| Total Comments | Sum of comments across posts |
142| Avg Engagement | Average reactions+comments per post |
143| Top Post URL | Highest engagement post |
144| Top Post Preview | First 100 chars of top post |
145| Source | post-data / web-research / both |
146 
147## Tools Required
148 
149- **Apify API token** — set as `APIFY_API_TOKEN` in `.env`
150- **Apify actors used:**
151 - `apimaestro/linkedin-posts-search-scraper-no-cookies` (keyword search)
152 
153## Example Usage
154 
155**Trigger phrases:**
156- "Find KOLs in the freight/logistics space"
157- "Who are the influencers in [industry]?"
158- "Discover thought leaders for [domain]"
159- "Run KOL discovery for [client]"
160 
161**With existing config:**
162```bash
163python3 skills/kol-discovery/scripts/kol_discovery.py \
164 --config clients/example/configs/kol-discovery.json \
165 --output-dir clients/example/leads --yes
166```
167 

Discussion

Alternatives

Also in Company & contact data