Positive reply scoring skill

Pulls replies from a Smartlead campaign, classifies each as positive/neutral/negative/OOO/bounce/unsubscribe using Claude, and reports the positive reply rate — the north-star metric for cold email.

by growthenginenowoslawski·MIT license·★ 736 Stars on the repo·GitHub ↗

Use now

Files of Positive reply scoring

growthenginenowoslawski/main1 file shown
SKILL.md
Show the full text197 lines

Positive Reply Scoring

Reply rate tells you if people are paying attention. Positive reply rate tells you if they want what you're selling. This skill computes the second.

Why this exists

A campaign can get 5% reply rate and still be a disaster. If 90% of those replies are "unsubscribe" and "not a fit," you're burning your domains for nothing.

The metric that matters is:

positive_reply_rate = positive_replies / total_sent

Compared side-by-side:

  • Campaign A: 1% reply rate, 70% positive → 0.7% positive reply rate
  • Campaign B: 5% reply rate, 10% positive → 0.5% positive reply rate
  • Campaign A wins.

Classification schema

Every reply is classified into exactly one bucket:

Label Meaning Count as "positive"?
positive_interested "Yes, tell me more" or booked a meeting ✅
positive_soft "Send more info" / "reach out in Q3" / info request ✅
positive_referral "Not me, but talk to X" ✅ (referral is high-value)
neutral_question Clarifying question, no commitment yet ❌ (optional — some score as half)
negative_notnow "Not right now, maybe later" ❌
negative_notfit "Not a fit" / "we don't need this" ❌
negative_hostile Angry reply, complaint, report ❌ (and track separately as risk signal)
unsubscribe Explicit opt-out ❌
ooo Out-of-office auto-reply ❌ (exclude from denominators)
bounce Technical bounce ❌ (exclude from denominators)
other Can't tell ❌

Positive reply rate = (positive_interested + positive_soft + positive_referral) / total_sent

Inputs

  • Smartlead API key (env: SMARTLEAD_API_KEY)
  • Campaign ID to score
  • Optional: client_id (if using a sub-client setup)
  • Optional: date range (defaults to full campaign)

Steps

1. Fetch all leads + replies from the campaign

Run the fetch script:

npx tsx scripts/fetch-campaign-replies.ts --campaign-id=12345 --out=/tmp/replies.json

This walks /campaigns/{id}/leads paginated, identifies leads with replies (has_reply = true), then fetches /campaigns/{id}/leads/{lead_id}/message-history for each, and writes them to a JSON file with one object per reply.

Output schema per reply:

{
  "lead_id": "...",
  "email": "...",
  "lead_first_name": "...",
  "company": "...",
  "reply_time": "ISO timestamp",
  "reply_subject": "...",
  "reply_body": "... full text ...",
  "sequence_step": 1
}
2. Classify replies in the Claude Code conversation

Once the JSON is written, Claude (the one running this skill) reads the file and classifies each reply. For speed, fan out in batches of 20-30 via the Task tool (see personalization-subagent-pattern skill for fan-out mechanics).

Classification prompt (per batch):

Classify each reply as one of:
- positive_interested, positive_soft, positive_referral
- neutral_question
- negative_notnow, negative_notfit, negative_hostile
- unsubscribe, ooo, bounce, other

For each reply, output: { lead_id, label, confidence: 0.0-1.0, one_line_reason }

Rules:
- OOO auto-replies ("I'm out of office") → ooo
- Bounces (delivery failure messages) → bounce
- "Take me off your list", "unsubscribe", "STOP" → unsubscribe
- "Not interested", "not a fit" → negative_notfit
- "Not right now, circle back in Q3" → negative_notnow
- "Try [other person]" → positive_referral
- "Send more info" or "Tell me more" → positive_soft
- "Yes, let's book a call", "what times work" → positive_interested
- Insults, reports, legal threats → negative_hostile

If confidence < 0.7, label as `other`.
3. Aggregate + compute rates

Run the aggregator:

npx tsx scripts/aggregate-scores.ts --replies=/tmp/classified-replies.json --campaign-id=12345

Output (to stdout + optional --out):

Campaign 12345 — Positive Reply Scoring

Total sent:              5,284
Total replies:              212 (4.01%)
  ooo/bounce (excluded):    34
  Net replies:             178

Breakdown:
  positive_interested:    22
  positive_soft:          31
  positive_referral:       8
  neutral_question:       14
  negative_notnow:        28
  negative_notfit:        52
  negative_hostile:        3
  unsubscribe:            20
  other:                   0

Positive reply rate:     1.15% (61 / 5,284)
Positive % of replies:   34.3% (61 / 178)
Negative hostile risk:    0.06% (3 / 5,284)
Unsub rate:              0.38% (20 / 5,284)

Benchmarks (B2B cold email):
  Good positive reply rate: ≥1%
  Great: ≥2%
  Hostile >0.3% or unsub >2% → deliverability risk, pause campaign
4. Save to disk

Write aggregate results to:

~/cold-email-ai-skills/profiles/<business-slug>/scores/<campaign-id>-<YYYY-MM-DD>.json

This builds a history so you can trend positive reply rate over campaigns.

5. Flag action items

At the end, surface:

  • Positive replies that need a human response — list the top 10 positive_interested leads and their reply bodies. The user should reply to these within 30 seconds of seeing this report.
  • Referrals that need follow-up — positive_referral labels. Add the referred contacts to a new outreach list.
  • Hostile flags — any negative_hostile replies. Read them manually; consider pausing the inbox if someone is genuinely angry.
  • Unsubscribes — confirm they're globally suppressed (Smartlead does this automatically, but double-check).

When to use this skill

  • After a campaign has run for at least 14 days (otherwise sample is too small)
  • When comparing two campaigns in an experiment (use the same cutoff date for both)
  • Weekly as a quality check on running campaigns
  • Before deciding to kill or scale a campaign

Common gotchas

  • Don't trust reply rate alone. A 5% reply rate from spam-trap replies and unsubscribes is worse than a 2% reply rate from real buyers.
  • Exclude OOO + bounce from denominators. They're not real replies. The script does this automatically.
  • Smartlead's built-in AI categorization exists but is less controllable. This skill uses Claude directly for transparency and prompt-tunable classification.
  • Small samples lie. Below ~500 sent, the positive reply rate has too much noise. Wait for more volume before declaring winners/losers.
  • Classify only FIRST reply per lead. If a lead replied, you replied, they replied again — only the first reply is the signal. Later messages are the conversation, not the scoring.

What to do next

Respond to every positive_interested reply within 30 seconds of seeing it. Then /experiment-design to plan the next iteration based on what worked.

If positive reply rate is <1% after 200+ sends: the 1% rule failed. Run /email-deliverability-audit (are you reaching the inbox?) (check for vague CTAs, generic first lines, em dashes).

Or wait: this skill is the Wednesday task in /cold-email-weekly-rhythm. Run it weekly going forward.

  • /experiment-design — uses positive reply rate as the success metric
  • /email-deliverability-audit — if hostile + unsub are elevated, run this next
  • /cold-email-starter-kit → 10-reply-handling.md for what to do with the positive replies once flagged

Scripts

  • scripts/fetch-campaign-replies.ts — pulls replies via Smartlead API
  • scripts/aggregate-scores.ts — computes rates from classified JSON
1---
2name: positive-reply-scoring
3description: Pulls replies from a Smartlead campaign, classifies each as positive/neutral/negative/OOO/bounce/unsubscribe using Claude, and reports the positive reply rate — the north-star metric for cold email. Use when the user wants to know if a campaign is actually working (not just getting replies, but getting the RIGHT replies). Triggers on "score my replies", "how's campaign X doing", "positive reply rate", "is this campaign working".
4---
5 
6# Positive Reply Scoring
7 
8**Reply rate tells you if people are paying attention. Positive reply rate tells you if they want what you're selling.** This skill computes the second.
9 
10## Why this exists
11 
12A campaign can get 5% reply rate and still be a disaster. If 90% of those replies are "unsubscribe" and "not a fit," you're burning your domains for nothing.
13 
14The metric that matters is:
15```
16positive_reply_rate = positive_replies / total_sent
17```
18 
19Compared side-by-side:
20- Campaign A: 1% reply rate, 70% positive → 0.7% positive reply rate
21- Campaign B: 5% reply rate, 10% positive → 0.5% positive reply rate
22- Campaign A wins.
23 
24## Classification schema
25 
26Every reply is classified into exactly one bucket:
27 
28| Label | Meaning | Count as "positive"? |
29|---|---|---|
30| `positive_interested` | "Yes, tell me more" or booked a meeting | ✅ |
31| `positive_soft` | "Send more info" / "reach out in Q3" / info request | ✅ |
32| `positive_referral` | "Not me, but talk to X" | ✅ (referral is high-value) |
33| `neutral_question` | Clarifying question, no commitment yet | ❌ (optional — some score as half) |
34| `negative_notnow` | "Not right now, maybe later" | ❌ |
35| `negative_notfit` | "Not a fit" / "we don't need this" | ❌ |
36| `negative_hostile` | Angry reply, complaint, report | ❌ (and track separately as risk signal) |
37| `unsubscribe` | Explicit opt-out | ❌ |
38| `ooo` | Out-of-office auto-reply | ❌ (exclude from denominators) |
39| `bounce` | Technical bounce | ❌ (exclude from denominators) |
40| `other` | Can't tell | ❌ |
41 
42Positive reply rate = (positive_interested + positive_soft + positive_referral) / total_sent
43 
44## Inputs
45 
46- Smartlead API key (env: `SMARTLEAD_API_KEY`)
47- Campaign ID to score
48- Optional: client_id (if using a sub-client setup)
49- Optional: date range (defaults to full campaign)
50 
51## Steps
52 
53### 1. Fetch all leads + replies from the campaign
54 
55Run the fetch script:
56 
57```bash
58npx tsx scripts/fetch-campaign-replies.ts --campaign-id=12345 --out=/tmp/replies.json
59```
60 
61This walks `/campaigns/{id}/leads` paginated, identifies leads with replies (`has_reply = true`), then fetches `/campaigns/{id}/leads/{lead_id}/message-history` for each, and writes them to a JSON file with one object per reply.
62 
63Output schema per reply:
64```json
65{
66 "lead_id": "...",
67 "email": "...",
68 "lead_first_name": "...",
69 "company": "...",
70 "reply_time": "ISO timestamp",
71 "reply_subject": "...",
72 "reply_body": "... full text ...",
73 "sequence_step": 1
74}
75```
76 
77### 2. Classify replies in the Claude Code conversation
78 
79Once the JSON is written, Claude (the one running this skill) reads the file and classifies each reply. For speed, fan out in batches of 20-30 via the Task tool (see `personalization-subagent-pattern` skill for fan-out mechanics).
80 
81Classification prompt (per batch):
82 
83```
84Classify each reply as one of:
85- positive_interested, positive_soft, positive_referral
86- neutral_question
87- negative_notnow, negative_notfit, negative_hostile
88- unsubscribe, ooo, bounce, other
89 
90For each reply, output: { lead_id, label, confidence: 0.0-1.0, one_line_reason }
91 
92Rules:
93- OOO auto-replies ("I'm out of office") → ooo
94- Bounces (delivery failure messages) → bounce
95- "Take me off your list", "unsubscribe", "STOP" → unsubscribe
96- "Not interested", "not a fit" → negative_notfit
97- "Not right now, circle back in Q3" → negative_notnow
98- "Try [other person]" → positive_referral
99- "Send more info" or "Tell me more" → positive_soft
100- "Yes, let's book a call", "what times work" → positive_interested
101- Insults, reports, legal threats → negative_hostile
102 
103If confidence < 0.7, label as `other`.
104```
105 
106### 3. Aggregate + compute rates
107 
108Run the aggregator:
109 
110```bash
111npx tsx scripts/aggregate-scores.ts --replies=/tmp/classified-replies.json --campaign-id=12345
112```
113 
114Output (to stdout + optional `--out`):
115 
116```
117Campaign 12345 — Positive Reply Scoring
118 
119Total sent: 5,284
120Total replies: 212 (4.01%)
121 ooo/bounce (excluded): 34
122 Net replies: 178
123 
124Breakdown:
125 positive_interested: 22
126 positive_soft: 31
127 positive_referral: 8
128 neutral_question: 14
129 negative_notnow: 28
130 negative_notfit: 52
131 negative_hostile: 3
132 unsubscribe: 20
133 other: 0
134 
135Positive reply rate: 1.15% (61 / 5,284)
136Positive % of replies: 34.3% (61 / 178)
137Negative hostile risk: 0.06% (3 / 5,284)
138Unsub rate: 0.38% (20 / 5,284)
139 
140Benchmarks (B2B cold email):
141 Good positive reply rate: ≥1%
142 Great: ≥2%
143 Hostile >0.3% or unsub >2% → deliverability risk, pause campaign
144```
145 
146### 4. Save to disk
147 
148Write aggregate results to:
149```
150~/cold-email-ai-skills/profiles/<business-slug>/scores/<campaign-id>-<YYYY-MM-DD>.json
151```
152 
153This builds a history so you can trend positive reply rate over campaigns.
154 
155### 5. Flag action items
156 
157At the end, surface:
158 
159- **Positive replies that need a human response** — list the top 10 `positive_interested` leads and their reply bodies. The user should reply to these within 30 seconds of seeing this report.
160- **Referrals that need follow-up** — `positive_referral` labels. Add the referred contacts to a new outreach list.
161- **Hostile flags** — any `negative_hostile` replies. Read them manually; consider pausing the inbox if someone is genuinely angry.
162- **Unsubscribes** — confirm they're globally suppressed (Smartlead does this automatically, but double-check).
163 
164## When to use this skill
165 
166- After a campaign has run for at least 14 days (otherwise sample is too small)
167- When comparing two campaigns in an experiment (use the same cutoff date for both)
168- Weekly as a quality check on running campaigns
169- Before deciding to kill or scale a campaign
170 
171## Common gotchas
172 
173- **Don't trust reply rate alone.** A 5% reply rate from spam-trap replies and unsubscribes is worse than a 2% reply rate from real buyers.
174- **Exclude OOO + bounce from denominators.** They're not real replies. The script does this automatically.
175- **Smartlead's built-in AI categorization** exists but is less controllable. This skill uses Claude directly for transparency and prompt-tunable classification.
176- **Small samples lie.** Below ~500 sent, the positive reply rate has too much noise. Wait for more volume before declaring winners/losers.
177- **Classify only FIRST reply per lead.** If a lead replied, you replied, they replied again — only the first reply is the signal. Later messages are the conversation, not the scoring.
178 
179## What to do next
180 
181**Respond to every `positive_interested` reply within 30 seconds** of seeing it. Then `/experiment-design` to plan the next iteration based on what worked.
182 
183**If positive reply rate is <1% after 200+ sends:** the 1% rule failed. Run `/email-deliverability-audit` (are you reaching the inbox?) (check for vague CTAs, generic first lines, em dashes).
184 
185**Or wait:** this skill is the Wednesday task in `/cold-email-weekly-rhythm`. Run it weekly going forward.
186 
187## Related skills
188 
189- `/experiment-design` — uses positive reply rate as the success metric
190- `/email-deliverability-audit` — if hostile + unsub are elevated, run this next
191- `/cold-email-starter-kit` → `10-reply-handling.md` for what to do with the positive replies once flagged
192 
193## Scripts
194 
195- `scripts/fetch-campaign-replies.ts` — pulls replies via Smartlead API
196- `scripts/aggregate-scores.ts` — computes rates from classified JSON
197 

Discussion