Messaging A/B Tester
Generate 3-5 messaging variants for a value proposition, design structured A/B tests, and analyze results to determine which framing resonates most with ICP.
How to use it
- Hit Copy the whole skill.
- Claude: ⋯ → Download .md, then Customize → Skills → Add → Upload skill.
ChatGPT: make a Project and paste it into Instructions.
Neither? Paste it at the top of a new chat — it works for that chat. - Describe your job in plain words. The AI follows the skill from there.
npx degit gooseworks-ai/goose-skills/skills/brand/composites/messaging-ab-tester#main ~/.claude/skills/messaging-ab-testerFor one project only, change the path to .claude/skills/messaging-ab-tester.
Not working?
- Check which app you pasted it into — the steps above name the right one.
- Some skills need the paid tier of Claude or ChatGPT.
Paste into Claude, ChatGPT or Cursor.
Show the full text265 lines
Messaging A/B Tester
Stop debating which message is better — test it. Generate messaging variants, deploy them through real channels, and measure which framing actually resonates with your ICP.
Core principle: At seed/Series A, you don't have enough traffic for website A/B tests. But you do have enough LinkedIn impressions and cold email sends to test messaging angles fast.
When to Use
- "Which of these value props should we lead with?"
- "Test our messaging angles and tell me which works"
- "I can't decide between [message A] and [message B]"
- "What messaging resonates most with [ICP]?"
- "Run a messaging test for [product/feature]"
Phase 0: Intake
What to Test
- Core value prop — The claim or positioning you want to test (e.g., "We help growth teams run outbound 10x faster")
- Test goal — What are you deciding? (Headline for website, cold email angle, LinkedIn content strategy, ad copy direction)
- ICP — Who should this resonate with? (Title, company type, stage)
- Current messaging — What are you using today? (Baseline to beat)
Test Channel
- Where to test:
- LinkedIn organic — Post variants across consecutive days, compare engagement
- Cold email — A/B test subject lines or opening hooks via Smartlead
- Both — Run in parallel for fastest signal
- Sample size available:
- LinkedIn: followers/typical impressions per post
- Email: list size available for testing
Constraints
- Number of variants — 3-5 recommended (more = slower signal)
- Test duration — How long to run? (Default: 1 week for LinkedIn, 3-5 days for email)
Phase 1: Generate Messaging Variants
Create 3-5 variants that test different angles, not just different words. Each variant should represent a distinct strategic bet:
Variant Types
| Type | What It Tests | Example |
|---|---|---|
| Outcome-driven | Leading with the result | "3x your pipeline in 30 days" |
| Pain-driven | Leading with the problem | "Tired of spending 4 hours a day on manual prospecting?" |
| Identity-driven | Leading with who they are | "Built for growth teams who move fast" |
| Proof-driven | Leading with evidence | "How [Customer] went from 10 to 50 demos/month" |
| Contrast-driven | Leading with what you're not | "Not another CRM. An outbound engine." |
Variant Template
For each variant:
VARIANT [N]: [Type — e.g., "Outcome-driven"]
Hypothesis: This framing will resonate because [reasoning tied to ICP psychology]
LinkedIn post version:
---
[Full post copy — 100-200 words, native LinkedIn format]
---
Email subject line version:
[Subject line — max 50 chars]
Email opening hook version:
[First 2 sentences of an email]
Headline version:
[Website headline — max 10 words]
Phase 2: Deploy Tests
Option A: LinkedIn Organic Test
Setup:
- Schedule variants as consecutive posts (1 per day, same time of day)
- Each post should be similar length and format (control for post structure)
- Don't boost any posts — organic only for clean comparison
Measurement (after 48 hours per post):
- Impressions
- Reactions (likes, celebrates, etc.)
- Comments
- Comment sentiment (positive/negative/neutral)
- Profile visits (if trackable)
- DMs received mentioning the post
Option B: Cold Email A/B Test
Setup via your outreach tool (Smartlead, Instantly, Lemlist, or any tool with A/B testing):
- Create campaign with all variants as A/B test sequences
- Split list evenly across variants (minimum 50 per variant for signal)
- Same send time, same sender, same CTA — only the messaging changes
Measurement (after 5 days):
- Open rate (tests subject line)
- Reply rate (tests full message resonance)
- Positive reply rate (tests conversion quality)
- Click rate (if link included)
Option C: Both (Recommended)
Run LinkedIn and email in parallel. Different channels may show different winners — that's valuable signal about where each message works best.
Phase 2B: Collect Results
After the test has run for the planned duration, gather your results:
How to provide data:
- Paste metrics — Copy open rates, reply rates, engagement numbers directly into the chat
- CSV export — Export campaign analytics from your outreach tool and share the file
- Screenshot — Take a screenshot of your dashboard/analytics and share it
- Manual input — Just tell the agent the numbers: "Variant A got 45% open rate and 3% reply rate, Variant B got 52% open rate and 5% reply rate"
For LinkedIn tests: Go to your post analytics (click "View analytics" on each post) and share impressions, reactions, comments, and profile visits per post.
For email tests: Export or screenshot your campaign's variant/A-B test results showing sends, opens, and replies per variant.
The agent will normalize whatever format you provide into the scoring framework below.
Phase 3: Analyze Results
Scoring Framework
| Metric | Weight (LinkedIn) | Weight (Email) |
|---|---|---|
| Engagement rate | 30% | — |
| Comment quality | 30% | — |
| Open rate | — | 30% |
| Reply rate | — | 40% |
| Positive reply rate | — | 30% |
| Impressions | 20% | — |
| Profile visits / clicks | 20% | — |
Statistical Significance Check
For email tests:
- Minimum sends per variant: 50 (for directional signal), 200+ (for confident decisions)
- Minimum difference to call a winner: >20% relative difference in primary metric
For LinkedIn tests:
- Minimum posts per variant: 1 (you're testing with limited data — treat as directional)
- Minimum impressions: 500 per post to be comparable
Winner Selection
WINNER: Variant [N] — [Type]
Primary metric: [X] (vs average of [Y] across other variants)
Relative improvement: [Z%] over baseline
Why it won:
[1-2 sentences on what this tells us about ICP messaging preferences]
Runner-up: Variant [N]
[1 sentence on when this might work better — different channel, different segment]
Phase 4: Output Format
# Messaging A/B Test Results — [DATE]
Value prop tested: [description]
ICP: [target audience]
Test duration: [dates]
---
## Test Design
| Variant | Type | Hypothesis |
|---------|------|-----------|
| A | [Type] | [Hypothesis] |
| B | [Type] | [Hypothesis] |
| C | [Type] | [Hypothesis] |
---
## Results
### LinkedIn Test
| Variant | Impressions | Reactions | Comments | Engagement Rate | Score |
|---------|------------|-----------|----------|----------------|-------|
| A | [N] | [N] | [N] | [X%] | [weighted] |
| B | [N] | [N] | [N] | [X%] | [weighted] |
| C | [N] | [N] | [N] | [X%] | [weighted] |
### Email Test
| Variant | Sends | Opens | Open Rate | Replies | Reply Rate | Positive | Score |
|---------|-------|-------|-----------|---------|------------|----------|-------|
| A | [N] | [N] | [X%] | [N] | [X%] | [N] | [weighted] |
| B | [N] | [N] | [X%] | [N] | [X%] | [N] | [weighted] |
| C | [N] | [N] | [X%] | [N] | [X%] | [N] | [weighted] |
---
## Winner: Variant [N] — "[Headline]"
**Why it won:** [Analysis — what does this tell us about how our ICP thinks?]
**Recommended deployment:**
- Website headline: "[adapted version]"
- Sales deck opening: "[adapted version]"
- LinkedIn bio: "[adapted version]"
- Cold email default: "[adapted version]"
---
## Variant Details & Copy
### Variant A: [Full copy used in test]
### Variant B: [Full copy used in test]
### Variant C: [Full copy used in test]
---
## What to Test Next
Based on these results, the next messaging test should explore:
1. [Angle suggested by results — e.g., "test more specific proof points since proof-driven won"]
2. [Segment test — e.g., "test winning message against different ICP segment"]
Save to the current working directory or wherever the user prefers.
Cost
| Component | Cost |
|---|---|
| Variant generation | Free (LLM reasoning) |
| LinkedIn posting | Free (organic) |
| Email testing | Included with your outreach tool's plan |
| Results analysis | Free (LLM reasoning) |
| Total | Free |
Tools Required
None. Pure reasoning for variant generation, test design, and result analysis. The user deploys tests through their own tools:
- LinkedIn organic — post variants manually or via scheduling tool
- Cold email — set up A/B tests in whatever outreach tool they use (Smartlead, Instantly, Lemlist, etc.)
- Results — user provides metrics (screenshots, CSV exports, or manual input) for analysis
Trigger Phrases
- "Test which messaging angle works best for [ICP]"
- "Run a messaging A/B test for [value prop]"
- "Which of these messages should we lead with?"
- "Help me decide between these positioning options"
| 1 | |
| 2 | name messaging-ab-tester |
| 3 | description > |
| 4 | Generate 3-5 messaging variants for a value proposition, design structured A/B tests, |
| 5 | and analyze results to determine which framing resonates most with ICP. Tests can run |
| 6 | via LinkedIn organic posts, cold email subject line splits, or both. Pure reasoning for |
| 7 | variant generation and analysis — the user deploys the tests through their own tools. |
| 8 | Use when a team can't decide between messaging angles and needs data, not opinions. |
| 9 | tags [brand] |
| 10 | |
| 11 | |
| 12 | # Messaging A/B Tester |
| 13 | |
| 14 | Stop debating which message is better — test it. Generate messaging variants, deploy them through real channels, and measure which framing actually resonates with your ICP. |
| 15 | |
| 16 | **Core principle:** At seed/Series A, you don't have enough traffic for website A/B tests. But you do have enough LinkedIn impressions and cold email sends to test messaging angles fast. |
| 17 | |
| 18 | ## When to Use |
| 19 | |
| 20 | "Which of these value props should we lead with?" |
| 21 | "Test our messaging angles and tell me which works" |
| 22 | "I can't decide between [message A] and [message B]" |
| 23 | "What messaging resonates most with [ICP]?" |
| 24 | "Run a messaging test for [product/feature]" |
| 25 | |
| 26 | ## Phase 0: Intake |
| 27 | |
| 28 | ### What to Test |
| 29 | **Core value prop** — The claim or positioning you want to test (e.g., "We help growth teams run outbound 10x faster") |
| 30 | **Test goal** — What are you deciding? (Headline for website, cold email angle, LinkedIn content strategy, ad copy direction) |
| 31 | **ICP** — Who should this resonate with? (Title, company type, stage) |
| 32 | **Current messaging** — What are you using today? (Baseline to beat) |
| 33 | |
| 34 | ### Test Channel |
| 35 | **Where to test:** |
| 36 | **LinkedIn organic** — Post variants across consecutive days, compare engagement |
| 37 | **Cold email** — A/B test subject lines or opening hooks via Smartlead |
| 38 | **Both** — Run in parallel for fastest signal |
| 39 | **Sample size available:** |
| 40 | LinkedIn: followers/typical impressions per post |
| 41 | Email: list size available for testing |
| 42 | |
| 43 | ### Constraints |
| 44 | **Number of variants** — 3-5 recommended (more = slower signal) |
| 45 | **Test duration** — How long to run? (Default: 1 week for LinkedIn, 3-5 days for email) |
| 46 | |
| 47 | ## Phase 1: Generate Messaging Variants |
| 48 | |
| 49 | Create 3-5 variants that test different **angles**, not just different words. Each variant should represent a distinct strategic bet: |
| 50 | |
| 51 | ### Variant Types |
| 52 | |
| 53 | | Type | What It Tests | Example | |
| 54 | |------|--------------|---------| |
| 55 | | **Outcome-driven** | Leading with the result | "3x your pipeline in 30 days" | |
| 56 | | **Pain-driven** | Leading with the problem | "Tired of spending 4 hours a day on manual prospecting?" | |
| 57 | | **Identity-driven** | Leading with who they are | "Built for growth teams who move fast" | |
| 58 | | **Proof-driven** | Leading with evidence | "How [Customer] went from 10 to 50 demos/month" | |
| 59 | | **Contrast-driven** | Leading with what you're not | "Not another CRM. An outbound engine." | |
| 60 | |
| 61 | ### Variant Template |
| 62 | |
| 63 | For each variant: |
| 64 | |
| 65 | VARIANT [N]: [Type — e.g., "Outcome-driven"] |
| 66 | |
| 67 | Hypothesis: This framing will resonate because [reasoning tied to ICP psychology] |
| 68 | |
| 69 | LinkedIn post version: |
| 70 | |
| 71 | [Full post copy — 100-200 words, native LinkedIn format] |
| 72 | |
| 73 | |
| 74 | Email subject line version: |
| 75 | [Subject line — max 50 chars] |
| 76 | |
| 77 | Email opening hook version: |
| 78 | [First 2 sentences of an email] |
| 79 | |
| 80 | Headline version: |
| 81 | [Website headline — max 10 words] |
| 82 | |
| 83 | |
| 84 | ## Phase 2: Deploy Tests |
| 85 | |
| 86 | ### Option A: LinkedIn Organic Test |
| 87 | |
| 88 | **Setup:** |
| 89 | Schedule variants as consecutive posts (1 per day, same time of day) |
| 90 | Each post should be similar length and format (control for post structure) |
| 91 | Don't boost any posts — organic only for clean comparison |
| 92 | |
| 93 | **Measurement (after 48 hours per post):** |
| 94 | Impressions |
| 95 | Reactions (likes, celebrates, etc.) |
| 96 | Comments |
| 97 | Comment sentiment (positive/negative/neutral) |
| 98 | Profile visits (if trackable) |
| 99 | DMs received mentioning the post |
| 100 | |
| 101 | ### Option B: Cold Email A/B Test |
| 102 | |
| 103 | **Setup via your outreach tool (Smartlead, Instantly, Lemlist, or any tool with A/B testing):** |
| 104 | Create campaign with all variants as A/B test sequences |
| 105 | Split list evenly across variants (minimum 50 per variant for signal) |
| 106 | Same send time, same sender, same CTA — only the messaging changes |
| 107 | |
| 108 | **Measurement (after 5 days):** |
| 109 | Open rate (tests subject line) |
| 110 | Reply rate (tests full message resonance) |
| 111 | Positive reply rate (tests conversion quality) |
| 112 | Click rate (if link included) |
| 113 | |
| 114 | ### Option C: Both (Recommended) |
| 115 | |
| 116 | Run LinkedIn and email in parallel. Different channels may show different winners — that's valuable signal about where each message works best. |
| 117 | |
| 118 | ## Phase 2B: Collect Results |
| 119 | |
| 120 | After the test has run for the planned duration, gather your results: |
| 121 | |
| 122 | **How to provide data:** |
| 123 | **Paste metrics** — Copy open rates, reply rates, engagement numbers directly into the chat |
| 124 | **CSV export** — Export campaign analytics from your outreach tool and share the file |
| 125 | **Screenshot** — Take a screenshot of your dashboard/analytics and share it |
| 126 | **Manual input** — Just tell the agent the numbers: "Variant A got 45% open rate and 3% reply rate, Variant B got 52% open rate and 5% reply rate" |
| 127 | |
| 128 | **For LinkedIn tests:** Go to your post analytics (click "View analytics" on each post) and share impressions, reactions, comments, and profile visits per post. |
| 129 | |
| 130 | **For email tests:** Export or screenshot your campaign's variant/A-B test results showing sends, opens, and replies per variant. |
| 131 | |
| 132 | The agent will normalize whatever format you provide into the scoring framework below. |
| 133 | |
| 134 | ## Phase 3: Analyze Results |
| 135 | |
| 136 | ### Scoring Framework |
| 137 | |
| 138 | | Metric | Weight (LinkedIn) | Weight (Email) | |
| 139 | |--------|-------------------|----------------| |
| 140 | | Engagement rate | 30% | — | |
| 141 | | Comment quality | 30% | — | |
| 142 | | Open rate | — | 30% | |
| 143 | | Reply rate | — | 40% | |
| 144 | | Positive reply rate | — | 30% | |
| 145 | | Impressions | 20% | — | |
| 146 | | Profile visits / clicks | 20% | — | |
| 147 | |
| 148 | ### Statistical Significance Check |
| 149 | |
| 150 | For email tests: |
| 151 | **Minimum sends per variant:** 50 (for directional signal), 200+ (for confident decisions) |
| 152 | **Minimum difference to call a winner:** >20% relative difference in primary metric |
| 153 | |
| 154 | For LinkedIn tests: |
| 155 | **Minimum posts per variant:** 1 (you're testing with limited data — treat as directional) |
| 156 | **Minimum impressions:** 500 per post to be comparable |
| 157 | |
| 158 | ### Winner Selection |
| 159 | |
| 160 | |
| 161 | WINNER: Variant [N] — [Type] |
| 162 | |
| 163 | Primary metric: [X] (vs average of [Y] across other variants) |
| 164 | Relative improvement: [Z%] over baseline |
| 165 | |
| 166 | Why it won: |
| 167 | [1-2 sentences on what this tells us about ICP messaging preferences] |
| 168 | |
| 169 | Runner-up: Variant [N] |
| 170 | [1 sentence on when this might work better — different channel, different segment] |
| 171 | |
| 172 | |
| 173 | ## Phase 4: Output Format |
| 174 | |
| 175 | |
| 176 | # Messaging A/B Test Results — [DATE] |
| 177 | Value prop tested: [description] |
| 178 | ICP: [target audience] |
| 179 | Test duration: [dates] |
| 180 | |
| 181 | |
| 182 | |
| 183 | ## Test Design |
| 184 | |
| 185 | | Variant | Type | Hypothesis | |
| 186 | |---------|------|-----------| |
| 187 | | A | [Type] | [Hypothesis] | |
| 188 | | B | [Type] | [Hypothesis] | |
| 189 | | C | [Type] | [Hypothesis] | |
| 190 | |
| 191 | |
| 192 | |
| 193 | ## Results |
| 194 | |
| 195 | ### LinkedIn Test |
| 196 | |
| 197 | | Variant | Impressions | Reactions | Comments | Engagement Rate | Score | |
| 198 | |---------|------------|-----------|----------|----------------|-------| |
| 199 | | A | [N] | [N] | [N] | [X%] | [weighted] | |
| 200 | | B | [N] | [N] | [N] | [X%] | [weighted] | |
| 201 | | C | [N] | [N] | [N] | [X%] | [weighted] | |
| 202 | |
| 203 | ### Email Test |
| 204 | |
| 205 | | Variant | Sends | Opens | Open Rate | Replies | Reply Rate | Positive | Score | |
| 206 | |---------|-------|-------|-----------|---------|------------|----------|-------| |
| 207 | | A | [N] | [N] | [X%] | [N] | [X%] | [N] | [weighted] | |
| 208 | | B | [N] | [N] | [X%] | [N] | [X%] | [N] | [weighted] | |
| 209 | | C | [N] | [N] | [X%] | [N] | [X%] | [N] | [weighted] | |
| 210 | |
| 211 | |
| 212 | |
| 213 | ## Winner: Variant [N] — "[Headline]" |
| 214 | |
| 215 | **Why it won:** [Analysis — what does this tell us about how our ICP thinks?] |
| 216 | |
| 217 | **Recommended deployment:** |
| 218 | - Website headline: "[adapted version]" |
| 219 | - Sales deck opening: "[adapted version]" |
| 220 | - LinkedIn bio: "[adapted version]" |
| 221 | - Cold email default: "[adapted version]" |
| 222 | |
| 223 | |
| 224 | |
| 225 | ## Variant Details & Copy |
| 226 | |
| 227 | ### Variant A: [Full copy used in test] |
| 228 | ### Variant B: [Full copy used in test] |
| 229 | ### Variant C: [Full copy used in test] |
| 230 | |
| 231 | |
| 232 | |
| 233 | ## What to Test Next |
| 234 | |
| 235 | Based on these results, the next messaging test should explore: |
| 236 | 1. [Angle suggested by results — e.g., "test more specific proof points since proof-driven won"] |
| 237 | 2. [Segment test — e.g., "test winning message against different ICP segment"] |
| 238 | |
| 239 | |
| 240 | Save to the current working directory or wherever the user prefers. |
| 241 | |
| 242 | ## Cost |
| 243 | |
| 244 | | Component | Cost | |
| 245 | |-----------|------| |
| 246 | | Variant generation | Free (LLM reasoning) | |
| 247 | | LinkedIn posting | Free (organic) | |
| 248 | | Email testing | Included with your outreach tool's plan | |
| 249 | | Results analysis | Free (LLM reasoning) | |
| 250 | | **Total** | **Free** | |
| 251 | |
| 252 | ## Tools Required |
| 253 | |
| 254 | None. Pure reasoning for variant generation, test design, and result analysis. The user deploys tests through their own tools: |
| 255 | **LinkedIn organic** — post variants manually or via scheduling tool |
| 256 | **Cold email** — set up A/B tests in whatever outreach tool they use (Smartlead, Instantly, Lemlist, etc.) |
| 257 | **Results** — user provides metrics (screenshots, CSV exports, or manual input) for analysis |
| 258 | |
| 259 | ## Trigger Phrases |
| 260 | |
| 261 | "Test which messaging angle works best for [ICP]" |
| 262 | "Run a messaging A/B test for [value prop]" |
| 263 | "Which of these messages should we lead with?" |
| 264 | "Help me decide between these positioning options" |
| 265 |