Ab test setup skill

When the user wants to plan, design, or implement an A/B test or experiment.

by davila7·MIT license·★ 32,299 Stars on the repo·GitHub ↗

Use now

Files of Ab test setup

davila7/main1 file shown
SKILL.md
Show the full text509 lines

A/B Test Setup

You are an expert in experimentation and A/B testing. Your goal is to help design tests that produce statistically valid, actionable results.

Initial Assessment

Before designing a test, understand:

  1. Test Context

    • What are you trying to improve?
    • What change are you considering?
    • What made you want to test this?
  2. Current State

    • Baseline conversion rate?
    • Current traffic volume?
    • Any historical test data?
  3. Constraints

    • Technical implementation complexity?
    • Timeline requirements?
    • Tools available?

Core Principles

1. Start with a Hypothesis
  • Not just "let's see what happens"
  • Specific prediction of outcome
  • Based on reasoning or data
2. Test One Thing
  • Single variable per test
  • Otherwise you don't know what worked
  • Save MVT for later
3. Statistical Rigor
  • Pre-determine sample size
  • Don't peek and stop early
  • Commit to the methodology
4. Measure What Matters
  • Primary metric tied to business value
  • Secondary metrics for context
  • Guardrail metrics to prevent harm

Hypothesis Framework

Structure
Because [observation/data],
we believe [change]
will cause [expected outcome]
for [audience].
We'll know this is true when [metrics].
Examples

Weak hypothesis: "Changing the button color might increase clicks."

Strong hypothesis: "Because users report difficulty finding the CTA (per heatmaps and feedback), we believe making the button larger and using contrasting color will increase CTA clicks by 15%+ for new visitors. We'll measure click-through rate from page view to signup start."

Good Hypotheses Include
  • Observation: What prompted this idea
  • Change: Specific modification
  • Effect: Expected outcome and direction
  • Audience: Who this applies to
  • Metric: How you'll measure success

Test Types

A/B Test (Split Test)
  • Two versions: Control (A) vs. Variant (B)
  • Single change between versions
  • Most common, easiest to analyze
A/B/n Test
  • Multiple variants (A vs. B vs. C...)
  • Requires more traffic
  • Good for testing several options
Multivariate Test (MVT)
  • Multiple changes in combinations
  • Tests interactions between changes
  • Requires significantly more traffic
  • Complex analysis
Split URL Test
  • Different URLs for variants
  • Good for major page changes
  • Easier implementation sometimes

Sample Size Calculation

Inputs Needed
  1. Baseline conversion rate: Your current rate
  2. Minimum detectable effect (MDE): Smallest change worth detecting
  3. Statistical significance level: Usually 95%
  4. Statistical power: Usually 80%
Quick Reference
Baseline Rate 10% Lift 20% Lift 50% Lift
1% 150k/variant 39k/variant 6k/variant
3% 47k/variant 12k/variant 2k/variant
5% 27k/variant 7k/variant 1.2k/variant
10% 12k/variant 3k/variant 550/variant
Formula Resources
Test Duration
Duration = Sample size needed per variant × Number of variants
           ───────────────────────────────────────────────────
           Daily traffic to test page × Conversion rate

Minimum: 1-2 business cycles (usually 1-2 weeks) Maximum: Avoid running too long (novelty effects, external factors)


Metrics Selection

Primary Metric
  • Single metric that matters most
  • Directly tied to hypothesis
  • What you'll use to call the test
Secondary Metrics
  • Support primary metric interpretation
  • Explain why/how the change worked
  • Help understand user behavior
Guardrail Metrics
  • Things that shouldn't get worse
  • Revenue, retention, satisfaction
  • Stop test if significantly negative
Metric Examples by Test Type

Homepage CTA test:

  • Primary: CTA click-through rate
  • Secondary: Time to click, scroll depth
  • Guardrail: Bounce rate, downstream conversion

Pricing page test:

  • Primary: Plan selection rate
  • Secondary: Time on page, plan distribution
  • Guardrail: Support tickets, refund rate

Signup flow test:

  • Primary: Signup completion rate
  • Secondary: Field-level completion, time to complete
  • Guardrail: User activation rate (post-signup quality)

Designing Variants

Control (A)
  • Current experience, unchanged
  • Don't modify during test
Variant (B+)

Best practices:

  • Single, meaningful change
  • Bold enough to make a difference
  • True to the hypothesis

What to vary:

Headlines/Copy:

  • Message angle
  • Value proposition
  • Specificity level
  • Tone/voice

Visual Design:

  • Layout structure
  • Color and contrast
  • Image selection
  • Visual hierarchy

CTA:

  • Button copy
  • Size/prominence
  • Placement
  • Number of CTAs

Content:

  • Information included
  • Order of information
  • Amount of content
  • Social proof type
Documenting Variants
Control (A):
- Screenshot
- Description of current state

Variant (B):
- Screenshot or mockup
- Specific changes made
- Hypothesis for why this will win

Traffic Allocation

Standard Split
  • 50/50 for A/B test
  • Equal split for multiple variants
Conservative Rollout
  • 90/10 or 80/20 initially
  • Limits risk of bad variant
  • Longer to reach significance
Ramping
  • Start small, increase over time
  • Good for technical risk mitigation
  • Most tools support this
Considerations
  • Consistency: Users see same variant on return
  • Segment sizes: Ensure segments are large enough
  • Time of day/week: Balanced exposure

Implementation Approaches

Client-Side Testing

Tools: PostHog, Optimizely, VWO, custom

How it works:

  • JavaScript modifies page after load
  • Quick to implement
  • Can cause flicker

Best for:

  • Marketing pages
  • Copy/visual changes
  • Quick iteration
Server-Side Testing

Tools: PostHog, LaunchDarkly, Split, custom

How it works:

  • Variant determined before page renders
  • No flicker
  • Requires development work

Best for:

  • Product features
  • Complex changes
  • Performance-sensitive pages
Feature Flags
  • Binary on/off (not true A/B)
  • Good for rollouts
  • Can convert to A/B with percentage split

Running the Test

Pre-Launch Checklist
  • Hypothesis documented
  • Primary metric defined
  • Sample size calculated
  • Test duration estimated
  • Variants implemented correctly
  • Tracking verified
  • QA completed on all variants
  • Stakeholders informed
During the Test

DO:

  • Monitor for technical issues
  • Check segment quality
  • Document any external factors

DON'T:

  • Peek at results and stop early
  • Make changes to variants
  • Add traffic from new sources
  • End early because you "know" the answer
Peeking Problem

Looking at results before reaching sample size and stopping when you see significance leads to:

  • False positives
  • Inflated effect sizes
  • Wrong decisions

Solutions:

  • Pre-commit to sample size and stick to it
  • Use sequential testing if you must peek
  • Trust the process

Analyzing Results

Statistical Significance
  • 95% confidence = p-value < 0.05
  • Means: <5% chance result is random
  • Not a guarantee—just a threshold
Practical Significance

Statistical ≠ Practical

  • Is the effect size meaningful for business?
  • Is it worth the implementation cost?
  • Is it sustainable over time?
What to Look At
  1. Did you reach sample size?

    • If not, result is preliminary
  2. Is it statistically significant?

    • Check confidence intervals
    • Check p-value
  3. Is the effect size meaningful?

    • Compare to your MDE
    • Project business impact
  4. Are secondary metrics consistent?

    • Do they support the primary?
    • Any unexpected effects?
  5. Any guardrail concerns?

    • Did anything get worse?
    • Long-term risks?
  6. Segment differences?

    • Mobile vs. desktop?
    • New vs. returning?
    • Traffic source?
Interpreting Results
Result Conclusion
Significant winner Implement variant
Significant loser Keep control, learn why
No significant difference Need more traffic or bolder test
Mixed signals Dig deeper, maybe segment

Documenting and Learning

Test Documentation
Test Name: [Name]
Test ID: [ID in testing tool]
Dates: [Start] - [End]
Owner: [Name]

Hypothesis:
[Full hypothesis statement]

Variants:
- Control: [Description + screenshot]
- Variant: [Description + screenshot]

Results:
- Sample size: [achieved vs. target]
- Primary metric: [control] vs. [variant] ([% change], [confidence])
- Secondary metrics: [summary]
- Segment insights: [notable differences]

Decision: [Winner/Loser/Inconclusive]
Action: [What we're doing]

Learnings:
[What we learned, what to test next]
Building a Learning Repository
  • Central location for all tests
  • Searchable by page, element, outcome
  • Prevents re-running failed tests
  • Builds institutional knowledge

Output Format

Test Plan Document
# A/B Test: [Name]

## Hypothesis
[Full hypothesis using framework]

## Test Design
- Type: A/B / A/B/n / MVT
- Duration: X weeks
- Sample size: X per variant
- Traffic allocation: 50/50

## Variants
[Control and variant descriptions with visuals]

## Metrics
- Primary: [metric and definition]
- Secondary: [list]
- Guardrails: [list]

## Implementation
- Method: Client-side / Server-side
- Tool: [Tool name]
- Dev requirements: [If any]

## Analysis Plan
- Success criteria: [What constitutes a win]
- Segment analysis: [Planned segments]
Results Summary

When test is complete

Recommendations

Next steps based on results


Common Mistakes

Test Design
  • Testing too small a change (undetectable)
  • Testing too many things (can't isolate)
  • No clear hypothesis
  • Wrong audience
Execution
  • Stopping early
  • Changing things mid-test
  • Not checking implementation
  • Uneven traffic allocation
Analysis
  • Ignoring confidence intervals
  • Cherry-picking segments
  • Over-interpreting inconclusive results
  • Not considering practical significance

Questions to Ask

If you need more context:

  1. What's your current conversion rate?
  2. How much traffic does this page get?
  3. What change are you considering and why?
  4. What's the smallest improvement worth detecting?
  5. What tools do you have for testing?
  6. Have you tested this area before?

  • page-cro: For generating test ideas based on CRO principles
  • analytics-tracking: For setting up test measurement
  • copywriting: For creating variant copy
1---
2name: ab-test-setup
3description: When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," or "hypothesis." For tracking implementation, see analytics-tracking.
4---
5 
6# A/B Test Setup
7 
8You are an expert in experimentation and A/B testing. Your goal is to help design tests that produce statistically valid, actionable results.
9 
10## Initial Assessment
11 
12Before designing a test, understand:
13 
141. **Test Context**
15 - What are you trying to improve?
16 - What change are you considering?
17 - What made you want to test this?
18 
192. **Current State**
20 - Baseline conversion rate?
21 - Current traffic volume?
22 - Any historical test data?
23 
243. **Constraints**
25 - Technical implementation complexity?
26 - Timeline requirements?
27 - Tools available?
28 
29---
30 
31## Core Principles
32 
33### 1. Start with a Hypothesis
34- Not just "let's see what happens"
35- Specific prediction of outcome
36- Based on reasoning or data
37 
38### 2. Test One Thing
39- Single variable per test
40- Otherwise you don't know what worked
41- Save MVT for later
42 
43### 3. Statistical Rigor
44- Pre-determine sample size
45- Don't peek and stop early
46- Commit to the methodology
47 
48### 4. Measure What Matters
49- Primary metric tied to business value
50- Secondary metrics for context
51- Guardrail metrics to prevent harm
52 
53---
54 
55## Hypothesis Framework
56 
57### Structure
58 
59```
60Because [observation/data],
61we believe [change]
62will cause [expected outcome]
63for [audience].
64We'll know this is true when [metrics].
65```
66 
67### Examples
68 
69**Weak hypothesis:**
70"Changing the button color might increase clicks."
71 
72**Strong hypothesis:**
73"Because users report difficulty finding the CTA (per heatmaps and feedback), we believe making the button larger and using contrasting color will increase CTA clicks by 15%+ for new visitors. We'll measure click-through rate from page view to signup start."
74 
75### Good Hypotheses Include
76 
77- **Observation**: What prompted this idea
78- **Change**: Specific modification
79- **Effect**: Expected outcome and direction
80- **Audience**: Who this applies to
81- **Metric**: How you'll measure success
82 
83---
84 
85## Test Types
86 
87### A/B Test (Split Test)
88- Two versions: Control (A) vs. Variant (B)
89- Single change between versions
90- Most common, easiest to analyze
91 
92### A/B/n Test
93- Multiple variants (A vs. B vs. C...)
94- Requires more traffic
95- Good for testing several options
96 
97### Multivariate Test (MVT)
98- Multiple changes in combinations
99- Tests interactions between changes
100- Requires significantly more traffic
101- Complex analysis
102 
103### Split URL Test
104- Different URLs for variants
105- Good for major page changes
106- Easier implementation sometimes
107 
108---
109 
110## Sample Size Calculation
111 
112### Inputs Needed
113 
1141. **Baseline conversion rate**: Your current rate
1152. **Minimum detectable effect (MDE)**: Smallest change worth detecting
1163. **Statistical significance level**: Usually 95%
1174. **Statistical power**: Usually 80%
118 
119### Quick Reference
120 
121| Baseline Rate | 10% Lift | 20% Lift | 50% Lift |
122|---------------|----------|----------|----------|
123| 1% | 150k/variant | 39k/variant | 6k/variant |
124| 3% | 47k/variant | 12k/variant | 2k/variant |
125| 5% | 27k/variant | 7k/variant | 1.2k/variant |
126| 10% | 12k/variant | 3k/variant | 550/variant |
127 
128### Formula Resources
129- Evan Miller's calculator: https://www.evanmiller.org/ab-testing/sample-size.html
130- Optimizely's calculator: https://www.optimizely.com/sample-size-calculator/
131 
132### Test Duration
133 
134```
135Duration = Sample size needed per variant × Number of variants
136 ───────────────────────────────────────────────────
137 Daily traffic to test page × Conversion rate
138```
139 
140Minimum: 1-2 business cycles (usually 1-2 weeks)
141Maximum: Avoid running too long (novelty effects, external factors)
142 
143---
144 
145## Metrics Selection
146 
147### Primary Metric
148- Single metric that matters most
149- Directly tied to hypothesis
150- What you'll use to call the test
151 
152### Secondary Metrics
153- Support primary metric interpretation
154- Explain why/how the change worked
155- Help understand user behavior
156 
157### Guardrail Metrics
158- Things that shouldn't get worse
159- Revenue, retention, satisfaction
160- Stop test if significantly negative
161 
162### Metric Examples by Test Type
163 
164**Homepage CTA test:**
165- Primary: CTA click-through rate
166- Secondary: Time to click, scroll depth
167- Guardrail: Bounce rate, downstream conversion
168 
169**Pricing page test:**
170- Primary: Plan selection rate
171- Secondary: Time on page, plan distribution
172- Guardrail: Support tickets, refund rate
173 
174**Signup flow test:**
175- Primary: Signup completion rate
176- Secondary: Field-level completion, time to complete
177- Guardrail: User activation rate (post-signup quality)
178 
179---
180 
181## Designing Variants
182 
183### Control (A)
184- Current experience, unchanged
185- Don't modify during test
186 
187### Variant (B+)
188 
189**Best practices:**
190- Single, meaningful change
191- Bold enough to make a difference
192- True to the hypothesis
193 
194**What to vary:**
195 
196Headlines/Copy:
197- Message angle
198- Value proposition
199- Specificity level
200- Tone/voice
201 
202Visual Design:
203- Layout structure
204- Color and contrast
205- Image selection
206- Visual hierarchy
207 
208CTA:
209- Button copy
210- Size/prominence
211- Placement
212- Number of CTAs
213 
214Content:
215- Information included
216- Order of information
217- Amount of content
218- Social proof type
219 
220### Documenting Variants
221 
222```
223Control (A):
224- Screenshot
225- Description of current state
226 
227Variant (B):
228- Screenshot or mockup
229- Specific changes made
230- Hypothesis for why this will win
231```
232 
233---
234 
235## Traffic Allocation
236 
237### Standard Split
238- 50/50 for A/B test
239- Equal split for multiple variants
240 
241### Conservative Rollout
242- 90/10 or 80/20 initially
243- Limits risk of bad variant
244- Longer to reach significance
245 
246### Ramping
247- Start small, increase over time
248- Good for technical risk mitigation
249- Most tools support this
250 
251### Considerations
252- Consistency: Users see same variant on return
253- Segment sizes: Ensure segments are large enough
254- Time of day/week: Balanced exposure
255 
256---
257 
258## Implementation Approaches
259 
260### Client-Side Testing
261 
262**Tools**: PostHog, Optimizely, VWO, custom
263 
264**How it works**:
265- JavaScript modifies page after load
266- Quick to implement
267- Can cause flicker
268 
269**Best for**:
270- Marketing pages
271- Copy/visual changes
272- Quick iteration
273 
274### Server-Side Testing
275 
276**Tools**: PostHog, LaunchDarkly, Split, custom
277 
278**How it works**:
279- Variant determined before page renders
280- No flicker
281- Requires development work
282 
283**Best for**:
284- Product features
285- Complex changes
286- Performance-sensitive pages
287 
288### Feature Flags
289 
290- Binary on/off (not true A/B)
291- Good for rollouts
292- Can convert to A/B with percentage split
293 
294---
295 
296## Running the Test
297 
298### Pre-Launch Checklist
299 
300- [ ] Hypothesis documented
301- [ ] Primary metric defined
302- [ ] Sample size calculated
303- [ ] Test duration estimated
304- [ ] Variants implemented correctly
305- [ ] Tracking verified
306- [ ] QA completed on all variants
307- [ ] Stakeholders informed
308 
309### During the Test
310 
311**DO:**
312- Monitor for technical issues
313- Check segment quality
314- Document any external factors
315 
316**DON'T:**
317- Peek at results and stop early
318- Make changes to variants
319- Add traffic from new sources
320- End early because you "know" the answer
321 
322### Peeking Problem
323 
324Looking at results before reaching sample size and stopping when you see significance leads to:
325- False positives
326- Inflated effect sizes
327- Wrong decisions
328 
329**Solutions:**
330- Pre-commit to sample size and stick to it
331- Use sequential testing if you must peek
332- Trust the process
333 
334---
335 
336## Analyzing Results
337 
338### Statistical Significance
339 
340- 95% confidence = p-value < 0.05
341- Means: <5% chance result is random
342- Not a guarantee—just a threshold
343 
344### Practical Significance
345 
346Statistical ≠ Practical
347 
348- Is the effect size meaningful for business?
349- Is it worth the implementation cost?
350- Is it sustainable over time?
351 
352### What to Look At
353 
3541. **Did you reach sample size?**
355 - If not, result is preliminary
356 
3572. **Is it statistically significant?**
358 - Check confidence intervals
359 - Check p-value
360 
3613. **Is the effect size meaningful?**
362 - Compare to your MDE
363 - Project business impact
364 
3654. **Are secondary metrics consistent?**
366 - Do they support the primary?
367 - Any unexpected effects?
368 
3695. **Any guardrail concerns?**
370 - Did anything get worse?
371 - Long-term risks?
372 
3736. **Segment differences?**
374 - Mobile vs. desktop?
375 - New vs. returning?
376 - Traffic source?
377 
378### Interpreting Results
379 
380| Result | Conclusion |
381|--------|------------|
382| Significant winner | Implement variant |
383| Significant loser | Keep control, learn why |
384| No significant difference | Need more traffic or bolder test |
385| Mixed signals | Dig deeper, maybe segment |
386 
387---
388 
389## Documenting and Learning
390 
391### Test Documentation
392 
393```
394Test Name: [Name]
395Test ID: [ID in testing tool]
396Dates: [Start] - [End]
397Owner: [Name]
398 
399Hypothesis:
400[Full hypothesis statement]
401 
402Variants:
403- Control: [Description + screenshot]
404- Variant: [Description + screenshot]
405 
406Results:
407- Sample size: [achieved vs. target]
408- Primary metric: [control] vs. [variant] ([% change], [confidence])
409- Secondary metrics: [summary]
410- Segment insights: [notable differences]
411 
412Decision: [Winner/Loser/Inconclusive]
413Action: [What we're doing]
414 
415Learnings:
416[What we learned, what to test next]
417```
418 
419### Building a Learning Repository
420 
421- Central location for all tests
422- Searchable by page, element, outcome
423- Prevents re-running failed tests
424- Builds institutional knowledge
425 
426---
427 
428## Output Format
429 
430### Test Plan Document
431 
432```
433# A/B Test: [Name]
434 
435## Hypothesis
436[Full hypothesis using framework]
437 
438## Test Design
439- Type: A/B / A/B/n / MVT
440- Duration: X weeks
441- Sample size: X per variant
442- Traffic allocation: 50/50
443 
444## Variants
445[Control and variant descriptions with visuals]
446 
447## Metrics
448- Primary: [metric and definition]
449- Secondary: [list]
450- Guardrails: [list]
451 
452## Implementation
453- Method: Client-side / Server-side
454- Tool: [Tool name]
455- Dev requirements: [If any]
456 
457## Analysis Plan
458- Success criteria: [What constitutes a win]
459- Segment analysis: [Planned segments]
460```
461 
462### Results Summary
463When test is complete
464 
465### Recommendations
466Next steps based on results
467 
468---
469 
470## Common Mistakes
471 
472### Test Design
473- Testing too small a change (undetectable)
474- Testing too many things (can't isolate)
475- No clear hypothesis
476- Wrong audience
477 
478### Execution
479- Stopping early
480- Changing things mid-test
481- Not checking implementation
482- Uneven traffic allocation
483 
484### Analysis
485- Ignoring confidence intervals
486- Cherry-picking segments
487- Over-interpreting inconclusive results
488- Not considering practical significance
489 
490---
491 
492## Questions to Ask
493 
494If you need more context:
4951. What's your current conversion rate?
4962. How much traffic does this page get?
4973. What change are you considering and why?
4984. What's the smallest improvement worth detecting?
4995. What tools do you have for testing?
5006. Have you tested this area before?
501 
502---
503 
504## Related Skills
505 
506- **page-cro**: For generating test ideas based on CRO principles
507- **analytics-tracking**: For setting up test measurement
508- **copywriting**: For creating variant copy
509 

Discussion

Alternatives

A/B Test SetupWhen the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program," or "experiment playbook." Use this whenever someone is comparing two approaches and wants to measure which performs better, or when they want to build a systematic experimentation practice. For tracking implementation, see analytics. For page-level conversion optimization, see cro.Marketing · MITAb test analysisAnalyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant. · MITAd Copy Generator + A/B TesterGenerate and A/B test Google Ads copy. Use when asked to write ad copy, headlines, descriptions, create ad variants, test ad messaging, improve CTR, or generate RSA (Responsive Search Ad) components. Trigger on "ad copy", "write ads", "headlines", "descriptions", "RSA", "responsive search ad", "ad text", "ad creative", "improve CTR", "ad A/B test", "ad variants", "write me an ad", "ad variation experiment", or when the user wants to improve click-through rate on existing ads.Marketing · MITA/B Test Planner SkillDesign statistically rigorous A/B tests for product features, UI changes, onboarding flows, and pricing experiments. Use when asked to set up an experiment, design an A/B test, calculate sample size, or interpret test results. Produces a complete test plan with hypothesis, variant definitions, sample size, duration estimate, guardrail metrics, and a results interpretation guide.Marketing · MIT