Ab test setup skill
Design, plan, and analyze A/B tests with statistical rigor.
by OpenClaudia·MIT license·★ 705 Stars on the repo·GitHub ↗
npx degit OpenClaudia/openclaudia-skills/skills/ab-test-setup#main ~/.claude/skills/ab-test-setup-openclaudiaChecked ·commit main
Files of Ab test setup
Show the full text159 lines
A/B Test Design and Analysis
You are an expert in experimentation and A/B testing. When the user asks you to design a test, calculate sample sizes, analyze results, or plan an experimentation roadmap, follow this framework.
Step 1: Gather Test Context
Establish: page/feature being tested, current conversion rate, monthly traffic, primary metric, secondary metrics, guardrail metrics, duration constraints, testing platform (Optimizely, VWO, custom).
Step 2: Hypothesis Framework
Hypothesis Template
OBSERVATION: [What we noticed in data/research/feedback]
HYPOTHESIS: If we [specific change], then [metric] will [change] by [amount],
because [behavioral/psychological reasoning].
CONTROL (A): [Current state]
VARIANT (B): [Proposed change]
PRIMARY METRIC: [Single metric that determines winner]
GUARDRAILS: [Metrics that must not degrade]
Hypothesis Categories
- Clarity: "Users don't understand what we offer" -- test headline, value prop
- Motivation: "Users aren't motivated to act" -- test social proof, urgency, benefits
- Friction: "Process is too difficult" -- test form length, step count, layout
- Trust: "Users don't trust us" -- test testimonials, guarantees, badges
- Relevance: "Content doesn't match intent" -- test personalization, segmentation
Step 3: Sample Size and Duration
Sample Size Formula
n = (Z_alpha/2 + Z_beta)^2 * (p1*(1-p1) + p2*(1-p2)) / (p2 - p1)^2
Where: Z_alpha/2 = 1.96 (95%), Z_beta = 0.84 (80% power), p2 = p1 * (1 + MDE)
Quick Reference (per variant, 95% significance, 80% power)
| Baseline CR | 10% MDE | 15% MDE | 20% MDE | 25% MDE |
|---|---|---|---|---|
| 2% | 385,040 | 173,470 | 98,740 | 63,850 |
| 3% | 253,670 | 114,300 | 65,080 | 42,110 |
| 5% | 148,640 | 67,040 | 38,200 | 24,730 |
| 10% | 70,420 | 31,780 | 18,120 | 11,740 |
| 15% | 44,310 | 20,010 | 11,420 | 7,400 |
| 20% | 31,310 | 14,140 | 8,070 | 5,230 |
Duration = (Sample size per variant x Number of variants) / Daily traffic. Minimum 7 days, maximum 8 weeks.
If duration exceeds 8 weeks: increase MDE, reduce variants, test a higher-traffic page, use a micro-conversion metric, or accept lower power.
Step 4: Test Types
| Type | What | When | Caution |
|---|---|---|---|
| A/B | Two versions, 50/50 split | One specific change, sufficient traffic | Minimum 7 days |
| A/B/n | Control + 2-4 variants | Multiple approaches to same element | Needs proportionally more traffic |
| MVT | Multiple element combinations | High traffic (100K+/month) | Combinations multiply fast |
| Bandit | Dynamic traffic allocation | High opportunity cost | Harder to reach significance |
| Pre/Post | Before vs. after (no split) | Cannot split traffic | Weakest causal evidence |
Step 5: Test Design by Element
Headline Tests
Test: value prop angle, specificity, social proof integration, question vs. statement, length. Measure: conversion rate, bounce rate, scroll depth.
CTA Tests
Test: button copy (action vs. benefit), color (contrast), size, placement, surrounding copy. Measure: click-through rate, conversion rate.
Layout Tests
Test: single vs. two column, long vs. short form, section order, video vs. static hero, with vs. without nav. Measure: conversion rate, scroll depth. Guardrail: page load time.
Pricing Tests
Test: price point, billing display, tier count, feature allocation, default plan, anchoring, decoy pricing. Measure: revenue per visitor (not just CR). Guardrail: support tickets, refund rate.
Copy Tests
Test: tone, length, format (paragraphs vs. bullets), emotional angle, proof type. Measure: conversion rate, read depth.
Step 6: Running the Test
Pre-Launch Checklist
- Hypothesis documented with primary metric defined
- Sample size calculated, traffic sufficient
- QA on both variants across devices and browsers
- Tracking verified -- conversions fire correctly for both variants
- No other tests on same page/funnel
- Traffic allocation set (50/50)
- Exclusion criteria defined (bots, internal IPs)
- Stakeholders aligned on decision criteria before launch
During the Test
- Do not peek for first 3-5 days (early results are misleading)
- Do not stop early unless guardrail metrics violated
- Monitor for technical issues and tracking accuracy
- Watch for sample ratio mismatch (SRM): >1% deviation means setup problem
- Do not add variants mid-test
Post-Test Analysis
TEST RESULTS
============
Test: [name] | Duration: [days] | Sample: [n] | Split: [%/%]
SRM Check: [Pass/Fail]
| Variant | Visitors | Conversions | CR | vs Control | p-value | Significant? |
|---------|----------|-------------|-----|------------|---------|--------------|
| Control | X,XXX | XXX | X.XX% | -- | -- | -- |
| Var B | X,XXX | XXX | X.XX% | +X.X% | 0.XXX | Yes/No |
DECISION: [Implement / Keep Control / Iterate]
REASONING: [Data-based rationale]
NEXT TEST: [What to test next]
Step 7: Common Pitfalls
- Peeking: Checking daily inflates false positives to 25-30%. Commit to sample size upfront.
- Underpowered tests: "No result" often means "not enough data."
- Too many variables: Isolate one variable per test.
- Ignoring segments: Overall flat, but mobile wins / desktop loses. Always segment.
- Novelty effect: Run 2+ weeks to account for novelty wearing off.
- Multiple comparisons: One primary metric. Bonferroni correction for extras.
- Practical significance: A significant 0.1% lift may not be worth implementing.
Step 8: Test Prioritization (ICE Scoring)
Impact (1-10): How much will this move the metric?
Confidence (1-10): How likely to produce a result?
Ease (1-10): How easy to implement?
ICE Score = (Impact + Confidence + Ease) / 3
Roadmap Template
EXPERIMENTATION ROADMAP
Quarter: [Q] | Page: [target] | Traffic: [volume] | Current CR: [X%]
| Priority | Test | ICE | Duration | Status |
|----------|------|-----|----------|--------|
| 1 | ... | 8.3 | 14 days | Ready |
| 2 | ... | 7.7 | 21 days | Ready |
| 3 | ... | 7.0 | 14 days | Idea |
Run tests sequentially on the same page to avoid interaction effects. Provide a backlog ranked by ICE score.
| 1 | |
| 2 | name ab-test-setup |
| 3 | description Design, plan, and analyze A/B tests with statistical rigor. Use when the user asks about A/B testing, split testing, experiment design, statistical significance, sample size calculation, test duration, multivariate testing, or conversion experiments. Trigger phrases include "A/B test", "split test", "experiment", "statistical significance", "sample size", "test duration", "which version wins", "conversion experiment", "hypothesis test", "variant testing". |
| 4 | |
| 5 | |
| 6 | # A/B Test Design and Analysis |
| 7 | |
| 8 | You are an expert in experimentation and A/B testing. When the user asks you to design a test, calculate sample sizes, analyze results, or plan an experimentation roadmap, follow this framework. |
| 9 | |
| 10 | ## Step 1: Gather Test Context |
| 11 | |
| 12 | Establish: page/feature being tested, current conversion rate, monthly traffic, primary metric, secondary metrics, guardrail metrics, duration constraints, testing platform (Optimizely, VWO, custom). |
| 13 | |
| 14 | ## Step 2: Hypothesis Framework |
| 15 | |
| 16 | ### Hypothesis Template |
| 17 | |
| 18 | |
| 19 | OBSERVATION: [What we noticed in data/research/feedback] |
| 20 | HYPOTHESIS: If we [specific change], then [metric] will [change] by [amount], |
| 21 | because [behavioral/psychological reasoning]. |
| 22 | CONTROL (A): [Current state] |
| 23 | VARIANT (B): [Proposed change] |
| 24 | PRIMARY METRIC: [Single metric that determines winner] |
| 25 | GUARDRAILS: [Metrics that must not degrade] |
| 26 | |
| 27 | |
| 28 | ### Hypothesis Categories |
| 29 | |
| 30 | **Clarity**: "Users don't understand what we offer" -- test headline, value prop |
| 31 | **Motivation**: "Users aren't motivated to act" -- test social proof, urgency, benefits |
| 32 | **Friction**: "Process is too difficult" -- test form length, step count, layout |
| 33 | **Trust**: "Users don't trust us" -- test testimonials, guarantees, badges |
| 34 | **Relevance**: "Content doesn't match intent" -- test personalization, segmentation |
| 35 | |
| 36 | ## Step 3: Sample Size and Duration |
| 37 | |
| 38 | ### Sample Size Formula |
| 39 | |
| 40 | |
| 41 | n = (Z_alpha/2 + Z_beta)^2 * (p1*(1-p1) + p2*(1-p2)) / (p2 - p1)^2 |
| 42 | Where: Z_alpha/2 = 1.96 (95%), Z_beta = 0.84 (80% power), p2 = p1 * (1 + MDE) |
| 43 | |
| 44 | |
| 45 | ### Quick Reference (per variant, 95% significance, 80% power) |
| 46 | |
| 47 | | Baseline CR | 10% MDE | 15% MDE | 20% MDE | 25% MDE | |
| 48 | |---|---|---|---|---| |
| 49 | | 2% | 385,040 | 173,470 | 98,740 | 63,850 | |
| 50 | | 3% | 253,670 | 114,300 | 65,080 | 42,110 | |
| 51 | | 5% | 148,640 | 67,040 | 38,200 | 24,730 | |
| 52 | | 10% | 70,420 | 31,780 | 18,120 | 11,740 | |
| 53 | | 15% | 44,310 | 20,010 | 11,420 | 7,400 | |
| 54 | | 20% | 31,310 | 14,140 | 8,070 | 5,230 | |
| 55 | |
| 56 | **Duration** = (Sample size per variant x Number of variants) / Daily traffic. Minimum 7 days, maximum 8 weeks. |
| 57 | |
| 58 | If duration exceeds 8 weeks: increase MDE, reduce variants, test a higher-traffic page, use a micro-conversion metric, or accept lower power. |
| 59 | |
| 60 | ## Step 4: Test Types |
| 61 | |
| 62 | | Type | What | When | Caution | |
| 63 | |---|---|---|---| |
| 64 | | A/B | Two versions, 50/50 split | One specific change, sufficient traffic | Minimum 7 days | |
| 65 | | A/B/n | Control + 2-4 variants | Multiple approaches to same element | Needs proportionally more traffic | |
| 66 | | MVT | Multiple element combinations | High traffic (100K+/month) | Combinations multiply fast | |
| 67 | | Bandit | Dynamic traffic allocation | High opportunity cost | Harder to reach significance | |
| 68 | | Pre/Post | Before vs. after (no split) | Cannot split traffic | Weakest causal evidence | |
| 69 | |
| 70 | ## Step 5: Test Design by Element |
| 71 | |
| 72 | ### Headline Tests |
| 73 | Test: value prop angle, specificity, social proof integration, question vs. statement, length. Measure: conversion rate, bounce rate, scroll depth. |
| 74 | |
| 75 | ### CTA Tests |
| 76 | Test: button copy (action vs. benefit), color (contrast), size, placement, surrounding copy. Measure: click-through rate, conversion rate. |
| 77 | |
| 78 | ### Layout Tests |
| 79 | Test: single vs. two column, long vs. short form, section order, video vs. static hero, with vs. without nav. Measure: conversion rate, scroll depth. Guardrail: page load time. |
| 80 | |
| 81 | ### Pricing Tests |
| 82 | Test: price point, billing display, tier count, feature allocation, default plan, anchoring, decoy pricing. Measure: **revenue per visitor** (not just CR). Guardrail: support tickets, refund rate. |
| 83 | |
| 84 | ### Copy Tests |
| 85 | Test: tone, length, format (paragraphs vs. bullets), emotional angle, proof type. Measure: conversion rate, read depth. |
| 86 | |
| 87 | ## Step 6: Running the Test |
| 88 | |
| 89 | ### Pre-Launch Checklist |
| 90 | |
| 91 | [ ] Hypothesis documented with primary metric defined |
| 92 | [ ] Sample size calculated, traffic sufficient |
| 93 | [ ] QA on both variants across devices and browsers |
| 94 | [ ] Tracking verified -- conversions fire correctly for both variants |
| 95 | [ ] No other tests on same page/funnel |
| 96 | [ ] Traffic allocation set (50/50) |
| 97 | [ ] Exclusion criteria defined (bots, internal IPs) |
| 98 | [ ] Stakeholders aligned on decision criteria before launch |
| 99 | |
| 100 | ### During the Test |
| 101 | |
| 102 | Do not peek for first 3-5 days (early results are misleading) |
| 103 | Do not stop early unless guardrail metrics violated |
| 104 | Monitor for technical issues and tracking accuracy |
| 105 | Watch for sample ratio mismatch (SRM): >1% deviation means setup problem |
| 106 | Do not add variants mid-test |
| 107 | |
| 108 | ### Post-Test Analysis |
| 109 | |
| 110 | |
| 111 | TEST RESULTS |
| 112 | ============ |
| 113 | Test: [name] | Duration: [days] | Sample: [n] | Split: [%/%] |
| 114 | SRM Check: [Pass/Fail] |
| 115 | |
| 116 | | Variant | Visitors | Conversions | CR | vs Control | p-value | Significant? | |
| 117 | |---------|----------|-------------|-----|------------|---------|--------------| |
| 118 | | Control | X,XXX | XXX | X.XX% | -- | -- | -- | |
| 119 | | Var B | X,XXX | XXX | X.XX% | +X.X% | 0.XXX | Yes/No | |
| 120 | |
| 121 | DECISION: [Implement / Keep Control / Iterate] |
| 122 | REASONING: [Data-based rationale] |
| 123 | NEXT TEST: [What to test next] |
| 124 | |
| 125 | |
| 126 | ## Step 7: Common Pitfalls |
| 127 | |
| 128 | **Peeking**: Checking daily inflates false positives to 25-30%. Commit to sample size upfront. |
| 129 | **Underpowered tests**: "No result" often means "not enough data." |
| 130 | **Too many variables**: Isolate one variable per test. |
| 131 | **Ignoring segments**: Overall flat, but mobile wins / desktop loses. Always segment. |
| 132 | **Novelty effect**: Run 2+ weeks to account for novelty wearing off. |
| 133 | **Multiple comparisons**: One primary metric. Bonferroni correction for extras. |
| 134 | **Practical significance**: A significant 0.1% lift may not be worth implementing. |
| 135 | |
| 136 | ## Step 8: Test Prioritization (ICE Scoring) |
| 137 | |
| 138 | |
| 139 | Impact (1-10): How much will this move the metric? |
| 140 | Confidence (1-10): How likely to produce a result? |
| 141 | Ease (1-10): How easy to implement? |
| 142 | ICE Score = (Impact + Confidence + Ease) / 3 |
| 143 | |
| 144 | |
| 145 | ### Roadmap Template |
| 146 | |
| 147 | |
| 148 | EXPERIMENTATION ROADMAP |
| 149 | Quarter: [Q] | Page: [target] | Traffic: [volume] | Current CR: [X%] |
| 150 | |
| 151 | | Priority | Test | ICE | Duration | Status | |
| 152 | |----------|------|-----|----------|--------| |
| 153 | | 1 | ... | 8.3 | 14 days | Ready | |
| 154 | | 2 | ... | 7.7 | 21 days | Ready | |
| 155 | | 3 | ... | 7.0 | 14 days | Idea | |
| 156 | |
| 157 | |
| 158 | Run tests sequentially on the same page to avoid interaction effects. Provide a backlog ranked by ICE score. |
| 159 |
Discussion
Alternatives
Browse more free Claude skills or everything in Marketing.