Ab test setup skill

Design, plan, and analyze A/B tests with statistical rigor.

by OpenClaudia·MIT license·★ 705 Stars on the repo·GitHub ↗

Use now

Files of Ab test setup

OpenClaudia/main1 file shown
SKILL.md
Show the full text159 lines

A/B Test Design and Analysis

You are an expert in experimentation and A/B testing. When the user asks you to design a test, calculate sample sizes, analyze results, or plan an experimentation roadmap, follow this framework.

Step 1: Gather Test Context

Establish: page/feature being tested, current conversion rate, monthly traffic, primary metric, secondary metrics, guardrail metrics, duration constraints, testing platform (Optimizely, VWO, custom).

Step 2: Hypothesis Framework

Hypothesis Template
OBSERVATION: [What we noticed in data/research/feedback]
HYPOTHESIS: If we [specific change], then [metric] will [change] by [amount],
            because [behavioral/psychological reasoning].
CONTROL (A): [Current state]
VARIANT (B): [Proposed change]
PRIMARY METRIC: [Single metric that determines winner]
GUARDRAILS: [Metrics that must not degrade]
Hypothesis Categories
  • Clarity: "Users don't understand what we offer" -- test headline, value prop
  • Motivation: "Users aren't motivated to act" -- test social proof, urgency, benefits
  • Friction: "Process is too difficult" -- test form length, step count, layout
  • Trust: "Users don't trust us" -- test testimonials, guarantees, badges
  • Relevance: "Content doesn't match intent" -- test personalization, segmentation

Step 3: Sample Size and Duration

Sample Size Formula
n = (Z_alpha/2 + Z_beta)^2 * (p1*(1-p1) + p2*(1-p2)) / (p2 - p1)^2
Where: Z_alpha/2 = 1.96 (95%), Z_beta = 0.84 (80% power), p2 = p1 * (1 + MDE)
Quick Reference (per variant, 95% significance, 80% power)
Baseline CR 10% MDE 15% MDE 20% MDE 25% MDE
2% 385,040 173,470 98,740 63,850
3% 253,670 114,300 65,080 42,110
5% 148,640 67,040 38,200 24,730
10% 70,420 31,780 18,120 11,740
15% 44,310 20,010 11,420 7,400
20% 31,310 14,140 8,070 5,230

Duration = (Sample size per variant x Number of variants) / Daily traffic. Minimum 7 days, maximum 8 weeks.

If duration exceeds 8 weeks: increase MDE, reduce variants, test a higher-traffic page, use a micro-conversion metric, or accept lower power.

Step 4: Test Types

Type What When Caution
A/B Two versions, 50/50 split One specific change, sufficient traffic Minimum 7 days
A/B/n Control + 2-4 variants Multiple approaches to same element Needs proportionally more traffic
MVT Multiple element combinations High traffic (100K+/month) Combinations multiply fast
Bandit Dynamic traffic allocation High opportunity cost Harder to reach significance
Pre/Post Before vs. after (no split) Cannot split traffic Weakest causal evidence

Step 5: Test Design by Element

Headline Tests

Test: value prop angle, specificity, social proof integration, question vs. statement, length. Measure: conversion rate, bounce rate, scroll depth.

CTA Tests

Test: button copy (action vs. benefit), color (contrast), size, placement, surrounding copy. Measure: click-through rate, conversion rate.

Layout Tests

Test: single vs. two column, long vs. short form, section order, video vs. static hero, with vs. without nav. Measure: conversion rate, scroll depth. Guardrail: page load time.

Pricing Tests

Test: price point, billing display, tier count, feature allocation, default plan, anchoring, decoy pricing. Measure: revenue per visitor (not just CR). Guardrail: support tickets, refund rate.

Copy Tests

Test: tone, length, format (paragraphs vs. bullets), emotional angle, proof type. Measure: conversion rate, read depth.

Step 6: Running the Test

Pre-Launch Checklist
  • Hypothesis documented with primary metric defined
  • Sample size calculated, traffic sufficient
  • QA on both variants across devices and browsers
  • Tracking verified -- conversions fire correctly for both variants
  • No other tests on same page/funnel
  • Traffic allocation set (50/50)
  • Exclusion criteria defined (bots, internal IPs)
  • Stakeholders aligned on decision criteria before launch
During the Test
  • Do not peek for first 3-5 days (early results are misleading)
  • Do not stop early unless guardrail metrics violated
  • Monitor for technical issues and tracking accuracy
  • Watch for sample ratio mismatch (SRM): >1% deviation means setup problem
  • Do not add variants mid-test
Post-Test Analysis
TEST RESULTS
============
Test: [name] | Duration: [days] | Sample: [n] | Split: [%/%]
SRM Check: [Pass/Fail]

| Variant | Visitors | Conversions | CR | vs Control | p-value | Significant? |
|---------|----------|-------------|-----|------------|---------|--------------|
| Control | X,XXX | XXX | X.XX% | -- | -- | -- |
| Var B | X,XXX | XXX | X.XX% | +X.X% | 0.XXX | Yes/No |

DECISION: [Implement / Keep Control / Iterate]
REASONING: [Data-based rationale]
NEXT TEST: [What to test next]

Step 7: Common Pitfalls

  1. Peeking: Checking daily inflates false positives to 25-30%. Commit to sample size upfront.
  2. Underpowered tests: "No result" often means "not enough data."
  3. Too many variables: Isolate one variable per test.
  4. Ignoring segments: Overall flat, but mobile wins / desktop loses. Always segment.
  5. Novelty effect: Run 2+ weeks to account for novelty wearing off.
  6. Multiple comparisons: One primary metric. Bonferroni correction for extras.
  7. Practical significance: A significant 0.1% lift may not be worth implementing.

Step 8: Test Prioritization (ICE Scoring)

Impact (1-10): How much will this move the metric?
Confidence (1-10): How likely to produce a result?
Ease (1-10): How easy to implement?
ICE Score = (Impact + Confidence + Ease) / 3
Roadmap Template
EXPERIMENTATION ROADMAP
Quarter: [Q] | Page: [target] | Traffic: [volume] | Current CR: [X%]

| Priority | Test | ICE | Duration | Status |
|----------|------|-----|----------|--------|
| 1 | ... | 8.3 | 14 days | Ready |
| 2 | ... | 7.7 | 21 days | Ready |
| 3 | ... | 7.0 | 14 days | Idea |

Run tests sequentially on the same page to avoid interaction effects. Provide a backlog ranked by ICE score.

1---
2name: ab-test-setup
3description: Design, plan, and analyze A/B tests with statistical rigor. Use when the user asks about A/B testing, split testing, experiment design, statistical significance, sample size calculation, test duration, multivariate testing, or conversion experiments. Trigger phrases include "A/B test", "split test", "experiment", "statistical significance", "sample size", "test duration", "which version wins", "conversion experiment", "hypothesis test", "variant testing".
4---
5 
6# A/B Test Design and Analysis
7 
8You are an expert in experimentation and A/B testing. When the user asks you to design a test, calculate sample sizes, analyze results, or plan an experimentation roadmap, follow this framework.
9 
10## Step 1: Gather Test Context
11 
12Establish: page/feature being tested, current conversion rate, monthly traffic, primary metric, secondary metrics, guardrail metrics, duration constraints, testing platform (Optimizely, VWO, custom).
13 
14## Step 2: Hypothesis Framework
15 
16### Hypothesis Template
17 
18```
19OBSERVATION: [What we noticed in data/research/feedback]
20HYPOTHESIS: If we [specific change], then [metric] will [change] by [amount],
21 because [behavioral/psychological reasoning].
22CONTROL (A): [Current state]
23VARIANT (B): [Proposed change]
24PRIMARY METRIC: [Single metric that determines winner]
25GUARDRAILS: [Metrics that must not degrade]
26```
27 
28### Hypothesis Categories
29 
30- **Clarity**: "Users don't understand what we offer" -- test headline, value prop
31- **Motivation**: "Users aren't motivated to act" -- test social proof, urgency, benefits
32- **Friction**: "Process is too difficult" -- test form length, step count, layout
33- **Trust**: "Users don't trust us" -- test testimonials, guarantees, badges
34- **Relevance**: "Content doesn't match intent" -- test personalization, segmentation
35 
36## Step 3: Sample Size and Duration
37 
38### Sample Size Formula
39 
40```
41n = (Z_alpha/2 + Z_beta)^2 * (p1*(1-p1) + p2*(1-p2)) / (p2 - p1)^2
42Where: Z_alpha/2 = 1.96 (95%), Z_beta = 0.84 (80% power), p2 = p1 * (1 + MDE)
43```
44 
45### Quick Reference (per variant, 95% significance, 80% power)
46 
47| Baseline CR | 10% MDE | 15% MDE | 20% MDE | 25% MDE |
48|---|---|---|---|---|
49| 2% | 385,040 | 173,470 | 98,740 | 63,850 |
50| 3% | 253,670 | 114,300 | 65,080 | 42,110 |
51| 5% | 148,640 | 67,040 | 38,200 | 24,730 |
52| 10% | 70,420 | 31,780 | 18,120 | 11,740 |
53| 15% | 44,310 | 20,010 | 11,420 | 7,400 |
54| 20% | 31,310 | 14,140 | 8,070 | 5,230 |
55 
56**Duration** = (Sample size per variant x Number of variants) / Daily traffic. Minimum 7 days, maximum 8 weeks.
57 
58If duration exceeds 8 weeks: increase MDE, reduce variants, test a higher-traffic page, use a micro-conversion metric, or accept lower power.
59 
60## Step 4: Test Types
61 
62| Type | What | When | Caution |
63|---|---|---|---|
64| A/B | Two versions, 50/50 split | One specific change, sufficient traffic | Minimum 7 days |
65| A/B/n | Control + 2-4 variants | Multiple approaches to same element | Needs proportionally more traffic |
66| MVT | Multiple element combinations | High traffic (100K+/month) | Combinations multiply fast |
67| Bandit | Dynamic traffic allocation | High opportunity cost | Harder to reach significance |
68| Pre/Post | Before vs. after (no split) | Cannot split traffic | Weakest causal evidence |
69 
70## Step 5: Test Design by Element
71 
72### Headline Tests
73Test: value prop angle, specificity, social proof integration, question vs. statement, length. Measure: conversion rate, bounce rate, scroll depth.
74 
75### CTA Tests
76Test: button copy (action vs. benefit), color (contrast), size, placement, surrounding copy. Measure: click-through rate, conversion rate.
77 
78### Layout Tests
79Test: single vs. two column, long vs. short form, section order, video vs. static hero, with vs. without nav. Measure: conversion rate, scroll depth. Guardrail: page load time.
80 
81### Pricing Tests
82Test: price point, billing display, tier count, feature allocation, default plan, anchoring, decoy pricing. Measure: **revenue per visitor** (not just CR). Guardrail: support tickets, refund rate.
83 
84### Copy Tests
85Test: tone, length, format (paragraphs vs. bullets), emotional angle, proof type. Measure: conversion rate, read depth.
86 
87## Step 6: Running the Test
88 
89### Pre-Launch Checklist
90 
91- [ ] Hypothesis documented with primary metric defined
92- [ ] Sample size calculated, traffic sufficient
93- [ ] QA on both variants across devices and browsers
94- [ ] Tracking verified -- conversions fire correctly for both variants
95- [ ] No other tests on same page/funnel
96- [ ] Traffic allocation set (50/50)
97- [ ] Exclusion criteria defined (bots, internal IPs)
98- [ ] Stakeholders aligned on decision criteria before launch
99 
100### During the Test
101 
102- Do not peek for first 3-5 days (early results are misleading)
103- Do not stop early unless guardrail metrics violated
104- Monitor for technical issues and tracking accuracy
105- Watch for sample ratio mismatch (SRM): >1% deviation means setup problem
106- Do not add variants mid-test
107 
108### Post-Test Analysis
109 
110```
111TEST RESULTS
112============
113Test: [name] | Duration: [days] | Sample: [n] | Split: [%/%]
114SRM Check: [Pass/Fail]
115 
116| Variant | Visitors | Conversions | CR | vs Control | p-value | Significant? |
117|---------|----------|-------------|-----|------------|---------|--------------|
118| Control | X,XXX | XXX | X.XX% | -- | -- | -- |
119| Var B | X,XXX | XXX | X.XX% | +X.X% | 0.XXX | Yes/No |
120 
121DECISION: [Implement / Keep Control / Iterate]
122REASONING: [Data-based rationale]
123NEXT TEST: [What to test next]
124```
125 
126## Step 7: Common Pitfalls
127 
1281. **Peeking**: Checking daily inflates false positives to 25-30%. Commit to sample size upfront.
1292. **Underpowered tests**: "No result" often means "not enough data."
1303. **Too many variables**: Isolate one variable per test.
1314. **Ignoring segments**: Overall flat, but mobile wins / desktop loses. Always segment.
1325. **Novelty effect**: Run 2+ weeks to account for novelty wearing off.
1336. **Multiple comparisons**: One primary metric. Bonferroni correction for extras.
1347. **Practical significance**: A significant 0.1% lift may not be worth implementing.
135 
136## Step 8: Test Prioritization (ICE Scoring)
137 
138```
139Impact (1-10): How much will this move the metric?
140Confidence (1-10): How likely to produce a result?
141Ease (1-10): How easy to implement?
142ICE Score = (Impact + Confidence + Ease) / 3
143```
144 
145### Roadmap Template
146 
147```
148EXPERIMENTATION ROADMAP
149Quarter: [Q] | Page: [target] | Traffic: [volume] | Current CR: [X%]
150 
151| Priority | Test | ICE | Duration | Status |
152|----------|------|-----|----------|--------|
153| 1 | ... | 8.3 | 14 days | Ready |
154| 2 | ... | 7.7 | 21 days | Ready |
155| 3 | ... | 7.0 | 14 days | Idea |
156```
157 
158Run tests sequentially on the same page to avoid interaction effects. Provide a backlog ranked by ICE score.
159 

Discussion

Alternatives

A/B Test SetupWhen the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program," or "experiment playbook." Use this whenever someone is comparing two approaches and wants to measure which performs better, or when they want to build a systematic experimentation practice. For tracking implementation, see analytics. For page-level conversion optimization, see cro.Marketing · MITAb test analysisAnalyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant. · MITAd Copy Generator + A/B TesterGenerate and A/B test Google Ads copy. Use when asked to write ad copy, headlines, descriptions, create ad variants, test ad messaging, improve CTR, or generate RSA (Responsive Search Ad) components. Trigger on "ad copy", "write ads", "headlines", "descriptions", "RSA", "responsive search ad", "ad text", "ad creative", "improve CTR", "ad A/B test", "ad variants", "write me an ad", "ad variation experiment", or when the user wants to improve click-through rate on existing ads.Marketing · MITA/B Test Planner SkillDesign statistically rigorous A/B tests for product features, UI changes, onboarding flows, and pricing experiments. Use when asked to set up an experiment, design an A/B test, calculate sample size, or interpret test results. Produces a complete test plan with hypothesis, variant definitions, sample size, duration estimate, guardrail metrics, and a results interpretation guide.Marketing · MIT