19 ab test setup global skill

Use when the user wants a VALID experiment instead of a guess — hypothesis, one variable, sample size and runtime math, statistical significance, primary versus secondary metrics, multi-arm designs, and a results template, across Optimizely, VWO, and native Meta and Google tests.

by minhnv0807·MIT license·★ 599 Stars on the repo·GitHub ↗

Use now

Files of 19 ab test setup global

minhnv0807/master1 file shown
SKILL.md
Show the full text400 lines

A/B Test Setup (Global)

Run experiments that produce decisions, not noise. Most "A/B tests" in marketing are underpowered, peeked-at, and badly hypothesized — meaning the team learns nothing and ships the louder variant.


For Newbies

A valid A/B test answers one question: "Did this change cause a real improvement, or am I seeing noise?"

To answer it credibly you need four things:

  1. A specific hypothesis with a numeric prediction
  2. One variable changed (everything else identical)
  3. Enough sample to detect the effect you care about
  4. Statistical significance before you call a winner (typically p < 0.05)

If any one of these is missing, you don't have an A/B test — you have a coin flip with extra steps.

Common newbie mistake: running a test for 3 days, seeing variant B 40% higher, declaring victory, and shipping. Three days is too short to absorb day-of-week effects, and small samples produce wild swings. Variant B may revert (or reverse) by day 14.


Step 0 — Read Context

Read .agents/product-marketing-context.md if it exists. Audience size, average traffic, and current conversion rate determine whether a test is even feasible.


Step 1 — Information Gathering

Ask up to 4 questions:

  1. What are you testing? (Ad headline / Landing page section / Email subject / Pricing display / CTA button / Creative video)
  2. Primary metric? (CTR / Conversion rate / CPM / CPA / Revenue / Open rate / Reply rate)
  3. Daily traffic to the test surface? (Needed for sample size and duration)
  4. Goal of the test? (Lift X% on primary metric / Pick a winner among N candidates / Validate a strategic hypothesis)

The 7 Principles of a Valid A/B Test

1. Test exactly one variable

The cardinal rule. Change two things at once and you cannot attribute the result.

  • Bad: "I changed the headline, the hero image, and the CTA color." → You learn nothing about which element drove the lift.
  • Good: Change only the headline. Image, CTA, layout, traffic source, and audience targeting are identical.

If you must test multiple changes, use a multivariate test (MVT) — but those need much more traffic (often 4×–8× a single A/B).

2. Hypothesize with a number

Format: "If we [change X], [metric Y] will increase by [Z%] because [reason]."

  • Good: "If we change the CTA from 'Sign up' to 'Get my free demo,' conversion rate will increase by 15% because action-specific language reduces ambiguity."
  • Bad: "The new copy will be better." (No metric, no number, no causal reasoning — un-testable.)

The "because" matters: if your hypothesis is wrong but the reasoning was sound, you've still learned something generalizable.

3. Sufficient sample size

Don't stop early. Statistical tests need adequate data to distinguish signal from noise.

  • Minimum rule of thumb: 100 conversions per variant (not 100 visitors)
  • Better: Calculate sample size up front based on baseline conversion rate and minimum detectable effect (formula below)
4. Sufficient duration

Run for whole weeks, not 3 days, not 10 days. Different weekdays produce different audience behavior — Monday B2B traffic is not Saturday DTC traffic.

  • Minimum: 7 days
  • Recommended: 14 days
  • Watch for: holidays, paydays, monthly billing cycles, ad spend ramp-ups
5. Don't peek

Looking at results every hour and stopping when "B looks good" is the most common error in marketing experimentation. Early peeks combined with early stops dramatically inflate false positive rates.

  • Define the end date in advance. Honor it.
  • If you must monitor, use sequential testing methods designed for it (Bayesian frameworks like Optimizely's Stats Engine, or platforms with built-in sequential controls).
6. Statistical significance: p < 0.05

Most marketing teams use 95% confidence (p-value < 0.05) as the bar.

  • p-value < 0.05 → less than 5% chance the observed difference is random
  • p-value 0.05–0.10 → suggestive but inconclusive — extend the test
  • p-value > 0.10 → no evidence of an effect — keep control or test something else

For high-stakes tests (pricing, branding) consider 99% confidence (p < 0.01).

7. Document everything

Write down:

  • Hypothesis (with number)
  • Start date / end date
  • Sample size achieved
  • Primary metric, secondary metrics
  • Result + p-value
  • Decision + reasoning
  • What you'd test next

A documented test history prevents your team from re-testing things that already failed and from forgetting why you made past decisions.


Sample Size Calculation

Quick formula
Sample size per variant ≈ 16 × p × (1 − p) / MDE²

where:
  p   = baseline conversion rate (e.g. 0.03 = 3%)
  MDE = minimum detectable effect, in absolute terms
        (e.g. 0.006 = lift from 3% to 3.6%)

This produces sample size for 80% power, 95% confidence, 50/50 split — sensible defaults for most marketing tests.

Worked example A — landing page CRO

Current conversion rate is 3%. You want to detect a 20% relative lift (from 3% to 3.6%).

  • p = 0.03
  • MDE (absolute) = 0.20 × 0.03 = 0.006
  • Sample size per variant = 16 × 0.03 × 0.97 / 0.006² = 12,933 visitors
  • Total: ~25,866 visitors. At 500 visitors/day → ~52 days.

That's slow. Either run it (if the change matters), test something with a bigger expected lift, or get more traffic on the test surface.

Worked example B — email subject line

Current open rate is 25%. You want to detect a 10% relative lift (to 27.5%).

  • p = 0.25
  • MDE = 0.025
  • Sample size per variant = 16 × 0.25 × 0.75 / 0.025² = 4,800 sends
  • Total: 9,600 sends per email — usually achievable in one campaign.
Feasibility quick-reference
Daily volume Conv. rate Days needed Test feasibility
< 100 any 2+ months Skip — focus on traffic first
100–500 2–5% 3–6 weeks Yes, but be patient
500–2K 2–5% 2–3 weeks Yes — ideal range
2K–10K 2–5% 1–2 weeks Yes — rapid iteration
10K+ any days Yes — multi-arm tests possible

If volume is below 100/day, A/B testing is statistically wasted — concentrate on increasing traffic before running experiments.


Multi-Arm and Multivariate

Beyond simple A vs B:

  • Multi-arm (A/B/C/D): test 3+ variants at once. Sample size grows roughly linearly with arms.
  • Multivariate (MVT): test multiple elements simultaneously (headline × image × CTA = 8 combinations). Sample size grows multiplicatively. Only viable with very high traffic.
  • Sequential / Bayesian (Thompson Sampling): dynamically allocate more traffic to better-performing variants. Optimizely, Google Optimize successors, and Meta's auto-optimization use this.

For most teams: stick to A/B until traffic exceeds ~10K/day on the test surface.


What to Test (in priority order)

1. Headline — highest impact

Roughly 80% of visitors read the headline; 20% read the body. Optimizing the headline gives the largest expected lift per unit of effort.

Variations to try:

  • Question vs statement
  • Specific number vs generic ("3,247 founders trust us" vs "Trusted by founders")
  • Outcome-focused vs feature-focused
  • Short (5–7 words) vs long (12–15 words)
2. CTA button

Easy to change, often 5–25% lift potential.

Variations:

  • Text: "Sign up" vs "Get free demo" vs "Start free trial" vs "See pricing"
  • Color: brand primary vs contrast (high-contrast usually wins)
  • Size: standard vs large
  • Position: above-the-fold vs sticky vs end-of-page
3. Hero visual
  • Product shot vs lifestyle shot
  • Static image vs video
  • Founder face vs anonymous model
  • Demo screencast vs testimonial clip
4. Pricing display
  • Monthly vs annual primary
  • Strikethrough discount vs clean price
  • Number formatting ($299 vs $299.00 vs $299/mo)
  • Anchor pricing (3-tier with middle highlighted)
5. Social proof placement
  • Numbers vs detailed reviews
  • Logo wall vs customer count
  • Above CTA vs below CTA
  • Video testimonial vs text testimonial
6. Form fields
  • 3 vs 5 vs 7 fields (fewer fields almost always wins on conversion, but lead quality may drop)
  • Label position (above vs left)
  • Single-step vs multi-step
  • Optional fields marked vs required marked
7. Email subject lines
  • Question vs benefit
  • Personalization vs generic
  • Emoji vs no emoji (varies by region/audience)
  • Length (under 40 chars vs 60+)
8. Ad creative
  • First 3-second hook variants
  • UGC style vs polished brand
  • Format: single image vs carousel vs video vs Reels-native
  • Headline on creative vs in copy field

Tool Recommendations

Tool Best for Cost
Meta Ads built-in A/B test Creative, audience, placement on Meta Free
TikTok Ads Split Test TikTok ad creative and audience tests Free
Google Ads Experiments Google Ads campaigns and ad copy Free
Optimizely Web Enterprise web experimentation, sequential testing $$$ enterprise
VWO Mid-market web A/B + heatmaps $199+/mo
Convert.com Privacy-first web testing $99+/mo
PostHog Product feature flags + experiments + analytics Free tier, generous
GrowthBook Open-source A/B testing platform Free / hosted plans
Statsig Product experimentation with feature flags Free tier
AB Tasty Web experimentation + personalization $$$
Unbounce / Instapage Built-in A/B for landing pages $90+/mo
Custom (split URL) Two pages, 50/50 redirect, GA4/Pixel attribution Free

Note: Google Optimize was sunset in September 2023. Migration paths: GA4 + a third-party platform (Optimizely, VWO, Convert) or PostHog/GrowthBook for product-led teams.


Setup Without Dedicated Tools

1. Build two versions of the page: /landing-a and /landing-b
2. Split traffic 50/50:
   - Meta Ads: 2 ad sets, identical audience, different destination URLs
   - Google Ads: 2 ads in the same ad group, identical targeting, different URLs
   - Email: list-split feature in your ESP
3. Track conversions per variant:
   - Meta Pixel custom event with parameter: page_version = "A" / "B"
   - GA4 event with custom dimension
   - PostHog feature flag exposure event
4. Run for the planned duration. Don't peek mid-test.
5. Export raw counts. Run significance test (calculator below).

Result Analysis

Statistical significance

Use a calculator. Recommended:

  • Evan Miller's calculator — evanmiller.org/ab-testing/chi-squared.html
  • AB Testguide — abtestguide.com/calc/
  • Optimizely's calculator — built into platform
  • Survey Monkey calculator — for sample size pre-test

Inputs:

  • Variant A: visitors + conversions
  • Variant B: visitors + conversions

Outputs:

  • p-value (need < 0.05 for 95% confidence)
  • Confidence interval on the lift
  • Lift % (relative or absolute)
Decision matrix
p-value Lift size Decision
< 0.05 > 5% B wins — implement and document
< 0.05 < 5% Significant but small — weigh implementation cost
0.05–0.10 > 10% Borderline — extend test if feasible
> 0.10 any No evidence — keep A or design a stronger test
Common Pitfalls
  1. Peeking and stopping early. Most common cause of false positives. If the platform shows "B is winning" on day 3, the platform is misleading you (unless it's specifically designed for sequential testing).
  2. Uneven splits. If split is 30/70 instead of intended 50/50, your delivery infrastructure has a bug. Investigate before trusting results.
  3. Seasonality. Tests run only on weekdays vs weekends produce different results. Always run for whole weeks.
  4. Novelty effect. New variants attract attention for the first 2–3 days, then performance regresses. Long enough tests absorb this.
  5. Sample ratio mismatch (SRM). Even with 50/50 intent, if visitor counts diverge significantly (e.g. 8,400 vs 11,600), there's likely a tracking or assignment bug. Tools like PostHog and Optimizely flag this automatically.
  6. Mixing traffic sources mid-test. Don't add a new ad campaign halfway through — it changes audience composition.
  7. Multiple comparison problem. Running 20 simultaneous tests means ~1 will look "significant" by chance alone. Adjust thresholds (Bonferroni correction) or pre-register hypotheses.

Output Template

# A/B Test: [test name]
Created: [YYYY-MM-DD]
Owner: [name]

## 1. Hypothesis
"If we [change X], [metric Y] will increase by [Z%] because [reason]."

## 2. Variants
- Variant A (Control): [current state description]
- Variant B (Challenger): [changed state description]
- Single change: [the one element that differs]

## 3. Metrics
- Primary: [e.g. conversion rate]
- Secondary (guardrails): [e.g. bounce rate, time on page, AOV]

## 4. Sample size & duration
- Baseline (p): [%]
- Minimum detectable effect (MDE): [%]
- Sample needed per variant: [N]
- Daily traffic to test surface: [N]
- Estimated days to complete: [N]

## 5. Setup
- Tool: [Optimizely / VWO / PostHog / Meta built-in / custom]
- Variant A URL or asset: [...]
- Variant B URL or asset: [...]
- Tracking events: [list]
- Split ratio: 50/50

## 6. Timeline
- Start: [date]
- End (planned): [date]
- Review meeting: [date]

## 7. Results (filled in after test ends)
| Variant | Visitors | Conversions | Rate | Lift vs A |
|---------|----------|-------------|------|-----------|
| A | | | | — |
| B | | | | +X% |

p-value: [x]
95% CI on lift: [lower%, upper%]
Significant (p < 0.05): [Yes / No]

## 8. Decision
[Ship B / Keep A / Inconclusive — extend or redesign]

## 9. Action
[Implement variant B globally / Roll back / Schedule next iteration]

## 10. Lessons
[What this teaches generalizable for future tests]

Quality Checklist

  • Exactly one variable changed
  • Hypothesis is specific and includes a numeric prediction
  • Sample size calculated up front (minimum 100 conv/variant)
  • Duration is at least 1 full week, ideally 2 weeks
  • No peeking — end date defined and honored
  • p-value calculated before declaring a winner (target p < 0.05)
  • Result documented (winner or not — both are learning)
  • Secondary metrics checked (no guardrail violations)
  • Sample ratio verified (no SRM red flags)
  • Next test identified based on what this one taught
1---
2name: 19-ab-test-setup-global
3description: "Use when the user wants a VALID experiment instead of a guess — hypothesis, one variable, sample size and runtime math, statistical significance, primary versus secondary metrics, multi-arm designs, and a results template, across Optimizely, VWO, and native Meta and Google tests. Trigger on 'A/B test', 'split test', 'how long should I run the test', 'is this result significant', 'test two versions', 'which creative is actually better'. Also use when a winner was declared after two days on tiny numbers. Not for — scaling the proven winner, see `55-scaling-ads-global`; analyzing data already collected, see `13-data-analysis-global`; auditing the account, see `21-ads-audit-global`."
4metadata:
5 version: 1.0.1
6 category: performance
7 language: en
8license: MIT
9triggers:
10 - "A/B test"
11 - "split test"
12 - "multivariate test"
13 - "experiment design"
14 - "statistical significance"
15 - "sample size calculator"
16output: A .md file containing hypothesis, sample size calculation, primary/secondary metrics, test setup, timeline, and a results template ready for analysis
17related:
18 - product-marketing-context-global
19 - 13-data-analysis-global
20 - 03-performance-eval-global
21 - 21-ads-audit-global
22---
23 
24# A/B Test Setup (Global)
25 
26> Run experiments that produce decisions, not noise. Most "A/B tests" in marketing are underpowered, peeked-at, and badly hypothesized — meaning the team learns nothing and ships the louder variant.
27 
28---
29 
30## For Newbies
31 
32A valid A/B test answers one question: "Did this change cause a real improvement, or am I seeing noise?"
33 
34To answer it credibly you need four things:
351. A **specific hypothesis** with a numeric prediction
362. **One variable changed** (everything else identical)
373. **Enough sample** to detect the effect you care about
384. **Statistical significance** before you call a winner (typically p < 0.05)
39 
40If any one of these is missing, you don't have an A/B test — you have a coin flip with extra steps.
41 
42**Common newbie mistake:** running a test for 3 days, seeing variant B 40% higher, declaring victory, and shipping. Three days is too short to absorb day-of-week effects, and small samples produce wild swings. Variant B may revert (or reverse) by day 14.
43 
44---
45 
46## Step 0 — Read Context
47 
48Read `.agents/product-marketing-context.md` if it exists. Audience size, average traffic, and current conversion rate determine whether a test is even feasible.
49 
50---
51 
52## Step 1 — Information Gathering
53 
54Ask up to 4 questions:
55 
561. **What are you testing?** (Ad headline / Landing page section / Email subject / Pricing display / CTA button / Creative video)
572. **Primary metric?** (CTR / Conversion rate / CPM / CPA / Revenue / Open rate / Reply rate)
583. **Daily traffic to the test surface?** (Needed for sample size and duration)
594. **Goal of the test?** (Lift X% on primary metric / Pick a winner among N candidates / Validate a strategic hypothesis)
60 
61---
62 
63## The 7 Principles of a Valid A/B Test
64 
65### 1. Test exactly one variable
66 
67The cardinal rule. Change two things at once and you cannot attribute the result.
68 
69- **Bad:** "I changed the headline, the hero image, and the CTA color." → You learn nothing about which element drove the lift.
70- **Good:** Change only the headline. Image, CTA, layout, traffic source, and audience targeting are identical.
71 
72If you must test multiple changes, use a **multivariate test (MVT)** — but those need much more traffic (often 4×–8× a single A/B).
73 
74### 2. Hypothesize with a number
75 
76**Format:** "If we [change X], [metric Y] will increase by [Z%] because [reason]."
77 
78- **Good:** "If we change the CTA from 'Sign up' to 'Get my free demo,' conversion rate will increase by 15% because action-specific language reduces ambiguity."
79- **Bad:** "The new copy will be better." (No metric, no number, no causal reasoning — un-testable.)
80 
81The "because" matters: if your hypothesis is wrong but the reasoning was sound, you've still learned something generalizable.
82 
83### 3. Sufficient sample size
84 
85Don't stop early. Statistical tests need adequate data to distinguish signal from noise.
86 
87- **Minimum rule of thumb:** 100 conversions per variant (not 100 visitors)
88- **Better:** Calculate sample size up front based on baseline conversion rate and minimum detectable effect (formula below)
89 
90### 4. Sufficient duration
91 
92Run for **whole weeks**, not 3 days, not 10 days. Different weekdays produce different audience behavior — Monday B2B traffic is not Saturday DTC traffic.
93 
94- **Minimum:** 7 days
95- **Recommended:** 14 days
96- **Watch for:** holidays, paydays, monthly billing cycles, ad spend ramp-ups
97 
98### 5. Don't peek
99 
100Looking at results every hour and stopping when "B looks good" is the most common error in marketing experimentation. Early peeks combined with early stops dramatically inflate false positive rates.
101 
102- Define the end date in advance. Honor it.
103- If you must monitor, use **sequential testing** methods designed for it (Bayesian frameworks like Optimizely's Stats Engine, or platforms with built-in sequential controls).
104 
105### 6. Statistical significance: p < 0.05
106 
107Most marketing teams use **95% confidence (p-value < 0.05)** as the bar.
108 
109- p-value < 0.05 → less than 5% chance the observed difference is random
110- p-value 0.05–0.10 → suggestive but inconclusive — extend the test
111- p-value > 0.10 → no evidence of an effect — keep control or test something else
112 
113For high-stakes tests (pricing, branding) consider 99% confidence (p < 0.01).
114 
115### 7. Document everything
116 
117Write down:
118- Hypothesis (with number)
119- Start date / end date
120- Sample size achieved
121- Primary metric, secondary metrics
122- Result + p-value
123- Decision + reasoning
124- What you'd test next
125 
126A documented test history prevents your team from re-testing things that already failed and from forgetting why you made past decisions.
127 
128---
129 
130## Sample Size Calculation
131 
132### Quick formula
133 
134```
135Sample size per variant ≈ 16 × p × (1 − p) / MDE²
136 
137where:
138 p = baseline conversion rate (e.g. 0.03 = 3%)
139 MDE = minimum detectable effect, in absolute terms
140 (e.g. 0.006 = lift from 3% to 3.6%)
141```
142 
143This produces sample size for **80% power, 95% confidence, 50/50 split** — sensible defaults for most marketing tests.
144 
145### Worked example A — landing page CRO
146 
147Current conversion rate is 3%. You want to detect a 20% relative lift (from 3% to 3.6%).
148 
149- p = 0.03
150- MDE (absolute) = 0.20 × 0.03 = 0.006
151- Sample size per variant = 16 × 0.03 × 0.97 / 0.006² = **12,933 visitors**
152- Total: ~25,866 visitors. At 500 visitors/day → ~52 days.
153 
154That's slow. Either run it (if the change matters), test something with a bigger expected lift, or get more traffic on the test surface.
155 
156### Worked example B — email subject line
157 
158Current open rate is 25%. You want to detect a 10% relative lift (to 27.5%).
159 
160- p = 0.25
161- MDE = 0.025
162- Sample size per variant = 16 × 0.25 × 0.75 / 0.025² = **4,800 sends**
163- Total: 9,600 sends per email — usually achievable in one campaign.
164 
165### Feasibility quick-reference
166 
167| Daily volume | Conv. rate | Days needed | Test feasibility |
168|--------------|-----------|------------|------------------|
169| < 100 | any | 2+ months | Skip — focus on traffic first |
170| 100–500 | 2–5% | 3–6 weeks | Yes, but be patient |
171| 500–2K | 2–5% | 2–3 weeks | Yes — ideal range |
172| 2K–10K | 2–5% | 1–2 weeks | Yes — rapid iteration |
173| 10K+ | any | days | Yes — multi-arm tests possible |
174 
175If volume is below 100/day, A/B testing is statistically wasted — concentrate on increasing traffic before running experiments.
176 
177---
178 
179## Multi-Arm and Multivariate
180 
181Beyond simple A vs B:
182 
183- **Multi-arm (A/B/C/D):** test 3+ variants at once. Sample size grows roughly linearly with arms.
184- **Multivariate (MVT):** test multiple elements simultaneously (headline × image × CTA = 8 combinations). Sample size grows multiplicatively. Only viable with very high traffic.
185- **Sequential / Bayesian (Thompson Sampling):** dynamically allocate more traffic to better-performing variants. Optimizely, Google Optimize successors, and Meta's auto-optimization use this.
186 
187For most teams: stick to A/B until traffic exceeds ~10K/day on the test surface.
188 
189---
190 
191## What to Test (in priority order)
192 
193### 1. Headline — highest impact
194Roughly 80% of visitors read the headline; 20% read the body. Optimizing the headline gives the largest expected lift per unit of effort.
195 
196Variations to try:
197- Question vs statement
198- Specific number vs generic ("3,247 founders trust us" vs "Trusted by founders")
199- Outcome-focused vs feature-focused
200- Short (5–7 words) vs long (12–15 words)
201 
202### 2. CTA button
203Easy to change, often 5–25% lift potential.
204 
205Variations:
206- Text: "Sign up" vs "Get free demo" vs "Start free trial" vs "See pricing"
207- Color: brand primary vs contrast (high-contrast usually wins)
208- Size: standard vs large
209- Position: above-the-fold vs sticky vs end-of-page
210 
211### 3. Hero visual
212- Product shot vs lifestyle shot
213- Static image vs video
214- Founder face vs anonymous model
215- Demo screencast vs testimonial clip
216 
217### 4. Pricing display
218- Monthly vs annual primary
219- Strikethrough discount vs clean price
220- Number formatting ($299 vs $299.00 vs $299/mo)
221- Anchor pricing (3-tier with middle highlighted)
222 
223### 5. Social proof placement
224- Numbers vs detailed reviews
225- Logo wall vs customer count
226- Above CTA vs below CTA
227- Video testimonial vs text testimonial
228 
229### 6. Form fields
230- 3 vs 5 vs 7 fields (fewer fields almost always wins on conversion, but lead quality may drop)
231- Label position (above vs left)
232- Single-step vs multi-step
233- Optional fields marked vs required marked
234 
235### 7. Email subject lines
236- Question vs benefit
237- Personalization vs generic
238- Emoji vs no emoji (varies by region/audience)
239- Length (under 40 chars vs 60+)
240 
241### 8. Ad creative
242- First 3-second hook variants
243- UGC style vs polished brand
244- Format: single image vs carousel vs video vs Reels-native
245- Headline on creative vs in copy field
246 
247---
248 
249## Tool Recommendations
250 
251| Tool | Best for | Cost |
252|------|----------|------|
253| **Meta Ads built-in A/B test** | Creative, audience, placement on Meta | Free |
254| **TikTok Ads Split Test** | TikTok ad creative and audience tests | Free |
255| **Google Ads Experiments** | Google Ads campaigns and ad copy | Free |
256| **Optimizely Web** | Enterprise web experimentation, sequential testing | $$$ enterprise |
257| **VWO** | Mid-market web A/B + heatmaps | $199+/mo |
258| **Convert.com** | Privacy-first web testing | $99+/mo |
259| **PostHog** | Product feature flags + experiments + analytics | Free tier, generous |
260| **GrowthBook** | Open-source A/B testing platform | Free / hosted plans |
261| **Statsig** | Product experimentation with feature flags | Free tier |
262| **AB Tasty** | Web experimentation + personalization | $$$ |
263| **Unbounce / Instapage** | Built-in A/B for landing pages | $90+/mo |
264| **Custom (split URL)** | Two pages, 50/50 redirect, GA4/Pixel attribution | Free |
265 
266> **Note:** Google Optimize was sunset in September 2023. Migration paths: GA4 + a third-party platform (Optimizely, VWO, Convert) or PostHog/GrowthBook for product-led teams.
267 
268---
269 
270## Setup Without Dedicated Tools
271 
272```
2731. Build two versions of the page: /landing-a and /landing-b
2742. Split traffic 50/50:
275 - Meta Ads: 2 ad sets, identical audience, different destination URLs
276 - Google Ads: 2 ads in the same ad group, identical targeting, different URLs
277 - Email: list-split feature in your ESP
2783. Track conversions per variant:
279 - Meta Pixel custom event with parameter: page_version = "A" / "B"
280 - GA4 event with custom dimension
281 - PostHog feature flag exposure event
2824. Run for the planned duration. Don't peek mid-test.
2835. Export raw counts. Run significance test (calculator below).
284```
285 
286---
287 
288## Result Analysis
289 
290### Statistical significance
291 
292Use a calculator. Recommended:
293- **Evan Miller's calculator** — `evanmiller.org/ab-testing/chi-squared.html`
294- **AB Testguide** — `abtestguide.com/calc/`
295- **Optimizely's calculator** — built into platform
296- **Survey Monkey calculator** — for sample size pre-test
297 
298**Inputs:**
299- Variant A: visitors + conversions
300- Variant B: visitors + conversions
301 
302**Outputs:**
303- p-value (need < 0.05 for 95% confidence)
304- Confidence interval on the lift
305- Lift % (relative or absolute)
306 
307### Decision matrix
308 
309| p-value | Lift size | Decision |
310|---------|-----------|----------|
311| < 0.05 | > 5% | B wins — implement and document |
312| < 0.05 | < 5% | Significant but small — weigh implementation cost |
313| 0.05–0.10 | > 10% | Borderline — extend test if feasible |
314| > 0.10 | any | No evidence — keep A or design a stronger test |
315 
316### Common Pitfalls
317 
3181. **Peeking and stopping early.** Most common cause of false positives. If the platform shows "B is winning" on day 3, the platform is misleading you (unless it's specifically designed for sequential testing).
3192. **Uneven splits.** If split is 30/70 instead of intended 50/50, your delivery infrastructure has a bug. Investigate before trusting results.
3203. **Seasonality.** Tests run only on weekdays vs weekends produce different results. Always run for whole weeks.
3214. **Novelty effect.** New variants attract attention for the first 2–3 days, then performance regresses. Long enough tests absorb this.
3225. **Sample ratio mismatch (SRM).** Even with 50/50 intent, if visitor counts diverge significantly (e.g. 8,400 vs 11,600), there's likely a tracking or assignment bug. Tools like PostHog and Optimizely flag this automatically.
3236. **Mixing traffic sources mid-test.** Don't add a new ad campaign halfway through — it changes audience composition.
3247. **Multiple comparison problem.** Running 20 simultaneous tests means ~1 will look "significant" by chance alone. Adjust thresholds (Bonferroni correction) or pre-register hypotheses.
325 
326---
327 
328## Output Template
329 
330```markdown
331# A/B Test: [test name]
332Created: [YYYY-MM-DD]
333Owner: [name]
334 
335## 1. Hypothesis
336"If we [change X], [metric Y] will increase by [Z%] because [reason]."
337 
338## 2. Variants
339- Variant A (Control): [current state description]
340- Variant B (Challenger): [changed state description]
341- Single change: [the one element that differs]
342 
343## 3. Metrics
344- Primary: [e.g. conversion rate]
345- Secondary (guardrails): [e.g. bounce rate, time on page, AOV]
346 
347## 4. Sample size & duration
348- Baseline (p): [%]
349- Minimum detectable effect (MDE): [%]
350- Sample needed per variant: [N]
351- Daily traffic to test surface: [N]
352- Estimated days to complete: [N]
353 
354## 5. Setup
355- Tool: [Optimizely / VWO / PostHog / Meta built-in / custom]
356- Variant A URL or asset: [...]
357- Variant B URL or asset: [...]
358- Tracking events: [list]
359- Split ratio: 50/50
360 
361## 6. Timeline
362- Start: [date]
363- End (planned): [date]
364- Review meeting: [date]
365 
366## 7. Results (filled in after test ends)
367| Variant | Visitors | Conversions | Rate | Lift vs A |
368|---------|----------|-------------|------|-----------|
369| A | | | | — |
370| B | | | | +X% |
371 
372p-value: [x]
37395% CI on lift: [lower%, upper%]
374Significant (p < 0.05): [Yes / No]
375 
376## 8. Decision
377[Ship B / Keep A / Inconclusive — extend or redesign]
378 
379## 9. Action
380[Implement variant B globally / Roll back / Schedule next iteration]
381 
382## 10. Lessons
383[What this teaches generalizable for future tests]
384```
385 
386---
387 
388## Quality Checklist
389 
390- [ ] Exactly one variable changed
391- [ ] Hypothesis is specific and includes a numeric prediction
392- [ ] Sample size calculated up front (minimum 100 conv/variant)
393- [ ] Duration is at least 1 full week, ideally 2 weeks
394- [ ] No peeking — end date defined and honored
395- [ ] p-value calculated before declaring a winner (target p < 0.05)
396- [ ] Result documented (winner or not — both are learning)
397- [ ] Secondary metrics checked (no guardrail violations)
398- [ ] Sample ratio verified (no SRM red flags)
399- [ ] Next test identified based on what this one taught
400 

Discussion

Alternatives

A/B Test SetupWhen the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program," or "experiment playbook." Use this whenever someone is comparing two approaches and wants to measure which performs better, or when they want to build a systematic experimentation practice. For tracking implementation, see analytics. For page-level conversion optimization, see cro.Marketing · MITAb test analysisAnalyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant. · MITAd Copy Generator + A/B TesterGenerate and A/B test Google Ads copy. Use when asked to write ad copy, headlines, descriptions, create ad variants, test ad messaging, improve CTR, or generate RSA (Responsive Search Ad) components. Trigger on "ad copy", "write ads", "headlines", "descriptions", "RSA", "responsive search ad", "ad text", "ad creative", "improve CTR", "ad A/B test", "ad variants", "write me an ad", "ad variation experiment", or when the user wants to improve click-through rate on existing ads.Marketing · MITA/B Test Planner SkillDesign statistically rigorous A/B tests for product features, UI changes, onboarding flows, and pricing experiments. Use when asked to set up an experiment, design an A/B test, calculate sample size, or interpret test results. Produces a complete test plan with hypothesis, variant definitions, sample size, duration estimate, guardrail metrics, and a results interpretation guide.Marketing · MIT