Ab test analysis skill
Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations.
by phuryn·MIT license·★ 26,557 Stars on the repo·GitHub ↗
npx degit phuryn/pm-skills/pm-data-analytics/skills/ab-test-analysis#main ~/.claude/skills/ab-test-analysis-2Checked ·commit main
Files of Ab test analysis
Show the full text83 lines
A/B Test Analysis
Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.
Context
You are analyzing A/B test results for $ARGUMENTS.
If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.
Instructions
Understand the experiment:
- What was the hypothesis?
- What was changed (the variant)?
- What is the primary metric? Any guardrail metrics?
- How long did the test run?
- What is the traffic split?
Validate the test setup:
- Sample size: Is the sample large enough for the expected effect size?
- Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
- Flag if the test is underpowered (<80% power)
- Duration: Did the test run for at least 1-2 full business cycles?
- Randomization: Any evidence of sample ratio mismatch (SRM)?
- Novelty/primacy effects: Was there enough time to wash out initial behavior changes?
- Sample size: Is the sample large enough for the expected effect size?
Calculate statistical significance:
- Conversion rate for control and variant
- Relative lift: (variant - control) / control × 100
- p-value: Using a two-tailed z-test or chi-squared test
- Confidence interval: 95% CI for the difference
- Statistical significance: Is p < 0.05?
- Practical significance: Is the lift meaningful for the business?
If the user provides raw data, generate and run a Python script to calculate these.
Check guardrail metrics:
- Did any guardrail metrics (revenue, engagement, page load time) degrade?
- A winning primary metric with degraded guardrails may not be a true win
Interpret results:
Outcome Recommendation Significant positive lift, no guardrail issues Ship it — roll out to 100% Significant positive lift, guardrail concerns Investigate — understand trade-offs before shipping Not significant, positive trend Extend the test — need more data or larger effect Not significant, flat Stop the test — no meaningful difference detected Significant negative lift Don't ship — revert to control, analyze why Provide the analysis summary:
## A/B Test Results: [Test Name] **Hypothesis**: [What we expected] **Duration**: [X days] | **Sample**: [N control / M variant] | Metric | Control | Variant | Lift | p-value | Significant? | |---|---|---|---|---|---| | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No | | [Guardrail] | ... | ... | ... | ... | ... | **Recommendation**: [Ship / Extend / Stop / Investigate] **Reasoning**: [Why] **Next steps**: [What to do]
Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.
Further Reading
| 1 | |
| 2 | name ab-test-analysis |
| 3 | description "Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant." |
| 4 | |
| 5 | |
| 6 | ## A/B Test Analysis |
| 7 | |
| 8 | Evaluate A/B test results with statistical rigor and translate findings into clear product decisions. |
| 9 | |
| 10 | ### Context |
| 11 | |
| 12 | You are analyzing A/B test results for **$ARGUMENTS**. |
| 13 | |
| 14 | If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed. |
| 15 | |
| 16 | ### Instructions |
| 17 | |
| 18 | **Understand the experiment**: |
| 19 | What was the hypothesis? |
| 20 | What was changed (the variant)? |
| 21 | What is the primary metric? Any guardrail metrics? |
| 22 | How long did the test run? |
| 23 | What is the traffic split? |
| 24 | |
| 25 | **Validate the test setup**: |
| 26 | **Sample size**: Is the sample large enough for the expected effect size? |
| 27 | Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE² |
| 28 | Flag if the test is underpowered (<80% power) |
| 29 | **Duration**: Did the test run for at least 1-2 full business cycles? |
| 30 | **Randomization**: Any evidence of sample ratio mismatch (SRM)? |
| 31 | **Novelty/primacy effects**: Was there enough time to wash out initial behavior changes? |
| 32 | |
| 33 | **Calculate statistical significance**: |
| 34 | **Conversion rate** for control and variant |
| 35 | **Relative lift**: (variant - control) / control × 100 |
| 36 | **p-value**: Using a two-tailed z-test or chi-squared test |
| 37 | **Confidence interval**: 95% CI for the difference |
| 38 | **Statistical significance**: Is p < 0.05? |
| 39 | **Practical significance**: Is the lift meaningful for the business? |
| 40 | |
| 41 | If the user provides raw data, generate and run a Python script to calculate these. |
| 42 | |
| 43 | **Check guardrail metrics**: |
| 44 | Did any guardrail metrics (revenue, engagement, page load time) degrade? |
| 45 | A winning primary metric with degraded guardrails may not be a true win |
| 46 | |
| 47 | **Interpret results**: |
| 48 | |
| 49 | | Outcome | Recommendation | |
| 50 | |---|---| |
| 51 | | Significant positive lift, no guardrail issues | **Ship it** — roll out to 100% | |
| 52 | | Significant positive lift, guardrail concerns | **Investigate** — understand trade-offs before shipping | |
| 53 | | Not significant, positive trend | **Extend the test** — need more data or larger effect | |
| 54 | | Not significant, flat | **Stop the test** — no meaningful difference detected | |
| 55 | | Significant negative lift | **Don't ship** — revert to control, analyze why | |
| 56 | |
| 57 | **Provide the analysis summary**: |
| 58 | |
| 59 | ## A/B Test Results: [Test Name] |
| 60 | |
| 61 | **Hypothesis**: [What we expected] |
| 62 | **Duration**: [X days] | **Sample**: [N control / M variant] |
| 63 | |
| 64 | | Metric | Control | Variant | Lift | p-value | Significant? | |
| 65 | |---|---|---|---|---|---| |
| 66 | | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No | |
| 67 | | [Guardrail] | ... | ... | ... | ... | ... | |
| 68 | |
| 69 | **Recommendation**: [Ship / Extend / Stop / Investigate] |
| 70 | **Reasoning**: [Why] |
| 71 | **Next steps**: [What to do] |
| 72 | |
| 73 | |
| 74 | Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided. |
| 75 | |
| 76 | |
| 77 | |
| 78 | ### Further Reading |
| 79 | |
| 80 | [A/B Testing 101 + Examples] |
| 81 | [Testing Product Ideas: The Ultimate Validation Experiments Library] |
| 82 | [Are You Tracking the Right Metrics?] |
| 83 |
Discussion
Alternatives
Browse more free Claude skills or everything in Marketing.