Files of CRO testing methodology
wondelai/
Show the full text344 lines
CRO Testing Methodology
Deep-dive into experiment design, statistical rigor, and test prioritization from the CRE Methodology.
Table of Contents
- The Philosophy of Bold Testing
- A/B Testing vs. Multivariate Testing
- Statistical Significance
- ICE Prioritization Framework
- Test Documentation
- When Tests Fail
- CRO Team Dynamics
- Testing Platform Comparison
The Philosophy of Bold Testing
Why "Meek Tweaks" Fail
Most A/B tests fail because they're too small to detect. Button color changes, minor copy tweaks, and micro-optimizations suffer from:
- Insufficient sample size - Small changes require massive traffic to reach significance
- Interaction effects - Minor changes get lost in noise
- Opportunity cost - Time spent on 2% wins could find 200% wins
Rule: Test big changes that could double conversion, not small changes that might move it 5%.
The 10x Mindset
Before any test, ask: "Could this 10x our results?" If not, is it worth testing?
- Worth testing: Complete page redesign, new value proposition, fundamentally different offer
- Not worth testing: Button color, font size, image swap
A/B Testing vs. Multivariate Testing
A/B Testing (Split Testing)
Compare two (or more) complete versions against each other.
| Aspect | Details |
|---|---|
| Best for | Testing big concepts, page redesigns, offers |
| Traffic needed | Lower (split between 2-4 variants) |
| Insights | Which version wins overall |
| Limitation | Doesn't show which elements contributed |
When to use:
- You have a hypothesis about a major change
- Traffic is limited
- You're comparing conceptual approaches
Multivariate Testing (MVT)
Test multiple elements simultaneously to find optimal combination.
| Aspect | Details |
|---|---|
| Best for | Optimizing elements after winning concept proven |
| Traffic needed | Much higher (combinations multiply) |
| Insights | Which specific elements drive results |
| Limitation | Requires significant traffic |
When to use:
- You have a winning page to optimize further
- High traffic (100k+ monthly visitors)
- Clear, isolated elements to test
Traffic Requirements
A/B Test:
Minimum sample per variant = 250-500 conversions
For 2 variants with 5% conversion: 10,000-20,000 visitors needed
Multivariate Test:
Combinations = (Options for Element 1) × (Options for Element 2) × ...
Example: 3 headlines × 3 images × 2 CTAs = 18 combinations
Each combination needs 250+ conversions = 90,000+ conversions total
Recommendation: Start with A/B tests. Only move to MVT when you have:
- Proven winning page concept
- 100k+ monthly visitors
- Mature testing program
Statistical Significance
What It Means
Statistical significance tells you: "How likely is this result due to chance vs. a real effect?"
Industry standard: 95% confidence (p-value < 0.05)
- 95% confident the difference is real
- 5% chance it's random noise
Common Mistakes
1. Peeking and stopping early
Checking results daily and stopping when you see a winner leads to false positives.
- Wrong: "We're at 95% confidence after 3 days—ship it!"
- Right: Pre-determine sample size and test duration; don't stop early
2. Calling tests with insufficient data
| Visitors | Conversions | Can you call it? |
|---|---|---|
| 500 | 15 vs 20 | No |
| 5,000 | 150 vs 200 | Possibly |
| 50,000 | 1,500 vs 2,000 | Yes |
3. Ignoring practical significance
A statistically significant 0.1% lift isn't worth implementation complexity.
4. Multiple comparison problem
Testing 20 variants? One will show "significance" by chance alone.
Sample Size Calculation
Before testing, calculate required sample size:
Inputs needed:
- Baseline conversion rate
- Minimum detectable effect (MDE) you care about
- Statistical power (typically 80%)
- Significance level (typically 95%)
Rule of thumb:
For 5% baseline, 20% relative lift detection:
~25,000 visitors per variant needed
For 5% baseline, 50% relative lift detection:
~4,000 visitors per variant needed
Key insight: The smaller the effect you want to detect, the more traffic you need. This is why bold changes are better—they're detectable with less traffic.
Test Duration
Minimum test duration:
- At least 1 full business cycle (typically 1-2 weeks)
- Include weekdays AND weekends
- Account for seasonality
Why?
- Visitor behavior differs by day of week
- Friday buyers differ from Monday researchers
- Monthly cycles affect B2B especially
ICE Prioritization Framework
Prioritize test ideas using ICE scores:
Impact (1-10)
"If this wins, how big would the impact be?"
| Score | Impact Level |
|---|---|
| 10 | Could double conversion rate |
| 7-9 | Major improvement (30-50%+) |
| 4-6 | Moderate improvement (10-30%) |
| 1-3 | Minor improvement (<10%) |
Confidence (1-10)
"How confident are we this will work?"
| Score | Confidence Level |
|---|---|
| 10 | Proven in research, worked before |
| 7-9 | Strong research supports it |
| 4-6 | Reasonable hypothesis |
| 1-3 | Gut feeling, unvalidated |
Ease (1-10)
"How easy is this to implement and test?"
| Score | Ease Level |
|---|---|
| 10 | Text change only |
| 7-9 | Design change, no dev needed |
| 4-6 | Requires development |
| 1-3 | Major technical lift |
ICE Score Calculation
ICE Score = (Impact + Confidence + Ease) / 3
Or weighted:
ICE Score = (Impact × 2 + Confidence × 1.5 + Ease × 1) / 4.5
Sample Prioritization
| Test Idea | Impact | Confidence | Ease | Score |
|---|---|---|---|---|
| New headline from customer research | 8 | 9 | 10 | 9.0 |
| Add video testimonial | 7 | 7 | 6 | 6.7 |
| Redesign checkout flow | 9 | 6 | 3 | 6.0 |
| Change button color | 2 | 2 | 10 | 4.7 |
Test Documentation
Before the Test
Document:
- Hypothesis: "If we [change X], then [metric Y] will improve because [reason based on research]"
- Primary metric: One metric that determines winner
- Secondary metrics: Additional metrics to monitor
- Guardrail metrics: Metrics that shouldn't decrease
- Sample size requirement
- Test duration
- Traffic allocation
After the Test
Document:
- Results: Raw numbers, conversion rates, confidence interval
- Statistical significance: p-value, confidence level
- Practical significance: Is the lift worth implementing?
- Learnings: What does this teach us about our customers?
- Next steps: Ship winner, iterate, or abandon?
Learnings Database
Every test should add to organizational knowledge:
| Test | Hypothesis | Result | Learning | Applicable to |
|---|---|---|---|---|
| Homepage headline A/B | Customer language converts better | Winner: +27% | Customers care about outcomes, not features | All landing pages |
| Form length test | Shorter forms convert better | Loser: no diff | Our audience expects detailed forms | Lead gen pages |
When Tests Fail
Types of "Failure"
1. No winner (inconclusive)
- Sample size too small
- Effect size too small to detect
- Test needed to run longer
2. Control wins
- New version is worse
- Hypothesis was wrong
- Still a learning!
3. Technical problems
- Tracking broke
- Experience differed from plan
- Sample contamination
What to Do
- Document the learning - "We learned customers prefer X"
- Investigate why - Go back to research
- Don't give up on the page - The opportunity exists, you just haven't found the solution
- Try a bolder change - Maybe the change wasn't big enough
Critical insight: A failed test that teaches you something is more valuable than a winning test you don't understand.
CRO Team Dynamics
Roles in a CRO Program
| Role | Responsibility |
|---|---|
| CRO Lead | Strategy, prioritization, stakeholder management |
| Researcher | User research, surveys, analytics analysis |
| Designer | Wireframes, mockups, user flows |
| Developer | Test implementation, technical QA |
| Analyst | Results analysis, statistical rigor |
Getting Stakeholder Buy-In
Common objections:
| Objection | Counter |
|---|---|
| "We already know what works" | "Then testing will confirm it quickly" |
| "Testing takes too long" | "Shipping wrong things costs more" |
| "Our traffic is too low" | "Then we test bigger changes" |
| "The CEO wants X" | "Let's test to validate the idea" |
Building credibility:
- Start with quick wins (high-traffic pages, obvious problems)
- Document and share learnings widely
- Quantify impact in revenue terms
- Build testing into the culture, not just a project
Test Velocity
Goal: Increase valid tests per month over time.
| Maturity | Tests/Month | Characteristics |
|---|---|---|
| Beginner | 1-2 | Manual processes, ad-hoc |
| Developing | 4-6 | Established backlog, regular cadence |
| Advanced | 10-20 | Parallel testing, mature process |
| Expert | 20+ | Multiple simultaneous tests, automated |
Testing Platform Comparison
| Platform | Best For | Limitations |
|---|---|---|
| Google Optimize | Beginners, free tier | Sunsetting, limited features |
| VWO | Mid-market, visual editor | Can be slow, limited targeting |
| Optimizely | Enterprise, complex tests | Expensive, learning curve |
| LaunchDarkly | Dev-centric, feature flags | Not optimized for marketing |
| Custom | Full control | Development cost |
Key Features to Look For
- Visual editor for non-developers
- Robust statistical engine
- Segment targeting
- Integrations (analytics, CDP, etc.)
- Flicker prevention
- Mutually exclusive experiments
| 1 | # CRO Testing Methodology |
| 2 | |
| 3 | Deep-dive into experiment design, statistical rigor, and test prioritization from the CRE Methodology. |
| 4 | |
| 5 | |
| 6 | ## Table of Contents |
| 7 | [The Philosophy of Bold Testing] |
| 8 | [A/B Testing vs. Multivariate Testing] |
| 9 | [Statistical Significance] |
| 10 | [ICE Prioritization Framework] |
| 11 | [Test Documentation] |
| 12 | [When Tests Fail] |
| 13 | [CRO Team Dynamics] |
| 14 | [Testing Platform Comparison] |
| 15 | |
| 16 | |
| 17 | |
| 18 | ## The Philosophy of Bold Testing |
| 19 | |
| 20 | ### Why "Meek Tweaks" Fail |
| 21 | |
| 22 | Most A/B tests fail because they're too small to detect. Button color changes, minor copy tweaks, and micro-optimizations suffer from: |
| 23 | |
| 24 | **Insufficient sample size** - Small changes require massive traffic to reach significance |
| 25 | **Interaction effects** - Minor changes get lost in noise |
| 26 | **Opportunity cost** - Time spent on 2% wins could find 200% wins |
| 27 | |
| 28 | **Rule: Test big changes that could double conversion, not small changes that might move it 5%.** |
| 29 | |
| 30 | ### The 10x Mindset |
| 31 | |
| 32 | Before any test, ask: "Could this 10x our results?" If not, is it worth testing? |
| 33 | |
| 34 | **Worth testing:** Complete page redesign, new value proposition, fundamentally different offer |
| 35 | **Not worth testing:** Button color, font size, image swap |
| 36 | |
| 37 | |
| 38 | |
| 39 | ## A/B Testing vs. Multivariate Testing |
| 40 | |
| 41 | ### A/B Testing (Split Testing) |
| 42 | |
| 43 | Compare two (or more) complete versions against each other. |
| 44 | |
| 45 | | Aspect | Details | |
| 46 | |--------|---------| |
| 47 | | **Best for** | Testing big concepts, page redesigns, offers | |
| 48 | | **Traffic needed** | Lower (split between 2-4 variants) | |
| 49 | | **Insights** | Which version wins overall | |
| 50 | | **Limitation** | Doesn't show which elements contributed | |
| 51 | |
| 52 | **When to use:** |
| 53 | You have a hypothesis about a major change |
| 54 | Traffic is limited |
| 55 | You're comparing conceptual approaches |
| 56 | |
| 57 | ### Multivariate Testing (MVT) |
| 58 | |
| 59 | Test multiple elements simultaneously to find optimal combination. |
| 60 | |
| 61 | | Aspect | Details | |
| 62 | |--------|---------| |
| 63 | | **Best for** | Optimizing elements after winning concept proven | |
| 64 | | **Traffic needed** | Much higher (combinations multiply) | |
| 65 | | **Insights** | Which specific elements drive results | |
| 66 | | **Limitation** | Requires significant traffic | |
| 67 | |
| 68 | **When to use:** |
| 69 | You have a winning page to optimize further |
| 70 | High traffic (100k+ monthly visitors) |
| 71 | Clear, isolated elements to test |
| 72 | |
| 73 | ### Traffic Requirements |
| 74 | |
| 75 | **A/B Test:** |
| 76 | |
| 77 | Minimum sample per variant = 250-500 conversions |
| 78 | For 2 variants with 5% conversion: 10,000-20,000 visitors needed |
| 79 | |
| 80 | |
| 81 | **Multivariate Test:** |
| 82 | |
| 83 | Combinations = (Options for Element 1) × (Options for Element 2) × ... |
| 84 | Example: 3 headlines × 3 images × 2 CTAs = 18 combinations |
| 85 | Each combination needs 250+ conversions = 90,000+ conversions total |
| 86 | |
| 87 | |
| 88 | **Recommendation:** Start with A/B tests. Only move to MVT when you have: |
| 89 | Proven winning page concept |
| 90 | 100k+ monthly visitors |
| 91 | Mature testing program |
| 92 | |
| 93 | |
| 94 | |
| 95 | ## Statistical Significance |
| 96 | |
| 97 | ### What It Means |
| 98 | |
| 99 | Statistical significance tells you: "How likely is this result due to chance vs. a real effect?" |
| 100 | |
| 101 | **Industry standard:** 95% confidence (p-value < 0.05) |
| 102 | 95% confident the difference is real |
| 103 | 5% chance it's random noise |
| 104 | |
| 105 | ### Common Mistakes |
| 106 | |
| 107 | **1. Peeking and stopping early** |
| 108 | |
| 109 | Checking results daily and stopping when you see a winner leads to false positives. |
| 110 | |
| 111 | **Wrong:** "We're at 95% confidence after 3 days—ship it!" |
| 112 | **Right:** Pre-determine sample size and test duration; don't stop early |
| 113 | |
| 114 | **2. Calling tests with insufficient data** |
| 115 | |
| 116 | | Visitors | Conversions | Can you call it? | |
| 117 | |----------|-------------|------------------| |
| 118 | | 500 | 15 vs 20 | No | |
| 119 | | 5,000 | 150 vs 200 | Possibly | |
| 120 | | 50,000 | 1,500 vs 2,000 | Yes | |
| 121 | |
| 122 | **3. Ignoring practical significance** |
| 123 | |
| 124 | A statistically significant 0.1% lift isn't worth implementation complexity. |
| 125 | |
| 126 | **4. Multiple comparison problem** |
| 127 | |
| 128 | Testing 20 variants? One will show "significance" by chance alone. |
| 129 | |
| 130 | ### Sample Size Calculation |
| 131 | |
| 132 | Before testing, calculate required sample size: |
| 133 | |
| 134 | **Inputs needed:** |
| 135 | Baseline conversion rate |
| 136 | Minimum detectable effect (MDE) you care about |
| 137 | Statistical power (typically 80%) |
| 138 | Significance level (typically 95%) |
| 139 | |
| 140 | **Rule of thumb:** |
| 141 | |
| 142 | For 5% baseline, 20% relative lift detection: |
| 143 | ~25,000 visitors per variant needed |
| 144 | |
| 145 | For 5% baseline, 50% relative lift detection: |
| 146 | ~4,000 visitors per variant needed |
| 147 | |
| 148 | |
| 149 | **Key insight:** The smaller the effect you want to detect, the more traffic you need. This is why bold changes are better—they're detectable with less traffic. |
| 150 | |
| 151 | ### Test Duration |
| 152 | |
| 153 | **Minimum test duration:** |
| 154 | At least 1 full business cycle (typically 1-2 weeks) |
| 155 | Include weekdays AND weekends |
| 156 | Account for seasonality |
| 157 | |
| 158 | **Why?** |
| 159 | Visitor behavior differs by day of week |
| 160 | Friday buyers differ from Monday researchers |
| 161 | Monthly cycles affect B2B especially |
| 162 | |
| 163 | |
| 164 | |
| 165 | ## ICE Prioritization Framework |
| 166 | |
| 167 | Prioritize test ideas using ICE scores: |
| 168 | |
| 169 | ### Impact (1-10) |
| 170 | "If this wins, how big would the impact be?" |
| 171 | |
| 172 | | Score | Impact Level | |
| 173 | |-------|--------------| |
| 174 | | 10 | Could double conversion rate | |
| 175 | | 7-9 | Major improvement (30-50%+) | |
| 176 | | 4-6 | Moderate improvement (10-30%) | |
| 177 | | 1-3 | Minor improvement (<10%) | |
| 178 | |
| 179 | ### Confidence (1-10) |
| 180 | "How confident are we this will work?" |
| 181 | |
| 182 | | Score | Confidence Level | |
| 183 | |-------|------------------| |
| 184 | | 10 | Proven in research, worked before | |
| 185 | | 7-9 | Strong research supports it | |
| 186 | | 4-6 | Reasonable hypothesis | |
| 187 | | 1-3 | Gut feeling, unvalidated | |
| 188 | |
| 189 | ### Ease (1-10) |
| 190 | "How easy is this to implement and test?" |
| 191 | |
| 192 | | Score | Ease Level | |
| 193 | |-------|------------| |
| 194 | | 10 | Text change only | |
| 195 | | 7-9 | Design change, no dev needed | |
| 196 | | 4-6 | Requires development | |
| 197 | | 1-3 | Major technical lift | |
| 198 | |
| 199 | ### ICE Score Calculation |
| 200 | |
| 201 | |
| 202 | ICE Score = (Impact + Confidence + Ease) / 3 |
| 203 | |
| 204 | |
| 205 | Or weighted: |
| 206 | |
| 207 | ICE Score = (Impact × 2 + Confidence × 1.5 + Ease × 1) / 4.5 |
| 208 | |
| 209 | |
| 210 | ### Sample Prioritization |
| 211 | |
| 212 | | Test Idea | Impact | Confidence | Ease | Score | |
| 213 | |-----------|--------|------------|------|-------| |
| 214 | | New headline from customer research | 8 | 9 | 10 | 9.0 | |
| 215 | | Add video testimonial | 7 | 7 | 6 | 6.7 | |
| 216 | | Redesign checkout flow | 9 | 6 | 3 | 6.0 | |
| 217 | | Change button color | 2 | 2 | 10 | 4.7 | |
| 218 | |
| 219 | |
| 220 | |
| 221 | ## Test Documentation |
| 222 | |
| 223 | ### Before the Test |
| 224 | |
| 225 | Document: |
| 226 | **Hypothesis:** "If we [change X], then [metric Y] will improve because [reason based on research]" |
| 227 | **Primary metric:** One metric that determines winner |
| 228 | **Secondary metrics:** Additional metrics to monitor |
| 229 | **Guardrail metrics:** Metrics that shouldn't decrease |
| 230 | **Sample size requirement** |
| 231 | **Test duration** |
| 232 | **Traffic allocation** |
| 233 | |
| 234 | ### After the Test |
| 235 | |
| 236 | Document: |
| 237 | **Results:** Raw numbers, conversion rates, confidence interval |
| 238 | **Statistical significance:** p-value, confidence level |
| 239 | **Practical significance:** Is the lift worth implementing? |
| 240 | **Learnings:** What does this teach us about our customers? |
| 241 | **Next steps:** Ship winner, iterate, or abandon? |
| 242 | |
| 243 | ### Learnings Database |
| 244 | |
| 245 | Every test should add to organizational knowledge: |
| 246 | |
| 247 | | Test | Hypothesis | Result | Learning | Applicable to | |
| 248 | |------|------------|--------|----------|---------------| |
| 249 | | Homepage headline A/B | Customer language converts better | Winner: +27% | Customers care about outcomes, not features | All landing pages | |
| 250 | | Form length test | Shorter forms convert better | Loser: no diff | Our audience expects detailed forms | Lead gen pages | |
| 251 | |
| 252 | |
| 253 | |
| 254 | ## When Tests Fail |
| 255 | |
| 256 | ### Types of "Failure" |
| 257 | |
| 258 | **1. No winner (inconclusive)** |
| 259 | Sample size too small |
| 260 | Effect size too small to detect |
| 261 | Test needed to run longer |
| 262 | |
| 263 | **2. Control wins** |
| 264 | New version is worse |
| 265 | Hypothesis was wrong |
| 266 | Still a learning! |
| 267 | |
| 268 | **3. Technical problems** |
| 269 | Tracking broke |
| 270 | Experience differed from plan |
| 271 | Sample contamination |
| 272 | |
| 273 | ### What to Do |
| 274 | |
| 275 | **Document the learning** - "We learned customers prefer X" |
| 276 | **Investigate why** - Go back to research |
| 277 | **Don't give up on the page** - The opportunity exists, you just haven't found the solution |
| 278 | **Try a bolder change** - Maybe the change wasn't big enough |
| 279 | |
| 280 | **Critical insight:** A failed test that teaches you something is more valuable than a winning test you don't understand. |
| 281 | |
| 282 | |
| 283 | |
| 284 | ## CRO Team Dynamics |
| 285 | |
| 286 | ### Roles in a CRO Program |
| 287 | |
| 288 | | Role | Responsibility | |
| 289 | |------|----------------| |
| 290 | | **CRO Lead** | Strategy, prioritization, stakeholder management | |
| 291 | | **Researcher** | User research, surveys, analytics analysis | |
| 292 | | **Designer** | Wireframes, mockups, user flows | |
| 293 | | **Developer** | Test implementation, technical QA | |
| 294 | | **Analyst** | Results analysis, statistical rigor | |
| 295 | |
| 296 | ### Getting Stakeholder Buy-In |
| 297 | |
| 298 | **Common objections:** |
| 299 | |
| 300 | | Objection | Counter | |
| 301 | |-----------|---------| |
| 302 | | "We already know what works" | "Then testing will confirm it quickly" | |
| 303 | | "Testing takes too long" | "Shipping wrong things costs more" | |
| 304 | | "Our traffic is too low" | "Then we test bigger changes" | |
| 305 | | "The CEO wants X" | "Let's test to validate the idea" | |
| 306 | |
| 307 | **Building credibility:** |
| 308 | Start with quick wins (high-traffic pages, obvious problems) |
| 309 | Document and share learnings widely |
| 310 | Quantify impact in revenue terms |
| 311 | Build testing into the culture, not just a project |
| 312 | |
| 313 | ### Test Velocity |
| 314 | |
| 315 | **Goal:** Increase valid tests per month over time. |
| 316 | |
| 317 | | Maturity | Tests/Month | Characteristics | |
| 318 | |----------|-------------|-----------------| |
| 319 | | Beginner | 1-2 | Manual processes, ad-hoc | |
| 320 | | Developing | 4-6 | Established backlog, regular cadence | |
| 321 | | Advanced | 10-20 | Parallel testing, mature process | |
| 322 | | Expert | 20+ | Multiple simultaneous tests, automated | |
| 323 | |
| 324 | |
| 325 | |
| 326 | ## Testing Platform Comparison |
| 327 | |
| 328 | | Platform | Best For | Limitations | |
| 329 | |----------|----------|-------------| |
| 330 | | Google Optimize | Beginners, free tier | Sunsetting, limited features | |
| 331 | | VWO | Mid-market, visual editor | Can be slow, limited targeting | |
| 332 | | Optimizely | Enterprise, complex tests | Expensive, learning curve | |
| 333 | | LaunchDarkly | Dev-centric, feature flags | Not optimized for marketing | |
| 334 | | Custom | Full control | Development cost | |
| 335 | |
| 336 | ### Key Features to Look For |
| 337 | |
| 338 | Visual editor for non-developers |
| 339 | Robust statistical engine |
| 340 | Segment targeting |
| 341 | Integrations (analytics, CDP, etc.) |
| 342 | Flicker prevention |
| 343 | Mutually exclusive experiments |
| 344 |
Discussion
Alternatives
Browse more free Claude skills or everything in Product.