CRO testing methodology skill

Deep-dive into experiment design, statistical rigor, and test prioritization from the CRE Methodology.

by wondelai·MIT license·★ 2,235 Stars on the repo·GitHub ↗

Use now

Files of CRO testing methodology

wondelai/main1 file
testing-methodology.md
Show the full text344 lines

CRO Testing Methodology

Deep-dive into experiment design, statistical rigor, and test prioritization from the CRE Methodology.

Table of Contents

  1. The Philosophy of Bold Testing
  2. A/B Testing vs. Multivariate Testing
  3. Statistical Significance
  4. ICE Prioritization Framework
  5. Test Documentation
  6. When Tests Fail
  7. CRO Team Dynamics
  8. Testing Platform Comparison

The Philosophy of Bold Testing

Why "Meek Tweaks" Fail

Most A/B tests fail because they're too small to detect. Button color changes, minor copy tweaks, and micro-optimizations suffer from:

  1. Insufficient sample size - Small changes require massive traffic to reach significance
  2. Interaction effects - Minor changes get lost in noise
  3. Opportunity cost - Time spent on 2% wins could find 200% wins

Rule: Test big changes that could double conversion, not small changes that might move it 5%.

The 10x Mindset

Before any test, ask: "Could this 10x our results?" If not, is it worth testing?

  • Worth testing: Complete page redesign, new value proposition, fundamentally different offer
  • Not worth testing: Button color, font size, image swap

A/B Testing vs. Multivariate Testing

A/B Testing (Split Testing)

Compare two (or more) complete versions against each other.

Aspect Details
Best for Testing big concepts, page redesigns, offers
Traffic needed Lower (split between 2-4 variants)
Insights Which version wins overall
Limitation Doesn't show which elements contributed

When to use:

  • You have a hypothesis about a major change
  • Traffic is limited
  • You're comparing conceptual approaches
Multivariate Testing (MVT)

Test multiple elements simultaneously to find optimal combination.

Aspect Details
Best for Optimizing elements after winning concept proven
Traffic needed Much higher (combinations multiply)
Insights Which specific elements drive results
Limitation Requires significant traffic

When to use:

  • You have a winning page to optimize further
  • High traffic (100k+ monthly visitors)
  • Clear, isolated elements to test
Traffic Requirements

A/B Test:

Minimum sample per variant = 250-500 conversions
For 2 variants with 5% conversion: 10,000-20,000 visitors needed

Multivariate Test:

Combinations = (Options for Element 1) × (Options for Element 2) × ...
Example: 3 headlines × 3 images × 2 CTAs = 18 combinations
Each combination needs 250+ conversions = 90,000+ conversions total

Recommendation: Start with A/B tests. Only move to MVT when you have:

  • Proven winning page concept
  • 100k+ monthly visitors
  • Mature testing program

Statistical Significance

What It Means

Statistical significance tells you: "How likely is this result due to chance vs. a real effect?"

Industry standard: 95% confidence (p-value < 0.05)

  • 95% confident the difference is real
  • 5% chance it's random noise
Common Mistakes

1. Peeking and stopping early

Checking results daily and stopping when you see a winner leads to false positives.

  • Wrong: "We're at 95% confidence after 3 days—ship it!"
  • Right: Pre-determine sample size and test duration; don't stop early

2. Calling tests with insufficient data

Visitors Conversions Can you call it?
500 15 vs 20 No
5,000 150 vs 200 Possibly
50,000 1,500 vs 2,000 Yes

3. Ignoring practical significance

A statistically significant 0.1% lift isn't worth implementation complexity.

4. Multiple comparison problem

Testing 20 variants? One will show "significance" by chance alone.

Sample Size Calculation

Before testing, calculate required sample size:

Inputs needed:

  • Baseline conversion rate
  • Minimum detectable effect (MDE) you care about
  • Statistical power (typically 80%)
  • Significance level (typically 95%)

Rule of thumb:

For 5% baseline, 20% relative lift detection:
~25,000 visitors per variant needed

For 5% baseline, 50% relative lift detection:
~4,000 visitors per variant needed

Key insight: The smaller the effect you want to detect, the more traffic you need. This is why bold changes are better—they're detectable with less traffic.

Test Duration

Minimum test duration:

  • At least 1 full business cycle (typically 1-2 weeks)
  • Include weekdays AND weekends
  • Account for seasonality

Why?

  • Visitor behavior differs by day of week
  • Friday buyers differ from Monday researchers
  • Monthly cycles affect B2B especially

ICE Prioritization Framework

Prioritize test ideas using ICE scores:

Impact (1-10)

"If this wins, how big would the impact be?"

Score Impact Level
10 Could double conversion rate
7-9 Major improvement (30-50%+)
4-6 Moderate improvement (10-30%)
1-3 Minor improvement (<10%)
Confidence (1-10)

"How confident are we this will work?"

Score Confidence Level
10 Proven in research, worked before
7-9 Strong research supports it
4-6 Reasonable hypothesis
1-3 Gut feeling, unvalidated
Ease (1-10)

"How easy is this to implement and test?"

Score Ease Level
10 Text change only
7-9 Design change, no dev needed
4-6 Requires development
1-3 Major technical lift
ICE Score Calculation
ICE Score = (Impact + Confidence + Ease) / 3

Or weighted:

ICE Score = (Impact × 2 + Confidence × 1.5 + Ease × 1) / 4.5
Sample Prioritization
Test Idea Impact Confidence Ease Score
New headline from customer research 8 9 10 9.0
Add video testimonial 7 7 6 6.7
Redesign checkout flow 9 6 3 6.0
Change button color 2 2 10 4.7

Test Documentation

Before the Test

Document:

  1. Hypothesis: "If we [change X], then [metric Y] will improve because [reason based on research]"
  2. Primary metric: One metric that determines winner
  3. Secondary metrics: Additional metrics to monitor
  4. Guardrail metrics: Metrics that shouldn't decrease
  5. Sample size requirement
  6. Test duration
  7. Traffic allocation
After the Test

Document:

  1. Results: Raw numbers, conversion rates, confidence interval
  2. Statistical significance: p-value, confidence level
  3. Practical significance: Is the lift worth implementing?
  4. Learnings: What does this teach us about our customers?
  5. Next steps: Ship winner, iterate, or abandon?
Learnings Database

Every test should add to organizational knowledge:

Test Hypothesis Result Learning Applicable to
Homepage headline A/B Customer language converts better Winner: +27% Customers care about outcomes, not features All landing pages
Form length test Shorter forms convert better Loser: no diff Our audience expects detailed forms Lead gen pages

When Tests Fail

Types of "Failure"

1. No winner (inconclusive)

  • Sample size too small
  • Effect size too small to detect
  • Test needed to run longer

2. Control wins

  • New version is worse
  • Hypothesis was wrong
  • Still a learning!

3. Technical problems

  • Tracking broke
  • Experience differed from plan
  • Sample contamination
What to Do
  1. Document the learning - "We learned customers prefer X"
  2. Investigate why - Go back to research
  3. Don't give up on the page - The opportunity exists, you just haven't found the solution
  4. Try a bolder change - Maybe the change wasn't big enough

Critical insight: A failed test that teaches you something is more valuable than a winning test you don't understand.


CRO Team Dynamics

Roles in a CRO Program
Role Responsibility
CRO Lead Strategy, prioritization, stakeholder management
Researcher User research, surveys, analytics analysis
Designer Wireframes, mockups, user flows
Developer Test implementation, technical QA
Analyst Results analysis, statistical rigor
Getting Stakeholder Buy-In

Common objections:

Objection Counter
"We already know what works" "Then testing will confirm it quickly"
"Testing takes too long" "Shipping wrong things costs more"
"Our traffic is too low" "Then we test bigger changes"
"The CEO wants X" "Let's test to validate the idea"

Building credibility:

  1. Start with quick wins (high-traffic pages, obvious problems)
  2. Document and share learnings widely
  3. Quantify impact in revenue terms
  4. Build testing into the culture, not just a project
Test Velocity

Goal: Increase valid tests per month over time.

Maturity Tests/Month Characteristics
Beginner 1-2 Manual processes, ad-hoc
Developing 4-6 Established backlog, regular cadence
Advanced 10-20 Parallel testing, mature process
Expert 20+ Multiple simultaneous tests, automated

Testing Platform Comparison

Platform Best For Limitations
Google Optimize Beginners, free tier Sunsetting, limited features
VWO Mid-market, visual editor Can be slow, limited targeting
Optimizely Enterprise, complex tests Expensive, learning curve
LaunchDarkly Dev-centric, feature flags Not optimized for marketing
Custom Full control Development cost
Key Features to Look For
  • Visual editor for non-developers
  • Robust statistical engine
  • Segment targeting
  • Integrations (analytics, CDP, etc.)
  • Flicker prevention
  • Mutually exclusive experiments
1# CRO Testing Methodology
2 
3Deep-dive into experiment design, statistical rigor, and test prioritization from the CRE Methodology.
4 
5 
6## Table of Contents
71. [The Philosophy of Bold Testing](#the-philosophy-of-bold-testing)
82. [A/B Testing vs. Multivariate Testing](#ab-testing-vs-multivariate-testing)
93. [Statistical Significance](#statistical-significance)
104. [ICE Prioritization Framework](#ice-prioritization-framework)
115. [Test Documentation](#test-documentation)
126. [When Tests Fail](#when-tests-fail)
137. [CRO Team Dynamics](#cro-team-dynamics)
148. [Testing Platform Comparison](#testing-platform-comparison)
15 
16---
17 
18## The Philosophy of Bold Testing
19 
20### Why "Meek Tweaks" Fail
21 
22Most A/B tests fail because they're too small to detect. Button color changes, minor copy tweaks, and micro-optimizations suffer from:
23 
241. **Insufficient sample size** - Small changes require massive traffic to reach significance
252. **Interaction effects** - Minor changes get lost in noise
263. **Opportunity cost** - Time spent on 2% wins could find 200% wins
27 
28**Rule: Test big changes that could double conversion, not small changes that might move it 5%.**
29 
30### The 10x Mindset
31 
32Before any test, ask: "Could this 10x our results?" If not, is it worth testing?
33 
34- **Worth testing:** Complete page redesign, new value proposition, fundamentally different offer
35- **Not worth testing:** Button color, font size, image swap
36 
37---
38 
39## A/B Testing vs. Multivariate Testing
40 
41### A/B Testing (Split Testing)
42 
43Compare two (or more) complete versions against each other.
44 
45| Aspect | Details |
46|--------|---------|
47| **Best for** | Testing big concepts, page redesigns, offers |
48| **Traffic needed** | Lower (split between 2-4 variants) |
49| **Insights** | Which version wins overall |
50| **Limitation** | Doesn't show which elements contributed |
51 
52**When to use:**
53- You have a hypothesis about a major change
54- Traffic is limited
55- You're comparing conceptual approaches
56 
57### Multivariate Testing (MVT)
58 
59Test multiple elements simultaneously to find optimal combination.
60 
61| Aspect | Details |
62|--------|---------|
63| **Best for** | Optimizing elements after winning concept proven |
64| **Traffic needed** | Much higher (combinations multiply) |
65| **Insights** | Which specific elements drive results |
66| **Limitation** | Requires significant traffic |
67 
68**When to use:**
69- You have a winning page to optimize further
70- High traffic (100k+ monthly visitors)
71- Clear, isolated elements to test
72 
73### Traffic Requirements
74 
75**A/B Test:**
76```
77Minimum sample per variant = 250-500 conversions
78For 2 variants with 5% conversion: 10,000-20,000 visitors needed
79```
80 
81**Multivariate Test:**
82```
83Combinations = (Options for Element 1) × (Options for Element 2) × ...
84Example: 3 headlines × 3 images × 2 CTAs = 18 combinations
85Each combination needs 250+ conversions = 90,000+ conversions total
86```
87 
88**Recommendation:** Start with A/B tests. Only move to MVT when you have:
89- Proven winning page concept
90- 100k+ monthly visitors
91- Mature testing program
92 
93---
94 
95## Statistical Significance
96 
97### What It Means
98 
99Statistical significance tells you: "How likely is this result due to chance vs. a real effect?"
100 
101**Industry standard:** 95% confidence (p-value < 0.05)
102- 95% confident the difference is real
103- 5% chance it's random noise
104 
105### Common Mistakes
106 
107**1. Peeking and stopping early**
108 
109Checking results daily and stopping when you see a winner leads to false positives.
110 
111- **Wrong:** "We're at 95% confidence after 3 days—ship it!"
112- **Right:** Pre-determine sample size and test duration; don't stop early
113 
114**2. Calling tests with insufficient data**
115 
116| Visitors | Conversions | Can you call it? |
117|----------|-------------|------------------|
118| 500 | 15 vs 20 | No |
119| 5,000 | 150 vs 200 | Possibly |
120| 50,000 | 1,500 vs 2,000 | Yes |
121 
122**3. Ignoring practical significance**
123 
124A statistically significant 0.1% lift isn't worth implementation complexity.
125 
126**4. Multiple comparison problem**
127 
128Testing 20 variants? One will show "significance" by chance alone.
129 
130### Sample Size Calculation
131 
132Before testing, calculate required sample size:
133 
134**Inputs needed:**
135- Baseline conversion rate
136- Minimum detectable effect (MDE) you care about
137- Statistical power (typically 80%)
138- Significance level (typically 95%)
139 
140**Rule of thumb:**
141```
142For 5% baseline, 20% relative lift detection:
143~25,000 visitors per variant needed
144 
145For 5% baseline, 50% relative lift detection:
146~4,000 visitors per variant needed
147```
148 
149**Key insight:** The smaller the effect you want to detect, the more traffic you need. This is why bold changes are better—they're detectable with less traffic.
150 
151### Test Duration
152 
153**Minimum test duration:**
154- At least 1 full business cycle (typically 1-2 weeks)
155- Include weekdays AND weekends
156- Account for seasonality
157 
158**Why?**
159- Visitor behavior differs by day of week
160- Friday buyers differ from Monday researchers
161- Monthly cycles affect B2B especially
162 
163---
164 
165## ICE Prioritization Framework
166 
167Prioritize test ideas using ICE scores:
168 
169### Impact (1-10)
170"If this wins, how big would the impact be?"
171 
172| Score | Impact Level |
173|-------|--------------|
174| 10 | Could double conversion rate |
175| 7-9 | Major improvement (30-50%+) |
176| 4-6 | Moderate improvement (10-30%) |
177| 1-3 | Minor improvement (<10%) |
178 
179### Confidence (1-10)
180"How confident are we this will work?"
181 
182| Score | Confidence Level |
183|-------|------------------|
184| 10 | Proven in research, worked before |
185| 7-9 | Strong research supports it |
186| 4-6 | Reasonable hypothesis |
187| 1-3 | Gut feeling, unvalidated |
188 
189### Ease (1-10)
190"How easy is this to implement and test?"
191 
192| Score | Ease Level |
193|-------|------------|
194| 10 | Text change only |
195| 7-9 | Design change, no dev needed |
196| 4-6 | Requires development |
197| 1-3 | Major technical lift |
198 
199### ICE Score Calculation
200 
201```
202ICE Score = (Impact + Confidence + Ease) / 3
203```
204 
205Or weighted:
206```
207ICE Score = (Impact × 2 + Confidence × 1.5 + Ease × 1) / 4.5
208```
209 
210### Sample Prioritization
211 
212| Test Idea | Impact | Confidence | Ease | Score |
213|-----------|--------|------------|------|-------|
214| New headline from customer research | 8 | 9 | 10 | 9.0 |
215| Add video testimonial | 7 | 7 | 6 | 6.7 |
216| Redesign checkout flow | 9 | 6 | 3 | 6.0 |
217| Change button color | 2 | 2 | 10 | 4.7 |
218 
219---
220 
221## Test Documentation
222 
223### Before the Test
224 
225Document:
2261. **Hypothesis:** "If we [change X], then [metric Y] will improve because [reason based on research]"
2272. **Primary metric:** One metric that determines winner
2283. **Secondary metrics:** Additional metrics to monitor
2294. **Guardrail metrics:** Metrics that shouldn't decrease
2305. **Sample size requirement**
2316. **Test duration**
2327. **Traffic allocation**
233 
234### After the Test
235 
236Document:
2371. **Results:** Raw numbers, conversion rates, confidence interval
2382. **Statistical significance:** p-value, confidence level
2393. **Practical significance:** Is the lift worth implementing?
2404. **Learnings:** What does this teach us about our customers?
2415. **Next steps:** Ship winner, iterate, or abandon?
242 
243### Learnings Database
244 
245Every test should add to organizational knowledge:
246 
247| Test | Hypothesis | Result | Learning | Applicable to |
248|------|------------|--------|----------|---------------|
249| Homepage headline A/B | Customer language converts better | Winner: +27% | Customers care about outcomes, not features | All landing pages |
250| Form length test | Shorter forms convert better | Loser: no diff | Our audience expects detailed forms | Lead gen pages |
251 
252---
253 
254## When Tests Fail
255 
256### Types of "Failure"
257 
258**1. No winner (inconclusive)**
259- Sample size too small
260- Effect size too small to detect
261- Test needed to run longer
262 
263**2. Control wins**
264- New version is worse
265- Hypothesis was wrong
266- Still a learning!
267 
268**3. Technical problems**
269- Tracking broke
270- Experience differed from plan
271- Sample contamination
272 
273### What to Do
274 
2751. **Document the learning** - "We learned customers prefer X"
2762. **Investigate why** - Go back to research
2773. **Don't give up on the page** - The opportunity exists, you just haven't found the solution
2784. **Try a bolder change** - Maybe the change wasn't big enough
279 
280**Critical insight:** A failed test that teaches you something is more valuable than a winning test you don't understand.
281 
282---
283 
284## CRO Team Dynamics
285 
286### Roles in a CRO Program
287 
288| Role | Responsibility |
289|------|----------------|
290| **CRO Lead** | Strategy, prioritization, stakeholder management |
291| **Researcher** | User research, surveys, analytics analysis |
292| **Designer** | Wireframes, mockups, user flows |
293| **Developer** | Test implementation, technical QA |
294| **Analyst** | Results analysis, statistical rigor |
295 
296### Getting Stakeholder Buy-In
297 
298**Common objections:**
299 
300| Objection | Counter |
301|-----------|---------|
302| "We already know what works" | "Then testing will confirm it quickly" |
303| "Testing takes too long" | "Shipping wrong things costs more" |
304| "Our traffic is too low" | "Then we test bigger changes" |
305| "The CEO wants X" | "Let's test to validate the idea" |
306 
307**Building credibility:**
3081. Start with quick wins (high-traffic pages, obvious problems)
3092. Document and share learnings widely
3103. Quantify impact in revenue terms
3114. Build testing into the culture, not just a project
312 
313### Test Velocity
314 
315**Goal:** Increase valid tests per month over time.
316 
317| Maturity | Tests/Month | Characteristics |
318|----------|-------------|-----------------|
319| Beginner | 1-2 | Manual processes, ad-hoc |
320| Developing | 4-6 | Established backlog, regular cadence |
321| Advanced | 10-20 | Parallel testing, mature process |
322| Expert | 20+ | Multiple simultaneous tests, automated |
323 
324---
325 
326## Testing Platform Comparison
327 
328| Platform | Best For | Limitations |
329|----------|----------|-------------|
330| Google Optimize | Beginners, free tier | Sunsetting, limited features |
331| VWO | Mid-market, visual editor | Can be slow, limited targeting |
332| Optimizely | Enterprise, complex tests | Expensive, learning curve |
333| LaunchDarkly | Dev-centric, feature flags | Not optimized for marketing |
334| Custom | Full control | Development cost |
335 
336### Key Features to Look For
337 
338- Visual editor for non-developers
339- Robust statistical engine
340- Segment targeting
341- Integrations (analytics, CDP, etc.)
342- Flicker prevention
343- Mutually exclusive experiments
344 

Discussion

Alternatives

Research methodology design for health literacy and medication adherence in aotearoa new zealandExplore the methodological design for researching health literacy and its impact on medication adherence among adults with chronic diseases in Aotearoa New Zealand.Business & ops · CC0-1.0Scientific critical thinkingEvaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.Science · MITAcademic research synthesizerAcademic research synthesis specialist. Use PROACTIVELY for comprehensive research on academic topics, literature reviews, technical investigations, and well-cited analysis combining multiple sources. <example>Context: A podcast episode needs a segment grounded in peer-reviewed evidence with formal citations. user: "Research the current state of transformer efficiency techniques for the episode, with proper academic citations." assistant: "I'll use the academic-research-synthesizer agent to search arXiv and Semantic Scholar, extract full-text findings via WebFetch, and produce a cited literature synthesis with confidence levels." <commentary>Use academic-research-synthesizer (not comprehensive-researcher) when the episode segment needs peer-reviewed sourcing, formal citation format, and explicit confidence tagging rather than general-purpose multi-source coverage.</commentary></example> <example>Context: The episode-orchestrator has routed a "literature review" request for a technical deep-dive segment. user: "Summarize the research landscape on federated learning privacy guarantees." assistant: "I'll invoke academic-research-synthesizer to systematically search academic sources, note peer-review status per source, and synthesize consensus vs. open debates."</example>Business & ops · MITAcademic researcherAcademic research specialist for scholarly sources, peer-reviewed papers, and academic literature. Use PROACTIVELY for research paper analysis, literature reviews, citation tracking, and academic methodology evaluation. <example>Context: The research-orchestrator has kicked off Phase 4 parallel research on 'efficacy of intermittent fasting' and needs peer-reviewed evidence. user: "Find the academic evidence on intermittent fasting outcomes." assistant: "I'll use the academic-researcher agent to search Semantic Scholar, PubMed, and OpenAlex for peer-reviewed studies and write structured findings to academic-research.md." <commentary>The request is specifically for scholarly/peer-reviewed evidence rather than general web coverage or code, so academic-researcher (not web-researcher or technical-researcher) is the right specialist.</commentary></example> <example>Context: The user wants a literature review comparing methodologies across studies on a topic. user: "Can you review the literature on transformer model interpretability and identify research gaps?" assistant: "Let me invoke the academic-researcher agent to pull foundational and recent papers, extract methodologies, and surface open research gaps." <commentary>Literature review, methodology extraction, and research-gap identification are core academic-researcher capabilities, distinct from technical-researcher's focus on code repositories and implementations.</commentary></example>Business & ops · MIT