Experiment patterns for lean UX skill

Experiments are the engine of learning in Lean UX.

by wondelai·MIT license·★ 2,235 Stars on the repo·GitHub ↗

Use now

Files of Experiment patterns for lean UX

wondelai/main1 file
experiment-patterns.md
Show the full text316 lines

Experiment Patterns for Lean UX

Experiments are the engine of learning in Lean UX. The right experiment answers the hypothesis with the least effort. Choosing the wrong experiment type wastes time, money, or both. This reference covers the full spectrum of UX experiments, from napkin sketches to coded A/B tests.

Table of Contents

  1. Types of UX Experiments
  2. Choosing the Right Experiment
  3. Experiment Design Template
  4. Minimum Viable Tests
  5. Running Experiments in Practice
  6. Experiment Cheat Sheet

Types of UX Experiments

1. Paper Prototypes

What it is: Hand-drawn screens on paper or index cards. A facilitator plays "computer," swapping screens as the user taps or points.

Best for: Early concept validation, flow testing, information architecture.

Effort: Very low (30 minutes to create). Fidelity: Very low. Confidence: Low-medium. Validates flow and concept, not visual design or interaction details.

When to use:

  • You have multiple competing concepts and need to narrow down
  • The hypothesis is about flow or content, not aesthetics
  • You need to test today, not next week

When NOT to use:

  • The hypothesis depends on visual design, animation, or micro-interactions
  • Users need to interact with real data
  • Stakeholders will not trust low-fidelity evidence

How to run:

  1. Sketch each screen on a separate sheet or card
  2. Write a realistic task scenario for the participant
  3. Ask the participant to "tap" or point at what they would interact with
  4. Swap screens manually based on their choices
  5. Note where they hesitate, get confused, or go off-script
2. Clickable Prototypes

What it is: Interactive mockups built in tools like Figma, Sketch, or InVision. Users click through a realistic-looking interface, but no backend logic exists.

Best for: Usability testing, flow validation, stakeholder buy-in, developer communication.

Effort: Medium (1-3 days). Fidelity: Medium-high. Confidence: Medium-high. Validates flow, layout, and basic usability.

When to use:

  • The hypothesis involves user navigation or task completion
  • You need to test with users who expect a realistic experience
  • The prototype will also serve as a design reference for developers

When NOT to use:

  • A paper prototype would suffice (over-investing)
  • The hypothesis is about performance, load times, or real data behavior
  • You need to test with hundreds of users (use coded experiments instead)

How to run:

  1. Build the key screens and link hotspots in your prototyping tool
  2. Write 3-5 task scenarios
  3. Recruit 5-8 participants matching your persona
  4. Run moderated sessions (20-30 minutes each)
  5. Track task completion, time on task, errors, and qualitative feedback
3. Concierge MVP

What it is: Deliver the service or experience manually, person-to-person, without building any technology. The user receives the full value, but the backend is entirely human-powered.

Best for: Validating that the solution genuinely solves the problem before investing in automation.

Effort: Medium (ongoing manual work per user). Fidelity: High (the experience is real). Confidence: High. Real behavior with real value delivery.

When to use:

  • You are unsure whether the solution concept works at all
  • The cost of building the automated version is high
  • You want to deeply understand the user's experience and edge cases

When NOT to use:

  • The hypothesis is about scale or technology performance
  • You need to test with more than 10-20 users simultaneously
  • The value proposition depends on speed that only automation can provide

Example: A meal-planning app manually emails personalized weekly meal plans and shopping lists to 10 users based on their dietary preferences, before building the algorithm.

4. Wizard of Oz

What it is: The user interacts with what appears to be a functioning product, but a human behind the scenes is performing the work the technology would eventually do.

Best for: Testing the user experience of an automated feature before building the automation.

Effort: Medium (build the frontend; humans operate the backend). Fidelity: High from the user's perspective. Confidence: High. Users interact with what feels like a real product.

When to use:

  • The hypothesis depends on the user experience of an AI, algorithm, or automation feature
  • Building the actual technology is expensive or risky
  • You want to learn what the "right" output looks like before training a model

When NOT to use:

  • The hypothesis is about system performance or response time
  • Manual operation cannot replicate the technology's speed
  • Ethical issues arise from deception (always disclose if legally required)

Example: A "smart" scheduling assistant that appears to use AI but is actually a team member reading requests and sending calendar invites manually.

5. Landing Page / Smoke Test

What it is: A single web page describing a product or feature that does not yet exist, with a call to action (sign up, pre-order, request access). Measures demand by tracking how many people take the action.

Best for: Demand validation before building anything.

Effort: Low (half a day to create). Fidelity: Low (no product). Confidence: Medium. Measures stated intent, not actual usage.

When to use:

  • You need to validate demand before committing development resources
  • The hypothesis is about whether people want this at all
  • You want to build an early-access list for future testing

When NOT to use:

  • You already know there is demand and need to validate usability
  • The product concept is hard to explain without a demo
  • Your audience is internal (use interviews instead)

How to run:

  1. Create a landing page with a clear value proposition, 2-3 key benefits, and a CTA
  2. Drive targeted traffic (ads, social media, email, communities)
  3. Measure conversion rate (visitors to CTA clicks or sign-ups)
  4. Set success threshold before launch (e.g., 5% sign-up rate from 500 visitors)
  5. Follow up with sign-ups for qualitative interviews
6. A/B Test (Coded Experiment)

What it is: Two or more versions of a live feature are shown to different user segments. Statistical analysis determines which version performs better on a target metric.

Best for: Optimizing existing features, validating specific design changes with statistical rigor.

Effort: High (requires code, traffic, and statistical analysis). Fidelity: Production-level. Confidence: Very high (if properly powered).

When to use:

  • You have sufficient traffic to reach statistical significance
  • The hypothesis involves a measurable behavior change in an existing product
  • You need high-confidence evidence to justify a significant investment

When NOT to use:

  • Traffic is too low for statistical significance (fewer than 1,000 users per variant)
  • The concept is entirely new (test with prototypes first)
  • The change is too small to produce a detectable effect

Choosing the Right Experiment

The decision depends on three factors: what you need to learn, how much confidence you need, and how much you can invest.

Experiment Selection Matrix
Question to Answer Best Experiment Fidelity Time Confidence
"Does anyone want this?" Landing page smoke test Low 1-2 days Medium
"Does the flow make sense?" Paper prototype Very low 1 day Low-Medium
"Can users complete this task?" Clickable prototype Medium 3-5 days Medium-High
"Does this solution actually work?" Concierge MVP High 1-2 weeks High
"Will the automated version feel right?" Wizard of Oz High 1-2 weeks High
"Which version performs better?" A/B test Production 2-4 weeks Very High
The Fidelity Ladder

Start at the lowest rung that can answer your question. Only climb higher when lower fidelity cannot provide the needed confidence.

Level 1: Paper prototype / Sketches
  ↓ (if concept validated, test usability)
Level 2: Clickable prototype (Figma, Sketch)
  ↓ (if usability validated, test real value)
Level 3: Concierge MVP / Wizard of Oz
  ↓ (if value validated, test at scale)
Level 4: Coded experiment / A/B test
  ↓ (if optimized, ship)
Level 5: Production release

Experiment Design Template

Use this template for every experiment, regardless of type:

EXPERIMENT DESIGN
=================
Date: _______________
Hypothesis ID: _______________
Experimenter: _______________

HYPOTHESIS
We believe _______________
will happen if _______________
achieves _______________
with _______________.

EXPERIMENT TYPE
[ ] Paper prototype  [ ] Clickable prototype  [ ] Concierge MVP
[ ] Wizard of Oz     [ ] Landing page test    [ ] A/B test
[ ] Other: _______________

AUDIENCE
Target persona: _______________
Sample size: _______________
Recruitment method: _______________

DESIGN
What we will build/prepare: _______________
What the participant will do: _______________
What we will observe/measure: _______________

SUCCESS CRITERIA
Primary metric: _______________
Success threshold: _______________
Failure threshold: _______________

TIME BOX
Build time: _______________
Run time: _______________
Analysis time: _______________
Total: _______________

RESULTS (fill after experiment)
Primary metric result: _______________
Qualitative observations: _______________
Surprises: _______________
Decision: [ ] Validate  [ ] Iterate  [ ] Pivot  [ ] Kill
Next step: _______________

Minimum Viable Tests

A minimum viable test (MVT) is the simplest possible experiment that can answer a specific question. The goal is to learn before you build, not to test what you have already built.

MVT Examples by Question
Question Minimum Viable Test Time Cost
"Do people understand our value proposition?" Show landing page to 5 people, ask them to explain it back 2 hours Free
"Will users find this navigation intuitive?" Card sort with 10 users using index cards 3 hours Free
"Is this onboarding flow clear?" Clickable prototype with 5 users 2 days Free
"Do users prefer layout A or B?" First-click test on UsabilityHub 4 hours $50-100
"Will users pay for this feature?" Add pricing page with "buy" button that leads to waitlist 1 day $50 for ads
"Is this workflow faster than the current one?" Time-on-task comparison: 5 users on old flow, 5 on prototype 1 day Free
The 5-User Rule

Jakob Nielsen's research shows that 5 users uncover approximately 85% of usability problems. For Lean UX experiments:

  • 5 users for qualitative usability tests (prototype tests, task analysis)
  • 20+ users for quantitative surveys or preference tests
  • 1,000+ users per variant for statistically significant A/B tests

Do not over-recruit for qualitative tests. Five users, tested quickly, are better than 50 users tested slowly.

Running Experiments in Practice

Weekly Experiment Cadence

A mature Lean UX team runs experiments every week. Here is a sample cadence:

Day Activity
Monday Review last week's results. Write new hypotheses. Design this week's experiment.
Tuesday Build experiment artifact (prototype, landing page, test script).
Wednesday Recruit participants (or launch ad traffic for smoke tests).
Thursday Run experiment sessions (usability tests, interviews).
Friday Synthesize results. Update hypothesis log. Plan next experiment.
Remote Experiment Tips
  • Use screen-sharing tools (Zoom, Lookback) for moderated prototype tests
  • Unmoderated tools (Maze, UserTesting) scale to more participants but lose qualitative depth
  • Record sessions (with consent) so the full team can watch asynchronously
  • Use virtual whiteboards (FigJam, Miro) for collaborative synthesis
Common Experiment Failures
Failure Cause Prevention
Leading questions Facilitator hints at the "right" answer Use neutral prompts: "What would you do next?" not "Would you click here?"
Confirmation bias Team sees only evidence that supports their idea Assign a devil's advocate; review raw data before discussing
Too few participants Results are unreliable Minimum 5 for qualitative, 1,000+ per variant for A/B
No success criteria Any result is interpreted as success Define thresholds before running the experiment
Testing too late Feature is already built; team is reluctant to change Test early with low-fidelity artifacts; never skip the prototype stage
Wrong audience Testing with colleagues instead of real users Recruit external participants matching the target persona

Experiment Cheat Sheet

Quick reference for choosing and running experiments:

If you need to learn... Use this experiment Minimum time Participants
Does the concept resonate? Landing page smoke test 2 days 200+ visitors
Does the flow work? Paper or clickable prototype 1-2 days 5 users
Is the solution valuable? Concierge MVP 1-2 weeks 5-10 users
Does the "smart" feature feel right? Wizard of Oz 1-2 weeks 5-10 users
Which design wins? A/B test 2-4 weeks 1,000+ per variant
What do users really need? Customer interview 1 day 5-8 users
How do users organize information? Card sort 3 hours 10-15 users
What do users notice first? First-click or 5-second test 4 hours 20+ users
1# Experiment Patterns for Lean UX
2 
3Experiments are the engine of learning in Lean UX. The right experiment answers the hypothesis with the least effort. Choosing the wrong experiment type wastes time, money, or both. This reference covers the full spectrum of UX experiments, from napkin sketches to coded A/B tests.
4 
5 
6## Table of Contents
71. [Types of UX Experiments](#types-of-ux-experiments)
82. [Choosing the Right Experiment](#choosing-the-right-experiment)
93. [Experiment Design Template](#experiment-design-template)
104. [Minimum Viable Tests](#minimum-viable-tests)
115. [Running Experiments in Practice](#running-experiments-in-practice)
126. [Experiment Cheat Sheet](#experiment-cheat-sheet)
13 
14---
15 
16## Types of UX Experiments
17 
18### 1. Paper Prototypes
19 
20**What it is:** Hand-drawn screens on paper or index cards. A facilitator plays "computer," swapping screens as the user taps or points.
21 
22**Best for:** Early concept validation, flow testing, information architecture.
23 
24**Effort:** Very low (30 minutes to create).
25**Fidelity:** Very low.
26**Confidence:** Low-medium. Validates flow and concept, not visual design or interaction details.
27 
28**When to use:**
29- You have multiple competing concepts and need to narrow down
30- The hypothesis is about flow or content, not aesthetics
31- You need to test today, not next week
32 
33**When NOT to use:**
34- The hypothesis depends on visual design, animation, or micro-interactions
35- Users need to interact with real data
36- Stakeholders will not trust low-fidelity evidence
37 
38**How to run:**
391. Sketch each screen on a separate sheet or card
402. Write a realistic task scenario for the participant
413. Ask the participant to "tap" or point at what they would interact with
424. Swap screens manually based on their choices
435. Note where they hesitate, get confused, or go off-script
44 
45### 2. Clickable Prototypes
46 
47**What it is:** Interactive mockups built in tools like Figma, Sketch, or InVision. Users click through a realistic-looking interface, but no backend logic exists.
48 
49**Best for:** Usability testing, flow validation, stakeholder buy-in, developer communication.
50 
51**Effort:** Medium (1-3 days).
52**Fidelity:** Medium-high.
53**Confidence:** Medium-high. Validates flow, layout, and basic usability.
54 
55**When to use:**
56- The hypothesis involves user navigation or task completion
57- You need to test with users who expect a realistic experience
58- The prototype will also serve as a design reference for developers
59 
60**When NOT to use:**
61- A paper prototype would suffice (over-investing)
62- The hypothesis is about performance, load times, or real data behavior
63- You need to test with hundreds of users (use coded experiments instead)
64 
65**How to run:**
661. Build the key screens and link hotspots in your prototyping tool
672. Write 3-5 task scenarios
683. Recruit 5-8 participants matching your persona
694. Run moderated sessions (20-30 minutes each)
705. Track task completion, time on task, errors, and qualitative feedback
71 
72### 3. Concierge MVP
73 
74**What it is:** Deliver the service or experience manually, person-to-person, without building any technology. The user receives the full value, but the backend is entirely human-powered.
75 
76**Best for:** Validating that the solution genuinely solves the problem before investing in automation.
77 
78**Effort:** Medium (ongoing manual work per user).
79**Fidelity:** High (the experience is real).
80**Confidence:** High. Real behavior with real value delivery.
81 
82**When to use:**
83- You are unsure whether the solution concept works at all
84- The cost of building the automated version is high
85- You want to deeply understand the user's experience and edge cases
86 
87**When NOT to use:**
88- The hypothesis is about scale or technology performance
89- You need to test with more than 10-20 users simultaneously
90- The value proposition depends on speed that only automation can provide
91 
92**Example:** A meal-planning app manually emails personalized weekly meal plans and shopping lists to 10 users based on their dietary preferences, before building the algorithm.
93 
94### 4. Wizard of Oz
95 
96**What it is:** The user interacts with what appears to be a functioning product, but a human behind the scenes is performing the work the technology would eventually do.
97 
98**Best for:** Testing the user experience of an automated feature before building the automation.
99 
100**Effort:** Medium (build the frontend; humans operate the backend).
101**Fidelity:** High from the user's perspective.
102**Confidence:** High. Users interact with what feels like a real product.
103 
104**When to use:**
105- The hypothesis depends on the user experience of an AI, algorithm, or automation feature
106- Building the actual technology is expensive or risky
107- You want to learn what the "right" output looks like before training a model
108 
109**When NOT to use:**
110- The hypothesis is about system performance or response time
111- Manual operation cannot replicate the technology's speed
112- Ethical issues arise from deception (always disclose if legally required)
113 
114**Example:** A "smart" scheduling assistant that appears to use AI but is actually a team member reading requests and sending calendar invites manually.
115 
116### 5. Landing Page / Smoke Test
117 
118**What it is:** A single web page describing a product or feature that does not yet exist, with a call to action (sign up, pre-order, request access). Measures demand by tracking how many people take the action.
119 
120**Best for:** Demand validation before building anything.
121 
122**Effort:** Low (half a day to create).
123**Fidelity:** Low (no product).
124**Confidence:** Medium. Measures stated intent, not actual usage.
125 
126**When to use:**
127- You need to validate demand before committing development resources
128- The hypothesis is about whether people want this at all
129- You want to build an early-access list for future testing
130 
131**When NOT to use:**
132- You already know there is demand and need to validate usability
133- The product concept is hard to explain without a demo
134- Your audience is internal (use interviews instead)
135 
136**How to run:**
1371. Create a landing page with a clear value proposition, 2-3 key benefits, and a CTA
1382. Drive targeted traffic (ads, social media, email, communities)
1393. Measure conversion rate (visitors to CTA clicks or sign-ups)
1404. Set success threshold before launch (e.g., 5% sign-up rate from 500 visitors)
1415. Follow up with sign-ups for qualitative interviews
142 
143### 6. A/B Test (Coded Experiment)
144 
145**What it is:** Two or more versions of a live feature are shown to different user segments. Statistical analysis determines which version performs better on a target metric.
146 
147**Best for:** Optimizing existing features, validating specific design changes with statistical rigor.
148 
149**Effort:** High (requires code, traffic, and statistical analysis).
150**Fidelity:** Production-level.
151**Confidence:** Very high (if properly powered).
152 
153**When to use:**
154- You have sufficient traffic to reach statistical significance
155- The hypothesis involves a measurable behavior change in an existing product
156- You need high-confidence evidence to justify a significant investment
157 
158**When NOT to use:**
159- Traffic is too low for statistical significance (fewer than 1,000 users per variant)
160- The concept is entirely new (test with prototypes first)
161- The change is too small to produce a detectable effect
162 
163## Choosing the Right Experiment
164 
165The decision depends on three factors: what you need to learn, how much confidence you need, and how much you can invest.
166 
167### Experiment Selection Matrix
168 
169| Question to Answer | Best Experiment | Fidelity | Time | Confidence |
170|-------------------|-----------------|----------|------|------------|
171| "Does anyone want this?" | Landing page smoke test | Low | 1-2 days | Medium |
172| "Does the flow make sense?" | Paper prototype | Very low | 1 day | Low-Medium |
173| "Can users complete this task?" | Clickable prototype | Medium | 3-5 days | Medium-High |
174| "Does this solution actually work?" | Concierge MVP | High | 1-2 weeks | High |
175| "Will the automated version feel right?" | Wizard of Oz | High | 1-2 weeks | High |
176| "Which version performs better?" | A/B test | Production | 2-4 weeks | Very High |
177 
178### The Fidelity Ladder
179 
180Start at the lowest rung that can answer your question. Only climb higher when lower fidelity cannot provide the needed confidence.
181 
182```
183Level 1: Paper prototype / Sketches
184 ↓ (if concept validated, test usability)
185Level 2: Clickable prototype (Figma, Sketch)
186 ↓ (if usability validated, test real value)
187Level 3: Concierge MVP / Wizard of Oz
188 ↓ (if value validated, test at scale)
189Level 4: Coded experiment / A/B test
190 ↓ (if optimized, ship)
191Level 5: Production release
192```
193 
194## Experiment Design Template
195 
196Use this template for every experiment, regardless of type:
197 
198```
199EXPERIMENT DESIGN
200=================
201Date: _______________
202Hypothesis ID: _______________
203Experimenter: _______________
204 
205HYPOTHESIS
206We believe _______________
207will happen if _______________
208achieves _______________
209with _______________.
210 
211EXPERIMENT TYPE
212[ ] Paper prototype [ ] Clickable prototype [ ] Concierge MVP
213[ ] Wizard of Oz [ ] Landing page test [ ] A/B test
214[ ] Other: _______________
215 
216AUDIENCE
217Target persona: _______________
218Sample size: _______________
219Recruitment method: _______________
220 
221DESIGN
222What we will build/prepare: _______________
223What the participant will do: _______________
224What we will observe/measure: _______________
225 
226SUCCESS CRITERIA
227Primary metric: _______________
228Success threshold: _______________
229Failure threshold: _______________
230 
231TIME BOX
232Build time: _______________
233Run time: _______________
234Analysis time: _______________
235Total: _______________
236 
237RESULTS (fill after experiment)
238Primary metric result: _______________
239Qualitative observations: _______________
240Surprises: _______________
241Decision: [ ] Validate [ ] Iterate [ ] Pivot [ ] Kill
242Next step: _______________
243```
244 
245## Minimum Viable Tests
246 
247A minimum viable test (MVT) is the simplest possible experiment that can answer a specific question. The goal is to learn before you build, not to test what you have already built.
248 
249### MVT Examples by Question
250 
251| Question | Minimum Viable Test | Time | Cost |
252|----------|-------------------|------|------|
253| "Do people understand our value proposition?" | Show landing page to 5 people, ask them to explain it back | 2 hours | Free |
254| "Will users find this navigation intuitive?" | Card sort with 10 users using index cards | 3 hours | Free |
255| "Is this onboarding flow clear?" | Clickable prototype with 5 users | 2 days | Free |
256| "Do users prefer layout A or B?" | First-click test on UsabilityHub | 4 hours | $50-100 |
257| "Will users pay for this feature?" | Add pricing page with "buy" button that leads to waitlist | 1 day | $50 for ads |
258| "Is this workflow faster than the current one?" | Time-on-task comparison: 5 users on old flow, 5 on prototype | 1 day | Free |
259 
260### The 5-User Rule
261 
262Jakob Nielsen's research shows that 5 users uncover approximately 85% of usability problems. For Lean UX experiments:
263 
264- **5 users** for qualitative usability tests (prototype tests, task analysis)
265- **20+ users** for quantitative surveys or preference tests
266- **1,000+ users per variant** for statistically significant A/B tests
267 
268Do not over-recruit for qualitative tests. Five users, tested quickly, are better than 50 users tested slowly.
269 
270## Running Experiments in Practice
271 
272### Weekly Experiment Cadence
273 
274A mature Lean UX team runs experiments every week. Here is a sample cadence:
275 
276| Day | Activity |
277|-----|----------|
278| **Monday** | Review last week's results. Write new hypotheses. Design this week's experiment. |
279| **Tuesday** | Build experiment artifact (prototype, landing page, test script). |
280| **Wednesday** | Recruit participants (or launch ad traffic for smoke tests). |
281| **Thursday** | Run experiment sessions (usability tests, interviews). |
282| **Friday** | Synthesize results. Update hypothesis log. Plan next experiment. |
283 
284### Remote Experiment Tips
285 
286- Use screen-sharing tools (Zoom, Lookback) for moderated prototype tests
287- Unmoderated tools (Maze, UserTesting) scale to more participants but lose qualitative depth
288- Record sessions (with consent) so the full team can watch asynchronously
289- Use virtual whiteboards (FigJam, Miro) for collaborative synthesis
290 
291### Common Experiment Failures
292 
293| Failure | Cause | Prevention |
294|---------|-------|------------|
295| Leading questions | Facilitator hints at the "right" answer | Use neutral prompts: "What would you do next?" not "Would you click here?" |
296| Confirmation bias | Team sees only evidence that supports their idea | Assign a devil's advocate; review raw data before discussing |
297| Too few participants | Results are unreliable | Minimum 5 for qualitative, 1,000+ per variant for A/B |
298| No success criteria | Any result is interpreted as success | Define thresholds before running the experiment |
299| Testing too late | Feature is already built; team is reluctant to change | Test early with low-fidelity artifacts; never skip the prototype stage |
300| Wrong audience | Testing with colleagues instead of real users | Recruit external participants matching the target persona |
301 
302## Experiment Cheat Sheet
303 
304Quick reference for choosing and running experiments:
305 
306| If you need to learn... | Use this experiment | Minimum time | Participants |
307|------------------------|-------------------|-------------|-------------|
308| Does the concept resonate? | Landing page smoke test | 2 days | 200+ visitors |
309| Does the flow work? | Paper or clickable prototype | 1-2 days | 5 users |
310| Is the solution valuable? | Concierge MVP | 1-2 weeks | 5-10 users |
311| Does the "smart" feature feel right? | Wizard of Oz | 1-2 weeks | 5-10 users |
312| Which design wins? | A/B test | 2-4 weeks | 1,000+ per variant |
313| What do users really need? | Customer interview | 1 day | 5-8 users |
314| How do users organize information? | Card sort | 3 hours | 10-15 users |
315| What do users notice first? | First-click or 5-second test | 4 hours | 20+ users |
316 

Discussion

Alternatives

Make UI/UX better of an already Created ApplicationGenerate a comprehensive, actionable development plan to enhance the existing web application.Coding · CC0-1.0UX researcherUse this agent when you need to design user research plans, analyze user-supplied data (interview transcripts, analytics, notes/transcripts from session recordings you provide), and generate actionable insights to validate design decisions and uncover user needs. Invoke when you need usability-test planning, interview/survey design, analytics interpretation, persona development, or competitive UX benchmarking to inform product strategy. Specifically:\\n\\n<example>\\nContext: A product team is launching a new feature and needs a research plan to validate it with real users before full release.\\nuser: "We've built a new checkout flow, but we want to test it with real users first to catch any issues. Can you help us plan usability testing?"\\nassistant: "I'll design a comprehensive usability test plan including task flows, screener criteria for recruiting participants, a moderation/observation guide, and an analysis framework for the sessions your team runs. Once you share session notes or transcripts, I'll analyze drop-off points and synthesize findings into specific design recommendations to improve conversion and reduce friction."\\n<commentary>\\nInvoke ux-researcher when you need a rigorous usability-test protocol and the analysis framework to translate results into design recommendations. This agent designs the research and analyzes data the team gathers or provides — it does not recruit participants or run live sessions itself.\\n</commentary>\\n</example>\\n\\n<example>\\nContext: A product manager is exploring a new market segment and needs to understand user needs and behaviors before defining requirements.\\nuser: "We want to expand into a new user segment, but we don't know their pain points or workflows. How can we understand what they need?"\\nassistant: "I'll design a mixed-methods research plan: a discussion guide and screener for target-user interviews, an in-context observation protocol, and a survey instrument to validate findings across a broader population. Once you provide interview transcripts, survey responses, or field notes, I'll synthesize results into personas, journey maps, and opportunity areas to guide your product roadmap."\\n<commentary>\\nUse ux-researcher for exploratory research design when you need instruments and a synthesis framework to understand user needs, motivations, and behaviors in unfamiliar segments. The agent designs the protocols and analyzes the data you supply.\\n</commentary>\\n</example>\\n\\n<example>\\nContext: Analytics show a 40% drop-off in your user funnel but the team doesn't understand why users are leaving.\\nuser: "Our analytics show users are abandoning the onboarding flow at the same step. What's causing this and how do we fix it?"\\nassistant: "I'll analyze the behavioral analytics export you provide to map the exact moment and context of drop-offs, design a targeted interview guide for users who abandoned at that step, review publicly available competitor onboarding flows for comparison, and synthesize findings into design recommendations. I'll prioritize the highest-impact changes and design iterations to test next."\\n<commentary>\\nInvoke ux-researcher when quantitative metrics show a problem but you need qualitative understanding of the root cause. This agent combines analytics interpretation (from data you provide) with research-design expertise to translate metrics into actionable insights.\\n</commentary>\\n</example>Design & UI · MITUX researcher designerUX research and design toolkit for Senior UX Designer/Researcher including data-driven persona generation, journey mapping, usability testing frameworks, and research synthesis. Use for user research, persona creation, journey mapping, and design validation.Design & UI · MITCs UX researcherUX research agent for research planning, persona generation, journey mapping, and usability test analysis. Use when product decisions need user evidence — e.g., planning interview scripts and recruiting criteria for a discovery study, or synthesizing usability-test sessions into prioritized findings and updated personas.Design & UI · MIT