Usability testing

Plan and run usability tests on existing or prototype designs including test design, task scripts, moderation, observation, and findings synthesis.

Usability testing — Creative Direction skill highlight diagram. Navy header card reads 'Impactful Creative Direction' with the subtitle… (from the rampstackco/claude-skills README)

From the rampstackco/claude-skills README — shows the whole collection, not only this skill. · view on GitHub

How to use it

Claude Code
  1. Run the line below. It pulls the whole folder into ~/.claude/skills/usability-testing.
  2. Describe your job in plain words. Claude Code follows the skill from there.
Claude Code — installs the whole folder, not just SKILL.md
npx degit rampstackco/claude-skills/skills/usability-testing#main ~/.claude/skills/usability-testing

For one project only, change the path to .claude/skills/usability-testing.

Claude (web or desktop app)
  1. On this page open ⋯ → Download .md.
  2. Save it as SKILL.md in a folder, zip the folder, then Customize → Skills → + → Create skill → Upload a skill.
  3. Pick the file and Save. Claude shows the name and description and runs a security scan.
  4. Check the skill is switched on.
  5. Start a new chat and describe your job in plain words. The AI follows the skill from there.
ChatGPT or another app
  1. ChatGPT: make a Project and paste it into Instructions.
  2. Neither? Paste it at the top of a new chat — it works for that chat.
Not working?
  • Check which app you pasted it into — the steps above name the right one.
  • Some skills need the paid tier of Claude or ChatGPT.
Step-by-step guide with screenshots · Ask in the forum

Paste into Claude, ChatGPT or Cursor.

Source of Usability testing

Show the full text245 lines
namedescriptioncategorycatalog_summarydisplay_order
usability-testingPlan and run usability tests on existing or prototype designs including test design, task scripts, moderation, observation, and findings synthesis. Use this skill whenever the user wants to test usability, run a moderated test, run an unmoderated test, validate a design, find usability issues, or improve task completion. Triggers on usability test, usability testing, moderated test, unmoderated test, task script, think aloud, prototype testing, user testing, design validation, task completion. Also triggers when the user has built something and wants to know if real users can use it before shipping.researchTest design, moderation, findings reports2

Usability Testing

Plan and run tests that find usability problems before users hit them in production. Stack-agnostic. Tool-agnostic.

This skill is for testing existing designs or prototypes. For broader discovery research, use ux-research. For conversion testing in production, use cro-optimization.


When to use

  • Before launching a new flow or major redesign
  • After a redesign to verify it doesn't introduce new problems
  • When analytics show drop-off but you don't know why
  • When customer support tickets pattern around specific UI areas
  • Pre-launch user validation
  • Comparing two design directions

When NOT to use

  • Discovery / generative research (use ux-research)
  • Live conversion optimization (use cro-optimization)
  • Mapping the broader experience (use journey-mapping)
  • Pure quantitative measurement (use analytics-strategy)

Required inputs

  • The design or prototype to test (functional or near-functional)
  • Specific tasks users would do
  • The audience (who should be tested)
  • Testing infrastructure (moderated tool, unmoderated tool, in-person setup)

The framework: 5 phases

1. Define what to test

Don't test the whole product. Test specific tasks.

Task selection criteria:

  • The task represents a real user goal (not "click around and explore")
  • The task has a clear start and end
  • The task is achievable in 2 to 10 minutes
  • The task is one of: most common, most strategic, most problematic

Examples of testable tasks:

"You want to find a contractor near you who can install a fence. Show me how you'd do that on this site."

"You're a first-time visitor. You want to understand if this product fits your needs. Walk me through how you'd evaluate it."

"Your team needs a new tool to manage projects. Use this site to figure out which plan is right for a 12-person team."

Task framing rules:

  • State the user goal, not the system action ("find a place to stay" not "click the search button")
  • Provide context (why are you doing this?)
  • Don't reveal the path
  • Don't use product terminology in the task framing
2. Choose moderated or unmoderated

Moderated (live, with researcher):

  • Researcher observes and probes in real time
  • Best for early-stage prototypes, complex tasks, novel concepts
  • Higher cost, smaller sample (5 to 8 participants typical)
  • Catches surprises and probe deeper

Unmoderated (recorded, asynchronous):

  • Participant completes alone, often via tool (UserTesting, Maze, Lookback)
  • Best for stable designs, simple tasks, larger sample
  • Lower cost, larger sample (15 to 30 participants typical)
  • Catches patterns at scale, less depth per session

For most teams: moderated for early/critical decisions, unmoderated for ongoing validation.

3. Recruit

Target audience - not just convenience.

Recruit criteria:

  • Match real users (target audience, not just "anyone")
  • Mix of experience levels with the product (new and existing if applicable)
  • Mix of relevant device types (mobile, desktop, tablet if relevant)
  • Exclude friends, family, employees

Sample size:

  • Moderated: 5 to 8 participants (Nielsen's "5 users find 85% of usability issues" for the most common segment)
  • Unmoderated: 15 to 30 participants (more participants compensate for less probing)
  • Multi-segment testing: 5 to 8 per segment
4. Run the test

Pre-task setup:

  • Confirm recording works
  • Brief participant (purpose, anonymity, recording, "no wrong answers")
  • Get verbal consent
  • Have participant share screen if remote

Moderated session structure:

  1. Warm-up (2 to 3 min). Easy questions to put participant at ease.
  2. Pre-test questions (3 to 5 min). Background context, current behavior with similar products.
  3. Task 1 (5 to 10 min). Describe task. Have participant attempt while thinking aloud.
  4. Post-task questions (1 to 2 min). What was easy/hard? Anything confusing?
  5. Repeat for tasks 2, 3, 4 (typically 3 to 5 tasks per 60-minute session).
  6. Overall debrief (5 to 10 min). General reactions, comparisons to alternatives, anything else.
  7. Close (2 min).

Moderation principles:

  • Encourage think-aloud ("What's going through your mind?")
  • Don't help unless they're truly stuck (and even then, only after a long pause)
  • Don't lead ("Are you looking for the menu?" - bad)
  • Note where they hesitate, scroll, or backtrack
  • Note their language vs the product's language
  • Note emotional reactions

Anti-patterns:

  • Talking too much (researcher should talk maybe 20% of the time)
  • Defending the design when participants struggle
  • Helping prematurely
  • Asking participants to predict their future behavior
  • Treating participant suggestions as features ("Users want X" - test demand for X separately)
5. Synthesize and report

Patterns across participants are signal. Single-participant complaints are weaker (but worth investigating).

Synthesis steps:

  1. Issue inventory. Every issue observed, with which participant, which task, severity.
  2. Cluster. Issues that are the same root problem.
  3. Severity.
    • Critical: Blocks task completion. Most users hit this.
    • Major: Significantly slows task. Many users hit this.
    • Minor: Friction. Some users hit this. Workaround exists.
    • Cosmetic: Polish. Doesn't affect task.
  4. Recommendations. For each issue, propose specific fixes.
  5. Prioritize. By severity and effort.

Report structure:

# Usability Test: [Design / flow]

## Summary
[2 to 3 paragraphs covering: what was tested, headline findings, top 3 priorities]

## Method
[Moderated/unmoderated, sample size, audience, dates, tasks]

## Critical findings
[Each with description, frequency, supporting evidence (quotes/clips), recommendation. Where evidence was not obtained, state the gap per the data-availability rule]

## Major findings
[Same structure]

## Minor findings
[Brief]

## Cosmetic findings
[Briefest]

## What worked well
[Calibration: capture successes too]

## Recommendations
[Prioritized list with effort estimates]

## Next steps
[Test re-run schedule, design iteration plan]

Workflow

  1. Define the goals. What decisions hinge on this? What tasks matter most?
  2. Design tasks. 3 to 5 specific, realistic, goal-framed tasks.
  3. Choose moderated vs unmoderated. Match to stage and depth needed.
  4. Recruit. Specific to audience.
  5. Pilot. 1 to 2 sessions before main batch. Refine tasks if needed.
  6. Run. Follow the protocol. Stay disciplined.
  7. Synthesize during, not just after. Patterns emerge by session 4 or 5.
  8. Report. Multiple formats - written report + highlight clips.
  9. Track fixes. Every critical issue should have an owner and date.
  10. Re-test after fixes. Verify the fix worked, didn't introduce new issues.

Failure patterns

  • Testing the whole product instead of specific tasks. Vague results.
  • Tasks that reveal the path. ("Click the menu and find...")
  • Friends and family as participants. Biased, not representative.
  • Researcher leading the participant. Findings reflect the researcher.
  • Defending the design when participants struggle. Misses real issues.
  • Helping too quickly. Participant doesn't experience the friction.
  • Treating participant suggestions as features. Users solve their problem; product team designs the solution.
  • One participant = data point. A single strong opinion isn't a finding.
  • Skipping severity scoring. All findings treated equally; team can't prioritize.
  • Reports no one reads. Highlight clips and live walkthroughs work better than 80-page decks.
  • Testing once, never re-testing. Fixes that introduce new problems go undetected.

Output format

Default outputs:

  1. Test plan (before testing) - usability-test-plan-[topic].md
  2. Task script (per session) - usability-tasks-[topic].md
  3. Findings report (after synthesis) - usability-findings-[topic].md
  4. Highlight clips (separately produced)

If required data is unavailable

This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.


Reference files

1---
2name: usability-testing
3description: "Plan and run usability tests on existing or prototype designs including test design, task scripts, moderation, observation, and findings synthesis. Use this skill whenever the user wants to test usability, run a moderated test, run an unmoderated test, validate a design, find usability issues, or improve task completion. Triggers on usability test, usability testing, moderated test, unmoderated test, task script, think aloud, prototype testing, user testing, design validation, task completion. Also triggers when the user has built something and wants to know if real users can use it before shipping."
4category: research
5catalog_summary: "Test design, moderation, findings reports"
6display_order: 2
7---
8 
9# Usability Testing
10 
11Plan and run tests that find usability problems before users hit them in production. Stack-agnostic. Tool-agnostic.
12 
13This skill is for testing existing designs or prototypes. For broader discovery research, use `ux-research`. For conversion testing in production, use `cro-optimization`.
14 
15---
16 
17## When to use
18 
19- Before launching a new flow or major redesign
20- After a redesign to verify it doesn't introduce new problems
21- When analytics show drop-off but you don't know why
22- When customer support tickets pattern around specific UI areas
23- Pre-launch user validation
24- Comparing two design directions
25 
26## When NOT to use
27 
28- Discovery / generative research (use `ux-research`)
29- Live conversion optimization (use `cro-optimization`)
30- Mapping the broader experience (use `journey-mapping`)
31- Pure quantitative measurement (use `analytics-strategy`)
32 
33---
34 
35## Required inputs
36 
37- The design or prototype to test (functional or near-functional)
38- Specific tasks users would do
39- The audience (who should be tested)
40- Testing infrastructure (moderated tool, unmoderated tool, in-person setup)
41 
42---
43 
44## The framework: 5 phases
45 
46### 1. Define what to test
47 
48Don't test the whole product. Test specific tasks.
49 
50**Task selection criteria:**
51 
52- The task represents a real user goal (not "click around and explore")
53- The task has a clear start and end
54- The task is achievable in 2 to 10 minutes
55- The task is one of: most common, most strategic, most problematic
56 
57**Examples of testable tasks:**
58 
59> "You want to find a contractor near you who can install a fence. Show me how you'd do that on this site."
60 
61> "You're a first-time visitor. You want to understand if this product fits your needs. Walk me through how you'd evaluate it."
62 
63> "Your team needs a new tool to manage projects. Use this site to figure out which plan is right for a 12-person team."
64 
65**Task framing rules:**
66 
67- State the user goal, not the system action ("find a place to stay" not "click the search button")
68- Provide context (why are you doing this?)
69- Don't reveal the path
70- Don't use product terminology in the task framing
71 
72### 2. Choose moderated or unmoderated
73 
74**Moderated** (live, with researcher):
75 
76- Researcher observes and probes in real time
77- Best for early-stage prototypes, complex tasks, novel concepts
78- Higher cost, smaller sample (5 to 8 participants typical)
79- Catches surprises and probe deeper
80 
81**Unmoderated** (recorded, asynchronous):
82 
83- Participant completes alone, often via tool (UserTesting, Maze, Lookback)
84- Best for stable designs, simple tasks, larger sample
85- Lower cost, larger sample (15 to 30 participants typical)
86- Catches patterns at scale, less depth per session
87 
88For most teams: moderated for early/critical decisions, unmoderated for ongoing validation.
89 
90### 3. Recruit
91 
92Target audience - not just convenience.
93 
94**Recruit criteria:**
95 
96- Match real users (target audience, not just "anyone")
97- Mix of experience levels with the product (new and existing if applicable)
98- Mix of relevant device types (mobile, desktop, tablet if relevant)
99- Exclude friends, family, employees
100 
101**Sample size:**
102 
103- Moderated: 5 to 8 participants (Nielsen's "5 users find 85% of usability issues" for the most common segment)
104- Unmoderated: 15 to 30 participants (more participants compensate for less probing)
105- Multi-segment testing: 5 to 8 per segment
106 
107### 4. Run the test
108 
109**Pre-task setup:**
110 
111- Confirm recording works
112- Brief participant (purpose, anonymity, recording, "no wrong answers")
113- Get verbal consent
114- Have participant share screen if remote
115 
116**Moderated session structure:**
117 
1181. **Warm-up** (2 to 3 min). Easy questions to put participant at ease.
1192. **Pre-test questions** (3 to 5 min). Background context, current behavior with similar products.
1203. **Task 1** (5 to 10 min). Describe task. Have participant attempt while thinking aloud.
1214. **Post-task questions** (1 to 2 min). What was easy/hard? Anything confusing?
1225. **Repeat for tasks 2, 3, 4** (typically 3 to 5 tasks per 60-minute session).
1236. **Overall debrief** (5 to 10 min). General reactions, comparisons to alternatives, anything else.
1247. **Close** (2 min).
125 
126**Moderation principles:**
127 
128- Encourage think-aloud ("What's going through your mind?")
129- Don't help unless they're truly stuck (and even then, only after a long pause)
130- Don't lead ("Are you looking for the menu?" - bad)
131- Note where they hesitate, scroll, or backtrack
132- Note their language vs the product's language
133- Note emotional reactions
134 
135**Anti-patterns:**
136 
137- Talking too much (researcher should talk maybe 20% of the time)
138- Defending the design when participants struggle
139- Helping prematurely
140- Asking participants to predict their future behavior
141- Treating participant suggestions as features ("Users want X" - test demand for X separately)
142 
143### 5. Synthesize and report
144 
145Patterns across participants are signal. Single-participant complaints are weaker (but worth investigating).
146 
147**Synthesis steps:**
148 
1491. **Issue inventory.** Every issue observed, with which participant, which task, severity.
1502. **Cluster.** Issues that are the same root problem.
1513. **Severity.**
152 - **Critical:** Blocks task completion. Most users hit this.
153 - **Major:** Significantly slows task. Many users hit this.
154 - **Minor:** Friction. Some users hit this. Workaround exists.
155 - **Cosmetic:** Polish. Doesn't affect task.
1564. **Recommendations.** For each issue, propose specific fixes.
1575. **Prioritize.** By severity and effort.
158 
159**Report structure:**
160 
161```markdown
162# Usability Test: [Design / flow]
163 
164## Summary
165[2 to 3 paragraphs covering: what was tested, headline findings, top 3 priorities]
166 
167## Method
168[Moderated/unmoderated, sample size, audience, dates, tasks]
169 
170## Critical findings
171[Each with description, frequency, supporting evidence (quotes/clips), recommendation. Where evidence was not obtained, state the gap per the data-availability rule]
172 
173## Major findings
174[Same structure]
175 
176## Minor findings
177[Brief]
178 
179## Cosmetic findings
180[Briefest]
181 
182## What worked well
183[Calibration: capture successes too]
184 
185## Recommendations
186[Prioritized list with effort estimates]
187 
188## Next steps
189[Test re-run schedule, design iteration plan]
190```
191 
192---
193 
194## Workflow
195 
1961. **Define the goals.** What decisions hinge on this? What tasks matter most?
1972. **Design tasks.** 3 to 5 specific, realistic, goal-framed tasks.
1983. **Choose moderated vs unmoderated.** Match to stage and depth needed.
1994. **Recruit.** Specific to audience.
2005. **Pilot.** 1 to 2 sessions before main batch. Refine tasks if needed.
2016. **Run.** Follow the protocol. Stay disciplined.
2027. **Synthesize during, not just after.** Patterns emerge by session 4 or 5.
2038. **Report.** Multiple formats - written report + highlight clips.
2049. **Track fixes.** Every critical issue should have an owner and date.
20510. **Re-test after fixes.** Verify the fix worked, didn't introduce new issues.
206 
207---
208 
209## Failure patterns
210 
211- **Testing the whole product instead of specific tasks.** Vague results.
212- **Tasks that reveal the path.** ("Click the menu and find...")
213- **Friends and family as participants.** Biased, not representative.
214- **Researcher leading the participant.** Findings reflect the researcher.
215- **Defending the design when participants struggle.** Misses real issues.
216- **Helping too quickly.** Participant doesn't experience the friction.
217- **Treating participant suggestions as features.** Users solve their problem; product team designs the solution.
218- **One participant = data point.** A single strong opinion isn't a finding.
219- **Skipping severity scoring.** All findings treated equally; team can't prioritize.
220- **Reports no one reads.** Highlight clips and live walkthroughs work better than 80-page decks.
221- **Testing once, never re-testing.** Fixes that introduce new problems go undetected.
222 
223---
224 
225## Output format
226 
227Default outputs:
228 
2291. **Test plan** (before testing) - `usability-test-plan-[topic].md`
2302. **Task script** (per session) - `usability-tasks-[topic].md`
2313. **Findings report** (after synthesis) - `usability-findings-[topic].md`
2324. **Highlight clips** (separately produced)
233 
234---
235 
236## If required data is unavailable
237 
238This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.
239 
240---
241 
242## Reference files
243 
244- [`references/task-script-patterns.md`](references/task-script-patterns.md) - Task framing patterns by common product type, with good and bad examples.
245 

Discussion

Alternatives

Also in User flowsSee all 106 in Design →
Make UI/UX better of an already Created ApplicationGenerate a comprehensive, actionable development plan to enhance the existing web application.Coding · CC0-1.0UX Researcher & DesignerUX research and design toolkit for Senior UX Designer/Researcher including data-driven persona generation, journey mapping, usability testing frameworks, and research synthesis. Use when conducting user research, creating personas, mapping user journeys, planning usability tests, or validating designs.Design & UI · MITUI UX testerUse this agent when you need exhaustive UI and UX functionality testing driven by documented user flows, with browser or desktop interaction tooling and structured defect reporting.Design & UI · MITImprove an AppGuided journey from a shipped app that works but feels rough to a product that fits the job, flows without friction, reads clearly, and persuades honestly. Orchestrates nine skills phase by phase - jobs-to-be-done, ux-heuristics, design-everyday-things, refactoring-ui, microinteractions, made-to-stick, influence-psychology, high-perf-browser, steve-jobs-design-review - asking the user questions at every decision point and recording results in the project docs/ folder (CUSTOMER.md, DESIGN.md, POSITIONING.md, IMPROVE-APP-PLAN.md) so the journey resumes across sessions. Use when the user wants to fix a clunky product, cut UX friction, sharpen in-app copy, or says ''the app works but feels rough''. Not for code, tests, or hardening - use improve-code-quality (fresh prototype) or remove-technical-debt (aged); no app yet, create-app; needs growth loops, grow-app; marketing-site friction, improve-website; one leaking flow, conversion-optimization. For one framework in isolation, invoke that skill directly.Design & UI · MIT