usability-testing — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited usability-testing (Agent Skill) and scored it 92/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 2 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 2 flagged
The text {match} tells the agent to skip the normal "ask the user first" gate. Used adversarially it removes the human-in-the-loop check before destructive or sensitive actions, turning a normally-gated agent into a fire-and-forget executor.
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Usability testing is the evaluative research practice of watching representative real people attempt realistic tasks on a prototype or product to surface the obstacles they encounter — making it empirical observation rather than expert opinion. Its core instrument is the think-aloud protocol (Ericsson & Simon), where participants narrate their reasoning so the divergence between their mental model and the design becomes visible at the moment it happens. Sessions are built around task scenarios that state a goal without prescribing steps ("find out how much you owe in taxes this quarter"), run by a moderator whose discipline is neutrality: frame the goal, prompt only with open questions, let silence and struggle sit, and never rescue or defend the design.
The craft has four hard gates. First, the artifact must exist at a fidelity where a real person can attempt the task. Second, the participant must match the study's target audience closely enough that the result transfers. Third, the task must be goal-framed rather than path-framed: "find out how much you owe this quarter," not "click Tax Summary, then View Details." Fourth, the claim type must match the study design: small formative rounds are excellent for discovering and diagnosing problems; summative benchmark claims require larger samples, metrics, and statistical justification.
Sample size is governed by purpose — the Nielsen/Landauer 5-user rule (~85% of major problems per homogeneous segment, per discrete task) holds for formative, diagnostic, iterative testing, but it is a long-run average with real per-round variance (Faulkner 2003), and summative benchmark claims require much larger statistical samples; conflating the two is a classic, invalidating error. Findings are triaged by Nielsen's 0–4 severity scale — itself a composite of frequency, impact, and persistence — complemented when needed by quantitative instruments (task success reported with a small-sample confidence interval, SEQ, SUS or UMUX-Lite). Current tooling — remote moderated sessions, unmoderated platforms, automatic transcripts, AI-assisted highlights, and LLM-agent simulations — has changed the logistics, not the epistemology: it can reduce operational cost, but it does not replace observing real humans. (For the standard practitioner reference behind these mechanics, see Rubin & Chisnell, Handbook of Usability Testing, 2008.)
Usability testing covers session planning, participant selection, task design, moderation or unmoderated setup, observation, severity triage, and claim-appropriate reporting.
| Purpose | Strong default | What it can claim | Watch out for |
|---|---|---|---|
| Find major usability problems early | Moderated formative test, often ~5 participants per homogeneous segment/task | Diagnostic findings and redesign priorities | Does not prove success rate, conversion lift, or statistical significance |
| Quickly check a simple, stable flow at scale | Remote unmoderated test with a pilot run | Behavioral recordings, task-success patterns, frequent friction | Weak for early brittle prototypes, emotional work, or unexpected follow-up |
| Evaluate IA, navigation labels, or first-path cues | Tree testing for hierarchy/findability; first-click testing for static-screen entry points | Directness, destination choice, first-click distribution, success patterns | Does not explain full interactive behavior; escalate to moderated testing when follow-up behavior matters |
| Compare designs or benchmark performance | Quantitative/summative usability study | Metrics such as success rate, time on task, SUS/SEQ, with confidence intervals | Usually needs many more participants, often dozens (commonly ~40), not five |
| Explore accessibility barriers with disabled users | Usability evaluation with relevant disabled participants and assistive technologies | Lived barriers and task evidence for the included user groups | Must be combined with WCAG conformance evaluation; it is not itself conformance |
| Prepare research with AI/synthetic users | Synthetic prompts, study-design dry runs, AI-assisted transcript review of real sessions | Hypotheses and analysis aids | Never report simulated users as real participants or final validation |
Sample size is only half the question; who you recruit decides whether the findings transfer to your real users. Participants must be representative of the actual user population for the task — recruit from the segments who will really use the product, not whoever is convenient (colleagues, friends, or other designers carry insider mental models and invalidate the test). Write a screener: a short set of qualifying questions on the behaviors, domain familiarity, and tool experience that define each segment, plus disqualifiers (e.g. people who work in UX or for a competitor). When the product genuinely serves distinct segments (first-time vs. power users, administrators vs. end users, mobile-only vs. desktop, assistive-technology users, or different roles), treat each as a separate cohort and apply the 5-user heuristic per segment — one mixed group of five does not cover two segments. Document the realized participant profile, and who was excluded and why, so a reader can judge transferability rather than assuming it.
Task scenarios state a goal with enough context for the participant to care, while withholding the path the design team hopes they will take ("you want to find out how much you owe in taxes this quarter"). Good scenarios avoid UI labels, feature names, insider language, jokes, offensive or sensitive assumptions, and hypothetical "how would you…" wording. "Find a way to receive upcoming events by email" is stronger than "click Sign up for newsletter"; "find the symptoms of the flu" is stronger than "how would you find the symptoms?" — because the point is observed behavior, not a verbal plan. A task whose wording names the destination has already given away the finding you came to observe.
The classic Ericsson & Simon protocol is concurrent think-aloud (CTA): the participant narrates while performing the task, and the moderator says almost nothing beyond "keep talking." CTA captures reasoning at the moment of struggle but introduces reactivity — narrating can slow users down or change how they work, and some participants fall silent exactly when a task gets hard (the moment you most want to hear). Retrospective think-aloud (RTA) addresses this: the participant works in silence, then narrates while reviewing a recording. RTA removes reactivity and often yields fuller explanations, but it relies on memory, can introduce post-hoc rationalization, and roughly doubles session time. Hertzum's 2024 meta-analytic review (ACM TOCHI) found CTA and RTA surface comparable sets of problems but reveal them differently — choose CTA by default, RTA when the task is too cognitively demanding to narrate live.
Moderator neutrality is active discipline. The strict Ericsson & Simon model forbids any acknowledgment (no "mm-hmm," no nodding) to avoid contaminating the cognitive trace. In practice, Boren & Ramey (2000) showed that real usability moderation works better as speech communication — the participant is the primary speaker and the moderator is a listener who keeps the channel open with minimal, non-leading backchannel tokens ("mm-hmm," "keep going," "what are you thinking now?") rather than stony silence, which keeps participants talking without steering them toward an answer. The moderator frames the session, gives tasks one at a time, clarifies the task only when needed, and lets silence sit. They do not teach the interface, defend the design, explain what was intended, praise a path, or rescue a participant as soon as friction appears. Both models agree on the hard rule: if a participant gets stuck, the stuckness is the finding — never explain, defend, or rescue.
Moderated testing gives richer data because a researcher can notice surprise, ask neutral follow-up, and recover when a prototype glitch or participant misunderstanding would otherwise derail the session; it costs more scheduling effort. Unmoderated testing scales faster because the platform administers tasks asynchronously and records screen, voice, clicks, transcripts, and sometimes clips; it requires more meticulous setup because no human can adapt mid-session. Pick moderated for early, complex, or exploratory flows where probing matters; pick unmoderated for fast validation of a simple, well-scoped flow at volume. A common mature stack pairs one of each. Both are now predominantly remote (moderated over video with screen sharing; unmoderated async on panel-recruiting platforms), though in-person still matters when the environment, device setup, physical product, assistive technology, privacy, or context is part of the behavior being tested. For obvious blockers, the RITE method (Rapid Iterative Testing and Evaluation) fixes a confirmed problem immediately after the participant who hit it, then tests the fix with the next participant — trading a clean fixed-sample comparison for faster convergence on the worst issues.
Beyond task-based observation, two lightweight quantitative methods answer narrower findability questions and are usually run unmoderated at scale. First-click testing measures where users click first to start a task — first-click success correlates strongly with overall task success, so it is a fast signal for whether an entry point or label is discoverable; its key metrics are first-click success and time to first click. Tree testing (reverse card sorting) evaluates an information hierarchy in the abstract, stripped of visual design, by asking users where they would look to find an item — isolating navigation-label and structure problems from layout problems; its key metrics are success and directness (reaching the target without backtracking). Treat their claims narrowly: they can show that people pick the wrong label or miss the expected entry point, but they do not test the full interactive sequence, state changes, error recovery, or emotional reaction. Both are evaluative usability methods; when the question is designing the hierarchy or labels rather than evaluating an attempt, that is adjacent information-architecture work, not this skill.
The widely-cited Nielsen/Landauer "5-user rule" estimates that 5 users surface ~85% of major usability problems for a homogeneous user group on a discrete task, with steeply diminishing returns afterward. Read it as "small repeated rounds often have the best learning-to-cost ratio," and treat the 85% as a long-run average, not a guarantee for any one round. Faulkner's 2003 study re-sampled sets from a 60-user pool and found individual sets of 5 caught anywhere from 55% to 99% of known problems; raising the sample to 10 lifted the worst-case floor to ~80%, and 20 users to ~95%. The practical reading: 5 is a defensible minimum for a fast formative round, but if a round can only be run once or the cost of missing a major problem is high, 8–12 users buys a meaningfully tighter floor. For problem discovery the real question is what problem probability you need to detect, and with what discovery confidence — rarer problems require more participants.
The rule has further limits: it applies per distinct user segment, per discrete task scope, and to formative (iterative diagnostic) testing — not to summative (benchmark) studies. Summative or quantitative claims need a separate sample-size rationale; many quantitative UX studies need dozens of participants (commonly ~40 under specific assumptions). Do not present "3 of 5" as proof of a population rate.
Findings are organized by a severity rating (Nielsen's 0–4 scale: 0 = not a problem, 1 = cosmetic, 2 = minor, 3 = major, 4 = catastrophic) so the team can triage under time pressure. Severity is not a single gut impression — Nielsen defines it as a composite of three factors: frequency (common or rare?), impact (easy or hard for the user to overcome?), and persistence (a one-time hurdle once learned, or a repeated annoyance?). Rate the factors and combine them, rather than assigning a number by feel; a rare but catastrophic-impact, persistent problem (e.g. a blocker that stops checkout) outranks a frequent cosmetic one. Each finding should carry enough evidence to be falsifiable: task, participant segment, observed behavior, quote or timestamp when available, expected path or mental model, impact, severity, and recommended next step.
Qualitative observation is the primary signal, but numbers help when a stakeholder needs one. Task success rate (binary or with partial-success levels) is the most direct behavioral metric — but with the small samples usability tests use, report it as a confidence interval, not a bare percentage: "3 of 5 succeeded" is 60% with a 95% interval roughly spanning 23%–88% (adjusted-Wald method, Sauro & Lewis 2005, which the standard Wald grossly understates at small n), so the point estimate alone overstates certainty. Time on task and error counts measure efficiency. For subjective ease, the Single Ease Question (SEQ) — a 7-point post-task difficulty rating (Sauro & Dumas 2009) — is fast and well-validated; for whole-product perceived usability, the System Usability Scale (SUS) (Brooke 1996) or the two-item UMUX-Lite (Lewis et al., which correlates strongly with SUS) work post-test; NASA-TLX fits complex workload-heavy tasks. Use each at the right scope (SEQ post-task, SUS/UMUX-Lite post-test). Quantitative instruments belong to summative or mixed studies; they do not rescue an underpowered formative round, and a number without the path, hesitation, misinterpretation, or workaround behind it is weak evidence for redesign.
Recruiting representative users includes recruiting people with disabilities who use assistive technologies (screen readers, switch access, screen magnification, voice control). The W3C WAI guidance ("Involving Users in Web Projects / Evaluation") is explicit on two points: usability testing with disabled users surfaces real barriers that automated checks and expert review miss, and it does not prove WCAG conformance — a clean session with three screen-reader users is empirical evidence of usability, not a conformance certificate, and does not generalize to all disabled users. Recruit for the relevant disability/AT profiles, set up the participant's technology, make facilities or remote tooling accessible, adapt facilitation style, and report scope precisely (e.g. "two NVDA users, one VoiceOver user, one switch-access user") rather than a vague "we tested with accessibility users," because the specific assistive technology determines which barriers are observable. This complements the a11y skill's conformance audit; it does not replace it (see Do NOT Use When).
Always run a pilot session (a dry run with one participant or a colleague) before the real rounds — it catches broken task wording, prototype dead-ends, login/technical failures, misleading follow-up questions, and timing problems while they are still cheap to fix; in an unmoderated study, where no human can adapt mid-session, an un-piloted broken task can corrupt an entire batch. Treat participants ethically: obtain informed consent before recording, explain the purpose and how the recording will be used and stored, make clear the participant is free to stop at any time, and that you are testing the design, not them. Protect the resulting data: de-identify transcripts and recordings where feasible and store sensitive participant data under access control. These are not optional courtesies — a recorded session without consent is an ethics (and often legal) failure.
The skill also covers session anti-patterns: leading prompts ("don't you think the menu should be here?"), defending the design, explaining how the design "is supposed to work" when the participant gets stuck, praising a particular path, and over-fitting interpretations to a single dramatic finding from one participant.
The upstream-displacement check is clear: a 2024–2025 wave of AI tooling has improved the logistics and analysis around usability testing, but it has not displaced the method. Three honest uses, and one hard line:
The hard line: synthetic users do not replace observing real users, and presenting their output as real-user findings is dishonest (Nielsen Norman Group is explicit on this). The mechanism is the one that justifies the whole skill: a model predicts the statistically likely path, while the value of a usability test is the surprising, off-path behavior a real human produces. Empirical checks bear this out — synthetic users show an over-optimistic completion bias (reporting success real users never achieve), and LLM agents in usability tasks (e.g. the UXAgent CHI 2025 work) follow neat, direct paths while real users wander, second-guess, and abandon. Label synthetic output as synthetic, treat it as a hypothesis, and verify it with real users before any product decision. When a request asks to "use synthetic users instead of real participants," this skill owns the answer — it is the explanation of why direct observation is irreplaceable.
Usability testing is built on a humbling claim: designers and engineers cannot reliably predict where users will struggle. The mental models that make a design feel obvious to its creators are exactly the models a fresh user lacks, and only direct observation closes that gap. The discipline rejects "I think users will understand this" in favor of "we watched users; here is what happened." Each session that confirms the design entirely is mildly suspicious — either the tasks were too easy or the moderator was unintentionally helping.
The practice is opinionated about moderator behavior. The moderator's job is to be uninteresting — to let the silence sit, to let the participant struggle long enough for the obstacle to become visible, to not rescue. This is hard because the social instinct is to help and the design instinct is to defend. A moderator who explains the design after a participant gets stuck has destroyed the evidence; the obstacle the participant just encountered is the finding, and it cannot be re-observed in that session.
The discipline is also anti-theater. A polished report with no contradictions, no behavioral evidence, no representative participants, and no painful findings is not a successful usability test — it is probably a leading script, a weak task set, a mismatched sample, or a moderator who rescued the design. Useful testing produces specific changes the team would not have made from opinion alone.
<!-- skill-graph-context:start (generated — do not edit by hand) -->
Classification
designtrueWhen to use
usability test, think aloud, test this prototype, task scenarios, test with users, synthetic users for usability, tree testing, first-click testingNot for
Related skills
a11yprototyping, user-research, research-synthesis, design-thinking, testing-strategy, a11y, information-architectureConcept
Grounding
universalhttps://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/, https://www.nngroup.com/articles/how-many-test-users/, https://measuringu.com/specific-sample-sizes-in-problem-discovery-studies/, https://www.nngroup.com/articles/summary-quant-sample-sizes/, https://link.springer.com/article/10.3758/BF03195514, https://measuringu.com/article/estimating-completion-rates-from-small-samples-using-binomial-confidence-intervals-comparisons-and-recommendations/, https://www.nngroup.com/articles/task-scenarios-usability-testing/, https://www.nngroup.com/articles/better-usability-tasks/, https://www.nngroup.com/articles/thinking-aloud-the-1-usability-tool/, https://doi.org/10.1109/47.867942, https://forskning.ruc.dk/en/publications/concurrent-or-retrospective-thinking-aloud-in-usability-tests-a-m/, https://www.nngroup.com/articles/navigation-ia-tests/, https://www.nngroup.com/articles/usability-checklist/, https://www.nngroup.com/articles/unmoderated-usability-testing/, https://www.nngroup.com/articles/unmoderated-user-testing-tools/, https://www.nngroup.com/articles/how-to-rate-the-severity-of-usability-problems/, https://www.nngroup.com/articles/measuring-perceived-usability/, https://www.nngroup.com/articles/success-rate-the-simplest-usability-metric/, https://measuringu.com/umux-lite/, https://www.nngroup.com/articles/screening-participants/, https://www.nngroup.com/articles/informed-consent/, https://www.w3.org/WAI/test-evaluate/involving-users/, https://www.nngroup.com/articles/synthetic-users/, https://help.usertesting.com/hc/en-us/articles/13268691111453-AI-insight-summary, https://www.amazon.science/publications/uxagent-an-llm-agent-based-usability-testing-framework-for-web-designKeywords
think aloud protocol, task scenario, moderated usability test, unmoderated test, severity rating, five user rule, formative testing, summative testing, tree testing, first-click testing<!-- skill-graph-context:end -->
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.