Usability testing with five users: what it catches and what it misses

By Meet Patel · 2026-10-03 · 5 min read

Summary

Nielsen Norman Group's 2000 analysis found one test user uncovers about 31% of usability problems, so five find roughly 85%. That holds for common problems in one user group. Rare problems, several groups and quantitative studies (20+ users) need more participants.

Key Metrics & Takeaways

31 percent
typical share of usability problems found by a single test user, averaged across projects Nielsen studied (Nielsen Norman Group, 2000)
about 85 percent
share of problems found by five users, using Nielsen's formula with L = 31 percent (Nielsen Norman Group, 2000)
20 users
minimum Nielsen Norman Group advises for statistically significant quantitative usability studies (2012)

Jakob Nielsen's argument for five test users rests on one number. In “Why You Only Need to Test with 5 Users” (Nielsen Norman Group, 18 March 2000), he reports that the typical share of usability problems found by a single test user is 31 percent, averaged across a large number of projects he studied. From that, the share found by n users is 1 minus (1 minus 0.31) to the power n. At five users the formula gives roughly 85 percent.

That result is real and useful, and it is also widely misread. Five users is a good size for a qualitative round on one user group, when the problems you are hunting are common. It is too few for rare problems, for several distinct groups and for any study that needs statistics. This post sets out what five users catch, what they miss, and how to run a round in a week.

What the 85 percent means

Using Nielsen's formula with his average of 31 percent, the cumulative share of problems found looks like this.

These are my calculations from his formula, and they show the shape of the curve. The first user delivers the most. In Nielsen's words, “As soon as you collect data from a single test user, your insights shoot up and you have already learned almost a third of all there is to know.” By the fifth user the sessions repeat what earlier users showed. He puts it bluntly: “After the fifth user, you are wasting your time by observing the same findings repeatedly but not learning much new.”

His practical advice follows from the curve. “The best results come from testing no more than 5 users and running as many small tests as you can afford.” Three rounds of five with a redesign between each is his preferred plan over one round of fifteen. My reading is that the fixes from round one clear the obstacles that hid later problems from round two.

What five users miss

Problems that affect a minority of people. The 31 percent figure is an average. A problem that affects only 1 user in 10 has a 41 percent chance of appearing among five participants, by the same formula. A problem hitting 1 in 5 has a 67 percent chance. To have an 85 percent chance of seeing a problem that affects 1 in 10, you need about 18 participants. Jeff Sauro makes the same point in his history of the magic number 5 (MeasuringU, 21 July 2010): rather than guess an average, choose the minimum problem frequency you want to detect and size the test from that.

Problems that surface in open-ended tasks. Sauro's history also summarizes Jared Spool and Will Schroeder's 2001 paper, “Testing Web Sites: Five Users Is Nowhere Near Enough.” Their participants were given open-ended tasks, browsing up to four websites for CDs they wanted, and serious problems were still being found after dozens of users. The lesson I draw is that the more freedom the task allows, the more distinct paths users take, and the more users it takes to see them.

Several user groups. In a 2012 update, Nielsen Norman Group advises representatives of each target group, with 3 to 4 users per group when behaviors overlap and around 5 per group when they differ.

Anything quantitative. The same update says to test at least 20 users to get statistically significant numbers. Five users tell you what is broken. They do not tell you how many customers are affected.

A week-long round

The schedule below assumes one product area, one user group and a team of two people who can give most of their week to it.

  1. Monday: define the round. Write three to five tasks in the customer's words. A task is a realistic activity, which Nielsen Norman Group lists as one of three components of a test alongside the facilitator and the participant. Draw the tasks from job statements, so they describe an outcome in place of a screen, for example “invite a teammate and give them read-only access.”
  2. Tuesday: recruit. Book five participants who match the group, plus one spare because someone will cancel. Real customers or people from the same segment work better than colleagues.
  3. Wednesday and Thursday: run the sessions. Plan five sessions of 45 to 60 minutes. One person facilitates and one takes notes. Observers stay silent and keep their questions for the debrief.
  4. Thursday afternoon: debrief. Each observer writes down the problems they saw before anyone speaks, then the team merges the lists.
  5. Friday: decide. Rank problems, assign fixes, and book round two for the following week with the fixes in place.

Running the session

The facilitator's job is to keep the participant doing the task. Read the task aloud, ask them to think aloud, and stay quiet when they struggle. If they ask “should I click here?”, answer with “what would you do if I were not here?” Avoid explaining the interface, defending a design or hinting at the answer. A participant who gives up on a task has told you something that no survey will.

Record the task, the outcome (completed, completed with difficulty, failed or abandoned) and the moment where they hesitated. These three facts per task per participant are enough to build the summary. Frustration is data too, and I made the case for mapping emotional states in emotion-first product design. I wrote about how product work removes friction users never mention in the invisible PM, and a usability session is the cheapest way to watch that friction happen.

Turning observations into decisions (hypothetical)

Suppose a 20-person SaaS company tests its invite flow with five users. Four of the five could not find the “Invite teammate” button, which is now in a settings menu. Two chose the wrong permission level because the labels were ambiguous. One disliked the color of the confirmation message.

Four of five is a problem that affects most users, and it gets fixed before round two. Two of five is likely to be common as well, and it gets fixed. One of five is logged and watched: if it appears again in round two, it moves up. The color complaint is noted and takes no time from the other two.

This triage works because of the curve above. A problem seen by several of five participants is probably frequent, and frequent problems are what a small round is designed to find. A problem seen by one participant might be rare or might be an accident of that person, and a second round tells you which.

When to use more than five

Use more than five when the stakes of a rare failure are high, such as a payment or a safety step, when the product serves several groups, or when you need numbers. Otherwise the cheaper plan is a small round every week or two. The product I would trust is the one that has been tested five users at a time, repeatedly, with every round's fixes tested in the next. The number of rounds matters more than the number of participants in any one of them, and a team that runs a round a week for a quarter will have watched sixty people work through its product and fixed what they found.

Perspectives

“The best results come from testing no more than 5 users and running as many small tests as you can afford.”

— Jakob Nielsen, Co-founder, Nielsen Norman Group

“After the fifth user, you are wasting your time by observing the same findings repeatedly but not learning much new.”

— Jakob Nielsen, Co-founder, Nielsen Norman Group

Frequently asked questions

Why do you only need five users for usability testing?

Jakob Nielsen's March 2000 Nielsen Norman Group article reports that a typical test user finds 31 percent of usability problems on average, so by the formula 1 minus (1 minus 0.31) to the power n, five users find about 85 percent. After that, sessions mostly repeat earlier findings. His advice is to run no more than five users per test and as many small tests as you can afford.

When are five users not enough?

Five is too few when you need to detect problems that affect a small share of users (about 18 participants give an 85 percent chance of seeing a problem affecting 1 in 10), when you test several distinct user groups, when tasks are open-ended, or when you need statistics. Nielsen Norman Group advises at least 20 users for quantitative studies.

How long does a five-user usability test take?

A focused round fits in a week for a team of two: define three to five tasks on Monday, recruit on Tuesday, run five sessions of 45 to 60 minutes on Wednesday and Thursday, debrief on Thursday afternoon, and rank problems and assign fixes on Friday. Then book round two for the following week with the fixes in place.

Sources

Written by Meet Patel — startup operator and growth strategist in Dubai.

Read on themeetpatel.com