How Do You Run a Usability Test With Just Five People?
How Do You Run a Usability Test With Just Five People?
Recruit five people who resemble your users, give them realistic tasks, and stay quiet while they try. That is genuinely the method. The value comes from watching someone fail at something you thought was obvious, and five people is enough for that to happen repeatedly.
Teams delay this work because they imagine a lab, a research plan, and a statistician. None of that is required for the kind of testing that improves a product, and waiting for the proper version usually means never doing it at all.
Here is how we run these, including the parts that are easy to get wrong.
Why Is Five the Number?
Because of diminishing returns, and the maths is public. Jakob Nielsen published the argument on 18 March 2000, modelling the problems found as the total number of usability problems multiplied by one minus L, raised to the power of the number of users, where L is the proportion one user finds.
The value that makes it work is stated plainly. Nielsen gives a typical L of 31 percent, averaged across a large number of projects they studied. Run that curve and five users surface roughly 85 percent of the usability problems in the thing you are testing.
The sixth user mostly repeats what you already saw. That is the whole argument, and it is an argument about efficiency rather than rigour. You are not trying to measure anything. You are trying to find problems, and problems repeat.
What Should You Do With the Rest of the Budget?
Run more rounds. Nielsen's recommendation is explicit: rather than one elaborate study, spend the budget on three studies with five users each. Fix what you find between rounds and test the redesign, which is how the method compounds.
The reason that beats one big study is that a large test tells you about one version of the design. Three small tests tell you whether your fixes worked, which is the question you will actually have after the first round.
It also changes the politics. A single expensive study becomes an event with findings that must justify the cost. Three cheap ones become a routine, and routines survive budget conversations far better than events do.
What Should You Test and What Should You Skip?
Test the paths where money or trust is at stake. Signup, first use, the pricing decision, the moment someone has to give you data. Skip the pages people rarely visit and the flows your team already argues about for aesthetic reasons.
The best candidate is usually whatever your team disagrees about most confidently. If two people are certain about opposite things, neither has evidence, and a five-person test settles it in an afternoon for less than the cost of the meeting.
You do not need a finished product. Thinking aloud works at any stage, from paper prototypes through to live systems, which means you can test a design before anyone builds it. Our notes on SaaS onboarding flow design cover the flow that most repays this.
What Is Thinking Aloud and Why Does It Work?
You ask participants to say what they are thinking continuously while they use the thing. Nielsen called it a window on the soul in a piece published on 15 January 2012, because it exposes not just what people do but what they believed was going to happen.
Its strengths are practical. It is cheap, needing very little equipment, and you can collect data from several users in a day. It is robust, producing reasonable findings even when run imperfectly, which is not true of quantitative studies. And it is persuasive, because a team that watches a user struggle argues less afterwards.
It has honest limits too. Talking continuously is unnatural, some people narrate a tidied-up version of their thoughts, and a facilitator who prompts too much changes the behaviour they came to observe. Knowing that is most of the defence against it.
How Do You Write Tasks That Do Not Lead People?
Describe a goal, never an interface. Find out whether this plan includes single sign-on is a task. Click the pricing link and then open the comparison table is a script, and it tests nothing except whether the person can follow instructions.
Avoid your own vocabulary. If your navigation says Workspaces and your task says find your workspace settings, you have handed them the answer. Use the words a customer would use, which is usually a clue that your labels need work anyway.
Give each task a clear finish line so both of you know when it is over. Ending a task on I think that is it tells you something important: that the interface never confirmed success, which is a finding in itself.
Who Should You Recruit?
People who resemble the users you care about in the ways that matter for the task. For a B2B product that usually means role and familiarity, not demographics. Someone who has bought software like yours behaves differently from someone who never has.
Do not test with colleagues, and be careful with existing customers who love you. Both groups have learned your interface and your language, and their fluency will hide exactly the problems a new visitor hits.
Five is per audience, not per study. If you genuinely serve two distinct groups, such as an admin and an end user, that is two sets of five, because their tasks and their mental models barely overlap.
What Do You Actually Do in the Session?
Very little, on purpose. Nielsen reduces the method to three steps: recruit representative users, give them realistic tasks, and then, in his words, shut up and let the users do the talking. Most of the skill is in resisting the urge to help.
When someone goes quiet, prompt with a neutral question rather than a hint. What are you thinking, or what did you expect to happen, keeps them talking without steering. Never ask do you see the button, because now they do.
When someone gets stuck, let the silence run longer than feels polite. The stuck moment is the finding, and rescuing them destroys the data you came for. Thirty uncomfortable seconds teaches you more than the other twenty minutes.
How Do You Turn Notes Into Decisions?
Write observations, not opinions, then count repeats. An observation is that three of five people looked for pricing in the top navigation and did not find it. An opinion is that the navigation is confusing. Only the first one tells the team what to change.
Sort by how many people hit it and how badly it blocked them. Anything that stopped three or more people is a fix now. Anything one person hit and recovered from goes on a list and waits for a second sighting in the next round.
Then write down what you changed and why, so the next round can tell whether it worked. Without that record the third study repeats the first, and nobody notices. Our notes on running a design review process cover where these findings should land.
What Are the Most Common Mistakes?
Testing too late is the big one. A usability test run the week before launch produces findings nobody has time to act on, which converts research into a source of guilt. Test when changing your mind is still cheap.
The second is asking for opinions. Would you use this and do you like the design produce polite, useless answers. People are bad at predicting their own behaviour and good at being kind to whoever is in the room with them.
The third is testing alone. Get one other person from the team to watch at least one session live. Second-hand findings get argued with; watching a real person fail ends the argument in about ninety seconds. Our piece on pre-launch website testing covers the wider checklist.
How Do You Make This a Habit?
Put it in the calendar rather than in the plan. A standing slot, same day every fortnight, with whatever is ready at the time, beats a research phase that gets cut when the timeline tightens. Nielsen suggests weekly sessions for teams that want to get good at it.
Keep a standing list of recruits so scheduling is never the blocker. The reason most teams stop is not that testing failed, it is that finding five people took two weeks and by then the design had shipped.
If you want help setting this up, or you want someone outside your team to moderate so your own assumptions stay out of the room, we are happy to help. We do this alongside build work at phoenix.studio, and the first round usually pays for itself in arguments avoided.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.