Usability testing services range from self-serve participant platforms to fully managed research agencies. The right model depends on whether you need recordings against a scenario you already understand or a researcher to design and interpret the study.
Most usability testing that disappoints was designed badly rather than run badly. The participants did what was asked, the recordings arrived, and nothing in them could be acted on because the session asked for opinions instead of setting tasks.
This covers what each method costs, how many people you actually need, and how to write a session that produces findings rather than commentary.
Use RentHuman when
You already know the website or app flow you want tested, can write the tasks, and need recordings or structured answers from people matching countries, languages or devices. You review the evidence and decide what it means.
Use a research agency when
You need somebody to define the research question, recruit a difficult professional audience, moderate live sessions, probe follow-up questions and turn the evidence into recommendations. That is a managed study, not a distributed task.
The methods, and what each is for
Unmoderated remote testing. A participant works through tasks alone, recording their screen and speaking aloud. Fast, cheap, scales to any number, and no scheduling. You lose the ability to ask why.
Moderated remote sessions. A researcher watches live and probes. The richest signal available short of being in the room, and the most expensive — scheduling, a researcher's time, and a higher incentive.
Guerrilla testing. Approaching people informally with a prototype. Nearly free, uncontrolled sample, useful for catching obvious problems early.
Real-device field testing. Participants use their own hardware in their own conditions, which surfaces problems that never appear on a developer's machine. How that differs from a device farm is a related question.
Full-service research. An agency designs, recruits, moderates and analyses. Several thousand per study, and correct when the question is strategic rather than "does this flow work".
Cost
| Method | Per participant | Setup |
|---|---|---|
| Unmoderated, 10 minutes | $8 – $20 | Hours |
| Unmoderated, 20–30 minutes | $20 – $60 | Hours |
| Moderated, 45–60 minutes | $60 – $150 | Days |
| Real-device field test | $12 – $40 | Hours |
| Full-service study | $3,000 – $25,000 total | Weeks |
How many participants
Five per distinct user group, per round.
Usability problems cluster heavily. The serious ones appear in the first three or four sessions, and beyond about eight you are paying to watch people encounter issues you have already logged. The classic finding that a handful of users surfaces the large majority of problems has held up well in practice for single-flow testing.
The exception is distinct groups. If new users and returning users experience genuinely different journeys, that is two groups and two sets of five.
The right response to having budget for twenty sessions is three rounds of five with fixes in between, not one round of twenty. The second round measures whether the fix worked, which is the round that earns its cost.
Writing tasks instead of asking questions
This is where a session succeeds or fails.
Set a goal, not a route. "Buy a size medium in blue and have it delivered to a work address" is a task. "Click the size selector, then add to basket" is a script that tests whether people can follow instructions.
Give a reason, briefly. "You have seen a jacket a friend recommended and want to check whether it comes in your size" puts somebody in a frame of mind. Context changes behaviour more than most teams expect.
One goal per task. Compound tasks produce a report you cannot attribute to a step.
Do not name the interface elements. The moment you say "the filter menu", you have told them where to look and destroyed the finding.
Ask what they expected before they act. "What do you think will happen when you press that?" catches mismatches between the label and the behaviour, which is the most common class of usability defect.
Ask for recall afterwards. What do they remember, without looking. What survives a minute of distance is what the interface actually communicated.
Include a task you expect to fail. A round where everything succeeds usually means the tasks were too easy, and you have paid for reassurance.
A scenario template
Copy and adapt. Six blocks, twenty minutes.
Context, 1 minute
Who they are in this scenario and why they are here. No product explanation.
First impression, 1 minute
Landing on the page: what is this, who is it for, what can you do here? Ask before they scroll.
Primary task, 6 minutes
The one flow that matters. Goal stated, route not.
Recovery task, 4 minutes
Something goes wrong — wrong item, needing to change something, an error state. Most products are tested only on the happy path and most support cost comes from the other one.
Secondary task, 4 minutes
A second flow, or the same one under a different constraint.
Debrief, 4 minutes
What was hardest, what was unexpected, what they would tell a friend about it. Then recall: what was the price, what was the delivery promise.
Ask participants to think aloud throughout, and to say when they feel stuck rather than pushing through in silence.
Running it on a task platform
Unmoderated rounds work well as structured tasks, and the mechanics are the same as any other brief.
You publish one brief with the scenario, the tasks in order, and the proof required — usually a screen recording with audio, plus written answers to the debrief questions. You set the reward and the number of participants, and target by country, language and where relevant the device.
On RentHuman funding is committed before the brief publishes, eligibility is checked before somebody reserves rather than after they deliver, review runs on a fixed 48-hour clock with one correction request, and a rejection opens a 72-hour appeal decided by an administrator. Participants see the reward and the exact requirements before committing. The task marketplace mechanics are identical for any other kind of brief.
What you do not get is analysis. The recordings and the answers arrive; deciding what they mean is yours. For a team that already knows what it is looking for, that is the cheap version of the same input. For a team that does not, an agency is the honest recommendation.
Where recruitment itself is the constraint, what participant recruitment costs and why incidence rate drives it covers the economics.
Recruiting the right participants
Who you test with changes the findings more than anything except the tasks, and teams routinely get this wrong in one of two directions.
People too close to the product
Colleagues, existing power users, anyone who already knows what the product does. They cannot un-know it, and they will sail through a flow that confuses everybody else.
People who would never use it
A general consumer sample for a specialist professional tool produces confusion that tells you nothing, because the confusion is about the domain rather than the interface.
The target is people who plausibly have the problem your product solves and have not seen your solution. That is usually a looser screen than teams write. Over-specifying the profile raises the cost, slows recruitment, and rarely changes what the session finds.
Three practical screens that matter more than demographics: have they used a competing product, how comfortable are they with this kind of software generally, and are they in the market where the product behaves as you expect.
Sessions that waste the budget
Six patterns, all common, all avoidable.
Testing a prototype that cannot fail. A clickable mock with only the happy path wired up. Participants cannot go wrong, so nothing is learned.
The team watching live and intervening. Someone unmutes to explain. The moment you explain, the session is over as measurement.
Leading questions. "Was that easy to find?" gets yes. "Talk me through how you found that" gets the truth.
Testing the whole product. A ninety-minute session across six flows produces shallow coverage of all of them. One flow, properly, beats six superficially.
Recruiting only in your own market for a product sold in several.
Not watching the recordings. More common than it should be. A round of eight unwatched sessions is a round of eight wasted incentives, and summaries written by somebody else lose exactly the hesitations you were paying to see.
Reading the results
Count, then rank. Six participants hesitating at the same point outranks one dramatic failure. Frequency first, severity second.
Separate three things. Defects go to engineering. Comprehension problems go to design. Preferences are noise unless several people independently share them.
Watch where they look before they act. Hesitation before a click is the signal; the click itself is the outcome. The pause is where the problem is.
Believe behaviour over commentary. Participants routinely say a thing was fine immediately after struggling with it, out of politeness and a sense of having failed a test.
What they did is the data.
Note the unprompted. Anything a participant mentions that nobody asked about is usually the most honest thing in the session.
Testing in other markets
The case where this has no substitute. A flow that works perfectly in your market fails elsewhere for reasons that are invisible from inside the team: a phone number format rejected by validation, an address form assuming a postcode, an absent payment method, a date order causing a misread, a translation that is grammatically correct and culturally wrong.
None of it appears in automated testing, because the code is correct. It appears when somebody in that country tries to use the thing with their own details.
Target by country and language, ask participants to complete the flow as though it were their own money, and include one question asking whether anything felt unusual for their market. That question surfaces more than the rest of the script combined.
Starting
One flow, five participants, a twenty-minute scenario with tasks rather than questions, and a debrief that asks for recall.
Watch all five recordings yourself before anybody summarises them. Then fix what three or more people hit, and run the same five tasks again. Two rounds of five with a fix between them is worth considerably more than one round of ten, and costs the same.
Where the question is whether people understood a message rather than whether they could complete a task, creative testing is the closer instrument and needs a different question set entirely.
Where the build has already been translated and the question is whether it survived contact with another language, a localization QA pass is the same mechanism aimed at a different class of defect.
For the wider question of which instrument answers which kind of question, a comparison of customer research methods sets out what each one cannot do.