A survey looks like the easy method. The tool is already licensed, the questions seem obvious, and the output arrives as a chart.
That impression is why so much of it is wasted. A survey is an instrument, and almost all of its error is built in before the first response arrives — in how questions were worded, which scale was used, what order things appeared in, and who was on the list.
What this method is for
Surveys measure prevalence. How many, how much, how often, and how one group differs from another. Within that scope they are the best instrument available and nothing else comes close on cost.
They are poor at explaining why, and unreliable about what people will do in future. Both limits are structural rather than fixable with better questions, and which method answers which shape of question works through the alternatives when your question is not a prevalence question.
Closed-ended questions
Fixed answer options. Countable, fast to answer, and limited to what you thought of when writing them.
Four common forms, with examples worth copying:
Dichotomous
Two options. "Have you bought running shoes in the last twelve months? Yes / No." Clean, unambiguous, and the right choice for a filter or a screener.
Multiple choice
A list, single or multi-select. "Which of these did you use most recently? [brands, randomised] / None of these / Cannot recall." The last two options are not optional — without them, people pick something.
Rating scale
A point scale with labelled ends. "How difficult was completing your order? Very easy 1 – 5 Very difficult." Label the ends explicitly rather than relying on numbers to carry meaning.
Ranking
Ordering a short list. "Rank these three by how much they influenced your choice." Cognitively expensive; keep it to four items or fewer or the later ranks become noise.
Open-ended questions
Answered in the respondent's own words, and the only part of a survey capable of telling you something you did not anticipate.
Examples that produce usable material: "What were you trying to do when you first looked for something like this?" — "What nearly stopped you from buying?" — "If you could change one thing about it, what would it be and why?" — "Why did you choose that answer?" placed immediately after a low rating.
Examples that do not: "Any other comments?" at the end of a long survey, which returns complaints about the survey; and "What do you like about us?", which returns the sentence already on your website.
Two or three open questions is usually the right number. They cost more to analyse than to collect, which is why programmes quietly stop including them, and the analysis is where the value was.
The closed questions confirm what you suspected. The open ones are the only place a survey can surprise you.
Wording, where most error enters
Six rules, and drafts routinely break the first three.
One idea per question. "How satisfied are you with the speed and accuracy of delivery?" cannot be answered by somebody who found it fast and wrong. Split it.
No presumption. "How useful was the new dashboard?" assumes usefulness and gets it. Ask what they used it for, and what happened.
No leading. "How much do you agree that our support is responsive?" is not a measurement. "How would you describe the response time when you last contacted support?" is.
Concrete over abstract. "How often do you use it?" produces guesses; "How many times did you use it last week?" produces a number people can actually retrieve.
Plain language and no internal vocabulary. If the term appears only in your product documentation, half your respondents will answer about something else.
And no double negatives. "Do you disagree that the process should not be simplified?" has a correct answer and nobody will find it.
Scales, and why they do not travel
The scale is a design decision with consequences most teams never examine.
Five or seven points both work; more points do not add precision, they add the appearance of it. Label every point where you can, because respondents interpret unlabelled numbers differently from each other.
Agreement scales — "strongly agree" to "strongly disagree" — carry acquiescence bias: a measurable tendency to agree, stronger in some cultures than others. Where you can, replace the agreement frame with a direct question about the thing itself.
Order, and the contamination nobody notices
Question order changes answers, reliably and by more than most campaigns move anything.
Ask unaided before aided. Once a respondent has seen your brand name in a list, every later question about awareness is contaminated, and the contamination always flatters you.
Randomise the order of options in lists, and randomise the order of blocks where the content permits. Fixed lists produce position effects that look like preferences.
Go from general to specific. A specific question early sets a frame the general question afterwards cannot escape — ask about a recent support problem and the subsequent "overall satisfaction" measures that problem.
And put demographic and sensitive questions at the end, once somebody is invested enough to answer them.
Length, paid for in representativeness
Completion falls steadily with length, and the fall is not random. The people who abandon a long survey are the busy and the indifferent, who are usually the people you needed most.
The practical discipline is to justify every question by naming the decision it informs. Questions that survive that test are few, and the ones that fail it are almost always the "while we have them" additions from three different teams.
Time your own draft, then add half again. Estimates by the person who wrote the questions are consistently short because they never have to re-read one. How the incentive interacts with length covers the other half of this trade — a longer survey needs a larger incentive, and past a point no incentive rescues it.
The response quality problem
A portion of any incentivised sample answers without engaging, and the design decides how much of that reaches your analysis.
Straightlining is the visible form — the same option selected down a grid, at a speed no genuine reader could manage. Set a minimum completion time from your own pilot rather than from a guess, and treat anything under about a third of the median as suspect rather than automatically invalid.
Attention checks work and are easy to overuse. One instructed-response item in a long survey is reasonable. Several read as distrust, and they annoy careful respondents more than they catch careless ones, which costs you good data to remove bad.
Better than either is a question whose answer you can verify from something other than the respondent's word — a detail only somebody with the described experience would know, or a fact you already hold about them. That is the survey version of the same principle that governs any distributed work: evidence produced as a by-product of the thing being true beats an assertion, and the requirement has to be set before fielding rather than applied afterwards.
Sampling, which decides more than the questionnaire
A perfect instrument administered to the wrong people produces a confident, precise, wrong answer.
Your own customer list is not a sample of the market. It contains people who chose you, are still subscribed, and still open your email. Missing from it: everybody who considered you and went elsewhere, everybody who left, and everybody who disengaged before leaving. Which populations a feedback programme never reaches covers that gap in detail.
Quota sampling — filling defined cells for age, market, category behaviour — is the workable approach for most commercial research, provided the quotas come from something real rather than from what was easy to fill. Watch fill rate by cell while in field; a cell that fills unusually fast is usually being filled by people who do not quite match it.
Subgroups need their own samples. A study with 1,000 responses reporting confidently on a segment of 60 has switched instruments halfway through the deck without saying so, and the arithmetic behind detecting small differences applies to every one of those cuts.
Where the profile is narrow, incidence dominates the budget: at one in fifty qualifying, you pay for forty-nine screeners per respondent. Why incidence drives recruitment cost sets out the numbers, and it is worth checking feasibility before writing a single question.
Pilot before you field
The cheapest insurance available, and routinely skipped under deadline.
Soft-launch to a small share of the sample and stop. Then look at four things: completion time distribution, straightlining, drop-off point, and the open-ended answers.
Completion times cluster oddly when a question is being misread. Straightlining — the same answer down a grid — tells you the grid is too long. The drop-off point names the question that lost you the sample. And ten open-ended answers will reveal a misunderstanding no amount of internal review caught, because everybody internal already knew what the question meant.
Multi-market
Translate properly rather than machine-translating, and have the translation checked by somebody who will actually administer it. A question rendered ambiguously in one language produces a systematic error in that market's data that looks like a genuine difference.
Build examples and answer options per market rather than translating one set. Brand lists, retailer names, income bands and job titles do not transfer, and a list containing four retailers nobody in that country uses tells you only that.
And treat each market as its own sample with its own quotas. What each layer of targeting costs sets out where narrowing gets expensive, which is usually the point where a study quietly drops the harder markets.
A short checklist
Before fielding: every question names a decision it informs; no question contains two ideas; every list has an escape option; unaided precedes aided; lists are randomised; the scale is labelled; the analysis plan and subgroups are written down; incidence and feasibility are confirmed; the incentive matches the length; and somebody outside the team has taken the survey and been asked what each question meant.
That last one takes twenty minutes and catches more than the rest of the list combined. For anything where the question turns out not to be a prevalence question after all, the comparison of research methods is the faster route back, and how a brief becomes reserved capacity and verified responses covers the recruitment side when your own list cannot supply the sample.