The method is one of the more defensible things in advertising measurement. Randomly withhold a campaign from part of the audience, survey both groups identically, and attribute the difference to exposure. That is a real experiment, and it answers a question attribution models cannot.
Most of what goes wrong happens in the execution, and two decisions carry nearly all of it: how the control group was built, and whether anybody checked that the sample could detect the effect being reported.
What is actually being measured
Not sales. Not revenue. A stated attitude, captured in a survey, at one moment.
The standard measures run in a rough sequence from easy to move to hard: unaided awareness, aided awareness, ad recall, favourability, consideration, purchase intent. Campaigns move the top of that list far more readily than the bottom, which is worth knowing before somebody sets a target on intent.
What lift is
The difference between the exposed group's answer and the control group's answer, expressed in percentage points. Ten percent of the control considers the brand, thirteen percent of the exposed group does, and the lift is three points.
What lift is not
A sales measurement, a causal account of why, or a durable property of the brand. It is a snapshot of an attitude, taken days after exposure, from people who overwhelmingly do not remember seeing the advertisement.
The last part deserves emphasis because it is frequently misunderstood as a flaw. Respondents not remembering the ad does not invalidate the method — the design does not depend on recall, it depends on randomisation. But it does explain why the effects are small, and small effects are where measurement problems hide.
The control group is the entire study
Everything rests on the two groups differing only in whether they saw the campaign.
The clean way to achieve that is a randomised holdout: before the campaign runs, a portion of the target audience is randomly excluded, and they are the control. Randomisation handles the variables you thought of and, more importantly, the ones you did not.
The compromised versions are common and worth recognising.
Post-hoc matching
Finding people who resemble the exposed group after the fact. Better than nothing and structurally weaker, because the reason somebody was exposed frequently correlates with the outcome being measured.
Unexposed-but-eligible
Using people who were targeted but never served an impression. They differ systematically — lighter platform users, different devices, different browsing patterns — and those differences are not random.
Geographic holdout
Excluding whole regions. Legitimate and coarse; regions differ from one another in ways a campaign is not responsible for, and the design needs enough of them to average that out.
Power, and why most studies cannot support their conclusions
This is the part that gets skipped and the part that decides whether the result means anything.
Lift effects are small. A campaign moving a consideration measure by three to five points is doing well. Detecting a difference that small, reliably, needs a sample sized for it, and the arithmetic is not gentle.
For a measure sitting around 30 percent, distinguishing a five-point difference from noise needs on the order of 1,300 completed responses in each group. Detecting two points needs roughly 8,000 per group — the requirement scales with the square of the effect you are trying to see, so halving the effect quadruples the sample.
| Effect to detect | Approx. responses per group | Practical read |
|---|---|---|
| 10 points | ~350 | Achievable, but this size of effect is rare |
| 5 points | ~1,300 | The realistic target for a well-run study |
| 3 points | ~3,700 | Expensive, and where most real effects live |
| 2 points | ~8,000 | Rarely funded, frequently claimed |
A study with four hundred people per group reporting a two-point lift has not measured a two-point lift. It has produced a number in that neighbourhood by chance.
The consequence is not that small studies are useless. It is that they can only support conclusions about large effects, and the honest write-up says so rather than reporting a point estimate to one decimal place. Slicing an underpowered study by age band, market or creative variant makes it worse in exactly the way that generates the most confident-sounding findings.
The platform conflict
Every major ad platform offers this measurement, usually free above a spend threshold, and the studies are generally competently run.
The structural problem is not competence. It is that the same party sells the media, defines what counts as exposure, constructs the control group, writes the questions, decides which results are reportable, and publishes the outcome. Each of those is a judgment call with a defensible answer, and every one of them sits with somebody who benefits from a positive result.
Two practical consequences follow, and both are worse than any question of good faith.
Cross-platform comparison becomes impossible. One platform's definition of an exposure may be an impression rendered, another's a video played to some threshold, another's an impression within a recency window. Comparing lift figures across platforms compares three different measurements wearing the same name.
And the counterfactual is invisible to you. You receive a result, not the design. What the control group actually was, how many responses backed each cell, and which questions were asked in what order are typically summarised rather than disclosed.
Running one independently
The alternative is to define the study yourself and recruit the respondents, which costs more and answers a question the platform version cannot.
The design requirement is the same as ever: a genuine holdout, identical questions across every arm and every platform, and a sample sized for the effect. What changes is that you control all three and can therefore compare across media rather than within one seller's measurement.
The hard part is not the survey. It is reaching a matched population in the markets that matter, at sufficient sample, including the markets where panels are thin. How professional and market-specific recruitment actually works covers that mechanism, and what each layer of targeting costs sets out where narrowing the audience definition starts to dominate the budget.
| Approach | Typical cost | What you get |
|---|---|---|
| Platform-run study | Free above a spend threshold | Directional, single-platform, opaque design |
| Research vendor | $15,000 – $60,000 | Full design control, panel access, analysis |
| Direct recruitment and short survey | $3,000 – $15,000 | Control at lower cost, you own the design work |
Writing questions that do not manufacture the answer
Small design choices move results by more than most campaigns do.
Ask unaided awareness before showing the brand name. Once a respondent has seen the name in an earlier question, every subsequent measure is contaminated and the contamination flatters you.
Randomise the order of brands in any list, and include competitors. A fixed list with your brand first produces a reliable lift that belongs to the position rather than the campaign.
Keep the questionnaire short. Attention decays quickly, and a long instrument produces straightlining — respondents selecting the same answer down the page — which mostly adds noise and occasionally adds bias.
And treat purchase intent with suspicion proportional to how much you want it. Stated intent has a weak and category-dependent relationship with behaviour, it is the measure most sensitive to question wording, and it is the one most often chosen as the headline because it sounds closest to money.
Timing
Field while the campaign is live or immediately after. Attitudinal effects decay quickly, and a study fielded three weeks later is measuring what survived rather than what happened.
For campaigns running over months, repeated smaller waves beat one large study at the end. Waves show you the shape of the response — where it started, whether it plateaued, whether it decayed between bursts — and that shape is more actionable than a single number, even though each wave individually supports weaker claims.
Reading a result somebody hands you
Most people encounter these as a finished slide rather than a study they commissioned, so it is worth knowing what to look for.
Find the base sizes per cell before reading any number. Not the total sample — the count behind the specific figure being shown. A deck reporting a headline from four thousand respondents and a market breakdown from two hundred has changed instrument halfway through without saying so.
Look for an interval or a significance note. A lift quoted as a bare number, particularly to one decimal place, is presenting an estimate as a measurement. The honest version says three points, plus or minus two, and lets you see how much room that leaves.
Check whether the subgroups were decided in advance. A study reporting that the campaign worked especially well among one age band in one market, where that cut appears nowhere in the original design, has found the most flattering slice of its own noise. Pre-registered cuts are a finding; discovered cuts are a hypothesis.
And ask what the flat results were. Every genuine study has cells that did not move, and a report containing only positive findings has been filtered somewhere between the field and the slide.
What to ask a supplier
Six questions, and the answers separate a real study from a report.
How is the control group constructed, and is it randomised before the campaign or matched afterwards?
What sample size per cell, and what effect size can that detect? A supplier who cannot answer the second half has not done the calculation.
What is the exposure definition, precisely?
Will we see the questionnaire in full, in field order, before it goes out?
What is the plan for subgroups, agreed in advance? Deciding which cuts to report after seeing the data is how noise becomes a finding.
And what happens if the result is flat? A supplier with a considered answer to that is one you can trust with a positive result.
What it will not tell you
Being clear about the boundary is more useful than widening the claim.
It does not measure sales. Incrementality of revenue is a different experiment with a different design, and a brand that moved attitudes without moving purchases has learned something real rather than nothing.
It does not measure durability. A lift measured during a campaign says nothing about the position in six months, and brand equity work operates on a timescale no single study captures.
And it does not explain why. A number tells you something moved; understanding the mechanism needs qualitative work alongside it, whether that is structured usability and comprehension testing on the creative itself or conversations with people in the category.
For the adjacent question of whether the creative is worth running at all before you spend the media budget measuring it, what human pre-testing answers that an in-platform test cannot is the earlier step, how a campaign brief becomes usable creative covers the production side, and how a brief becomes reserved capacity and verified responses describes the recruitment mechanism end to end.