Design a retargeting incrementality test that survives a finance review
A retargeting incrementality test randomises the retargeting audience into an exposed group and a held-out group, runs the flight, then compares registration rates with a two proportion test. Detecting a two point lift on an eighteen per cent base needs roughly 32,000 people at a ten per cent holdback.
The finance lead has one question about the retargeting line and it is a fair one. The platform says the campaign produced 5,982 registrations. How many of those people would have registered anyway?
A retargeting incrementality test is the only answer that holds up, because it is the only one that produces a comparison group. Everything else is a count of registrations by people who were already interested enough to visit your registration page, which is where the retargeting audience came from in the first place.
Why does the platform's retargeting number never survive scrutiny?
Retargeting is aimed at people who have already shown intent. They came to the site, they opened the form, some of them got as far as choosing a pass. That is the entire selection criterion, and it means the audience was going to register at a higher rate than the general population before a single impression was bought.
So when a platform reports 5,982 conversions, it is reporting registrations by people who saw an ad and then registered. The counterfactual, what those people would have done with no ad, is absent from the report and cannot be recovered from it.
Johnson, Lewis and Nubbemeyer set out the cleanest treatment of this in the Journal of Marketing Research in 2017. Their ghost ads method identifies the control group counterparts of exposed consumers inside a randomised experiment, so the comparison is between people the platform would have shown the ad to and people it did show it to. They report that, compared with public service announcement and intent-to-treat A/B tests, the approach measures lift just as precisely while spending at least an order of magnitude less, and their own implementation recorded more than 100 million predicted ghost ads a day. When they demonstrated it on an online retailer's display retargeting campaign, the ads lifted website visits by 17.2 per cent and purchases by 10.5 per cent.
Two things to take from that. Retargeting can work, measurably. And measuring it needs a control group that most organisers have to construct by hand, because ghost ad support is not available in every buying platform.
The holdout you can actually build
Without ghost ads, the practical design is a randomised holdout inside your own audience.
Take the retargeting audience as your registration system defines it, which for most shows means people who reached a given step of the form and have not completed. Assign each person a random number when they enter the audience. Everyone below the cut goes into the holdout and is never uploaded to any platform. Everyone above it is uploaded as normal.
Randomise at the person level, not the session level, and do it once. A person who re-enters the audience next week must land in the same group, otherwise the two groups leak into each other and the test measures nothing.
Then compare on assignment. Everyone in the exposed group counts, whether or not an impression was ever served to them. It is tempting to restrict the exposed side to people the platform confirms it reached, and it is the single most common way these tests get quietly broken, because the people who receive the most impressions are the people who browse the most, and they register at a higher rate for reasons that have nothing to do with your ads.
Building the audience by last completed step is a separate piece of work with its own definitions, covered in retargeting people who left the form part way, and this test assumes it is already running.
How big does the audience have to be to detect a 2 point lift?
This is where most event retargeting tests die, and it is better to find out before the flight than after.
Use the standard two proportion sample size formula. At five per cent two-sided significance and eighty per cent power, the multiplier is 7.84. Suppose the registration rate in the holdout is 18 per cent and you want to detect a lift to 20 per cent, a difference of 2 points. With an allocation ratio of nine exposed to one held out, the required holdout size is 7.84 multiplied by 0.1476 plus 0.16 divided by 9, all divided by 0.0004, which comes to 3,241. So you need about 3,250 people held out and 29,200 exposed, a retargeting audience of roughly 32,400 people.
Change the split and the arithmetic moves a long way. At an even fifty-fifty allocation the same test needs about 6,030 in each group, 12,060 in total. At a twenty per cent holdback it needs about 3,680 held out and 14,710 exposed, 18,390 in total.
That ordering surprises people, so it is worth stating plainly: a bigger holdout needs a smaller total audience. The precision of the test is limited by the smaller group, and a ten per cent holdout makes the smaller group very small indeed. My own preference on a show-sized audience is a twenty five per cent holdback, accepting the lost reach, because a ten per cent holdback on a typical audience gives a test that cannot detect anything worth acting on.
Run the calculation the other way when the audience is fixed. An audience of 8,000 split ninety-ten gives 800 held out, and the minimum detectable effect at that size is about 4 points, which means an 18 per cent base rate has to reach 22 per cent before the test can see it. If your expectation is a 2 point lift, that test will report no effect no matter what the ads do.
Running the two proportion test on one flight
Suppose you get the audience you need. The flight closes and the file looks like this.
Exposed group: 29,178 people, 5,982 registrations, a rate of 20.5 per cent. Holdout: 3,242 people, 592 registrations, a rate of 18.3 per cent. The observed difference is 2.24 points.
The pooled proportion is 6,574 divided by 32,420, which is 0.2028. The standard error is the square root of 0.2028 times 0.7972 times the sum of one over 29,178 and one over 3,242, which works out at 0.00744. The z statistic is 0.0224 divided by 0.00744, which is 3.01, and the two-sided p value is about 0.003.
That is a result you can take into a finance review, because every number in it is reproducible from two counts and two denominators, and the person checking it can redo the division on paper.
What the test does to your cost per registration
Here is the part that makes the meeting uncomfortable and makes the number trustworthy.
The incremental registrations are the lift applied to the exposed group: 2.24 per cent of 29,178, which is 654. If the flight cost 21,000 in media, the cost per incremental registration is 32.11. The platform's own figure, 21,000 divided by 5,982, is 3.51.
The tested number is roughly nine times the reported one, and the tested number is the one that answers the finance question. Both belong in the report, labelled, because the platform figure is still the right operational signal for judging creative and placement week to week.
Randall Lewis and Justin Rao, writing in the Quarterly Journal of Economics in 2015, are the necessary caution here. Across twenty five large field experiments representing 2.8 million dollars of digital advertising spend, they found the median confidence interval on return on investment was over 100 percentage points wide, and that individual-level sales are volatile enough that informative experiments can easily require more than 10 million person-weeks. Their finding is why this post tests a lift in registration rate and not a return on investment. Registration is a common event with a base rate near 18 per cent in this audience, so it is measurable at show scale. Revenue per registrant is not, and no honest test at your audience size will pin it down.
Where this stops
The design has three real limits and a finance reader will find all of them, so name them first.
Contamination is the first. A held-out person can still see your ads through a channel you did not hold back, receive the reminder email, or hear about the deadline from a colleague. Every one of those pushes the measured lift towards zero, so the test is conservative, and a null result means the ads did nothing detectable above everything else you were doing.
The second is that the answer expires. A lift measured in the final six weeks before doors, against a deadline, does not transfer to the quiet middle of a campaign, and a lift measured on one edition does not automatically hold for the next when the price, the venue or the calendar slot has moved.
The third is the opportunity cost of the holdout itself. Twenty five per cent of your audience receives no retargeting for the length of the flight, and if the ads work, that is real registrations you chose not to have in exchange for knowing. Decide it deliberately, and take it out of the reach forecast before anyone builds a target on it. Which channels deserve that trade in a given edition is a portfolio question covered in choosing where to spend your testing budget, and the same holdout logic applied to lapsed email cohorts has its own arithmetic in measuring a reactivation mailing.
The first step this week is a counting exercise. Pull the size of your current retargeting audience and its registration rate, then put both into the sample size formula above. If the audience is under about 12,000 people, no split will detect a 2 point lift this edition, and the useful decision is to pool the audience across two editions or across a portfolio of shows before spending anything on the test. That number, more than any argument about method, decides whether this work belongs in your acquisition and attribution plan this year.
Questions people ask about retargeting incrementality test
- How large does a retargeting audience need to be for an incrementality test?
- Detecting a two point lift on an eighteen per cent registration rate, at five per cent significance and eighty per cent power, needs about 3,250 people held out and 29,200 exposed. A fifty-fifty split needs about 12,100 people in total. Most single-show retargeting audiences are far smaller than either figure.
- Should the holdout be compared on exposure or on assignment?
- On assignment. Compare everyone randomised into the exposed group against everyone randomised into the holdout, whether or not an ad was actually served. Restricting the exposed side to people who saw an ad reintroduces the selection the experiment exists to remove, because heavier browsers see more impressions and register more often anyway.
- What does an incrementality test do to reported cost per registration?
- It usually multiplies it. In a flight where a platform claims 5,982 registrations on 21,000 of spend, the reported cost per registration is 3.51. If the tested lift is 2.24 points across 29,178 exposed people, the incremental count is 654 and the cost per incremental registration is 32.11.