Post show survey response bias and how to size it before publishing
Post show survey response bias is the non response rate multiplied by the difference between respondents and non respondents on whatever you are measuring. A nine per cent response rate is harmless if those two groups look alike and severe if they do not, so compare respondent composition against the full attendee file before publishing any mean.
Nine per cent responded. Two thousand one hundred and sixty replies against twenty-four thousand verified attendees, mean satisfaction 7.7, and the slide is due Thursday.
The instinct at that point is to argue about whether nine per cent is enough. That argument has no end, because post show survey response bias does not follow from the response rate. It follows from how different the people who replied are from the people who did not, and unlike the response rate, that difference is something you can partly measure with the files already on your desk.
Why does a low response rate not automatically mean a biased result?
Groves put the case directly in Public Opinion Quarterly in 2006, in a paper on nonresponse rates and nonresponse bias in household surveys. His abstract records that recent empirical findings illustrate cases where the link between nonresponse rates and nonresponse biases is absent, and the paper states that there is no minimum response rate below which survey estimates are necessarily subject to bias.
Groves and Peytcheva went further in the same journal in 2008, assembling fifty-nine methodological studies designed to estimate the size of nonresponse bias directly. The response rate turned out to be a weak guide to how much bias a survey carried.
That is liberating and slightly alarming at the same time. Liberating, because your nine per cent is not automatically worthless. Alarming, because a comfortable forty per cent response from a self-selecting group can be worse than a low response from a representative one, and nobody in the room will challenge the forty per cent.
The algebra is short enough to do here. Split the audience into respondents and non respondents. If r is the share who responded, the audience mean is r times the respondent mean plus one minus r times the non respondent mean. Rearrange and the error in the respondent mean comes out as the non response rate multiplied by the difference between the two group means. Two terms, and you have no control over the second one at the point of analysis.
The composition check you can run in an hour
Join the response file to the registration file on whatever key the survey platform returned, usually a token or an email hash. Then produce one table with two columns: the share of the full attendee file in each category, and the share of respondents in each category.
Run it on every attribute you hold. Badge type. Country or region. First visit or returning. Company size band. Registration source. Day of first scan. None of this requires new data collection and all of it comes out of the same attendee analytics join you use for the rest of the post-show pack.
What you are looking for is any category where the two shares differ by more than a few points. Those are the places where the bias term has something to work with.
Putting a number on the gap
Suppose the check comes back with one large discrepancy. Returning attendees are 30 per cent of the verified attendee file and 61 per cent of respondents.
Work out what the non respondents look like, because that is the group the formula needs. Returning attendees in the population number 0.30 times 24,000, which is 7,200. Returning attendees among respondents number 0.61 times 2,160, which is 1,318. The difference, 5,882, sits among the 21,840 non respondents, so their returning share is 5,882 divided by 21,840, or 26.9 per cent.
Now bring in what you know about satisfaction by group. Say returning attendees rate the show at 8.1 and first-time attendees at 7.2, a gap of 0.9 that you can compute from the respondents themselves.
The respondent mean is 0.61 times 8.1 plus 0.39 times 7.2, which is 7.75. The audience mean, if the group means hold, is 0.30 times 8.1 plus 0.70 times 7.2, which is 7.47. The reported figure is 0.28 too high.
Check it against the formula. The non respondent mean is 0.269 times 8.1 plus 0.731 times 7.2, or 7.44. The gap between respondents and non respondents is 0.31, the non response rate is 0.91, and 0.91 times 0.31 is 0.28. The two routes agree, which they should, and the agreement is worth confirming once so you trust the shortcut afterwards.
Three tenths of a point on a ten point scale is small enough that most boards would shrug and large enough to account for the entire year-on-year movement your report is about to attribute to the new hall layout.
The two dimensions worth checking first
Returning attendance is the one that skews almost everywhere, and the mechanism is obvious once stated. A person who has come to your show four times has a relationship with your brand, opens your email, and answers your survey. A first-time attendee who came for one meeting and left at lunch has none of that.
International share skews the other way and gets missed more often. The invitation lands while those attendees are in transit or back in a different time zone with a week of accumulated work, and the reply rate drops. If overseas visitors also rate the show differently, and at most shows they do, the direction of the resulting bias is predictable.
There is a third that is more of an outright defect than a skew. If exhibitor staff, press and speakers are on the same send as visitors, their responses land in the same file, and the mean you publish is a blend of four populations with different reply rates and different views. Split the send, or at minimum split the analysis, before you look at anything else.
What size of bias should worry you?
The threshold is not statistical. It depends on what the number will be used for.
If the satisfaction figure feeds a page in the show report that nobody acts on, an error of 0.28 changes nothing. If it feeds a year-on-year comparison where last year's figure came from a differently composed sample, then a bias of 0.28 in one edition and 0.11 in the other has manufactured a movement of 0.17 out of nothing at all. If it feeds a bonus calculation, someone will eventually audit it.
My own rule is that any bias estimate above about half the year-on-year movement you are reporting has to appear in the report. Below that, record it in the methodology note and move on.
Two things make the rule easier to apply than it sounds. The first is that composition tends to be stable across editions of the same show with the same send mechanics, so once you have measured the returning skew for one year you have a reasonable prior for the next, and a sudden change in the skew is itself worth investigating. The second is that the bias term scales with the group difference, which means an attribute where the two groups rate the show identically contributes nothing however lopsided the sample is on it. Country of residence is often like that: heavily skewed in the sample, and almost irrelevant to the mean, because overseas and domestic visitors rate the show within a tenth of each other. Check the gap in the outcome before you spend a week fixing a gap in the composition.
What to publish alongside the mean
Three lines, all short. The response rate with its numerator and denominator. The largest composition gap you found, named, with both percentages. And the estimated direction and size of the effect on the headline figure, if you were able to estimate it.
That last line is the one people resist, on the grounds that it invites doubt. It does the opposite in practice, because a reader who is told the figure is about 0.3 high and is still 7.5 has been given a number they can rely on. A reader given 7.75 with no qualification, who then discovers the sample was 61 per cent repeat attendees, stops believing the whole pack.
Correcting the mean rather than merely reporting the gap is the next step, and the post-stratification arithmetic for it is weighting survey results in D26.
Where this stops
Composition checks only reach the attributes you hold. If the real driver of who replies is something the registration record does not carry, such as how far the person walked to find the stand they came for, no join will surface it and the bias estimate you publish will be an underestimate of unknown size.
The second limit is that converting a composition gap into a bias estimate assumes the group means are themselves unbiased. In the arithmetic above, the 8.1 for returning attendees came from returning respondents, and if the returning attendees who reply are the happier ones, then 8.1 is too high and so is the correction built on it. The estimate is still worth having, because it is bounded and directional, and it beats the alternative of asserting that nine per cent is fine.
There is a design point here too. A shorter instrument reaches a wider slice of the audience, which shrinks the non response rate and therefore the first term in the formula, and the trade-off between item count and completion is survey length and completion rate in D24. Deleting questions the registration file already answers, as post event survey design in D22 sets out, does the same work from the other end.
This week, take the response file from your last edition and the registration file it came from, and produce a single two column comparison of returning share. If those two percentages are within three points of each other, you have a defensible sample and you can say so in the report. If they are thirty points apart, you have just found the reason your satisfaction score has looked stable for four years.
Questions people ask about post show survey response bias
- Does a low survey response rate mean the results are biased?
- Not on its own. Groves showed in Public Opinion Quarterly in 2006 that there is no minimum response rate below which estimates are necessarily biased, because bias depends on how different the non respondents are, not on how many of them there are. A nine per cent response from a group that resembles the whole audience beats a forty per cent response from a skewed one.
- How do I measure post show survey response bias?
- Join the response file to the registration file on the same key, then compare the two on every attribute you hold: badge type, country, first visit or returning, company size. Each gap is a candidate source of bias. Where you also know how the attribute relates to satisfaction, you can convert the gap into an estimate of how far the reported mean is off.
- Which attendee attributes usually skew a post show survey sample?
- Returning attendance is the reliable one, because people with a habit of coming back also have a habit of replying. International share usually skews the other way, since the invitation arrives while those attendees are still travelling. Badge type skews wherever exhibitor staff and press are on the same send as visitors, since their reply rates differ sharply.