Testing entrance counting accuracy against a manual clicker count
Testing entrance counting accuracy means running a manual clicker count alongside the automated feed at one door for two hours, splitting both into 15 minute bins, then plotting the difference against the mean for each bin. That gives a bias figure and a range of disagreement to publish beside the count.
The overhead sensor at the north entrance reported 1,842 people between 09:00 and 11:00. The steward standing under it with a hand counter reported 1,795. That is a gap of 47, and the meeting that followed spent forty minutes on which of the two was right.
Neither question is answerable and neither needs to be. Entrance counting accuracy is a property of the pair, and the useful output is a statement of how far apart the two methods run and how much that gap varies. Once you have that, the sensor becomes a measurement with a stated band around it, which is a considerably more useful object than a number people argue about.
Why does correlation tell you nothing here?
Because two counting methods can agree perfectly in their movements and disagree completely in their values.
Bland and Altman made this argument in the Lancet in 1986, in a paper on assessing agreement between two methods of clinical measurement. A correlation coefficient measures whether two sets of readings rise and fall together. Feed it a sensor that reads exactly 40 higher than the clicker in every single bin and it returns a correlation of 1.00, which reads as perfect agreement and describes a systematic error of 40 per bin.
Their proposal was a different plot. For each paired observation, take the difference between the two methods and the mean of the two methods, and plot difference against mean. The average of those differences is the bias. The spread around it, expressed as the mean difference plus and minus 1.96 standard deviations, gives limits of agreement, the range holding roughly 95 per cent of individual differences.
That structure suits door counting exactly, because the two things you want to know are whether the sensor runs high or low, and by how much it varies from one quarter hour to the next.
Running the count so it is worth running
Two hours at one door, split into 15 minute bins, gives eight paired observations. Small, and enough.
Choose the window deliberately. Agreement degrades as the door fills, because overhead sensors merge people walking abreast and human counters lose track at exactly the same moments. A two hour window spanning both a peak and a lull tells you far more than two hours of calm, and a sample taken only at 14:00 will produce a bias figure that flatters the sensor and fails in the morning.
Have the counter face the same direction the sensor is counting. If the sensor counts inbound only, the person clicks inbound only, and somebody needs to decide in advance what to do about a person who steps out and back. Write that rule down before the count starts, because deciding it afterwards is how the two datasets end up measuring different things.
Two people counting the same door for the same two hours is better still. Their disagreement is your reference method's own error, and it is usually larger than people expect.
Eight bins feels thin and it is defensible for a first pass. The width of the limits of agreement is itself estimated with uncertainty, and that uncertainty shrinks slowly, so doubling to four hours narrows the interval around the limits by roughly thirty per cent while doubling the demand on the person holding the counter. Two hours at one door on each of three show days beats six hours on one morning, because it samples three different crowds and lets you see whether the bias is stable. Stability is the property that decides whether you can correct for the bias later.
The failure modes differ by sensor type, and knowing which one you are testing changes what you look for. An overhead counter over-counts under crowding, merging and splitting bodies. A badge reader under-counts under crowding, because tailgating rises and hurried presentations fail. If your validation returns a positive bias from a badge reader or a negative bias from an overhead counter, look hard at the count itself before believing it, because both are the opposite of the expected direction.
The numbers from one two hour window
| Bin | Sensor | Clicker | Difference | Mean | Difference as per cent of mean |
|---|---|---|---|---|---|
| 09:00 | 120 | 124 | -4 | 122.0 | -3.3 |
| 09:15 | 208 | 195 | 13 | 201.5 | 6.5 |
| 09:30 | 265 | 262 | 3 | 263.5 | 1.1 |
| 09:45 | 320 | 297 | 23 | 308.5 | 7.5 |
| 10:00 | 288 | 282 | 6 | 285.0 | 2.1 |
| 10:15 | 238 | 236 | 2 | 237.0 | 0.8 |
| 10:30 | 222 | 216 | 6 | 219.0 | 2.7 |
| 10:45 | 181 | 183 | -2 | 182.0 | -1.1 |
The totals are 1,842 against 1,795, a difference of 47 on a mean of 1,818.5, which is an aggregate bias of 2.6 per cent. The sensor runs high across the window as a whole.
The bin level analysis says something slightly different. The eight percentage differences average 2.0 per cent, with a standard deviation of 3.6. Limits of agreement are 2.0 plus and minus 1.96 times 3.6, which is 2.0 plus and minus 7.0, giving a range from minus 5.0 to plus 9.1 per cent.
Those two bias figures, 2.6 and 2.0, are both correct and they answer different questions. The aggregate weights each person equally, so the busy bins dominate it. The bin mean weights each quarter hour equally, so the quiet bins count as much as the peak. Report the aggregate when you are talking about a daily total and the bin mean when you are talking about operational decisions taken bin by bin.
The pattern in the table is worth as much as the summary statistics. The two largest positive differences, 13 and 23, both fall in the busiest bins, and the two negative differences fall in the quietest. That is a sensor over-counting under crowding, which is the classic overhead counting failure where two people walking close together are resolved as three.
How wide can the limits be before the count is unusable?
It depends on the decision, and stating the decision first prevents an argument about whether 9 per cent is acceptable in the abstract.
For a daily attendance figure, wide limits matter less than the bias does, because the bin errors partly cancel across 32 bins. A 2.6 per cent bias on a daily count of 12,000 is roughly 310 people, and it is systematic, so it does not cancel at all. That is the number to correct for or disclose.
For an operational trigger, the limits are the whole story. If your rule is to open two more lanes when a bin exceeds 300 people, and your limits of agreement run to plus 9 per cent, then a true count of 280 can read as 305 and your rule fires on noise. Set the trigger with the band in mind, or use two consecutive bins to fire it.
For a year on year comparison, both matter and a third thing matters more, which is whether the sensor was the same and mounted in the same place. A 2.6 per cent bias that is stable across editions cancels in the comparison. A bias that changed because somebody moved the mounting height by 300 millimetres does not.
Lin published a single-number alternative in Biometrics in 1989, the concordance correlation coefficient, which combines how tightly the pairs sit around the line of equality with how far that line is displaced. It is useful for tracking one door across many editions in a single series. For the first conversation with an operations team, the plot works better, because it shows where in the day the disagreement lives.
What to do with the answer
Publish it. A door with a stated bias and a stated band is a door somebody can reason about, and it stops the next debrief repeating the same forty minute argument.
Correcting the count is a separate decision and I would be careful with it. A stable, well-measured bias can reasonably be corrected, provided the correction is applied to the published figure with a note saying so and the raw sensor number is retained. An unstable bias should be disclosed and left alone, because a correction derived from one two hour window in one weather condition will be wrong in a way nobody can trace later.
Where this test fits with everything else is straightforward. Gate coverage tells you which doors are measured at all. This test tells you how well the measured ones measure. The rules that turn reads into attendances sit in entry scan deduplication, and the exercise of making three systems agree on a total is reconciling badge scan counts. All four feed the same attendee analytics report and each one can quietly ruin it.
Where this stops
The clicker is a reference, and it is not truth. A person counting a busy door for two hours drifts, and two people on the same door will differ by one or two per cent from each other. Every limit above is really the disagreement between two imperfect methods, which is why the output is a band rather than a correction factor.
The test also measures one door for two hours. Doors differ, and the same door differs by day. A validation on the main entrance on Tuesday morning says nothing reliable about the side entrance on Thursday afternoon, and generalising it is exactly the move that produces a confident, wrong correction applied across a whole show.
Start this week by scheduling one two hour count at your busiest door on the next show morning, with a written rule about direction and re-entries agreed beforehand. Eight paired bins is enough to know whether your entrance counts are running high or low, and you will never get a cheaper answer.
Questions people ask about entrance counting accuracy
- How long should a manual validation count run?
- Two hours at one door gives eight 15 minute bins, which is enough to estimate a bias and a spread without exhausting the person holding the counter. Split the window so it covers both a busy period and a quiet one, because agreement usually degrades as the door fills and a quiet-only sample will flatter the sensor.
- Why is a correlation coefficient the wrong test for counting accuracy?
- Because correlation measures whether two sets of readings move together, and two counts can move together perfectly while one is consistently forty higher. Bland and Altman made this argument in the Lancet in 1986 and proposed plotting the difference between the two methods against their mean, which shows both the offset and its spread.
- What do limits of agreement mean for an attendance number?
- They give the range within which about 95 per cent of individual bin differences fall. A band from minus five to plus nine per cent says that any single quarter hour count could reasonably sit anywhere in that range, which matters for operational decisions taken bin by bin much more than for a daily total.
Related reading
- Gate coverage for attendance counting and the doors everyone forgets
- An entry scan deduplication rule that survives a busy hall door
- Reconciling badge scan counts when three systems each claim a different total