People counter accuracy validation you can run in a single morning
People counter accuracy validation means treating the sensor as a measurement instrument and separating its own repeatability from the reproducibility of whoever supplies the reference count. Two observers and three time windows on one door take a morning, and they tell you whether the disagreement you are seeing belongs to the sensor or to the manual count you are testing it against.
An operations lead stands at hall three with a clicker for twenty minutes, gets 402, pulls the sensor number for the same twenty minutes, gets 404, and reports that the counter is accurate. A month later somebody else does the same test at a different door, gets a 9 per cent gap, and the sensor gets a reputation.
Both tests were too small to support the conclusion drawn from them. Proper people counter accuracy validation costs about the same effort and answers a better question, because it measures the reference count as well as the sensor.
The reference count is a measurement too
Everyone treats the clicker total as truth. It is a reading produced by a tired person watching a busy doorway, and it has its own error, which is usually larger than anyone expects and almost never quantified.
Think about what an observer at a hall door is actually being asked to do at 09:20 on day one. Groups arrive four abreast. Two people step in, change their minds and step back out. A contractor pushes a trolley through the same doorway. Somebody stands in the threshold talking on the phone for ninety seconds. Every one of those is a judgement call, and two observers standing side by side will make different calls.
If you only ever run one observer, that variation is invisible and it silently gets assigned to the sensor.
The measurement world solved this problem long before anybody counted a trade show visitor. ISO 5725-1, first published in 1994, separates two conditions. Repeatability conditions hold the method, the item, the laboratory, the operator and the equipment fixed and vary only time. Reproducibility conditions vary the operator and the setting. The standard treats the two as the extremes of precision, the first giving the minimum variability you can expect and the second the maximum.
Applied to a hall door, repeatability is the counter doing the same thing twice. Reproducibility is two people disagreeing about what happened.
The design, which fits inside one morning
One door. Two observers. Three windows of 20 minutes, spread across the arrival profile so that the sensor is tested at different traffic levels.
The observers stand where they can both see the full width of the doorway, they do not confer, and they agree the counting rule in advance in writing. The rule matters more than the diligence. Write down what happens with a person who enters and immediately leaves, with staff and contractors, with a wheelchair and its pusher, and with somebody carrying a child. If the two observers apply different rules, you are measuring the rulebook rather than the door.
Then pull the sensor rows for those exact timestamps. Not the hour containing them. The minute-level rows, clipped to the same start and stop seconds the observers used, because a 60 second misalignment at 27 arrivals a minute is a 27 count error that has nothing to do with the sensor at all.
This is a simplified field version of the gauge repeatability and reproducibility study described in the AIAG Measurement Systems Analysis manual, fourth edition of 2010, which is the standard reference for this kind of work in manufacturing. A full study uses more parts, more trials and a proper variance decomposition. A door gives you three windows and two appraisers, and it still separates the two sources of error, which is the point of the exercise.
What does the arithmetic actually tell you?
Here is a set of results shaped the way real ones come out.
Window one, 09:00 to 09:20. Observer A counts 402, observer B counts 378, the sensor logs 404. Window two, 09:20 to 09:40. Observer A counts 551, observer B counts 519, the sensor logs 553. Window three, 09:40 to 10:00. Observer A counts 468, observer B counts 440, the sensor logs 470.
Start with the observers, because they are the part nobody usually measures. In window one the gap is 24 on a mean of 390, which is 6.2 per cent. Window two, 32 on 535, 6.0 per cent. Window three, 28 on 454, 6.2 per cent. Across the morning, A totals 1,421 and B totals 1,337, a gap of 84 on a mean of 1,379, which is 6.1 per cent.
Now the sensor, against the mean of the two observers. Window one, 404 over 390 is 1.036, so the sensor reads 3.6 per cent high. Window two, 553 over 535 is 1.034, 3.4 per cent high. Window three, 470 over 454 is 1.035, 3.5 per cent high.
Read those two lines together and the diagnosis writes itself. The sensor's bias sits at 3.5 per cent across a range of traffic levels, moving by two tenths of a percentage point between the quietest and busiest window. That is a stable instrument with a known offset. The observers disagree with each other by 6.1 per cent, nearly double the sensor's bias, and the disagreement does not shrink when the door gets quieter.
There is a trap buried in these numbers that is worth pointing at. When two observers differ by 6 per cent, the higher one sits about 3 per cent above their mean. A sensor with a 3.5 per cent high bias therefore lands almost exactly on observer A in all three windows. Had you run this test with observer A alone, you would have concluded the counter was accurate to within half a per cent and gone home.
Reading the result against a published band
The AIAG manual sets acceptance bands that transfer to this problem better than anything the sensor industry publishes. Variation under 10 per cent is acceptable. Between 10 and 30 per cent is acceptable depending on the application and the cost of the measurement. Above 30 per cent the measurement system is unacceptable. The manual pairs that with a second test, the number of distinct categories, which should be five or more, and both have to pass.
By those bands, a 3.5 per cent instrument bias and a 6.1 per cent appraiser spread are both inside the acceptable range, and the honest report says the counter is fit for the decisions you make with it. It also says something more useful. Your reference method is the weaker of the two instruments, so tightening the sensor further buys nothing until you tighten the counting rule the observers work to.
That conclusion changes how you spend the next hour of effort. Rewrite the rule, run two observers on it again, and see whether the 6.1 per cent comes down. If it does, you have improved your ability to audit every door you own. Choosing which door to instrument in the first place, and with what, is a different question with its own trade-offs across people counting sensors.
Why do two observers on the same door disagree?
Three causes account for most of it, and knowing which one you have tells you what to fix.
Group arrivals are the largest. A pair walking shoulder to shoulder through a 1.8 metre doorway gets counted as two by one observer and, in the split second the second person is hidden, as one by the other. This gets worse as traffic rises, which is why it damages exactly the windows you care about.
Threshold ambiguity is the second. Somebody steps in, turns around to wait for a colleague, steps out and comes back. That is one entry, three entries, or a net of one depending on the rule, and it is common at the top of the morning.
The third is fatigue. Observer accuracy falls after about fifteen to twenty minutes of continuous counting, which is why the design uses 20 minute windows with breaks between them rather than one continuous hour.
None of this applies to a passive method that never sees a person. Counting devices instead of bodies has a different failure entirely, since address randomisation broke the link between an identifier and a person, and the wifi footfall counting limits that result cannot be fixed by a better observer.
Where this stops
A morning at one door tells you about that door, at that traffic level, in that light, with that geometry. It does not transfer. A counter over a 3 metre entrance with an atrium behind it and a counter over a 1.2 metre side door behave differently, and the second one is usually worse because groups have to compress to get through.
The method also assumes the sensor's error is a proportional bias, which the worked example above happens to satisfy. Some counters instead saturate, holding a ceiling once flow passes some threshold. A validation run entirely inside quiet windows will report a clean 3 per cent bias and completely miss a ceiling that only bites at 09:15 on day one. Put at least one window inside your actual peak, even though that is the hardest one to staff.
And a single morning gives you three data points, which is enough to see a consistent bias and not enough to put an interval around it. Treat the output as a direction rather than a calibration constant, and repeat it at the next edition before you trust the number to move.
The step for this week is the rulebook. Write the counting rule for your doors on one page, covering groups, re-entries, staff, contractors and pushchairs, and have two people count the same twenty minutes against it at your next event of any size. If they land within 2 per cent of each other, you have a reference you can audit sensors with. If they land 10 per cent apart, fix that before you buy anything, because everything downstream in your attendee analytics inherits it.
Questions people ask about people counter accuracy validation
- How do you check whether a people counter is accurate?
- Put two independent observers on the same door for the same short windows, log the sensor for those exact windows, and compare all three. Running two observers instead of one lets you measure how much the reference count itself varies. Without that, any gap you find gets blamed on the sensor by default, which is often wrong.
- What counts as an acceptable error for a people counter?
- The AIAG Measurement Systems Analysis manual, fourth edition of 2010, sets bands used across manufacturing: under 10 per cent variation is acceptable, 10 to 30 per cent is acceptable only depending on the application and the cost of measuring, and above 30 per cent is unacceptable. Those bands transfer sensibly to door counting.
- How long does a people counter validation take?
- Three 20 minute windows on one door, staffed by two observers, is one hour of counting inside a two hour block. Add half an hour to pull the sensor data for exactly those timestamps and half an hour for the arithmetic. One morning covers one door properly, which is more useful than a shallow test across six.