Why wifi footfall counting limits have got worse since MAC randomisation
Wifi footfall counting infers presence from probe requests, which used to carry a stable hardware address per device. Apple randomised those addresses in iOS 8 during 2014 and Android made randomisation the default in version 10 during 2019, so one phone now emits many addresses across a day and unique device counts inflate accordingly.
Somebody presents a hall traffic chart built from wifi sensors and the y axis says 84,000 unique devices for a show that registered 9,840 people. Nobody in the room believes the number, and nobody can say by how much it is wrong. The wifi footfall counting limits behind that chart are well documented and entirely mechanical, and once you know the mechanism you can decide what the sensors are still good for.
The short version is that the method was built on an identifier that phones deliberately stopped providing.
What probe request counting was originally measuring
A phone looking for known networks broadcasts probe requests. For most of the last two decades those frames carried the device's real, globally unique hardware address. A passive sensor sitting in a hall could log every frame it heard, take the distinct addresses, and call that a device count. Divide by an assumed devices-per-person ratio and you had footfall.
The method was cheap, it needed no cooperation from the attendee, and it worked. It also worked rather too well from a privacy point of view, because the same stable identifier that lets you count a crowd lets you follow one person from a shopping centre to a station to an office over months.
That is the property the platform vendors removed.
What changed, and when
Martin and colleagues published the first wide-scale study of the practice in 2017, in the Proceedings on Privacy Enhancing Technologies. Their account of the timeline is the one to hold on to. In late 2014, Apple introduced MAC address randomisation with the release of iOS 8.0. Android added randomisation for probe requests with Android 6.0, plus an incremental patch to 5.0.
Android having the capability and Android using it turned out to be different questions. The same 2017 paper found that the overwhelming majority of Android devices were not implementing the randomisation capability already built into the operating system. Google closed that gap with a default. The Android developer documentation on privacy changes states it flatly: on devices that run Android 10 or higher, the system transmits randomised MAC addresses by default. Android 10 shipped in 2019.
So there are two dates on the chart, five years apart, and the effect on your counts came in two waves. A sensor deployment commissioned in 2016 and never revisited was measuring something quite different by 2021, and nothing in the reporting would have flagged the change, because the count kept going up.
Why does one phone now look like thirty-two people?
Work the arithmetic on a single attendee.
Suppose a phone rotates its randomised address every 15 minutes while it is unassociated and scanning. Your show floor is open for eight hours. Eight hours is 480 minutes, and 480 divided by 15 is 32. That one person contributed 32 distinct addresses to the day's list, and a naive distinct count records 32 devices.
Now scale it. Take 5,000 phones on the floor behaving that way. The distinct address count is 5,000 times 32, which is 160,000. Your actual device population was 5,000. The count is 32 times too high, and the multiplier is set by a rotation interval that the phone chooses and never tells you.
Those 15 minutes are an illustration rather than a measured constant, and that is the real problem. The inflation factor depends on the operating system, the version, whether the screen is on, whether the device is associated to an access point, and how long the person stayed. All of those vary across your audience and across editions of your show. A counting error you cannot bound is worse to work with than a large error you can, because you cannot even state a direction of travel year on year with confidence.
The devices-per-person assumption compounds it. Even before randomisation, converting devices into people meant dividing by an assumed ratio, and at a B2B exhibition that ratio has never been one. A sales engineer carries a phone, a tablet and a laptop that all probe. A visitor from a delegation carries a personal phone and a company phone. If you assume 1.4 devices per person and the true figure on your floor is 1.9, your headcount is 36 per cent high on that account alone, and it compounds with the randomisation multiplier rather than cancelling it. Two unmeasured factors multiplied together give you a number with no defensible interpretation.
There is a second, opposite bias sitting underneath the first. Martin and colleagues also showed that passive identification techniques defeated randomisation in roughly 96 per cent of the Android phones they examined, largely through information leaked in the probe frames themselves. Vendors have narrowed those leaks since. So a sensor platform doing sophisticated de-randomisation in 2017 was doing better than a naive count, and the same platform doing the same thing today recovers a smaller share, which means its counts drift over time in a direction nobody is monitoring.
The fixes people try, and what each one costs you
Four approaches show up in vendor documentation and in practice, and each trades one problem for another.
Fingerprinting the probe frame. Group addresses by the set of networks the device asks for, the information elements it includes, and the sequence numbers it uses. This is the technique the 2017 paper describes and it genuinely helps. It also degrades every time a platform closes another leak, so your historical series stops being comparable with your current one.
Counting only associated devices. Ask people to join the venue wifi, then count sessions on the controller. Addresses are stable within an association, so the count is clean. Association rates at exhibitions are low enough that you are now measuring the subset of people willing to join a captive portal, which skews young, skews international roaming, and skews towards anyone who wanted the free bandwidth.
Counting frames rather than devices. Give up on a unique count and report probe volume as an index. It moves with the crowd and it is internally consistent, so it is genuinely useful for shape. It cannot be turned into an attendance number, and somebody will eventually try.
Calibrating against a known count. Take your verified badge entries for the same period and fit a ratio. This is the most defensible of the four, because it anchors the sensor to a measurement, and it needs redoing every edition since the ratio moves with the device mix.
I would use the third and the fourth together, and I would refuse to publish a unique device count from probe requests at all. An index that honestly says relative traffic beats a number that claims to be people and is wrong by an unknown multiple.
What is wifi sensing still genuinely good for?
Three things survive the randomisation problem, and they are worth keeping.
Shape over time holds up well. If probe volume in hall two peaks at 11:15 and again at 14:30, that pattern is real, because the inflation factor is roughly constant within a day and cancels when you compare one bin against another. Relative comparison across zones in the same hall holds up for the same reason, provided the sensors are the same model at the same mounting height.
Dwell distribution partly survives, since a device seen across a long span of bins was probably present for a long span, even when you cannot name it consistently.
What does not survive is any absolute count, any year-on-year comparison of that count, and any per-person metric built on it. If you need people rather than an index, the instrument has to be one that identifies a person by design, which is what badge portals do and what UWB positioning for events does at much higher spatial resolution and much higher cost. Matching the instrument to the decision is the whole of the people counting sensor question.
Where this stops
Everything above assumes you are allowed to do this at all, and that assumption deserves a sentence of its own. Passive collection of device identifiers in a public venue sits inside data protection law in most jurisdictions your show runs in, and the analysis you can defend depends on what your privacy notice actually says. That is a legal question with its own owner, and I am not going to answer it here.
The narrower limit is that no amount of method fixes the underlying signal. Probe requests are emitted at the discretion of a device you do not control, running an operating system that treats being counted as a bug it is actively fixing. Every release closes a little more of the gap. Any pipeline you build on de-randomisation is built on the residue of somebody else's incomplete privacy work, and it will keep getting thinner.
There is also a floor problem that gets forgotten. Sensors hear what is in range, and range depends on the hall, the crowd density and the mounting height, so a sensor in a full hall hears proportionally less than the same sensor in an empty one. That biases your busiest bins downwards at precisely the moment you care most about them, which partly offsets the randomisation inflation and partly does not, in a ratio nobody can compute for you.
Do this before the next edition. Take one hour from your last show where you have both a wifi count and a verified badge entry count for the same doors, and divide one by the other. Write that ratio on the report next to the wifi number, with the date it was measured. If the ratio is 8.5 to 1, everyone reading the chart now knows what they are looking at, and next year you will be able to tell whether the instrument moved or the audience did. The way that count then feeds the rest of your attendee analytics becomes a documented conversion instead of a silent one.
Questions people ask about wifi footfall counting limits
- Does wifi footfall counting still work with MAC randomisation?
- It still detects that devices are present and it still tracks the shape of a day, because probe volume rises and falls with the crowd. It no longer produces a trustworthy count of unique devices, because a single phone emits many distinct randomised addresses over a few hours and each one looks like a separate device to the counter.
- When did phones start randomising MAC addresses?
- Martin and colleagues record that Apple introduced MAC address randomisation with the release of iOS 8.0 in late 2014, and that Android added randomisation for probe requests with Android 6.0. Google's own Android documentation states that devices running Android 10 or higher transmit randomised MAC addresses by default, which dates the shift to 2019.
- How much does MAC randomisation inflate a footfall count?
- It depends entirely on the rotation interval and how long people stay. A device that presents a fresh address every 15 minutes across an eight hour day produces 32 addresses. A crowd of 5,000 phones behaving that way would register as 160,000 unique devices unless the counting method groups addresses back together.