Why heat map normalization changes which part of the hall looks busy
Normalising a heat map means dividing each zone's count by the zone's floor area and by the minutes that zone was open, giving a rate per square metre per hour. Raw counts reward large zones with long opening hours, so the ranking they produce is mostly a ranking of zone size.
Two zones come out of the sensor export at the top of the table. Zone A logged 8,400 pings over the three days and zone B logged 3,100. The slide ranks them by count, zone A wins, and the sales director starts talking about what the zone A stands should pay next year.
Zone A covers 900 square metres of a hall entrance concourse. Zone B is a 200 square metre feature theatre that opens at 11:00 and shuts at 16:00. Heat map normalisation is the step that makes those two numbers comparable, and once it is done the ranking flips.
The two denominators a zone count needs
A zone count answers a question nobody asked: how many detections happened inside this boundary. Two things drive that number before any attendee behaviour does.
The first is floor area. A bigger zone contains more floor, more stands and more people at any instant, so it accumulates more detections. The second is open minutes. A zone that is accessible for eight hours a day collects roughly twice what an identical zone collects in four.
So the comparable figure has both denominators in it: detections per square metre per open hour. Store the area and the open minutes as attributes of the zone in the same place the zone boundary lives, and recompute nothing at report time. Zones get redrawn between editions and nobody remembers which version the areas came from.
Working two zones through
Zone A first. It holds 8,400 pings across 900 square metres, which is 8,400 divided by 900, or 9.33 pings per square metre. The hall opens 10:00 to 18:00 on all three days, so zone A was open for 24 hours. That gives 9.33 divided by 24, which is 0.389 pings per square metre per hour.
Zone B now. It holds 3,100 pings across 200 square metres, which is 15.5 pings per square metre. Already the smaller zone is denser by two thirds, on area alone. The theatre opened 11:00 to 16:00 on each of the three days, so 15 hours. That gives 15.5 divided by 15, which is 1.033 pings per square metre per hour.
Compare the two endpoints. On raw counts zone A beat zone B by 8,400 to 3,100, a factor of 2.71. On the normalised rate zone B beats zone A by 1.033 to 0.389, a factor of 2.66. The ordering reversed and the magnitude barely changed, which is what happens when the raw comparison was measuring size and hours the whole time.
Neither number is wrong. They answer different questions. The raw count answers how much footfall the concourse absorbed in total, which matters for stewarding and for cleaning. The rate answers how hard a square metre of that floor worked, which is the question a rate card conversation is actually about.
Why does the raw count always favour the big zone?
Because area sits in the numerator by stealth. Detections arrive in proportion to how many people are present, and how many people are present is roughly proportional to floor area at any given density. Rank by count and you have ranked by area with a small behavioural signal on top.
The size of that effect is easy to see in the same hall. Across all zones the hall holds 54,000 pings over 4,800 square metres of measured floor, so the hall runs at 11.25 pings per square metre, or 0.469 per square metre per hour over the 24 open hours. Zone A at 0.389 is below the hall rate. It was the busiest zone in the hall by count and it is a below-average zone by rate, and the only thing that changed is the denominator.
Open minutes do the same thing in the time dimension, and they bite hardest on feature areas, theatres, catering and anything that runs to a programme. Those are exactly the parts of the floor an organiser is trying to prove the value of.
Small zones produce unstable rates
Dividing by a small denominator buys the comparison at the price of precision, and this catches people out every year.
Anselin sets the problem out plainly in the rate mapping material for GeoDa, revised in 2018. A crude rate is the count of cases over the population at risk, and the variance of that estimate depends on the denominator: the larger the population, the smaller the variance, so rates estimated from small populations may carry a large standard error. Substitute floor area for population and the argument transfers without modification.
Take zone F, a 40 square metre meeting point holding 210 pings over the full 24 hours. That is 5.25 per square metre, or 0.219 per square metre per hour. Now suppose one group of eight people stands there talking for a few minutes and adds 40 pings. The zone reads 250, which is 6.25 per square metre, or 0.260 per hour. A 19 per cent move in the rate from one conversation.
For zone A to move 19 per cent you would need 0.19 times 8,400, which is 1,596 extra pings. One small group can move the small zone as far as sixteen hundred detections move the big one, and a table sorted by rate will put the small zone wherever that group happened to stand.
The standard correction is the one Anselin describes: shrink each zone's rate towards the overall rate, with the weight given to the zone's own rate rising with the size of its denominator. Set the shrinkage constant to 300 square metres for illustration, and zone F gets a weight of 40 divided by 340, which is 0.118. Its smoothed rate is 0.118 times 0.219 plus 0.882 times 0.469, which is 0.440, almost the hall rate. Zone A gets a weight of 900 divided by 1,200, which is 0.75, and its smoothed rate is 0.409 against a raw 0.389. The small zone is pulled a long way and the large one hardly moves. A real implementation estimates the constant from the data instead of picking it, and the behaviour is the same.
Is that hot zone real?
Rates and shrinkage tell you how confident to be in a single zone's number. They do not tell you whether a cluster of adjacent hot zones is a pattern or an accident, and that is a separate test with a separate tool.
Anselin published the general form in Geographical Analysis in 1995, in a paper on local indicators of spatial association running from page 93 to page 115 of volume 27. The idea is that a global measure of spatial association, such as Moran's I, can be decomposed into one contribution per observation, and each of those local contributions can be assessed on its own as an indicator of a local pocket of nonstationarity, which is what a hot spot is. Anselin's paper places the family alongside the Gi and G star i statistics that Getis and Ord published in 1992, and evaluates the local Moran in detail.
The practical shape of it on a hall is a permutation test. Hold zone B's value fixed, reshuffle the other zones' values across the hall at random, recompute the local statistic, and repeat 999 times. The pseudo p value is the number of reshuffles at least as extreme as the real one, plus one, over 1,000. If zone B's actual value beats every reshuffle, the pseudo p is 0.001, which is the smallest that test can produce.
Two cautions come with it. The test needs a neighbour definition, and on an exhibition floor that means deciding whether two zones separated by a solid stand build are neighbours at all. And with 22 zones each tested at the 5 per cent level, you would expect 22 times 0.05, which is 1.1 apparent hot spots, from randomness alone. Finding one hot zone in a hall of 22 is not a finding.
Where this stops
Normalisation makes zones comparable to each other. It does nothing about whether the underlying detection process is comparable, and on a real floor it often is not. A zone under a mezzanine, a zone with a receiver behind a truss, and a zone that happens to sit near the wifi rack all collect detections at different rates for reasons that have nothing to do with attendees. Dividing a biased count by a correct area gives you a precise biased rate.
The honest guard is a coverage attribute per zone, recorded on the day and reported next to every rate. If you cannot say what fraction of the zone had a clean line to a receiver, the rate is a number with an unknown multiplier in front of it.
The second limit is that the rate has no direction and no dwell in it. Two zones at 1.0 pings per square metre per hour can be a corridor everybody crosses quickly and a stand cluster where fewer people stay much longer, and nothing in the normalised rate separates those. If the decision turns on that difference, the rate is the wrong instrument.
Where this fits with its neighbours is straightforward. The gridded picture, with its bin size and colour scale, is a different object with its own four choices in C15. Ranking aisle runs by traffic per linear metre against the hall median, and diagnosing why the quiet ones are quiet, is the dead zone question in C16. And an instantaneous crowding decision needs persons per square metre over walkable area in C11, which is a different denominator again. All three want the same zone map, and fixing that map once is the highest-value hour in the whole attendee analytics setup.
Start with a spreadsheet of three columns: zone name, area in square metres, and open minutes. Fill it in for last edition from the floorplan and the programme, join it to your zone counts, and re-sort the table. If the top three zones change, every ranking you presented last year was a ranking of zone size.
Questions people ask about heat map normalization
- How do you normalise a show floor heat map?
- Divide each zone's total count by the zone's floor area in square metres, then divide again by the number of hours that zone was open to attendees. The result is a rate per square metre per hour, which is comparable across zones of different sizes and different opening patterns. Store the area and the open minutes as attributes of the zone.
- Why does a big zone always look busy on a heat map?
- Because a raw count is a count of everything that happened inside a boundary, and a larger boundary contains more floor, more stands and more people. A zone of 900 square metres holding 8,400 pings is at 9.3 per square metre, while a zone of 200 square metres holding 3,100 is at 15.5, so the smaller zone is denser by half again.
- How do you know a hot zone is a real hot spot?
- Test it against the pattern you would get by chance. Anselin's 1995 local indicators of spatial association compare each zone's value with a reference distribution built by reshuffling the values across zones many times. With 22 zones tested at a 5 per cent level you would expect about one apparent hot spot from randomness alone.
Related reading
- Show floor heat mapping that an operations lead can defend in a debrief
- Finding dead zones on a show floor before the exhibitors in them complain
- Turning aisle traffic density into a number your operations team can act on