Exhibitor lead volume distribution is skewed and your averages are hiding it
Exhibitor lead volume distribution is right-skewed on almost every show floor, because space, staffing and scanner counts compound together. A mean computed across it sits above most exhibitors, so publish the five-number summary, minimum through maximum, and the share of total scans held by the top decile of stands.
The number on the slide was 310. Average leads per exhibitor, second bullet of the post-show deck, and nobody in the room queried it.
Two weeks later a stand manager from a 36 square metre booth in the fluid power aisle emailed to ask why his team had come away with 74 when the show apparently averaged 310. He wanted to know what he had done wrong. He had done nothing wrong, and the number on the slide was never about him. Exhibitor lead volume distribution is skewed hard to the right on every floor I have looked at, so the mean sits above most of the stands that produced it.
What is that average leads per exhibitor made of?
Take a floor of 640 exhibiting companies and 198,400 unique leads across the four days. Divide one by the other and you get 310. That arithmetic is correct and the figure it produces is close to useless for the person reading it.
Lead counts across a show floor are not symmetric around a central value. They pile up at the low end and stretch a long way out at the high end, because the things that drive lead volume all compound. Space begets footfall, footfall begets scans, and the large stands that draw the most traffic are also the ones with the most staff holding the most scanners for the most hours. A 400 square metre pavilion with fourteen people on it does not produce eleven times the leads of a 36 square metre stand with two. It produces something more like twenty times, and it sits in the same average.
Sort the 640 stands by unique leads and split them into ten bands of 64. On a floor with this shape, the top band holds around 47 per cent of every unique lead captured. Sixty-four companies, just under half the total. The bottom band holds about half of one per cent, which works out at roughly 16 leads each.
That top band pulls the mean upward and leaves it stranded a long way above the middle of the floor. In this file the median is 168. The mean is 310, or 1.85 times the median, and the mean falls inside the eighth decile, which means roughly 460 of the 640 stands sit below the number you printed. Seventy-two per cent of your exhibitors read the average and correctly conclude they are behind it.
Five numbers instead of one
Tukey set out the five-number summary in his 1977 book on exploratory data analysis, and it is still the cheapest upgrade available to a post-show deck. Minimum, lower hinge, median, upper hinge, maximum.
For the floor above:
Minimum 1, lower quartile 74, median 168, upper quartile 355, maximum 4,912.
That takes the same amount of slide real estate as the mean and tells a reader something the mean cannot. The interquartile range is 281, which is wider than the median itself. Half your exhibitors landed somewhere between 74 and 355 unique leads, and the stand manager who wrote to you is sitting exactly on the lower quartile. He is at the boundary of the bottom quarter of the floor, which is a real finding and a very different conversation from being 236 leads below average.
Publish the quartiles and the arithmetic stops being adversarial. An exhibitor can locate themselves. Nobody has to take your word for what typical means.
The maximum deserves a word of caution. A single value of 4,912 is one stand, and printing it invites every other exhibitor to measure themselves against a company that bought 400 square metres and staffed it with fourteen people. If you show the maximum at all, show the space alongside it. Otherwise report the 90th percentile, which on this floor is around 690, and which at least describes a band of sixty companies instead of one.
The share held by the top decile
The decile share is worth computing explicitly, because it is the number that tells your commercial team how concentrated the floor is, and concentration is a portfolio risk that lead averages conceal entirely.
Lorenz published the cumulative share curve in the Publications of the American Statistical Association in 1905 to describe wealth distribution, and Gini published the concentration ratio derived from it in his 1912 book on variability and mutability. The machinery transfers to a show floor without modification, because the question is identical: what share of the total is held by what share of the population.
Work the cumulative shares from the bottom decile upward: 0.5, 1.7, 3.8, 7.0, 11.5, 17.6, 25.8, 36.8, 53.0, 100.0.
The grouped-data approximation to the Gini coefficient is one minus the average of each consecutive pair of cumulative shares. Add the ten pairs: 0.005, 0.022, 0.055, 0.108, 0.185, 0.291, 0.434, 0.626, 0.898, 1.530. That sums to 4.154. Divide by ten to get 0.415, and subtract from one.
The Gini coefficient is 0.58. Because the calculation groups exhibitors into deciles it loses the spread inside each band, so the true figure on the underlying rows will be a little higher.
A number in that region says something concrete. Roughly half your lead capture depends on sixty-odd companies, and if eight of them buy less space next edition, the floor total moves in a way that no amount of recruitment at the small end will offset. That is a sales planning fact, and it never appears in a mean.
The coefficient is also comparable across your own portfolio in a way that lead counts are not. A vertical show with 200 exhibitors and a national trade fair with 1,400 cannot be compared on average leads, because the audiences and the categories have nothing in common. They can be compared on concentration, because concentration is unit-free. If one show in the portfolio sits at 0.44 and another at 0.71, the second one is carrying a structural dependence on a small group of accounts, and that is worth knowing before the renewal season rather than after it. Track the figure across editions of the same show too. A concentration measure that climbs three years running is telling you the small end of the floor is thinning out, whatever the exhibitor count says.
Why does the mean keep moving when the median does not?
Run the same calculation on the previous edition and the mean will have jumped or dropped by an amount that has very little to do with how the floor performed.
One large pavilion switching from a single shared scanner to eight individual licences can add three or four thousand unique leads on its own. Divided across 640 exhibitors that is five or six points on the mean, and there is a real temptation to report it as growth. Nothing changed about the show. One exhibitor changed how it collected.
The median absorbs that shock because it only cares about the rank of the stand in the middle, and the stand in the middle is nowhere near the pavilion. The upper quartile moves a little. The mean moves a lot. For year-on-year reporting on lead volume I would put the median first, the quartiles second, and the mean nowhere at all, on the grounds that a statistic which can be moved five points by one exhibitor's procurement decision is not measuring the thing the sentence claims it is measuring.
The exception is when you genuinely want a total. Total unique leads across the floor is a real number and belongs in the report. The problem starts when you divide it by the exhibitor count and present the result as what an exhibitor can expect.
What to send an exhibitor
Whatever else sits in your exhibitor analytics pack, the per-stand page is the one that gets read. The CEIR 2015 study on exhibitor ROI and performance metric practices found lead generation ranking as the top objective for exhibiting, with brand awareness and reinforcement second, and it documented wide variation in how exhibiting companies measure any of it. That variation runs both ways. If the exhibitor's own measurement is inconsistent and yours is a single average, neither party has a defensible position in the renewal conversation.
What travels well in an exhibitor report is a rank statement with the population attached. Your stand captured 240 unique leads. Within the 47 stands in your product category, that places you 18th, at the 63rd percentile. The category median is 205.
Three properties make that work. It is checkable, because the exhibitor knows roughly who else was in their aisle. It is stable, because moving one large pavilion does not change a percentile. And it does not require you to publish anybody else's raw numbers, which matters, since exhibitors will ask and you cannot tell them.
The construction of those peer groups is its own problem, and a percentile taken over nine companies is not worth printing. That sits with E24.
Where this stops
Everything above assumes the lead counts you are ranking mean the same thing across stands, and they do not.
An exhibitor with two scanners at a demo station will re-scan the same badges repeatedly, which inflates total scans and, depending on how you deduplicate, unique leads too. What a healthy ratio of scans to unique leads looks like is E12's question. An exhibitor collecting business cards in a bowl and typing them up afterwards produces a lead count of zero in your system and several hundred in their CRM. Both sit in the same distribution, and the second one will be at the bottom of your ranking while genuinely having had a good show. Before any of this, count the stands that captured nothing at all, which is E3's ground, because a zero in the file and a zero on the floor are different findings.
So the distribution you are describing is the distribution of captured and recorded leads, which is a measurement of scanning behaviour at least as much as it is a measurement of commercial outcome. That does not make it worthless. It makes the header on the chart important. Label it captured leads, not leads, and say so out loud in the report, because the exhibitors who know the difference are the ones who will trust the rest of the pack.
The other limit is the tail itself. With 640 exhibitors, the top decile is 64 companies and the quantile estimates out at the 95th percentile and beyond rest on a handful of observations. Report the top of the distribution as a range or a count, and resist the urge to publish a 99th percentile figure computed from six stands.
Take last edition's export, sort every exhibiting company by unique leads, and write down five numbers: the minimum, the lower quartile, the median, the upper quartile and the maximum. Then count how many stands fall below the mean you published. If that count is more than half the floor, the mean should not have been in the deck.
Questions people ask about exhibitor lead volume distribution
- Why is average leads per exhibitor a misleading number?
- Because lead volume is not symmetric around a central value. On a floor of 640 stands capturing 198,400 unique leads, the mean is 310 while the median is 168, and the mean falls inside the eighth decile. Roughly 460 of the 640 exhibitors sit below the figure you printed, so most of your floor reads the average and correctly concludes it is behind.
- What should an exhibitor report show instead of the average?
- The five-number summary, which takes the same space on a slide. Minimum, lower quartile, median, upper quartile and maximum. An exhibitor on 74 unique leads can then see they are sitting exactly on the lower quartile of the floor rather than 236 leads below a mean that was never about them.
- How do you measure how concentrated lead capture is across a floor?
- Sort every exhibitor by unique leads, split them into deciles, and compute the cumulative share held by each. The top decile share is the headline number. A Gini coefficient computed from the same cumulative shares gives you one unit-free figure that is comparable across shows of very different sizes and across editions of the same show.
Related reading
- Lead capture coverage rate shows how many exhibitors captured nothing at all
- The scan to unique lead ratio and what a healthy number looks like