Skip to content

How to read industry benchmarks when your show sits outside the sample

Standards and researchUpdated 2026-08-239 min read

In short

A published industry benchmark describes the reference class its publisher sampled. Test your show against that class on geography, sector, size band and data vintage before quoting the figure. Where your show falls outside it, the comparable set is your own editions, and five of them yield a mean growth rate and a spread worth testing this year against.

The planning meeting has one slide with an outside number on it. Sector growth for the year, a single figure, no sample attached. Your show sells 4,240 square metres of net space in a regional hall, and the figure came from a study of events that fill three halls in a capital city. Somebody asks why you are behind.

Learning how to read industry benchmarks starts with a membership question, and the membership question is usually answerable in about fifteen minutes. Is a show like yours in the set that number was computed from? If the answer is no, the gap on the slide is measuring the distance between two populations, and no amount of work on your show will close it.

A published benchmark is somebody else's reference class

Bent Flyvbjerg set out the useful frame in Project Management Journal in 2006, writing about cost and demand forecasts for public infrastructure. His method, reference class forecasting, works by taking what he calls an outside view. "The outside view on a given project is based on knowledge about actual performance in a reference class of comparable projects" (Flyvbjerg, 2006).

The first of his three steps is the one that matters here. Flyvbjerg writes that the method requires "Identification of a relevant reference class of past, similar projects. The class must be broad enough to be statistically meaningful but narrow enough to be truly comparable with the specific project" (Flyvbjerg, 2006).

Read that as a test with two failure modes and it describes most of the benchmark arguments I have sat through. A class too narrow has four shows in it and no statistical content, so any average it produces moves with whichever of the four had a bad year. A class too broad averages your 4,240 square metre regional show with a 90,000 square metre international fair, and the result describes neither of them.

A published industry benchmark has already made that trade, for the publisher's purposes. The publisher was answering a question about a market. You are answering a question about one show in one city. Those two questions want different classes, and nothing guarantees that the published class contains you.

Is your show inside the sample or outside it?

Four tests, in the order that eliminates fastest.

Geography first, because it is quickest to check: which markets supplied the units. A national study is a national study whatever the cover says. Then sector, at whatever level of grouping the publisher used, since healthcare exhibitions and machine tool exhibitions behave differently enough that a combined average is a weak guide to either. Then size band, which is the test people skip and usually the one that disqualifies them. Last, vintage, meaning which year the underlying data describes as opposed to the year the report came out, a gap that is often two years wide and always worth pinning down before anyone builds a budget on it.

Size band deserves the attention because the publishers themselves take it seriously. CEIR, the official research division of the International Association of Exhibitions and Events, lists an Organizer Performance Benchmarking Study on its 2026 research agenda and splits the output into separate reports by size: large at 200,000 net square feet and above, midsized at 50,000 to 199,000, and small below 50,000, with healthcare and independent organiser events handled apart from those (CEIR, 2026). A research body that publishes one benchmark per size band has already concluded that one number will not cover all of them.

Put your own show through that scheme before assuming nothing fits. A show at 4,240 square metres is 4,240 times 10.7639, which is 45,639 square feet, so it lands inside CEIR's small band with about 4,400 square feet of headroom. The large band starts at 200,000 square feet, or 18,580 square metres, roughly 4.4 times your show. Any figure drawn from the large band describes a business that differs from yours in staffing, revenue mix and hall economics by design. It is worth checking what has actually been published for your own band before reaching for the global average, because a study scoped to shows under 50,000 net square feet is a different document from the one everybody quotes.

Judging whether a published sample is any good in the first place, including how it recruited and how many replied, is a separate job and P19 does it. This post starts one step later, from the finding that the sample does not include you.

Build the class you are actually in

Flyvbjerg's second step is "Establishing a probability distribution for the selected reference class", which he says "requires access to credible, empirical data for a sufficient number of projects within the reference class to make statistically meaningful conclusions" (Flyvbjerg, 2006).

For an exhibition organiser, the most comparable set of past, similar units available anywhere is your own show's editions. Same sector, same city, same venue in most years, same sales team, same definition of net space if you have been careful about it. No published sample will ever match your show that closely, because no published sample can.

That set is small, and I will come back to how small. It is still the right class to start from, because comparability is the binding constraint in Flyvbjerg's first step and sample size is the constraint you can partially relieve later by adding portfolio siblings.

A worked example on five year over year changes

Take six editions of net space sold, in square metres: 3,640 in 2021, 3,780 in 2022, 4,120 in 2023, 3,950 in 2024, 4,180 in 2025 and 4,240 this year.

The year over year changes are 3.8 per cent, 9.0 per cent, minus 4.1 per cent, 5.8 per cent and 1.4 per cent. Work the first one by hand and the rest follow the same way: 3,780 minus 3,640 is 140, and 140 divided by 3,640 is 0.0385.

The mean of those five is 3.2 per cent. The standard deviation is 4.9 points. Those two numbers are your reference class distribution, and having them changes what this year's 1.4 per cent means. One standard deviation either side of the mean runs from minus 1.7 per cent to 8.1 per cent, so 1.4 per cent is an ordinary year for this show, sitting closer to the middle of your own history than either 2023 or 2024 did.

Now run the comparison the slide was inviting. Suppose the outside figure is 6 per cent, from a study you have just established your show is not in. Six per cent falls inside your own one standard deviation band. Your show has produced a year above it and a year well below it within the last five. Being 4.6 points under it this year is inside the range your show generates unaided, and reading that gap as a performance finding hands the sales team blame for your own variance.

The reverse case is the one worth acting on. Had the show come in at minus 8 per cent, that is 2.3 standard deviations below your own mean, which is genuinely unusual against the only class that resembles you closely. That is the point at which to go looking for a cause.

What the outside number is still worth

Being excluded from a sample does not make the study useless to you. It changes what you take from it.

Direction and turning points survive the exclusion far better than levels do. A market that has stopped growing has usually stopped growing for reasons that reach a small regional show a year or two later, and the sign of a change carries across populations much more reliably than its size.

Definitions travel best of all. The 2026 CEIR Index Report, released by IAEE on 4 May 2026, covers the United States business to business exhibition industry across 14 sectors and measures year over year change on four things: net square feet of exhibit space sold, professional attendance, number of exhibiting companies, and gross revenue (CEIR, 2026). You can adopt those four definitions for your own series whatever market you sit in, and your internal benchmark is then at least built the same way as a published one. Using that Index from outside the United States raises its own set of questions that P35 works through.

The third thing an outside study gives you is a warning about arithmetic. If a published average was built by giving every show one vote, it describes the average show and says nothing about the average square metre, and that single choice moves the answer more than most readers expect. P22 has the method for redoing it.

What if you only have two editions?

Two editions give you one year over year change and no distribution at all, so borrow units from somewhere.

The first place to borrow is your own portfolio. Shows in adjacent sectors run by the same organiser share a sales approach, a pricing model and a calendar, which makes them a defensible class even where their subject matter differs. Six shows with four changes each gives you 24 observations, enough for a mean and a spread, provided you write in the methodology note that the class is mixed.

The second place is certified per-exhibition data, where somebody has already published the individual units and left the averaging to you. UFI's Euro Fair Statistics for 2024, published in November 2025, carries the certified statistics of 2,240 exhibitions from 16 countries (UFI, 2025). A dataset of individual exhibitions lets you pick out the ones near your own size and sector and assemble the class yourself, which is the thing a single aggregate figure never permits.

Both routes need the same discipline. Write down the selection rule before you look at any outcomes. A class chosen after seeing which comparators flatter you has stopped being a class and become an argument.

Where this stops

Your own history is a small sample too, and pretending otherwise swaps one error for another.

Take the five changes above. The standard error of their mean is the standard deviation divided by the square root of the count, so 4.9 divided by the square root of 5, which is 2.2 points. Your long run growth rate is therefore 3.2 per cent give or take roughly 2.2 points, and a self-benchmark built this way cannot resolve a difference of two points in either direction. Anyone presenting an internal baseline as precise is overselling it the same way the slide did.

The second limit is definitional drift inside your own file. If you changed the treatment of gangways in 2024, or folded a co-located event into the total, your series carries a break and part of the trend is an artefact of your own bookkeeping. An external sample at least has a written method behind it. Your internal series has whatever the last space planner decided and did not record.

The third limit is the honest one. An internal baseline cannot tell you whether the whole market moved. If your sector contracted 10 per cent and your show fell 4, you had a strong year, and nothing in your own five editions will ever reveal it. That is why the outside number keeps its place even when you are outside its sample. The working answer to a benchmark you are not in is a second series placed beside your own, each one labelled with the population it describes.

This week, take the last external figure your team quoted and write two lines under it: which population it was computed from, and whether a show of your size and market belongs to that population. Then pull your own net space for the last five editions, compute the four or five year over year changes, and put the mean and the standard deviation on the same slide as the external number. The comparison worth having is between this year and that spread.

Questions people ask about how to read industry benchmarks

How do I know if an industry benchmark applies to my show?
Check four things against the study's own description: which markets supplied the data, which sectors it groups together, what size of event it covers, and which year the underlying data describes. Size band is the test most readers skip. CEIR publishes separate organiser benchmarks for events above 200,000 net square feet, between 50,000 and 199,000, and below 50,000.
What should I benchmark against if my show is not in any published sample?
Your own editions. Five years of a single show give four or five year over year changes, which is enough for a mean and a standard deviation. Compare this year against that spread. Where two editions are all you have, pool shows from the same portfolio and write down the selection rule before you look at the outcomes.
Is an industry benchmark useless if my show is outside its sample?
No. Direction and turning points carry across populations better than levels do, so a market that has stopped growing is still information. The definitions carry too. The 2026 CEIR Index Report, released by IAEE on 4 May 2026, measures net square feet sold, professional attendance, exhibiting companies and gross revenue, and you can adopt those four for your own series.

Related reading

All standards and research articles