Skip to content

Minimum group size for benchmarks before a percentile reveals a competitor

BI and reportingUpdated 2026-08-238 min read

In short

A minimum group size for benchmarks is the number of members a comparison cell must hold before any statistic about it is shown. Below roughly ten, a reader who knows their own figure can bound their competitors' figures arithmetically. Suppress the tile, suppress a second cell to block differencing, and print the group count.

The exhibitor benchmark tile was the most requested feature in the portal and it shipped in a fortnight. Each stand sees where it sits against the median for its product category: leads per square metre, scan rate, follow up within seven days. Reception was good until an exhibitor in a thin category emailed to ask, politely, whether the organiser realised the tile was telling him roughly what the other three companies in his category had done.

He was right, and setting a minimum group size for benchmarks is the only thing that fixes it. The instinct is to treat this as a precision problem, where a small cell simply produces a noisy number. That is a real problem and it has a different answer. The problem here is disclosure. A statistic computed over four companies, shown to one of those four, is a description of the other three.

What a percentile gives away in a category of four

Take the tile as it was actually built. Product category, four exhibitors, and three summary figures: the count, the median and the mean. Add the top quartile because somebody asked for it. The exhibitor reading it knows a fourth number that nobody printed, which is his own.

Suppose the tile shows a count of 4, a median of 4.4 leads per square metre and a mean of 5.1. Our reader's own figure is 6.2.

The mean gives him the total: 4 times 5.1 is 20.4. The median of four values is the average of the second and third when they are sorted, so the middle pair sums to 8.8. His 6.2 is above the median, so he is either the third value or the fourth.

Test the fourth. If he is the largest, the smallest value must be 20.4 minus 8.8 minus 6.2, which is 5.4. But the smallest cannot exceed the second, and the second is at most 4.4 because the middle pair sums to 8.8 and the second is the lower of them. So 5.4 is impossible and he is the third value.

That settles the second value at 8.8 minus 6.2, which is 2.6 exactly. The remaining two sum to 20.4 minus 8.8, which is 11.6, and since the smallest is at most 2.6 the largest is at least 9.0.

From three published numbers and his own, he now knows that one competitor scores 2.6 or below and another scores at least 9.0, on a metric that maps directly onto how well a stand is being worked. He knows the category membership because it is printed in the show directory. Two phone calls and he knows which is which.

Why is a competitor a harder adversary than a researcher?

Sweeney formalised the underlying idea in 2002, defining k-anonymity as the property that a released record cannot be distinguished from at least k minus 1 others in the release. Statistical disclosure control practice builds on the same intuition, and the UK Government Analysis Function guidance (2026) sets out the threshold rule for magnitude tables as a requirement that at least n enterprise groups sit in a cell, naming 2, 3, 4 and 5 as typical values for n, and noting that cells of size 1 and 2 are usually risky in frequency tables.

Those thresholds are lower than what an exhibitor benchmark needs, and the reason is who is reading. A national statistics output is read by researchers and journalists who hold no rows in the table. An exhibitor benchmark is read by a member of the cell. That removes one unknown from every equation before the reader starts, and in a cell of four it turns an underdetermined system into the one worked above.

There is a second asymmetry. A researcher wanting to identify a firm has to guess at membership. An exhibitor already knows exactly who else is in the category, because they walked past their stands for three days.

Where the floor should sit

Ten is where I would start, and twelve where I would rather be.

The reasoning is arithmetic and you can redo it with your own metric. In a cell of n, a reader who knows their own value has n minus 1 unknowns and gets one equation from each published statistic. With three statistics on the tile and a cell of 4, that is three equations for three unknowns, which is why the example above resolved so cleanly. At n equal to 10 the same three statistics leave six unknowns underdetermined, and the bounds a reader can derive are wide enough that acting on them is guesswork.

The count of published statistics matters as much as the count of members, and it is the variable people forget. Every extra figure you put on the tile is another equation. A tile showing count, minimum, first quartile, median, third quartile, maximum and mean over eight companies amounts to a slightly obfuscated list of the eight values.

E24 reaches a floor near twelve from an entirely different argument, about how much one rank position moves a stated percentile in a thin cell. The two arguments are independent and they happen to meet, which is a comfortable place to set a rule. Where they disagree, take the higher number, because a tile that is stable and leaky is still a problem you have to explain to a lawyer.

Suppressing one cell is never enough

The Government Analysis Function guidance (2026) is blunt about the failure that follows a naive suppression: "disclosure by differencing one table from another can occur where the variable classifications are slightly different". A benchmark tile with filters on it produces exactly that condition, on demand, for free.

Work it. A category holds 12 exhibitors and passes the floor, so the tile renders with a mean of 5.1. The reader then applies the stand size filter and asks for stands over 30 square metres. That cell holds 10 exhibitors and also passes, with a mean of 5.6.

The two exhibitors under 30 square metres never had a tile of their own. Their combined total is 12 times 5.1 minus 10 times 5.6, which is 61.2 minus 56.0, or 5.2 across two companies. Their mean is 2.6. Both are identifiable from the floor plan, because small stands in a thin category are visible from the aisle.

The fix is complementary suppression, which means suppressing a second cell so the first cannot be recovered by subtraction. In practice, if any cell in a filtered breakdown falls below the floor, suppress the smallest cell above the floor as well, and suppress the total. It removes more from the report than anyone wants. That is the price of the tile being safe rather than approximately safe.

Working a floor of ten across 41 categories

Here is the sweep for a show with 640 exhibitors across 41 product categories.

Count the members of each category. Nine of the 41 come in below ten, holding 2, 3, 3, 4, 5, 5, 6, 6 and 7 exhibitors. That is 41 companies out of 640, or 6.4 per cent of the exhibitor base, who will see a suppressed tile.

Two responses are available and one of them is wrong. Widening the group by merging thin categories into a parent is defensible, as long as the merge is declared on the tile, because a benchmark against a group you were not told about is worse than no benchmark. Quietly dropping the floor for those nine because the complaints are annoying is the wrong one, and it is what happens when the floor lives in a report rather than in the query.

Put the suppression in the semantic layer, next to the tenant filter that scopes each exhibitor to their own rows, so a new report cannot render an unsuppressed cell. A rule that lives in one visual gets copied incorrectly the first time somebody builds a second visual. The isolation question V20 covers and this one are the same class of defect: a predicate that has to be everywhere and is only somewhere.

Should you tell exhibitors the floor?

Disclosure control guidance says to keep threshold parameters confidential, on the grounds that knowing them reduces protection. For an exhibitor benchmark I would publish the number anyway.

The protection here comes from the suppression itself, not from secrecy about where it triggers. An exhibitor who sees a blank tile with no explanation raises a ticket, and your support team then explains the rule one exhibitor at a time, badly. A line reading "not shown: fewer than 10 companies in this comparison" costs nothing and answers the question. The residual risk is that a reader learns their category has nine or fewer members, which they could have counted in the directory.

The wording matters more in an embedded tile than it would in an internal report, because a report living inside the exhibitor portal has no analyst standing next to it. V19 covers that embed. What belongs here is that the suppression message is part of the report, drafted once, rather than a blank rectangle and a hope. The same holds for every other page in the BI and reporting set that shows a peer comparison.

Where this stops

A floor protects against arithmetic. It does not protect against a reader who already knows the answer and is using your tile to confirm it, and it does nothing about disclosure across editions.

That second one is the gap worth watching. If a category held 11 members last edition and 9 this edition, and you published last edition's mean, the arrival or departure of two companies is itself information, and the earlier figure combined with this edition's category total narrows what the remaining nine did. Suppression rules applied one edition at a time do not see across time. If your benchmark shows edition on edition change, the floor has to be met in both editions or the change is suppressed too.

Start by running a count of members per comparison cell for your last edition, including every filter combination the tile actually offers rather than only the unfiltered view. Sort ascending and look at how many cells sit under ten. On most shows the number is larger than anyone expects, and it is the honest size of the decision you are making.

Questions people ask about minimum group size for benchmarks

What is a safe minimum group size for an exhibitor benchmark?
Ten to twelve members is a defensible floor for a benchmark shown to members of the group itself. Official statistics guidance uses lower thresholds, but those outputs are read by people outside the population. An exhibitor already knows one value in the cell, which removes an unknown from every equation and makes small cells far leakier.
Why does suppressing one small cell not solve the problem?
Because the suppressed cell can often be recovered by subtraction. If the category total and every other category are published, the missing one is the difference. The UK Government Analysis Function warns that disclosure by differencing occurs where two tables use slightly different classifications, which is exactly what a filterable benchmark produces.
Should the benchmark show the number of companies in the peer group?
Yes. A percentile with no group count behind it invites a reader to treat a four company comparison as an industry figure. Printing the count also explains a suppressed tile without a support call, and it lets a reader discount a thin cell themselves rather than acting on a number that moves by twenty points when one company joins.

Related reading

All bi and reporting articles