Where to set lead grade thresholds so an A actually means something
Lead grade thresholds should come from a measured distribution and a stated panel judgement, never a round number picked in the reporting layer. On a twelve point scale across 17,980 graded captures, a cut at ten and above yields 6.2 per cent A grades while a cut at five and above yields 51.5 per cent.
The grade mix comes back from the first edition where everybody used the same scale, and it is useless. Fifty eight per cent of the graded file is an A.
Nobody cheated. The scale was published, the axes were the same on every device, the four scores were stored. What was never decided, by anyone, in writing, was the number at which an A begins. Somebody put it in the reporting layer on a Thursday afternoon because the report would not render without it, and that number has now been printed on six hundred exhibitor scorecards.
Lead grade thresholds are the part of a grading scheme that gets the least attention and does the most work. The axes decide what you measure. The cut points decide what the answer looks like, and they are almost always set by feel.
Two ways to draw a line, and they answer different questions
Educational measurement worked through this argument sixty years ago and the vocabulary is worth borrowing, because it names the choice precisely.
Glaser, writing in American Psychologist in 1963, drew the distinction between criterion-referenced and norm-referenced measurement. A criterion-referenced cut point depends on an absolute standard of quality: this lead is an A because it meets a stated bar, and if every lead at the show meets it then every lead is an A. A norm-referenced cut point depends on a relative standard: this lead is an A because it sits in the top slice of the distribution, and the size of that slice is fixed in advance.
Both are defensible for lead grades and they behave very differently in the hands of an exhibitor.
A criterion cut is stable across editions, so an exhibitor can say their A count doubled and mean it. It is also capable of producing a show where 58 per cent of leads are A grades, or a show where none are, because nothing constrains the proportion.
A norm cut fixes the proportion, so the report always reads sensibly and the exhibitor at the ninetieth percentile is always at the ninetieth percentile. It also means that an exhibitor whose leads genuinely improved year on year can see their A count fall, because everyone else improved faster. That is the failure mode that gets you a phone call.
I would use a criterion cut for the letter that goes on the exhibitor's file and a norm cut as the annual sanity check on where the criterion sits. The rest of this post is about how to run both.
Look at the distribution before you argue about the number
Take the four axis scale, zero to three on each axis, summing to twelve. Why those axes and that range is a design question with its own answer, and what enters the graded pool at all depends on the show level definition of a qualified lead. Grade the whole show, then plot the totals before anyone proposes a cut point.
Here is a distribution from a single edition, 17,980 graded captures across the floor.
| Total score | Captures | Share |
|---|---|---|
| 0 | 640 | 3.6% |
| 1 | 1,180 | 6.6% |
| 2 | 1,930 | 10.7% |
| 3 | 2,410 | 13.4% |
| 4 | 2,560 | 14.2% |
| 5 | 2,300 | 12.8% |
| 6 | 1,980 | 11.0% |
| 7 | 1,640 | 9.1% |
| 8 | 1,290 | 7.2% |
| 9 | 940 | 5.2% |
| 10 | 620 | 3.4% |
| 11 | 340 | 1.9% |
| 12 | 150 | 0.8% |
Now run the three candidate cut points against it and watch what happens to the same file.
Cut A at ten and above, which is the tidy answer because it takes the top quarter of a twelve point range. That gives 620 plus 340 plus 150, so 1,110 captures, 6.2 per cent of the file. A very tight A, and an exhibitor with 180 graded leads will average eleven of them.
Cut A at eight and above. That gives the 1,290 that scored exactly eight plus the 2,050 above them, so 3,340 captures, 18.6 per cent.
Cut A at five and above, which is roughly where a stand grading generously would land. That gives 2,300 plus 1,980 plus 1,640 plus 3,340, so 9,260 captures, 51.5 per cent. Drop it one further point to four and above and you are at 11,820, or 65.7 per cent, which is the 58 per cent from the opening with a slightly different file behind it.
Three cut points, every one of them arguable in a meeting, producing A rates that differ by a factor of ten on identical data, which is the size of a decision currently being made by whoever needed the report to compile.
Why does the top quintile answer not land cleanly?
The norm-referenced version is easy to state. A is the top quintile of the graded distribution, B the next three, C below that.
Twenty per cent of 17,980 is 3,596 captures. Count down from the top: 150, then 490 at eleven and above, 1,110 at ten and above, 2,050 at nine and above, 3,340 at eight and above, 4,980 at seven and above.
The target of 3,596 sits between two bands. At eight and above you have 3,340, which is 18.6 per cent. At seven and above you have 4,980, which is 27.7 per cent. To hit 20 per cent exactly you would have to take 256 of the 1,640 captures that scored exactly seven, and there is nothing in the data to choose which 256.
An integer scale with thirteen possible values cannot produce arbitrary quantiles, and pretending otherwise by breaking ties on capture time or badge id would be inventing a distinction. Take the nearest band, take eight and above, and publish the actual figure of 18.6 per cent rather than the round number you asked for. Anyone who reads the report will trust 18.6 more than 20.
Borrowing a standard setting method from testing
The criterion cut needs a procedure, because otherwise it is one person's Thursday afternoon.
Angoff described the approach in 1971, in a chapter on scaling and norming in Thorndike's Educational Measurement, and modified versions of it are still the common way cut scores get set in licensing and certification. The mechanism is simple. Assemble a panel of people who know the domain. Have them describe the borderline candidate, the one who only just deserves to pass. Then, item by item, have each panellist estimate how that borderline candidate would perform. Sum the estimates and you have a cut score derived from stated judgements rather than from a feeling about round numbers.
The adaptation to lead grades is direct. Your panel is six or seven sales directors from exhibiting companies across different categories, in a room for ninety minutes. They describe the borderline A: the weakest lead a sales team would still drop everything for in week one. Then, for each of the four axes, each panellist writes down the score that borderline lead would most likely receive.
Average across the panel, axis by axis. Suppose it comes out at authority 2.4, stated need 2.1, timeline 2.0 and category fit 2.3. Sum is 8.8, so the panel's cut for an A is nine.
Nine and above is 2,050 captures, 11.4 per cent of the file.
That is the useful moment. The panel says nine, the distribution says eight, and the gap between them is the 1,290 captures that scored exactly eight, or 7.2 per cent of everything graded. One point of scale, one band, and a difference of nearly two thirds in how many A grades the show reports.
Which number should go on the exhibitor scorecard?
Take the panel's number for the letter and print the percentile beside it.
The letter is what the exhibitor's sales team acts on, and it has to mean the same thing in 2027 as it does now, so it needs to be criterion-referenced and it needs to hold still. The percentile is what tells the exhibitor whether their file is unusual, and it needs to be recomputed every edition from that edition's distribution.
So the line on the scorecard reads: 212 A grades, 14.8 per cent of your graded captures, against a show figure of 11.4 per cent. Both numbers, one sentence, and the exhibitor can see immediately that their bar was easier than the floor's or that their audience was better than the floor's. Which of those two it is takes a conversation, and the conversation is the point. Comparing them against a narrower set of similar stands rather than the whole floor is a peer group construction problem with its own small sample traps.
Then set a rule for moving the cut and write the rule down. Reconvene the panel every third edition, or sooner if the axis wording changes. When the cut moves, restate the previous edition under the new cut in the same report, so the year on year line does not break silently. A threshold that drifts by one point every year makes every trend in your exhibitor reporting a measurement artefact, and somebody will eventually work that out in front of a client.
There is a commercial reason to care about the stability as well as the arithmetic. The 2026 Marketing Spend Decision Report from CEIR found that sales metrics dominate how management evaluates exhibition return on investment, with lead volume and post-show closed deals ranking highest, and that exhibitions take 40.8 per cent of exhibitor marketing budgets. The A count is a lead volume figure that somebody defends to a finance director. Moving the definition underneath it without saying so is the kind of thing that ends a renewal, and it is the fastest way to lose trust in exhibitor analytics across a whole floor.
Where this stops
A threshold cannot be better than the scores it cuts. If a stand grades everything a three on stated need, no cut point rescues that file, and the only defence is the completion and consistency checks that belong with the scheme itself. A cut point also says nothing about ordering within a band, which is where a weighted score built from registration attributes does work that letters cannot.
The bigger limit is small exhibitors. A stand with 40 graded captures and a show A rate of 11.4 per cent expects between four and five A grades, and the ordinary variation around that number is wide enough that four versus seven means nothing at all. I would print the A count for every exhibitor and suppress the A percentage below about 60 graded captures, showing the raw counts instead. A percentage computed from 40 items reads as precise and is not, and it will be quoted back at you in a renewal meeting as though it were.
The third limit is that the panel is made of exhibitors, and exhibitors have a view about what a good lead looks like that reflects their own sales cycle. A panel drawn entirely from capital equipment sellers will set a bar that makes every consumables stand look excellent. Draw the panel across categories, record who was in the room, and publish the composition next to the cut point, because a cut score with named judges behind it survives an argument and an anonymous one does not.
Take last edition's graded captures, if you have any, and build the thirteen row table above for your whole floor. Then run your current A threshold against it and write down the percentage it produces. If that percentage is above about a quarter or below about a twentieth, the number in your reporting layer was set by whoever needed the report to render, and it is worth an hour of somebody senior's time before the next edition.
Questions people ask about lead grade thresholds
- Should lead grades be criterion-referenced or norm-referenced?
- Use a criterion cut for the letter on the exhibitor's file, so an A means the same thing in three years and a doubled A count means something real. Use a norm cut, recomputed each edition, as the annual sanity check on where the criterion sits. Print the letter and the percentile together.
- How do you set a lead grade cut point defensibly?
- Borrow the Angoff procedure from educational testing. Put six or seven sales directors from different categories in a room, have them describe the weakest lead a sales team would still drop everything for, then have each write the score that borderline lead would receive on every axis. Average across the panel and sum.
- Why not simply make an A the top twenty per cent of leads?
- An integer scale with thirteen possible values cannot produce arbitrary quantiles. Twenty per cent of 17,980 captures is 3,596, but eight and above gives 3,340 and seven and above gives 4,980. Hitting 20 per cent exactly means splitting the captures that scored seven, and nothing in the data chooses which ones.
Related reading
- Lead quality grading only works when the organiser defines the grades first
- Agreeing one qualified lead definition for exhibitors across an entire show
- Building a trade show lead scoring model an organiser can defend to exhibitors