Skip to content

Designing an exhibitor account health score a sales director will trust

Renewal intelligenceUpdated 2026-08-1810 min read

In short

Build an exhibitor account health score from five weighted components: commercial at 30, engagement at 25, service at 20, performance at 15 and relationship at 10. Score each component 0 to 100 using percentiles within the show, and require every input to be one an account manager can look up in under a minute. Publish the weights.

The health score goes live in February. Twenty-six inputs, a gradient boosted model underneath, a holdout that looks respectable. By November two of the eleven reps have it open during pipeline review and the other nine are working from a spreadsheet one of them maintains, which has four columns in it: last edition's space, this edition's space, days late on the final invoice, and a note.

Nobody killed the score. Arguing with the spreadsheet was easier, and arguing is most of what a renewal review consists of.

That is the thing to design for. An exhibitor account health score is a management artefact before it is a statistical one. It exists so a sales director can rank a book of eight hundred accounts, hand slices of it to eleven people, and then defend that allocation to a finance lead who wants to know why the team is spending September on forty exhibitors. Every one of those uses happens in a conversation, and a number nobody can take apart in a conversation loses to a worse number they can.

What does a health score have to do in a meeting?

Verbeke, Martens, Mues and Baesens made the same point from the modelling side in Expert Systems with Applications in 2011, in a paper on building comprehensible customer churn prediction models with advanced rule induction techniques. The premise in the title is the argument: they went to the trouble of developing rule induction methods so that a churn model would produce output a person can read, because accuracy on a holdout is not the only property that decides whether anyone uses the thing.

I would give up a couple of points of ranking accuracy for a score whose components an account manager can look up in systems they already have open. That is a trade about adoption. The output of this thing is a list of names, and the list only produces revenue if somebody picks up a phone, so a score that eleven people use at 0.72 beats a score that two people use at 0.79.

There is a second requirement that follows from the same logic. The score has to be stable enough that a rep who looked at an account in September recognises it in October. A model refitted weekly on twenty-six inputs will move an account from 61 to 48 because a competitor's registration file arrived and shifted a percentile somewhere. The rep will not be able to explain that to the exhibitor, so they will stop quoting it.

Five to seven components, and the test each one has to pass

The design I would ship has five components. Six is fine, seven is the ceiling. Past that nobody holds the structure in their head, and the whole point of the structure is that people hold it in their head.

Each component is scored 0 to 100 and carries a published weight. The weights sum to 100, so the score is out of 100 and every component's maximum possible contribution is visible on the page. That published structure is what makes the score usable as renewal intelligence instead of a number people work around.

Every input inside a component has to pass one test: an account manager can look the raw value up, in a system they can open, in under a minute, without asking the data team. That test kills anything derived from an embedding. It kills "propensity decile", because looking that up returns another model output. It keeps space contracted, invoice ageing, portal sessions, ticket reopen share, scans at the stand, meetings accepted.

The test sounds arbitrary until you watch what happens without it. An exhibitor rings to ask why their renewal call came early, the rep tries to explain the score, and the explanation terminates in something the rep cannot see. After that happens twice the rep stops opening the score.

What goes in each component

Commercial, weight 30. Space contracted this edition against the account's own three-edition mean, total contracted value including sponsorship and services, and days from invoice to payment. Heaviest because it is the only component made entirely of decisions the exhibitor has already taken with money.

Engagement, weight 25. Exhibitor-initiated portal sessions outside mandatory upload windows, staff registration completion against allocation, meetings accepted in the matchmaking programme, and marketing responses with opens excluded. Recency of exhibitor-initiated contact belongs here as well, and computing it properly is its own subject (G17).

Service, weight 20. Ticket volume, reopen share, median days to close, and escalations past the account manager, each normalised for stand size so the component does not quietly rank exhibitors by how much stand they bought. Volume is the weakest of the four and the one to weight lowest inside the component (G16).

Performance, weight 15. What the exhibitor got for the money. Scans at the stand per 100 square metres against their category peers, and share of allocated scanner licences activated.

Relationship, weight 10. Tenure in editions, live contacts on file, whether anyone with signing authority attended, and the exhibitor survey response where you have one.

The survey deserves a note, because teams keep trying to make it carry more weight than it can. UFI's May 2025 statistics release, using data from its partnership with Explori, reported exhibitor net promoter scores improving by between 20 and 29 points across regions after the pandemic. That is a real measurement and it is a show-level one. At account level the response rate defeats you: exhibitor surveys reach whoever staffed the stand, and the person who signs the contract rarely fills one in. If your coverage is under half the book, the component is missing for most accounts. Weight it at 10, treat a missing survey as the component median, and accept that it can never move a score more than ten points.

Scoring one account, end to end

Convert every input to a percentile within the show's exhibitor base for that edition, oriented so higher is better. Percentiles because the units are incompatible and because a percentile stays interpretable when the show grows by two hundred stands.

Take a 72 square metre exhibitor on a September show, sixth consecutive edition.

Commercial has three inputs. They contracted 72 square metres against their own three-edition mean of 96, a ratio of 0.75, which sits at the 18th percentile of the show. Contracted value including sponsorship is down 14 per cent, the 41st percentile. They paid at 38 days against a show median of 9, which puts them at the 12th percentile once the input is turned round so fast payment scores high. Averaging the three: 18 plus 41 plus 12 is 71, divided by three is 23.7. The commercial component is 24.

Run the same procedure on the rest and suppose engagement comes out at 28, service at 74, performance at 55, relationship at 88.

Now apply the weights. Commercial contributes 0.30 times 24, which is 7.2 points of a possible 30. Engagement contributes 0.25 times 28, which is 7.0 of 25. Service contributes 0.20 times 74, which is 14.8 of 20. Performance contributes 0.15 times 55, which is 8.25 of 15. Relationship contributes 0.10 times 88, which is 8.8 of 10.

Add them: 7.2 plus 7.0 plus 14.8 plus 8.25 plus 8.8 is 46.05. The account scores 46.

The shape is worth more than the total. Service and relationship are close to full marks, which says your operations team has looked after this exhibitor for six years and somebody senior still turns up. Commercial has delivered 7.2 of a possible 30. This is a well-treated, long-standing account that is quietly reducing its commitment and paying slowly, which is a specific and recognisable kind of exhibitor. Turning that shape into a sentence a rep can open a call with is what reason codes are for (G21), and whether 46 counts as amber or red is a separate decision with its own arithmetic (G19).

Where should the weights come from?

Not from the data, and this is where most builds go wrong.

The tempting move is to fit the weights by regressing renewal on the five components. Do that and you have built a churn model with five features, which is a legitimate thing to build and should then be compared against a proper one on proper terms, since a health score and a churn probability answer different questions (G20). What you no longer have is a description of the relationship, because every component has been rescaled by how well it happened to predict last year's outcome, and a component can matter to the relationship while carrying little predictive weight.

The OECD and the European Commission's Joint Research Centre published the Handbook on Constructing Composite Indicators in 2008, and its position on weighting is the one to take. "Regardless of which method is used, weights are essentially value judgements," the handbook says, and it asks constructors to select the weighting procedure with reference to the theoretical framework they hold. It then gives sensitivity analysis a step of its own, naming the choice of weights among the things that step has to test.

So set them in a room. The commercial director and the two longest-serving account managers, one session, numbers on a whiteboard, and a written sentence next to each explaining what it represents. Then publish the weights in the same place the score appears. A rep who can see that commercial carries 30 and relationship carries 10 can argue about the 30, and that argument is the process working.

Testing the weights instead of defending them

That sensitivity step has a cheap version that takes an afternoon.

Move one weight by 5 points, redistribute the difference proportionally across the other four, recompute every score, and count how many accounts move more than 20 places in a ranking of 900. Repeat for each component.

Suppose commercial moved from 30 to 35 shifts 61 accounts by more than 20 places, which is 6.8 per cent of the book. Relationship moved from 10 to 15 shifts 12 accounts, 1.3 per cent. That tells you the argument about the relationship weight is not worth having and the argument about the commercial weight is, which is useful to know before the meeting rather than during it.

The second test is about compensation. A weighted sum lets a strong component pay for a dead one, and the 46-point account above earned most of its points from components describing how well you have treated them. Report the minimum component next to the total. An account at 46 with a floor of 24 and an account at 46 with a floor of 41 are different situations, and one extra column carries that.

Where this stops

The honest limit is that a health score has no ground truth. There is no measurement of relationship health sitting in a table that you can validate against. You can check that the ranking predicts renewal, and if you tune the weights until it does, you have built a churn model and thrown away the description you wanted.

What you can do is check face validity, deliberately and on the record. Hand the top thirty and the bottom thirty to the people who have carried those accounts for years and ask them to mark the ones they think are wrong. Disagreements are the finding. Every one of them is either a missing input or a weight nobody actually believes.

The other limit is that any component your own team can move cheaply will get moved. Once people know staff registration completion feeds engagement, somebody starts chasing exhibitors to complete their staff registrations the week before scoring runs. The input rises, the score rises, the underlying situation is unchanged. Watch each component for a step change in the fortnight before your scoring date, and treat any step change you find as evidence about your own process, because that is where it came from.

Take last edition's closed book, build the five components from systems you already have, and score every account. Print the top thirty and the bottom thirty, put them in front of your two longest-serving account managers, and ask them to strike out the ones that look wrong. If they strike out more than five of the sixty, you have a missing input or a weight nobody believes, and you will learn which one in the same conversation.

Questions people ask about exhibitor account health score

What should go into an exhibitor account health score?
Commercial covers space against the account's own three edition mean, total contracted value and days from invoice to payment. Engagement covers voluntary portal sessions, staff registration completion and meetings accepted. Service covers tickets, reopens and close times normalised for stand size. Performance covers scans per 100 square metres. Relationship covers tenure, live contacts and senior attendance.
How many components should a health score have?
Five is a good target and seven is the ceiling. Past that nobody holds the structure in their head during a pipeline review, which is the only place the score gets used. Each component needs a published weight, and the weights should sum to 100 so every component's maximum possible contribution is visible on the same page as the total.
Should health score weights be fitted to renewal data?
No. Fit the weights to renewal outcomes and you have built a churn model with five features, which is a legitimate thing to build and a different thing from a description of the relationship. Set the weights in a room with the commercial director and your longest serving account managers, write a sentence next to each, then test how much the ranking depends on them.

Related reading

All renewals articles