Skip to content

Account health reason codes turn a score into a call agenda

Renewal intelligenceUpdated 2026-08-189 min read

In short

Account health reason codes are the individual component contributions that produce a health score, printed alongside it. For a weighted sum, each contribution is the component weight times the gap between the account's value and the book average. For a tree ensemble, Shapley values give the same additive breakdown a rep can read.

A sales director I know keeps a printout of her renewal book on the desk during the fortnight after the show closes. Every account has a health score on it. She hands the sheet to a rep, the rep looks at a row that says 38, and asks the only sensible question available. What do I say to them?

Nobody in the room can answer. The score came out of a model that nobody in the sales team built, it went down from 61 last edition, and the honest reply is that somebody would have to go and look. Which they will not do for four hundred accounts.

The score is doing half a job. It ranks the book and sorts it into bands somebody had to draw, which is worth something, and then it stops exactly where the work starts. Reason codes are the other half, and they are usually already computed and thrown away.

A score is a summary statistic, and summaries throw things away

Any health score compresses ten or twenty inputs into one number, and compression is the point. How those inputs get chosen and weighted is G18's subject. You cannot sort a book on twenty columns. You can sort it on one.

The cost of that compression is that two accounts scoring 38 can be at 38 for entirely different reasons. One has stopped paying on time and cut its stand from 54 square metres to 36. The other pays like clockwork, kept its space, and has not opened a single email since March while its named contact left the company. Those are two different calls. The first is a finance conversation and probably a payment plan. The second is a contact-refresh problem and a fifteen minute call to whoever has taken over.

Give a rep the number alone and both accounts get the same call, which is the vague one that starts with "I just wanted to check in on next year".

How do you compute a reason code from a weighted score?

If your health score is a weighted sum, the reason codes are already sitting inside it and nobody has bothered to print them. The contribution of a component is its weight multiplied by the gap between this account's value on that component and the book average.

Take a real-shaped example. The score runs 0 to 100, the book average score is 62, and the model has ten components. For account 41207 the four largest movers look like this.

Payment lateness index sits at 22 against a book average of 71, with a weight of 0.18. The contribution is 0.18 times minus 49, which is minus 8.82.

Space change sits at 30 against a book average of 69, weight 0.18. That is 0.18 times minus 39, or minus 7.02.

Staff registrations completed sits at 12 against 58, weight 0.11. That is 0.11 times minus 46, or minus 5.06.

Tenure sits at 92 against 63, weight 0.14. That one is positive: 0.14 times 29, or plus 4.06.

The remaining six components net to minus 7.16. Add it all up starting from the book average: 62 minus 8.82 minus 7.02 minus 5.06 plus 4.06 minus 7.16 gives 38.00.

Now the row reads differently. Payment minus 9, space minus 7, staff registrations minus 5, tenure plus 4. A rep can work with that in a way they cannot work with 38, because it names three things to ask about and one thing to lead with. Ten years on the show floor is the reason this account is worth a call at all.

The property that makes this honest is that the parts sum to the whole. If your reason codes do not reconstruct the score, they are commentary rather than explanation, and somebody will eventually notice.

When the model is a tree ensemble

Weighted sums are easy to pull apart. Gradient boosted trees are not, and most teams who build a second version of their health score end up with a tree ensemble because it handles interactions and missing values without much coaxing.

Lundberg and Lee, at NeurIPS in 2017, set out the approach that has become the default answer. They showed that a family of additive explanation methods, including several that had been invented independently, all point at the same unique solution when you require a small set of sensible properties, and that solution is the Shapley value from cooperative game theory. Each feature gets credit equal to its average marginal contribution across all the orders in which features could be added to the model.

The catch with Shapley values is that computing them exactly means considering every subset of features, and the number of subsets doubles with each feature. Lundberg and colleagues published the fix in Nature Machine Intelligence in 2020: an algorithm that exploits the structure of decision trees to compute exact Shapley values for tree ensembles in polynomial time. That paper is why this is a practical option for a renewal model over twenty features and three thousand accounts, and not a research curiosity.

What you get out is the same shape as the weighted-sum case. A base value, which is the average model output over the training set, plus one contribution per feature, summing to this account's prediction. The interpretation is different in an important way that I will come back to, but the format a rep sees is identical.

One detail catches people the first time. The base value is the average over whatever population you trained on, so if your training set is every exhibitor across eight shows and five years, the contributions for an account on your smallest show are measured against a book that account has nothing to do with. Every contribution then carries a slice of "this is a small show" in it. Training a separate model per show fixes the interpretation and usually costs you accuracy, because a show with four hundred exhibitors and five editions has two thousand rows to learn from. The compromise I would take is one model with the show as a feature, and reason codes computed against the base value of that show's own accounts, which keeps the comparison local while the model still learns across the portfolio.

What should the reason code row actually say?

Three negatives and one positive, each with a number, each in the language of the business rather than the model.

Not "payment_lateness_index: -8.82". Something closer to "Paid 41 days late on the last two invoices, against a book average of 6 days. Minus 9 points." The rep needs the underlying fact and the size of the effect, because the fact is what they raise on the call and the size is what tells them whether to raise it first or third.

The positive matters more than people expect. A call agenda built entirely from problems produces a call that sounds like an audit, and exhibitors who feel audited go quiet. Leading with ten editions of history and then asking about the invoices is the same information in an order somebody will answer.

I would cap the list at four rows. Reps do not read five, and beyond the third contribution the numbers are usually small enough that the ordering is noise. If your fourth-largest contribution is minus 1.2 on a score that moves twenty points between quartiles, printing it invites a conversation about something that does not matter.

The ordering question nobody asks

Sort by the size of the contribution and you get the largest movers. Sort by what the rep can actually do something about and you get a different list.

Tenure is usually near the top of the contribution list and there is nothing anyone can do about it. Neither can they do much about the account being in a category that has been shrinking for three years. A contribution list that leads with two immovable facts wastes the top of the page.

The version I would ship splits the row in two. The largest contributions overall, which explain the score, and then the largest contributions among components the sales team can influence, which is the agenda. Payment terms, unresolved service tickets, staff registration completion, meetings booked in the matchmaking tool, sponsorship attach. Those are things a call can change. Category decline is context for the forecast and not an action.

That split takes about a day to build once the contributions exist, and it is the difference between reps reading the reason codes and reps ignoring them by week two.

Where reason codes stop being true

A contribution is an attribution of the model's output, and the model is a correlational summary of your own history. It is not a causal account of why this exhibitor is unhappy, and the two come apart most sharply in exactly the situation the reason codes are trying to describe. This is also the point where the difference between a composite score and a calibrated probability starts to matter, which is G20's argument.

Suppose an exhibitor's marketing budget was frozen in February. That single event shows up in your data as late payment, a stand reduction, fewer staff registrations and a cancelled sponsorship. Four correlated symptoms of one cause. The model will split the credit between them, and how it splits depends on which correlations dominated the training data rather than on anything about this exhibitor. Present those four as four separate problems and you have manufactured an agenda with four items where the real agenda has one.

Correlated inputs also make the split unstable. Retrain on next year's data and the same account can show payment at minus 9 and space at minus 7 one month, then space at minus 11 and payment at minus 5 the next, with no change in the underlying facts. If reps notice that, they stop trusting the numbers, and they are right to.

The second failure is quieter. Once a team knows that staff registration completion moves the score, somebody will start chasing exhibitors to complete their staff registrations in the fortnight before scoring runs. The input moves, the score moves, and the underlying risk does not. Any component that is cheap for your own team to move is a component you should watch for drift, and probably one to keep out of the rep-facing agenda even if the model uses it.

Neither problem argues for going back to the bare score. Both argue for a health score whose components are things you would investigate anyway, and for keeping a record of what each account's reason codes said at the time so you can go back and check whether the calls that followed them worked.

Take last cycle's scored book and recompute it with contributions attached, then pull twenty accounts a rep already called and compare the top reason code against whatever they wrote in the CRM note. Where the two agree you have a model worth putting in front of the team. Where they disagree, the rep is usually right and you have found the feature you are missing.

Questions people ask about account health reason codes

What are account health reason codes?
They are the component contributions that add up to an account health score, shown next to it. Payment lateness minus 9, space change minus 7, staff registrations minus 5, tenure plus 4. Each one names a fact about the account and the number of points it moved, so a rep gets an agenda instead of a ranking.
How do you calculate reason codes for a health score?
If the score is a weighted sum, multiply each component's weight by the gap between this account's value and the book average. A weight of 0.18 on a payment index sitting 49 points below average contributes minus 8.82. The contributions have to sum back to the score, otherwise they are decoration.
Can you get reason codes out of a gradient boosted model?
Yes. Lundberg and Lee set out the Shapley value approach at NeurIPS in 2017, and Lundberg and colleagues published a polynomial time algorithm for tree ensembles in Nature Machine Intelligence in 2020. The output has the same shape as a weighted sum, a base value plus one contribution per feature, summing to the account's prediction.

Related reading

All renewals articles