Skip to content

Health score versus churn probability and why you need both numbers

Renewal intelligenceUpdated 2026-08-189 min read

In short

A health score describes the current state of a relationship so a book of accounts can be ranked and split between people. A churn probability forecasts one event at one date. Carry both as separate columns and cross tabulate them, because the accounts sitting in the contradicting corners are the ones your team is about to misread.

The renewal review stops on one row. Health score 78, churn probability 0.61. Green account, high risk. Somebody asks which number is right and the room splits, because the analyst who built the model says the model, and the account manager who has carried the account for seven years says the score, and neither of them can produce anything that settles it.

Both numbers are right. Health score versus churn probability is not a contest to settle: they were built to answer different questions and nobody said so on the page.

What is the difference between the two numbers?

A health score measures a state. It compresses what the relationship looks like right now into one figure so that a book of accounts can be ranked and split between people. It has no horizon in it. Ask what a health score of 78 means and the answer is a description: this exhibitor is committed, pays on time, uses what they bought, and somebody senior turns up.

A churn probability forecasts an event at a date. Ask what 0.61 means and the answer is a claim about one specific thing happening once: this account does not exhibit at the next edition, 61 times in 100.

Shmueli set this distinction out carefully in Statistical Science in 2010, arguing that explanatory modelling and predictive modelling are different activities routinely conflated, on an assumption that a model with high explanatory power must also predict well. Her point that carries here is that a variable can matter enormously to the description while adding nothing to the forecast, and that the model selection criteria for the two jobs genuinely differ. Tenure is a good example. Eleven editions of history tells you a great deal about what kind of relationship this is. It may add almost nothing to a forecast of whether they sign in the next four months, because everyone already renewed ten times and the thing that changed is not in their tenure.

Blattberg, Kim and Neslin's Database Marketing, published by Springer in 2008, keeps the two apart in its own structure, with customer lifetime value and RFM analysis in the analytics chapters and churn management given a chapter of its own much later, which is the shape a renewal book needs. Keep the description and the forecast as two columns and the disagreements become information. Collapse them and the disagreements become an argument in a meeting.

Why do they come apart?

Four mechanisms produce most of the divergence, and all four are design decisions somebody made deliberately.

The input sets differ on purpose. A health score built to be looked up excludes anything an account manager cannot open in a system, so category-level decline, macro conditions and interaction terms are out of it (G18). A churn model has no such constraint and will use all of them. An exhibitor in a category that has shrunk for three editions can hold a health score of 78 while the model quietly moves them to 0.61 on evidence the score was never allowed to see.

The horizons differ. The probability is conditioned on one renewal event at one date. The score describes a state that has no date attached, so an account carrying a signed commitment for the next edition is at low risk and can still be in poor health.

The score compensates and the model does not. A weighted sum lets a strong service component pay for a dead commercial one. A tree ensemble is free to let a single input dominate, so an account with a 34 per cent space reduction can be pushed to high risk regardless of everything else.

The objectives differ, which is Shmueli's point applied. The score is tuned until the people who know the accounts agree with the ranking. The model is tuned until it separates outcomes on a holdout. Two different targets, two different answers.

The cross tabulation, worked

Take one show's 900 exhibitor accounts, both numbers computed at show close. Health bands from the score, drawn where observed renewal says they belong: red below 41, amber 41 to 50, green 51 and above (G19). Risk bands from the model: low below 0.20, medium 0.20 to 0.45, high above 0.45.

Health bandLow riskMedium riskHigh riskAccounts
Green41211846576
Amber749658228
Red18324696
Accounts504246150900

Most of the book sits on the diagonal, which is what you would expect and what makes the off-diagonal cells worth reading. The two corners that contradict each other hold 46 plus 18, which is 64 accounts, one in fourteen of the book. On a show this size that is more than five per rep, and every one of them is an account somebody is about to misread.

Now follow all 900 to the next edition. Across the book, 709 renewed, 78.8 per cent. Inside the green band, 507 of 576 renewed, 88.0 per cent. Inside red, 46 of 96, 47.9 per cent.

The green and high risk corner renewed 21 of 46, which is 45.7 per cent. Those 46 accounts sat in a band that renews at 88 per cent and behaved like the red band. The model was right about them and the health score was wrong.

The red and low risk corner renewed 16 of 18, which is 88.9 per cent, against 47.9 per cent for the red band they sat in. The model was right about those too.

Read only that far and the conclusion is obvious and wrong: keep the probability, drop the score.

Following the quiet corner two editions out

Take the same accounts to the edition after next.

Across the book, of the 709 that renewed once, 572 renewed again, so 80.7 per cent of survivors came back a second time and two-edition survival is 572 of 900, 63.6 per cent.

Of the 21 green and high risk accounts that survived, 18 renewed again, 85.7 per cent of survivors. Better than the book. That is the signature of a one-off shock: a budget freeze upstream, a product launch moved to a different quarter, a parent company mid-acquisition. The relationship was in good shape, something outside it removed the money for one year, and the accounts that got through came back to normal behaviour.

Of the 16 red and low risk accounts that survived, 7 renewed again, 43.8 per cent of survivors, against 80.7 per cent for the book. Their next-edition renewal was almost certain and it bought you nothing, because the reason it was almost certain is that they had already committed. The commitment expires. The relationship underneath it was in the red band the whole time and the score was reading it correctly while the probability, correctly, said the next edition was safe.

That is the case for carrying both numbers. Two-edition survival of 43.8 per cent among accounts your model rated low risk is the most expensive blind spot in a renewal book, and the only column that saw it was the description. The way multi edition contracts distort every retention figure sitting underneath this is a separate subject (G38), and it is the mechanism behind that corner.

What each corner means and who works it

Green and high risk is usually a money problem outside the relationship. The call is a different call: it goes to whoever holds the marketing budget rather than the person who runs the stand, and the useful offer is a smaller footprint, a payment schedule, or a held option on the same position. Treating those 46 accounts like ordinary at-risk exhibitors and sending a discount misreads the situation, because the relationship was never the problem.

Red and low risk is a deferred problem, and the mistake is to close the row because the probability is 0.11. The action is to book the health conversation into the current year, while there is still a contract running and the exhibitor has a reason to talk to you. Waiting until the model flags them means waiting until the contract is nearly finished, which is the point at which the exhibitor has already run their own review.

Amber and medium risk is the crowded middle, 96 accounts here, and it is the cell where a grid stops helping and a ranked call list has to do the work (G22).

Do not average them into one number

Somebody will propose combining the two into a single figure, usually as a weighted average, usually to simplify a dashboard.

The result has no units and no interpretation. A health score is on a scale you defined, a probability is on a scale defined by observed frequencies, and averaging them gives a number that cannot be validated against anything, because there is no outcome that a blend of a description and a forecast is supposed to match. The probability could have been checked against realised renewal rates and now it cannot. The score could have been checked for face validity with the account managers and now it cannot.

The related move, which is quieter and more common, is fitting the health score's weights to churn. Do that and you have a churn model with five features, so the two columns now agree by construction, the disagreements vanish, and the 18 accounts in the red and low risk corner disappear into the middle of the book where nobody looks at them. The disagreement was the finding. Building it out is the one thing you should not do.

Where this stops

The cross tabulation is only as good as the calibration of the probability. If your model outputs scores that rank well and are systematically high, the high risk column fills up with accounts that were never at 0.45, the corner counts inflate, and the readings above become noise. Check that predicted probabilities match observed frequencies in bands before you build any grid on top of them, and before anyone sums them into a renewal revenue forecast, which is where miscalibration turns into a number in a board pack (G32). Calibration is the one property in a renewal intelligence stack that is cheap to check and expensive to skip.

The health score has no equivalent check, and that asymmetry is permanent. There is no measurement of relationship health to validate against, so when the two numbers disagree you can establish that the probability was right about the next edition and you can never establish that the score was right about the state. The two-edition follow-through above is the closest available substitute, and it falls well short of proof.

The cells also get small quickly. Eighteen accounts in the red and low risk corner is enough to notice and not enough to conclude much from, and 7 of 16 could easily have been 5 or 9. Pool three editions before you attach any weight to a corner rate, and pool across shows in the same vertical if a single show leaves you with single figures.

Take your last scored book, put the health band and the risk band side by side, and count the accounts in the two contradicting corners. Then pull the ones that are healthy and high risk, and ask the account managers what happened to each in the twelve months before scoring. If the answer is a budget freeze or a reorganisation more often than not, you have found a signal your churn model is reading and your health score is not, and you know which of the two columns to trust on those accounts next time.

Questions people ask about health score versus churn probability

Is a health score the same as a churn probability?
No. A health score compresses the present state of a relationship into one figure and has no horizon attached. A churn probability makes a claim about one specific event at one date, usually the next renewal. An account can be healthy and high risk when a budget freeze arrives from outside the relationship, and unhealthy and low risk when a contract still has a year to run.
What does a healthy but high risk exhibitor mean?
Usually a money problem upstream of the relationship: a budget freeze, a launch moved to another quarter, a parent company mid acquisition. The relationship was never the issue, so a discount misreads the situation. The productive call goes to whoever holds the marketing budget, and the useful offer is a smaller footprint, a payment schedule or a held option on the position.
Can you combine a health score and a churn probability into one number?
You can, and the result has no units and nothing to validate against. A probability can be checked against realised renewal rates and a score can be checked for face validity with account managers. Average them and you lose both checks. Fitting the score's weights to churn has the same effect more quietly, because the two columns then agree by construction.

Related reading

All renewals articles