Setting account health score bands so green actually means something
Set account health score bands from outcomes. Pool every scored account edition with whether it renewed, bucket by score, and put the green line where observed renewal crosses about 60 per cent. Judge candidate red lines on three numbers: the share of the book they hold, the share of churn they capture, and lift over the base rate.
The account health score bands arrived with the dashboard template. Green at 50 and above, amber from 25 to 49, red below 25, because those are the numbers the template shipped with and nobody had a better idea in the week the thing went live.
A year later somebody does the obvious check. Green holds 76 per cent of the book, and 54 per cent of last edition's churn came out of it. The sales director points out, reasonably, that a colour which covers three quarters of the accounts and more than half the losses is not telling anyone anything.
The score underneath is fine. The lines drawn across it were never derived from anything.
What is a band actually for?
A band is an instruction. Red means somebody calls this account in the next fortnight. Amber means it goes into the weekly review. Green means nobody touches it unless something changes. Those three sentences are the whole product of the exercise, and they are what the arithmetic has to serve.
Once you write it down that way, the question stops being where a natural break in the score sits and becomes what the consequence of each colour is. A line is a commitment of somebody's time. So it gets set from what happened to accounts on each side of it, and from how many calls your team can actually make.
The alternative, which is what most teams do, is to pick round numbers and then argue about individual accounts forever. Round numbers are defensible only if you have no outcome data. After three editions you have outcome data.
Build the evidence table before you draw a line
Pool every scored account-edition you have. Three editions of a 900-exhibitor show gives 2,700 rows. For each row you need the score as it stood at show close, and whether that account exhibited at the following edition.
The score has to come from a snapshot taken at the time. Recomputing history with today's model produces a table that describes a score nobody ever saw, and the bands you draw on it will not hold when the real score is applied.
Bucket the rows and compute observed renewal per bucket.
| Score band | Account editions | Renewed next edition | Renewal rate |
|---|---|---|---|
| 0 to 20 | 122 | 38 | 31% |
| 21 to 40 | 286 | 137 | 48% |
| 41 to 55 | 392 | 223 | 57% |
| 56 to 70 | 690 | 490 | 71% |
| 71 to 85 | 764 | 642 | 84% |
| 86 to 100 | 446 | 410 | 92% |
Across the whole table 1,940 of 2,700 renewed, so the base renewal rate is 71.9 per cent and 760 accounts churned. The curve is monotone, which is the first thing to check. If renewal does not rise with the score, no band will save you and the problem is upstream in how the score itself was built (G18).
Use wide buckets where the curve is flat and narrow ones where you expect to cut. A 15-point bucket is fine at the top and useless at the point where you are trying to locate a line to within a few points.
Where should the green line go?
Green is the promise that an account needs no attention, so set it where observed renewal is comfortably above the point at which you would want a conversation. Sixty per cent is a defensible place to put that, on the grounds that an account with a better than two-in-five chance of leaving is one you would want to have spoken to.
The 41 to 55 bucket renews at 57 per cent, which is below the line, and the 56 to 70 bucket renews at 71 per cent, which is above it. The crossing is hidden inside a 15-point range, so split it.
At 41 to 45 there are 118 accounts and 61 renewed, 52 per cent. At 46 to 50 there are 132 accounts and 74 renewed, 56 per cent. At 51 to 55 there are 142 accounts and 88 renewed, 62 per cent.
Green starts at 51. Above that line accounts renew at better than 60 per cent and the rate climbs steadily; below it they do not. That sentence is defensible in a meeting in a way that "green starts at 50 because it is half" never was.
Where the red line goes, and the three numbers that judge it
Red is the opposite promise: everything in here gets worked. So the line is set by what you can cover, which is a capacity question with its own arithmetic (G23), and judged by three numbers, all computed from the same table.
Size is the share of the book the band holds, because that is the workload. Capture is the share of all churn that falls inside it, because that is the point. Lift is the churn rate inside the band divided by the base churn rate, which tells you whether the band is doing better than picking names at random.
Take three candidate lines.
Red below 41 puts 408 accounts in the band, 15.1 per cent of the book. They contain 233 of the 760 churned accounts, so capture is 30.7 per cent. The churn rate inside the band is 233 over 408, which is 57.1 per cent, against a base of 28.1 per cent. Lift is 2.03.
Red below 46 puts 526 accounts in, 19.5 per cent of the book, containing 290 churners. Capture 38.2 per cent, band churn rate 55.1 per cent, lift 1.96.
Red below 51 puts 658 accounts in, 24.4 per cent of the book, containing 348 churners. Capture 45.8 per cent, band churn rate 52.9 per cent, lift 1.88.
Every step down the score buys capture and pays for it in lift, which is the trade you are actually making. Fawcett set out the geometry of this in Pattern Recognition Letters in 2006: each candidate line is one operating point, and the set of them traces a curve of true positive rate against false positive rate. Red below 41 catches 30.7 per cent of churners while wrongly flagging 175 of the 1,940 accounts that renewed, a false positive rate of 9.0 per cent. Red below 51 catches 45.8 per cent while wrongly flagging 310, a false positive rate of 16.0 per cent.
Fawcett's other point matters more than the picture. A single summary of the whole curve tells you about the score, and you will only ever run one operating point. Hand made the argument sharper in Machine Learning in 2009, showing that the area under an ROC curve implicitly uses a different cost distribution for different classifiers, which makes it incoherent as a way of choosing between them when costs are real. Compare candidate scores at the line you will actually use. The same caution applies when you put the band next to a calibrated churn probability, which is a different number doing a different job (G20).
The cut encodes a price nobody has written down
Look at the marginal step from red below 41 to red below 51. It adds 250 accounts to the call list and reaches 115 additional churners, so those extra calls cost 2.17 of them per churner reached.
Vickers and Elkin made this explicit in Medical Decision Making in 2006. The threshold at which somebody chooses to act carries information about how they weigh a false positive against a false negative, and you can read the implied exchange rate straight off the cut. Accepting 2.17 calls per additional churner reached is a statement that a churner reached is worth more than 2.17 calls. That figure belongs in the paragraph next to the band definition, because it is the assumption the whole band rests on and it is the one a finance lead will want to see.
The uncomfortable part is that reaching a churner is not the same as keeping one. Some of the 115 were going to leave whatever the call said, and a few would have renewed without it. Sorting persuadable accounts from doomed ones needs a held-out group and a different method (G24). Until you have that, treat capture as an upper bound on what the band can deliver and say so out loud.
How often to move the lines
Recompute the table every edition. Move the lines much less often than that.
The 51 to 55 bucket has 142 accounts renewing at 62 per cent. The standard error on that proportion is the square root of 0.62 times 0.38 divided by 142, which is 0.041, so the 95 per cent interval runs from roughly 54 to 70 per cent. The crossing point is not sharply located, and a bucket that reads 62 per cent this year and 58 per cent next year has told you nothing.
Set a rule before you have a reason to break it. Move a line when the crossing has shifted by more than two buckets, or when the base renewal rate has moved by more than five points, which usually means something structural happened to the show. Otherwise leave it, because a band that moves annually destroys any comparison of this year's green against last year's, and the comparison is a large part of what the colours are for.
Where this stops
The evidence table has a defect built into it, and it gets worse the longer the bands are in use.
Once red means somebody calls the account, red accounts get called. Their renewal rate rises. Next year's table shows red renewing better than it did, and a line drawn on that table will be drawn in the wrong place, because what you are now measuring is the outcome of your own intervention, with the underlying risk hidden behind it. The direction of the bias is predictable and the size of it is not. Recording which accounts were actually worked, and keeping a small unworked holdout where the commercial team will tolerate one, is the only clean fix. The wider problem of scored accounts changing the behaviour they are scored on has its own treatment (G39), and it is the main reason a renewal intelligence programme needs to log its own actions as data.
The second limit is sample size at the bottom. The 0 to 20 bucket has 122 rows pooled across three editions, so roughly forty a year. On a smaller show that becomes a dozen, and a band drawn on a dozen accounts moves whenever two of them behave unexpectedly. Pool across shows in the same vertical if you can, accept wider buckets if you cannot, and never draw a line on a bucket with fewer than about fifty rows behind it.
Pull three editions of scored accounts with their next-edition outcome, bucket them at five points, and plot renewal against score. Find where the curve crosses 60 per cent, then work out how many accounts sit below every candidate red line and how much of last year's churn each one would have caught. Take those three numbers to whoever owns the renewal target and let them pick the line, because they are the person who has to staff it.
Questions people ask about account health score bands
- Where should the green line on a health score sit?
- Where observed renewal is comfortably above the point at which you would want a conversation. Sixty per cent is defensible, on the grounds that an account with a better than two in five chance of leaving deserves a call. Find the crossing by splitting your buckets to five points near it, because a 15 point bucket hides the line inside itself.
- How do you know if a red band is set correctly?
- Compute three numbers for every candidate line. Size is the share of the book inside the band, which is the workload. Capture is the share of all churn that falls inside it. Lift is the band's churn rate divided by the base churn rate. Every step down the score buys capture and pays for it in lift.
- How often should you move score bands?
- Recompute the evidence table every edition and move the lines much less often. A bucket of 142 accounts renewing at 62 per cent carries a standard error of about 4 points, so a reading of 58 per cent next year means nothing. Move a line when the crossing shifts by more than two buckets or the base renewal rate moves five points.