Which sponsor renewal predictors actually appear in your own data
Sponsor renewal predictors should be tested on your own sponsor-edition history rather than borrowed from a list. Build one row per sponsor per edition, label it renewed or not, and fit a logistic model. With 175 rows and 49 non-renewals, the events-per-variable rule caps you at about four predictors.
Somebody forwards an article listing the seven warning signs that a sponsor is about to leave. It reads well. Reduced engagement, a new marketing director, questions about attribution, silence in the autumn. Your sponsorship lead nods at all seven and nothing changes, because a list of sponsor renewal predictors written from somebody else's portfolio has no weights on it and no way to tell you which of your accounts to call on Tuesday.
The alternative is small and unglamorous. You already have three to five editions of sponsor history sitting in a CRM and a finance system. That is enough to test candidate signals against your own outcomes and find out which ones carry information here, on this show, with these accounts.
Build one row per sponsor per edition
The unit of analysis is the sponsor-edition. One row for each company that sponsored a given edition, with the outcome recorded as whether that company sponsored the following edition.
The outcome column is where most of these projects quietly fail, because it inherits every ambiguity in how you define renewal at all. A sponsor who dropped from 60,000 dollars to 4,000 is labelled renewed by a logo rule and lapsed by a value rule, and a model fitted on one label learns something different from a model fitted on the other. Settle that first and write the rule into the extract, because a label that shifts meaning between editions produces coefficients that describe your reporting conventions.
Then assemble the predictor columns as they stood at the moment the renewal decision was open, which is usually the eight weeks after the show closes. A column that records the state of the world in February tells you about accounts you already know renewed. This is the leakage failure and it is easy to commit without noticing, because the CRM shows current values and the history you need is in the audit log.
A portfolio show with 35 sponsors and five editions gives you 175 sponsor-editions minus the most recent one, so roughly 140 usable rows. Two shows with similar sponsor profiles can be pooled if you add a show indicator, which costs you a variable and buys a lot of rows.
How many predictors can your data actually support?
This is the question that decides whether the exercise produces anything trustworthy, and it has a published answer.
Peduzzi, Concato, Kemper, Holford and Feinstein published a simulation study in the Journal of Clinical Epidemiology in 1996 that varied the number of events per variable in logistic regression from 2 up to 25 across 500 simulated analyses. Below roughly ten events per variable they found coefficients biased in both directions, unreliable variance estimates, and associations reaching significance with the wrong sign.
Apply that here. Events means the smaller of your two outcome classes. With 175 sponsor-editions and a 28 per cent non-renewal rate you have 49 non-renewals against 126 renewals, so the events count is 49. Divide by ten and the model supports about four predictors.
Four. Not the seven from the article, and not the fourteen fields somebody wants to throw at it.
The budget is also consumed faster than people expect. A categorical variable with four levels costs three degrees of freedom, so entering asset count as small, medium, large and title uses three quarters of your allowance in one column. Keep counts as numbers where the relationship is plausibly monotonic. An interaction term costs another. The show indicator for a pooled model costs one, which is why pooling has to buy more rows than it spends.
If you have 20 non-renewals, you do not have a model. You have a cross-tabulation, and a cross-tabulation of renewal against one variable at a time is a perfectly respectable thing to publish.
The five candidates worth testing first
Pick four of these five, based on which are cleanest in your systems.
Activation spend as a ratio of the rights fee. What the sponsor spent building out the thing they bought, divided by what they paid you for the rights. A sponsor who paid 30,000 dollars and spent 12,000 activating it has a ratio of 0.4 and has committed internal budget beyond your invoice. Capture it from booth services orders, custom build quotes and any spend that flowed through your own suppliers, and accept that you will miss what they spent elsewhere.
Number of distinct asset lines held. A sponsor with one lanyard deal and a sponsor with six lines across signage, email and education have different amounts of surface area, and different numbers of people internally who would notice if it went away.
Whether the named contact changed since the last edition. A single flag from the CRM. The most operationally useful of the five if it holds up, because it is knowable in advance and it tells the sales team exactly who to go and meet.
Days from invoice to payment. A finance field, already clean, and a reasonable proxy for how the sponsorship sits inside the sponsor's own budget process. A median of 34 days against a median of 71 is a real difference in enthusiasm.
Whether a fulfilment report was delivered, and when. ANA and MASB, in their July 2018 study of sponsorship accountability metrics, found that among the marketers with a standardised measurement process, 30 per cent audit or verify the metrics they receive from the sponsorship property and 14 per cent said they receive no metrics from the property at all. The second figure is the interesting one for an organiser, because it describes a group of sponsors renewing or leaving with no evidence from you either way. Whether that flag predicts anything in your file is a genuine empirical question, and tracking what was actually delivered is what makes the column exist.
Reading the coefficients without overreading them
Fit the model, then report signs and magnitudes before anyone asks for an accuracy figure.
Suppose the coefficient on contact-changed comes back at 0.9 on the log-odds scale. Exponentiate: e to the 0.9 is 2.46, so a changed contact multiplies the odds of non-renewal by about two and a half. Your baseline non-renewal odds are 49 over 126, which is 0.389. Multiply by 2.46 and you get 0.957, and converting back gives a probability of 0.489. A sponsor whose contact changed sits near a coin flip against a base rate of 28 per cent.
That is a number a sales director can use. It is also a number that needs a caveat attached in the same breath.
Shmueli made the distinction plainly in Statistical Science in 2010, arguing that explanatory modelling and predictive modelling are different activities and that high explanatory power does not imply high predictive power. Your coefficient on contact-changed is an association in 140 historical rows. Contact churn may cause non-renewal, or both may follow from a restructuring at the sponsor, or the flag may be recorded more diligently for accounts the team was already worried about. The model cannot separate those, and the third possibility is a measurement artefact you can check by asking when the CRM field was last edited.
Report the sign, the odds ratio and a confidence interval. If the interval on a predictor spans 1.0, say so and leave the variable in the model or drop it on a stated rule, without narrating it as a finding.
What should you actually do with the fitted model?
Ranking, and nothing more ambitious.
Score every current sponsor, sort descending by predicted non-renewal probability, and hand the top quintile to whoever runs renewal conversations. The value is in the ordering, which means a model with mediocre calibration can still be useful as long as it puts roughly the right accounts near the top.
Check that it does. Split the file by edition, fit on the earlier ones, score the most recent, and count how many of the actual non-renewals landed in your top quintile. If your top 28 accounts contain 21 of the 40 that left, the model is finding something. If it contains 12, which is roughly what random sorting would give, you have learned that these four variables do not carry the signal and the next step is different variables rather than a fancier model.
Rank instability across editions is worth watching too, since a coefficient that flips sign between fits on adjacent editions is telling you the sample is too small for that variable. Much of the raw variation will turn out to be tenure, and why first-year sponsors behave differently deserves handling on its own before you conclude that your other predictors are weak. Everything here sits alongside the rest of renewal and revenue analysis as one input to a conversation, never as a verdict.
Where this stops
Four predictors on 140 rows is a small model, and it will be beaten by a good salesperson who has been to dinner with the account.
The honest framing is that the model catches the accounts nobody was worried about. Your team already knows the title sponsor is wobbling. What they miss is the mid-tier account that paid late, lost its contact, held one asset line and never received a report, because none of those facts is dramatic enough to reach a pipeline review on its own.
There is also a censoring problem the arithmetic hides. Sponsors who left three editions ago are gone from the file, so the surviving population is progressively more loyal and your measured base rate drifts down for reasons that have nothing to do with anything you did. Fitting on a five-year window and reporting the base rate per edition alongside the model is the cheap defence.
This week, pull one column: for every sponsor in your last three editions, whether the named contact on the account changed since the prior edition. Cross-tabulate it against renewal. Two numbers, one afternoon, and you will know whether the most actionable of the five candidates is worth building the rest of the table for.
Questions people ask about sponsor renewal predictors
- How much data do you need to model sponsor renewal?
- Count the smaller of your two outcome classes, usually non-renewals, and divide by ten. Peduzzi and colleagues showed in 1996 that logistic regression coefficients become unstable below roughly ten events per variable. A file with 49 non-renewals therefore supports about four predictors, and each level of a categorical variable consumes part of that budget.
- Which sponsor signals are worth testing first?
- Activation spend as a ratio of the rights fee, the number of distinct asset lines held, whether the named contact changed since the last edition, days from invoice to payment, and whether a fulfilment report was actually delivered. All five come from systems you already run, which means you can build the history retrospectively rather than waiting a year.
- Can a renewal model tell you why a sponsor left?
- No. A fitted coefficient describes an association in your historical file, and the same variable can be a good forecaster while carrying no causal meaning. Shmueli set out that distinction in Statistical Science in 2010. Use the model to rank accounts for attention, and use conversations with the sponsors to work out the reasons.