Borrowing strength across shows when no single show has enough history
Borrowing strength across shows means estimating one pooled pattern plus a per-show deviation, then shrinking each deviation toward zero in proportion to how little data that show has. The weight on a show's own history is its edition count divided by that count plus the ratio of within-show to between-show variance.
A show launched three editions ago is now big enough that people care about its forecast. It has three registration curves. Somebody plots them, notices they do not agree with each other, and asks which one to use.
Borrowing strength across shows is the answer that stops you having to choose. The portfolio has eight shows and forty editions between them, the shape of a registration campaign is broadly similar across them, and a model can use all of it while still letting each show be itself.
The two bad answers a portfolio gives by default
Left alone, a team will do one of two things with a short-history show.
The first is to use only that show's own data, which Gelman and Hill call the no-pooling analysis in their 2007 book on regression and multilevel models. They are direct about what goes wrong. Writing about counties in a radon study, they note that "the estimates from the no-pooling model overstate the variation among counties and tend to make the individual counties look more different than they actually are." Swap counties for shows and the sentence needs no other edit. The show with three editions will look like the fastest grower or the worst performer in the portfolio, and it will look that way because it has three editions.
The second is to use the portfolio average for everything, which is complete pooling. That removes the noise and removes the show as well. If your kitchen and bath show genuinely registers later than your hospitality show, complete pooling hands both of them the same curve and both forecasts are wrong in opposite directions.
Partial pooling is the compromise, and it has an exact form rather than a judgement call.
How much is one show's own history actually worth?
Gelman and Hill write the multilevel estimate for a group as a weighted average of that group's own estimate and the estimate from the group-level model. Their weight on the group's own number is n_j divided by sigma y squared, over the sum of that term and one over sigma alpha squared, where n_j is the number of observations in the group, sigma y is the within-group standard deviation and sigma alpha the between-group standard deviation.
Multiply top and bottom by sigma y squared and the weight reads more usefully. It becomes n_j divided by the sum of n_j and the ratio of within-show variance to between-show variance.
That ratio is the number of editions the pooled pattern is worth. Gelman and Hill make the same reading of their own example: with a variance ratio of about a fifth, they say "the amount of information in this distribution is the same as that in 5 measurements within a county", and conclude that the multilevel estimate is closer to complete pooling below that count and closer to no pooling above it.
For a portfolio of shows, the crossover is the number you should know.
The weight, worked on a pacing curve
The quantity to pool is the share of final registrations a show holds at a fixed point in the campaign. Take day minus 60.
Across the eight shows the pooled share at day minus 60 is 58 per cent. Within a single show, the edition-to-edition standard deviation of that share is about 6 percentage points, so the within-show variance is 36. Across shows, the spread of true shares is tighter, around 3 points, so the between-show variance is 9.
The ratio is 36 divided by 9, which is 4. The pooled curve is worth four editions of a show's own data.
A show with three editions therefore gets a weight of 3 divided by 3 plus 4, which is 0.43. A show with twelve editions gets 12 divided by 12 plus 4, which is 0.75. The intraclass correlation, which is between-show variance over total variance, is 9 over 45, or 0.2, meaning a fifth of what you observe when you compare shares across shows is genuine difference between shows.
Note what the weight does not depend on. It does not depend on how big the show is, how much revenue it makes, or how confident its director sounds. Only edition count and the two variances.
The same measured share, two different forecasts
Both the three edition show and the twelve edition show have measured an average day-minus-60 share of 68 per cent, well above the pooled 58.
The three edition show is pulled to 0.43 times 68 plus 0.57 times 58, which is 62.3 per cent. The twelve edition show is pulled to 0.75 times 68 plus 0.25 times 58, which is 65.5 per cent. Same measurement, different answers, and the difference is entirely a statement about how much each measurement is worth.
Now turn those into registrations. Suppose both shows sit on 7,400 registrations at day minus 60.
Using its own share, the three edition show forecasts 7,400 divided by 0.68, which is 10,882. Using the pooled share it forecasts 7,400 divided by 0.58, which is 12,759. The two answers are 1,877 registrations apart, which is the size of the decision you are actually making. Partial pooling puts it at 7,400 divided by 0.623, which is 11,878. The twelve edition show, on the same count, forecasts 7,400 divided by 0.655, or 11,298.
The three edition show gets the higher forecast, which surprises people until they see why. Its own data said it registers early, the model only half believes that, and half-believing an early-registration claim means assuming more of the file is still to come.
Why does a show with twelve editions still get pooled at all?
Because 0.75 is not 1, and the remaining quarter is doing something useful.
Twelve editions of a share with a 6 point standard deviation gives a standard error on the show's own mean of 6 divided by the square root of 12, which is 1.7 points. That is good but finite, and the pooled curve carries independent information about where a show of this kind sits. Giving it a quarter weight costs almost nothing when the show's own estimate is right and helps when the show has had an odd run.
There is a second reason that matters more in a portfolio that keeps changing. A model with per-show weights handles a new acquisition, a show that skipped a year and a show in its twelfth edition with the same machinery. Nobody has to decide which shows are established enough to trust. The edition count decides, and it decides continuously.
Hyndman and Kostenko reached the same conclusion from the short-series side in Foresight in 2007. Their advice when data are scarce is to bring in other information alongside the available data, and they point specifically at analogous time series and Bayesian pooling as the route. A portfolio of shows is the cleanest supply of analogous series anyone is likely to have.
What the model needs from your data
Pool a quantity that is comparable across shows. A share, a growth rate, a curve parameter. Pooling raw registration counts across a 4,000 attendee show and a 40,000 attendee one produces a group mean that describes neither, and the shrinkage will drag the small show upward for no reason connected to its audience.
You also need enough shows for the between-show variance to be estimable. Estimating sigma alpha from three shows is guesswork, and the fitted value will often come back at zero, which silently collapses the model into complete pooling. Eight is workable. Four is where you should be checking the fitted variance components by hand before you believe anything.
Fitting all of this as one model across every series at once, with show identity as a feature the model reads, is a related but different construction, and a global forecasting model trained on the whole portfolio is O15's subject. Stacking the same data as show-year rows and running fixed effects is a third route, covered as pooled panel regression in O16.
Where this stops
Partial pooling assumes the shows are draws from one population. The moment that stops being true, the pooling is importing a pattern the show does not have.
A show that moved from a March slot to a June one has a registration campaign running through a different set of holidays, and pooling it with seven March shows will push its day-minus-60 share toward a number no June show would produce. Detecting that break and choosing an estimation window afterwards is O31 and O32's territory, and both belong upstream of any pooling you do.
The other limit is the one the arithmetic makes visible. The variance ratio of 4 was estimated from the same eight shows, and with eight groups the estimate of between-show variance is itself uncertain. If the true ratio is 2, the three edition show's weight rises from 0.43 to 0.60 and its forecast drops by about 300 registrations. That sensitivity is worth reporting alongside the forecast, because it is the honest width of the method rather than of the show.
A single weight for the whole portfolio, computed once and applied to every show, is the simpler cousin of this and is often good enough for a growth rate. The shrinkage estimator in O13 sets that version out, including how to get the two variances it needs.
Start by computing the day-minus-60 share for every closed edition you have, grouped by show. Take the standard deviation within each show and the standard deviation of the show means. Square both, divide the first by the second, and you have the number of editions your pooled curve is worth. Any show with fewer editions than that number is currently being forecast on evidence weaker than the portfolio it belongs to. The rest of what a system doing this needs sits with the forecasting methods behind it.
Questions people ask about borrowing strength across shows
- How is the pooling weight for a single show calculated?
- Gelman and Hill give it as n_j over sigma_y squared, divided by that quantity plus one over sigma_alpha squared. Cancelling gives a simpler reading: the weight is the show's edition count divided by that count plus the ratio of within-show variance to between-show variance.
- What does the variance ratio mean in practice?
- It is the number of editions the pooled pattern is worth. If within-show variance is 36 and between-show variance is 9, the ratio is 4, so the group-level distribution carries as much information as four editions of the show's own data. A show with fewer than four leans mostly on the group.
- Does partial pooling work when the shows are genuinely different sizes?
- Yes, provided you pool the shape and not the level. Pool a share, a growth rate or a curve parameter, all of which are comparable across shows of different sizes. Pooling raw registration counts across a 4,000 and a 40,000 attendee show produces a group mean that describes neither.
Related reading
- A shrinkage estimator pulls one noisy show forecast toward the portfolio mean
- A global forecasting model trained on every show in the portfolio at once
- Pooled panel regression turns eight shows and five years into forty rows