Skip to content

A shrinkage estimator pulls one noisy show forecast toward the portfolio mean

Forecasting methodsUpdated 2026-08-238 min read

In short

A shrinkage estimator replaces a single show's noisy growth estimate with a weighted average of that estimate and the portfolio mean. With a weight of 0.40, a show measured at 22 per cent growth against a portfolio mean of 6 per cent is reported at 12.4 per cent.

One show in the portfolio grew 22 per cent last edition. The rest of the group averaged 6 per cent. The sales director wants next year's target built on 22, the finance lead thinks 22 is a fluke, and neither of them has a way of settling it that is more than an opinion about the show.

A shrinkage estimator settles it with arithmetic. It reports a number between the two, and the position of that number is determined by how much of the spread across your shows is real difference and how much is measurement noise.

The estimate you have is the truth plus noise

Start with what the 22 per cent actually is. It is one year-on-year change, computed from two registration totals, on a show whose own history bounces around.

Write the measured growth of show i as the sum of its true underlying growth and an error term. The eight numbers you can see have a variance that combines both: the genuine spread between shows, and the noise in measuring each one. Call the first the between-show variance and the second the sampling variance.

If almost all the observed spread is genuine, the 22 should be taken at close to face value. If almost all of it is noise, the best estimate of every show is the portfolio mean, and 22 tells you nothing about this show that 6 does not tell you better. Everything in between is a weighted average, and the weight is a ratio of those two variances.

What weight belongs on one show's own number?

The empirical Bayes weight is the between-show variance divided by the total observed variance. Written as w, the shrunk estimate is w times the show's own number plus one minus w times the portfolio mean.

That is the whole method. The work is in getting the two variances, and both are available from data you already have.

Total observed variance comes from the eight numbers themselves. Sampling variance comes from each show's own history of year-on-year changes, because the spread of a show's growth rate around its own average is a direct measurement of how noisy a single year-on-year change is for that show.

The weight, worked on eight shows

Eight shows, year-on-year registration growth in per cent: 22, 9, 4, 1, minus 3, 7, 5 and 3. The mean is 48 divided by 8, which is 6.

Deviations from that mean are 16, 3, minus 2, minus 5, minus 9, 1, minus 1 and minus 3. Their squares sum to 386, and dividing by seven gives a sample variance of 55 and a standard deviation of 7.4 percentage points.

Now the noise. Across the portfolio, a single year-on-year change has historically varied by about 5.7 percentage points around a show's own average, so the sampling variance of one measured growth rate is roughly 33.

Between-show variance is what is left: 55 minus 33, which is 22. The weight is 22 divided by 55, which is 0.40.

Apply it to the show that grew 22 per cent. Nought point four times 22 is 8.8. Nought point six times 6 is 3.6. The shrunk estimate is 12.4 per cent.

Apply it to the show that shrank 3 per cent. Nought point four times minus 3 is minus 1.2, plus 3.6, which is 2.4 per cent. The worst performer in the portfolio is estimated to have grown slightly, because a single measured decline of 3 per cent against a noise level of 5.7 points is weak evidence of anything.

Two things fall out of this that are worth saying to a sales director. The gap between the best and worst show, 25 percentage points as measured, becomes 10 points after shrinkage. And the ordering does not change, so the show that grew fastest is still the show that grew fastest.

Two derivations, two weights, one bracket

The empirical Bayes weight is one route. James and Stein (1961), in a paper titled "Estimation with Quadratic Loss", give another, and on the same data it produces a different number, which is instructive.

That form keeps the portfolio mean and multiplies each deviation from it by one minus the quantity k minus 3 times the sampling variance, divided by the total sum of squared deviations, where k is the number of shows. Here that is one minus five times 33 over 386, which is one minus 0.43, or 0.57.

So one derivation says 0.40 and the other says 0.57. Show A comes out at 12.4 per cent under the first and 6 plus 0.57 times 16, which is 15.2 per cent, under the second.

Neither number is wrong. They make slightly different assumptions about what you know, and the useful reading is that the defensible answer for this show lies somewhere between 12 and 16, and nowhere near 22. If you want one number, take the more conservative of the two and say why.

What both derivations share is the guarantee James and Stein proved: estimating three or more quantities at once, the shrunken set has lower total squared error than the individual averages. Efron and Morris put the same result in plain language for Scientific American in May 1977, describing the case as one where "there are estimators better than the arithmetic average". The word doing the work in their sentence is "estimators", plural. The result is about the set.

From a shrunk growth rate to a shrunk forecast

The growth rate is an intermediate quantity. What operations wants is a registration number, and the translation is one multiplication that makes the size of the argument obvious.

Show A closed its last edition on 14,200 registrations. At the measured 22 per cent, next edition forecasts at 14,200 times 1.22, which is 17,324. At the shrunk 12.4 per cent it forecasts at 14,200 times 1.124, which is 15,961. The two numbers differ by 1,363 registrations, and that gap has to be carried by somebody: badge stock, catering covers, the shuttle schedule and the staffing roster all price off it.

Put both numbers on the sheet with the weight next to them. A show director who can see that the model kept 40 per cent of the show's own signal and 60 per cent of the group's will argue about the weight, which is the right argument to be having. A show director handed only 15,961 will argue about the model, which is not.

The same translation works in reverse for a decline. The show measured at minus 3 per cent closed on 8,600, so its raw forecast is 8,342 and its shrunk forecast, at plus 2.4 per cent, is 8,806. Shrinkage moves that show up by 464 registrations, and the reason is the same in both directions: one year of movement is thin evidence against a portfolio of eight.

Does shrinkage hide a real outlier?

It reduces one, and whether that is hiding depends on how much history the outlier has.

The weight is not a constant of nature. It moves with how well each show's own number is measured. Suppose instead of one year-on-year change you average three, which cuts the sampling variance by a factor of three, from 33 to about 11. The total observed variance falls to 22 plus 11, which is 33, and the weight rises to 22 over 33, or 0.67. The same measured 22 per cent now comes back as 0.67 times 22 plus 0.33 times 6, which is 16.7 per cent.

More evidence, less shrinkage. That relationship is exactly right, and it is also the reason a single weight applied to every show in the portfolio is a simplification. A show with twelve editions behind it deserves a weight near one and a show with three deserves one near zero, and giving each show its own weight is the multilevel version of this idea. That is O14's subject, and borrowing strength across shows when no single show has enough history sets out the per-show weight properly.

The genuine outlier case, where a show grew 22 per cent because it absorbed a competitor's audience or moved into a hall twice the size, is not a statistics problem. Shrinkage assumes the shows are interchangeable members of one group. If you can name the reason a show is different, take it out of the group before you compute anything, and forecast it on its own terms.

Where this stops

Shrinkage improves the total. It can make an individual show worse, and it will do so most often for the show everyone is watching.

If the show that grew 22 per cent really did grow 22 per cent, reporting 12.4 understates it by nearly ten points, and the person who owns that show will notice. The James and Stein result says the sum of squared errors across all eight shows goes down. It says nothing reassuring about the one show that mattered to the person reading the slide, and there is no version of this method that fixes that.

The second limit is the sampling variance, which is the input everybody eyeballs. If you set it too high you shrink everything to the mean and the portfolio looks flat. If you set it too low you barely shrink at all and the exercise was decoration. Estimating it from each show's own history of year-on-year changes is the defensible route, and it needs at least four or five changes per show to be worth anything, which is the same short-series problem that makes every model fit five points in O12.

There is also a scaling limit. Shrinking one summary statistic per show is a hand calculation. Doing the same thing to a whole forecasting model, so that every coefficient is partially pooled across the portfolio, means fitting one model to every series at once, which is a global forecasting model and O15's subject.

Take your last portfolio review, list the year-on-year growth of every show, and compute the variance of that list. Then take one show and compute the variance of its own year-on-year changes over the editions you have. If the second number is more than half the first, the differences your review is discussing are mostly noise, and the weight you should be putting on any single show's number is below 0.5. The broader question of which of these estimators a portfolio system should run sits with the forecasting methods it is built on.

Questions people ask about shrinkage estimator

How is the shrinkage weight calculated?
It is the ratio of between-show variance to total observed variance. If the eight measured growth rates have a variance of 55 and each measurement carries sampling variance of 33, the between-show variance is 22 and the weight on a show's own number is 22 divided by 55, which is 0.40.
Does shrinking a forecast make it less accurate for the individual show?
Sometimes, for one show. James and Stein proved in 1961 that for three or more quantities estimated together, the shrunken set has lower total squared error than the raw estimates. The guarantee is about the portfolio total, so an individual show can be moved in the wrong direction.
When should you not shrink a show's estimate?
When the show differs for a reason you can name. A venue move, a date change or a merged co-located event makes that show a poor member of the group whose mean you are shrinking toward. Shrinkage assumes the shows are exchangeable, and a known structural difference breaks that assumption.

Related reading

All forecasting methods articles