Prediction interval coverage is worse than the label claims with five editions
Prediction interval coverage falls well below the nominal level on a five edition series because the usual formula ignores parameter uncertainty and uses a normal multiplier. Replacing 1.96 with a t multiplier and adding the estimation terms widens a worked example by a factor of 2.35.
Somebody puts a forecast on a slide with a band around it and calls the band a 95 per cent interval. The number came out of a package, the package printed the interval, and nobody in the room has any reason to doubt it.
Prediction interval coverage is the property that band is claiming, and on a series of five annual editions the claim is usually wrong by a wide margin. The interval is too narrow, the reasons are arithmetic rather than mysterious, and you can compute the correct width by hand in about ten minutes.
Where the 1.96 came from and why it does not apply
Hyndman and Athanasopoulos give the standard construction in Forecasting: Principles and Practice, third edition: a 95 per cent interval is the point forecast plus or minus 1.96 times the estimated standard deviation of the forecast distribution, with the multiplier depending on the coverage probability you want.
That 1.96 comes from the normal distribution and it is correct when the standard deviation is known. On five editions it is not known. It is estimated, from a handful of residuals, and using a normal multiplier on an estimated standard deviation is the first of two understatements.
The second is that the standard error itself is usually the residual standard deviation and nothing else, which treats the fitted level and slope as if they were the truth. They are estimates too, and the forecast inherits their uncertainty.
Chatfield reviewed the whole territory in the Journal of Business and Economic Statistics in 1993, distinguishing theoretical formulas based on a fitted probability model with and without a correction for parameter uncertainty, approximate formulas that he argues should be avoided, and empirical, simulation and resampling procedures. The distinction between the corrected and uncorrected theoretical formula is the one that decides whether your band means anything.
The five editions, worked
Verified attendance across five editions: 7,300, 7,650, 7,500, 8,050 and 8,400. The mean is 38,900 divided by 5, which is 7,780.
Fit a straight line. Centring the edition index at zero, the slope is 2,600 divided by 10, or 260 per edition. The fitted values are 7,260, 7,520, 7,780, 8,040 and 8,300, so the residuals are 40, 130, minus 280, 10 and 100. Their squares sum to 107,000.
Three residual degrees of freedom, because two parameters went on the level and the slope. The residual variance is 107,000 divided by 3, which is 35,667, and the residual standard deviation is 189.
The forecast for edition six is 7,780 plus three times 260, which is 8,560.
Here is the interval most people publish. Take 1.96 times 189, which is 370, and report 8,190 to 8,930. Width 741 attendees.
Now the correct version. The standard error for a new observation is the residual standard deviation multiplied by the square root of one, plus one over n, plus the squared distance of the forecast point from the centre of the data divided by the sum of squared distances. That is the square root of 1 plus 0.2 plus 0.9, which is the square root of 2.1, or 1.449. So the standard error is 189 times 1.449, which is 274.
The multiplier is the t value at three degrees of freedom, which is 3.182 rather than 1.96.
The interval is 8,560 plus or minus 3.182 times 274, which is plus or minus 872, giving 7,688 to 9,432. Width 1,744 attendees.
The correct interval is 2.35 times as wide as the published one. Two thirds of that comes from the multiplier, which is 62 per cent wider on its own, and the rest from the estimation terms, which add another 45 per cent.
How much does the interval actually cover?
This is the number worth quoting internally, because it converts an abstract complaint into a rate.
The published interval was plus or minus 370. In units of the correct standard error of 274, that is plus or minus 1.35. The predictive distribution for a new observation from this fit is a t distribution on three degrees of freedom, and the probability that such a variable falls within 1.35 of zero is about 0.73.
So the band labelled 95 per cent covers about 73 per cent. Out of four future editions you would expect one to fall outside it, against the one in twenty the label promises.
That is with everything else assumed correct. If the show has a genuine chance of a step change, coverage is worse still, and the interval is silent about it.
What the standard formula still leaves out
Even the corrected interval is optimistic, and the textbook says so directly.
Hyndman and Athanasopoulos, discussing intervals from ARIMA models, write that they "tend to be too narrow. This occurs because only the variation in the errors has been accounted for." They then list what is missing: "There is also variation in the parameter estimates, and in the model order, that has not been included in the calculation. In addition, the calculation assumes that the historical patterns that have been modelled will continue into the forecast period."
The t correction above handles the first of those three for a simple regression. The second, uncertainty about which model is right, is not handled by anything in the standard formula. On five observations you chose between a mean, a line and a curve on evidence so thin that almost every candidate fits, which is O12's subject, and none of that uncertainty appears in the band around whichever one you picked.
The third is the assumption that the past describes the future. For a show that has kept the same venue, the same month and the same market, it is reasonable. For a show that just moved city it is not, and no widening of the interval repairs it, because the distribution being widened is the wrong distribution.
Can you just widen the interval until it is honest?
Partly, and it is better than the alternative, but the honest version has a cost people find hard to accept.
The corrected interval on this series runs from 7,688 to 9,432. Presented to an operations team planning badge stock and catering, that is a range of 1,744 on a base of 8,560, or plus and minus about 10 per cent. Somebody will say it is too wide to be useful. They are describing the state of the evidence accurately.
Narrowing it honestly requires more information. Hyndman and Kostenko put the scaling law plainly in Foresight in 2007: as sample size increases, the intervals "decrease at a rate proportional to the square root of n", and "if you quadruple the length of the series, you will cut in half the width of a PI". Four times five editions is twenty editions. On an annual show that is fifteen more years, which is a fine plan for 2041.
The route that works now is to stop treating the series as five observations. A model that estimates the shared pattern across the whole portfolio and a per-show deviation is fitted to forty annual observations, and its parameter uncertainty is correspondingly smaller. That is borrowing strength across shows in O14, and narrower intervals are the main practical reason to do it.
A distribution-free alternative exists as well, which builds the interval from the empirical distribution of past forecast errors instead of from a fitted model, and conformal prediction intervals are O21's subject.
Where this stops
Everything above assumes the residuals are independent draws from one distribution with constant variance. On five points you cannot test that, and the tests you might run have no power at that length.
The variance assumption is the one most likely to be wrong for a show. A 4,000 attendee edition and a 12,000 attendee edition do not have the same absolute error, so if your series grew substantially over the five editions, the residual standard deviation of 189 is an average of a smaller early error and a larger recent one, and the interval for edition six is too narrow again. Working in logs handles proportional errors and costs nothing, and it is worth doing whenever the series has moved by more than about half its own starting value.
There is also an uncomfortable point about what an interval is for. A band of plus and minus 10 per cent that is honestly derived will be ignored by the same people who acted on a band of plus and minus 4 per cent that was not. The defence is to publish the coverage rate alongside the interval, by recording every published forecast and checking afterwards how often the outcome landed inside. After four or five editions you have a measured hit rate, and a measured hit rate is an argument that survives contact with a budget meeting.
Counting what a method estimates before you fit it is the upstream discipline that keeps residual degrees of freedom from disappearing, and the parameter budget five observations leaves you covers it in O11.
This week, take the last forecast your team published with a band on it, find the residual standard deviation and the number of residual degrees of freedom behind it, and recompute the band with the t multiplier and the estimation terms. If the recomputed band is more than twice the published one, the published one was decoration, and the version you now have is the first honest statement of uncertainty that show has produced. What a system needs to hold to compute that automatically sits with the forecasting methods behind it.
Questions people ask about prediction interval coverage
- Why is a 95 per cent prediction interval too narrow on five observations?
- Because the 1.96 multiplier assumes the error variance is known, and it is estimated from three residual degrees of freedom. The correct multiplier is the t value at those degrees of freedom, 3.182, and the standard error must also carry the uncertainty in the fitted level and slope, which adds a further 45 per cent.
- What coverage does a nominal 95 per cent interval actually achieve?
- On the worked example here, about 73 per cent. The published interval of plus or minus 370 registrations corresponds to 1.35 standard errors of the correct predictive distribution, and for a t distribution on three degrees of freedom that spans roughly 0.73 of the probability. Roughly one edition in four falls outside.
- Can you make the interval narrower honestly?
- Only with more information. Hyndman and Kostenko note that interval width falls with the square root of the sample size, so halving it needs four times the data, which on an annual show is fifteen further years. Pooling across the portfolio adds observations now, which is the practical route.
Related reading
- Forecasting with five observations and the parameter budget that leaves you
- Overfitting short time series is easy when every model fits five points
- Borrowing strength across shows when no single show has enough history