Skip to content

Pickup model diagnostics that tell you which form to trust this year

Forecasting methodsUpdated 2026-08-238 min read

In short

Pickup model diagnostics compare the two forms empirically. For each closed edition, predict its final from the other editions using the additive rule and again using the multiplicative rule, then compare the standard deviations of the two residual sets. Five editions give five residuals per model, which is a weak but usable signal.

Two people bring two forecasts to the same Thursday meeting. One says 5,750, built by adding the average gain from the last four editions. The other says 7,200, built by dividing by the average share of final. The gap is 1,450 registrations, which is a hall, a badge run and a shuttle contract.

The meeting resolves this by seniority, or by whichever number the show director likes, and then everyone forgets that a defensible answer was available in twenty minutes. Pickup model diagnostics are that answer. You already have the data to decide which form fits, because every closed edition is a test case you never ran.

Make the choice a measurement

The procedure is a backtest on the editions you have already finished.

For each closed edition, pretend you did not know its final number. Predict it two ways, using only the other editions. The additive prediction is that edition's count at the chosen days-to-open point plus the average pickup of the other editions, which is O1's method. The multiplicative prediction is the same count divided by the average share of final of the other editions. Record both errors. Then compare the spread of the two error sets.

Leaving each edition out of its own prediction matters more here than it looks. Include an edition in the average that is used to predict it and both models get a free hint, the additive one especially, because a single unusual edition drags the mean toward itself and then gets rewarded for it. With five editions the leave-one-out version costs you five short calculations.

The procedure, worked on five editions

A packaging show, measured at 90 days before doors.

The counts at day 90 across the five editions were 1,400, 1,600, 1,900, 2,200 and 2,600. The finals were 3,980, 4,390, 5,240, 6,080 and 7,290. The absolute pickups are 2,580, 2,790, 3,340, 3,880 and 4,690. The shares of final are 0.3518, 0.3645, 0.3626, 0.3618 and 0.3567.

Start with the 2021 edition. The other four pickups average 3,675, so the additive prediction is 1,400 plus 3,675, which is 5,075 against an actual 3,980. The residual is minus 1,095. The other four shares average 0.3614, so the multiplicative prediction is 1,400 over 0.3614, which is 3,874. That residual is plus 106.

Repeat for the rest. The additive residuals come out at minus 1,095, minus 832, minus 145, plus 530 and plus 1,542. The multiplicative residuals come out at plus 106, minus 77, minus 57, minus 50 and plus 71.

The standard deviation of the additive residuals is 1,069 registrations. For the multiplicative residuals it is 84. That is a factor of about 12.8, and it is not a close call.

Look at the shape of the additive residuals as well as their size. In edition order they run from minus 1,095 up to plus 1,542, moving steadily in one direction as the show grows. That ordering is the tell. A model whose errors track the size of the thing being forecast has a scaling problem, and on a pickup model the scaling problem has exactly one cure, which is switching to the ratio form that O2 sets out.

What does the ratio of the two spreads tell you?

It tells you how much better one form is on the history you have, and it tells you nothing about whether that difference is real.

A factor of 12.8 is comfortable. Five residuals with that separation would be very hard to produce by accident, and I would switch forms on it without hesitating. A factor of 1.2, which is what most shows actually produce, tells you almost nothing, because the standard error of a standard deviation estimated from five numbers is enormous. Roughly, the relative standard error of a sample standard deviation on n observations is one over the square root of twice n minus two, which on five observations is about 0.35. A measured spread of 1,069 carries an implied range of something like 700 to 1,450 before you argue about anything else.

So the diagnostic has three outcomes and it is worth naming all three. A large gap decides it. A small gap means keep whichever form you are already using, since churn has its own cost and you have no evidence. A gap that flips direction depending on which days-to-open point you measure at means the two forms are equivalent for your purposes and the choice should be made on explainability.

Whether a measured gap is larger than noise has a formal treatment with a test statistic behind it, and comparing two forecasting methods on four error pairs is O27's subject. The diagnostic here is deliberately cruder and runs in a spreadsheet.

Why is five residuals such a weak signal?

Because model selection is itself a form of fitting, and you are spending the same scarce observations twice.

Green and Armstrong put the general case in the Journal of Business Research in 2015, reviewing 97 comparisons across 32 papers on simple against complex forecasting methods. None of the papers gave a balance of evidence that complexity improved accuracy, and in the 25 papers with quantitative comparisons complexity increased forecast error by 27 per cent on average. Their definition of simplicity is worth borrowing whole: a forecasting process is simple when the method, the way prior knowledge enters it, the relationships inside the model and the link between model, forecast and decision are all understandable to the people using the forecast.

Both pickup forms clear that bar. Anyone in the operations meeting can follow an average increment or an average share. That is the reason this diagnostic is allowed to be a comparison of two candidates instead of a search across twenty. Run a model selection loop over twenty variants on five editions and you will find one that fits beautifully and forecasts badly, and you will have no way of telling which you have.

Keep the candidate list at two, or at most three if you want to add a weighted version of the ratio form. Write the residuals down. Do the comparison once a year when an edition closes, rather than every time somebody dislikes a number.

The tell that does not need a backtest

There is a faster check that agrees with the backtest most of the time, and it takes five points on a chart.

Plot each closed edition's absolute pickup against the count it was holding at the same days-to-open value. On the packaging show those pairs are 1,400 with 2,580, 1,600 with 2,790, 1,900 with 3,340, 2,200 with 3,880 and 2,600 with 4,690. The correlation is 0.997. The additive model's central assumption is that the pickup is unrelated to the count in hand, and that plot has already falsified it before anybody computes a residual.

The same idea appears in the exponential smoothing literature as the choice between additive and multiplicative components. Hyndman and Athanasopoulos give the rule for seasonality in the third edition of Forecasting: Principles and Practice (2021). They prefer the additive form where the seasonal swings stay "roughly constant through the series", and the multiplicative form where those swings change in proportion to the level of the series. Substitute pickup for seasonal swings and the sentence becomes the diagnostic. Constant in absolute terms means additive. Proportional to the level means multiplicative.

The reason to run the residual version anyway is that the plot cannot tell you how much the choice costs. On the packaging show it was worth about 985 registrations of residual spread, and that is the number worth quoting when somebody asks why the forecast method changed.

Where this stops

The diagnostic answers a narrow question and people will try to make it answer a wider one.

It compares two specific forms at one specific days-to-open point. Run it at day 120 and at day 45 and you can easily get different winners, because the ratio form's weakness is concentrated early in the campaign and the additive form's weakness is concentrated on the shows that have grown most. If your operations calendar needs a number at both points, run the diagnostic at both and be prepared to use different forms at each, which sounds inconsistent and is simply what the evidence supports.

It also assumes the five editions are comparable, and that is doing quiet work. One venue move, one date shift, one edition that opened registration eleven weeks late, and the residuals you are comparing include a structural difference that neither model was ever going to catch. The diagnostic will still return a winner. It will be the form that happens to absorb that particular disruption better, which is a property of one historical accident rather than of your show.

The last limit is the one that applies to every backtest on annual data. Five residuals is the entire evidence base, and one more year adds one more residual. That is why the sensible use of this is a standing habit rather than a project. Add the new residual pair when an edition closes, keep the running comparison in the same file as the snapshots, and let the evidence accumulate at the only rate it can.

Take your last five closed editions, pick the days-to-open point your operations calendar actually needs, and compute the ten numbers: five additive residuals and five multiplicative ones, each leaving its own edition out. Put both standard deviations at the top of the page. That page is the answer to the next Thursday argument, and it belongs in the same file as the snapshots and the rest of the forecasting methods work.

Questions people ask about pickup model diagnostics

How do you choose between additive and multiplicative pickup?
Backtest both on your own closed editions. For each edition, forecast its final using the average increment from the other editions, then again using the average share of final from the other editions, and record both errors. Compare the standard deviation of the two error sets. The smaller spread wins, provided the gap is large.
How large does the gap between the two residual spreads need to be?
Large enough that four or five noisy observations could not have produced it by chance. A residual standard deviation of 1,069 registrations against 84 is a factor of twelve and settles the question. A gap of 1,069 against 890 does not, and the honest response there is to keep the simpler rule and revisit after another edition closes.
Can you tell which pickup form fits without running a backtest?
Partly. Plot each closed edition's absolute pickup against the count it held at the same days-to-open point. If the pickup grows with the count, the additive assumption of a constant increment is already contradicted. That plot takes five points and one minute, and it agrees with the backtest most of the time.

Related reading

All forecasting methods articles