Bottom up versus top down forecasting for a portfolio of eight shows
Bottom up sums individual show forecasts and keeps every show's detail while inheriting every show's error. Top down forecasts the smoother portfolio total then splits it, which is stable at the top and unreliable underneath when shares drift. Compute both and compare them on held-out editions.
Two people build the group registration forecast every year and they never build it the same way. The audience acquisition lead forecasts each show and adds them up. The finance lead forecasts the group total and allocates it by last year's mix. The two numbers differ, and the argument about which is right is really an argument about which level of the business carries the signal.
Bottom up versus top down is a choice about where the error goes. Both directions are defensible, both discard something, and the choice can be settled on your own data rather than by seniority.
What each direction actually assumes
Bottom up assumes each show can be forecast on its own history and that the sum of eight such forecasts is a sensible total. It keeps every show's own trend, its own pacing, its own venue change, and the total inherits all eight sets of errors.
Top down assumes the portfolio total is the reliable quantity and that the split between shows is stable enough to apply from history. It gets the smoother series to model and pays for it underneath, because whatever the split gets wrong lands entirely on individual shows.
Athanasopoulos, Ahmed and Hyndman put the aggregate case plainly in the International Journal of Forecasting in 2009, working on Australian domestic tourism. Of top-down approaches they write that "the simplicity of the application of these top-down approaches is their greatest attribute", since you only have to model the most aggregated series, and that these approaches "seem to produce quite reliable forecasts for the aggregate levels". Then the cost: "On the other hand, their greatest disadvantage is the loss of information due to aggregation."
They spell out what gets lost. Under a top-down approach, they say, "we are unable to capture and take advantage of individual series characteristics such as time dynamics, special events, etc." In an events business those individual characteristics are most of what the show team knows. A hall extension, a co-location, a competitor that folded, a programme that moved: every one of them is a show-level fact that a portfolio-level model cannot see and a historical share cannot carry.
The share that drifted, worked on eight shows
Portfolio registrations across five editions: 98,000, 103,000, 106,000, 111,000 and 115,000. The group forecast for the next edition is 118,000.
Eight shows, with average historical shares over those five editions of 0.24, 0.18, 0.14, 0.12, 0.11, 0.09, 0.07 and 0.05. They sum to one, so top down splits cleanly. Applied to 118,000 the smallest show gets 0.05 times 118,000, which is 5,900 registrations.
Now look at how that show got to an average share of 0.05. Its shares across the five editions were 0.031, 0.038, 0.048, 0.058 and 0.071. They average to 0.049, and they have been climbing by about one percentage point of share per edition without a single reversal.
Applying the average share to next year assigns this show 5,900 when its most recent edition alone held 7.1 per cent of a smaller portfolio. In registrations, its five editions ran at 3,038, 3,914, 5,088, 6,438 and 8,165.
Forecast that series on its own. The least squares slope through those five points is 1,278 registrations per edition and the mean is 5,329, so the next edition forecasts at 5,329 plus three times 1,278, which is 9,163. Bottom up says 9,163. Top down with historical proportions says 5,900. The gap is 3,263 registrations on a show that will run about nine thousand, and the top-down number is 36 per cent low.
Athanasopoulos, Ahmed and Hyndman describe exactly this mechanism. Basing disaggregation on historical and static proportions means, in their words, that "these proportions will miss any trends in the data". Their conclusion on the tourism hierarchies is blunter still: the two top-down approaches based on static historical proportions "are only useful for forecasting the very top level of the hierarchies".
Which version of top down are you actually running?
The word top-down covers several methods that behave very differently, and most event teams have only met the weakest one.
The first averages each show's historical share and applies the average. That is the 5,900 above.
The second takes the proportion of the historical averages, dividing each show's mean registrations by the portfolio's mean registrations. On a portfolio where the small show grew and the big ones did not, this gives a slightly different set of shares and the same basic problem.
The third forecasts the shares themselves. Extrapolate the small show's share sequence, which rose by 0.007, 0.010, 0.010 and 0.013, giving a mean increase of 0.010 and a forecast share of about 0.081. Apply that to 118,000 and the show gets 9,558, within four hundred registrations of the bottom-up answer.
That third version was the 2009 paper's own proposal, and it is the one that won. Their evaluation found that "the top-down approach based on forecast proportions and the optimal combination method perform best for the tourism hierarchies we consider". Top down beat bottom up, but only in a form that forecasts the split instead of assuming it.
One caveat travels with all three. Any top-down method, the same authors note, has the disadvantage that "these approaches do not produce unbiased revised forecasts, even if the base forecasts are unbiased". You inherit a bias at the show level in exchange for a stable total.
Does summing eight noisy forecasts beat forecasting the total?
This is the part of the argument that is usually asserted and rarely computed, and it turns entirely on whether your shows move together.
Suppose each of the eight show forecasts has a standard error of 900 registrations. If those errors are independent, the standard error of their sum is 900 times the square root of 8, which is 2,546, or 2.2 per cent of a 115,000 registration portfolio. Bottom up looks excellent at the top, because eight independent errors partly cancel.
Now suppose the errors are perfectly correlated, which is what a market-wide event produces: a travel disruption, a sector downturn, a competitor launching against your whole portfolio. The standard error of the sum becomes 8 times 900, which is 7,200, or 6.3 per cent. The cancellation vanishes and bottom up is now considerably worse at the top than a direct forecast of the total would be.
Real portfolios sit between the two, and you can measure where. Take your last several editions, compute each show's year-on-year growth, and correlate the eight series with each other. If the average pairwise correlation is near zero, bottom up is safe at the top. If it is 0.5 or more, your shows share a market and the aggregate deserves its own model.
That measurement takes an afternoon and it settles the argument better than any general claim about which direction is superior.
What to do with both numbers
Compute both. Then score both on the editions you have already closed.
For each of the last two or three editions, produce the bottom-up forecast and the top-down forecast as they would have looked at the time, and compare their absolute errors at the level you actually care about. Most portfolios care about two levels at once, the total for finance and the individual show for operations, and it is common to find that top down wins at the total and bottom up wins at the show. That result is not a contradiction. It is the reason a third option exists.
The third option keeps every base forecast and adjusts all of them by the smallest amount that makes the set add up, which is hierarchical forecast reconciliation and O17's subject. Hyndman, Ahmed, Athanasopoulos and Shang set out the combination version of it in Computational Statistics and Data Analysis in 2011, describing forecasts that add up appropriately across the hierarchy, are unbiased and have minimum variance among all combination forecasts under some simple assumptions. If both of your current methods are defensible, that is the argument for keeping both and combining them.
The individual show forecasts feeding a bottom-up sum are also the place where noise does most damage, and pulling a single volatile show toward the portfolio before summing is what a shrinkage estimator does in O13.
Where this stops
Neither direction rescues a portfolio whose composition changed. If you acquired two shows and divested one during the five editions you are averaging over, the historical shares describe a portfolio that no longer exists, and the top-down split is allocating a total across a membership list from three years ago. Rebuild the shares on a consistent set of shows before you compute anything, and accept that this shortens your history.
Bottom up has an equivalent failure that is easier to miss. A show forecast on its own five editions carries no information about the market, so eight bottom-up forecasts will all miss a shared downturn in the same direction at the same time, and the total will be wrong by the full amount. The correlation check above is the early warning for that, and it is worth running before every planning round rather than once.
The same choice appears on the time axis, where a weekly pacing forecast and an annual forecast of the same show disagree, and temporal hierarchy forecasting handles it in O19.
This week, take your smallest show and plot its share of the portfolio across every edition you have. If that line has a slope, your top-down forecast for it is wrong by a predictable amount and you can compute exactly how much before anyone else notices. The wider question of which of these a portfolio system should compute sits with the forecasting methods behind it.
Questions people ask about bottom up versus top down
- When does top down forecasting go wrong for a show portfolio?
- When a show's share of the portfolio is moving. Athanasopoulos, Ahmed and Hyndman note that disaggregating by historical and static proportions means those proportions will miss any trends in the data. A show whose share drifted from 3.1 to 7.1 per cent over five editions gets its average share of 5 per cent applied, and is understated badly.
- Does bottom up forecasting produce a better portfolio total?
- Only if the show-level errors are close to independent. Eight show forecasts each with a standard error of 900 registrations sum to a standard error of 2,546 if independent, and to 7,200 if they all move together. A shared market shock destroys the cancellation that makes bottom up attractive at the top.
- Is there a version of top down that handles drifting shares?
- Yes. Forecast each show's share forward instead of averaging its past shares, then apply the forecast proportions to the portfolio total. Athanasopoulos, Ahmed and Hyndman proposed this and found it among the best performers on their tourism hierarchies, ahead of both conventional top down and bottom up.
Related reading
- A shrinkage estimator pulls one noisy show forecast toward the portfolio mean
- Hierarchical forecast reconciliation makes show and portfolio numbers add up
- Temporal hierarchy forecasting reconciles the weekly pace with the annual number