Hierarchical forecast reconciliation makes show and portfolio numbers add up
Hierarchical forecast reconciliation takes forecasts produced independently at every level of a portfolio and projects them onto the set of forecasts that add up. Every number moves, including the aggregate, and the reconciled total lands between the sum of the shows and the independently forecast portfolio figure.
The portfolio review pack has two numbers on it that cannot both be right. The finance model forecasts 44,000 registrations across the group next year. The eight show forecasts, added together, come to 44,500. Nobody put the 500 there. It arrived because two different people forecast two different things well.
Hierarchical forecast reconciliation is the formal treatment of that gap. It takes the forecasts you already produced at every level and adjusts all of them by the smallest amount that makes the set add up.
Why do independently produced forecasts never add up?
Because each level is a different series with a different signal-to-noise ratio, and a good forecast of one is not the sum of good forecasts of the others.
An individual show's registration series is noisy. A booth relocation, a competitor moving dates, a single large association block booking, all of it lands on one series and none of it is predictable. The portfolio total is much steadier, because the idiosyncratic movements of eight shows partly cancel. A model fitted to the total will therefore find a cleaner trend, and it will produce a number that has no reason to equal the sum of eight noisy forecasts.
The instinct is to declare one level authoritative and derive the rest from it. Sum the shows and overwrite the total, or forecast the total and split it by share. Both are available, both discard information, and the comparison between them is O18's.
Reconciliation takes a different position: every level was informative, so use all of them.
The summing matrix, on a portfolio of four shows
Set up a small hierarchy so the arithmetic stays visible. Four shows, two regions, one portfolio total. Seven series in all, four of which are at the bottom.
Hyndman and Athanasopoulos define the structure in Forecasting: Principles and Practice, third edition, as an n by m matrix S, "the summing matrix", which "dictates the way in which the bottom-level series aggregate", so that the full vector of all series equals S times the vector of bottom-level series.
Here S has seven rows and four columns. The top row is one, one, one, one. The row for region A is one, one, zero, zero. The row for region B is zero, zero, one, one. The four show rows form an identity block. Read it as a set of instructions: the total is the sum of all four, region A is shows one and two, and each show is itself.
Any forecast vector you can write as S times something is coherent. Everything else is not, and the whole method is a projection back onto the space S describes. Reconciliation is written as S times G times the base forecasts, where G maps the base forecasts down to the bottom level. The same book gives the condition under which this leaves you unbiased: the reconciled forecasts are unbiased "provided S*G*S = S".
Reconciliation, worked
Base forecasts, produced independently at every level. Show one at 12,000, show two at 9,500, show three at 15,200, show four at 7,800. Region A forecast on its own history at 21,900, region B at 22,600. The portfolio total, forecast on its own history, at 44,000.
Nothing agrees. The shows in region A sum to 21,500 against a regional base forecast of 21,900. Region B's shows sum to 23,000 against a base of 22,600. The four shows together give 44,500 against a total base forecast of 44,000.
Take the ordinary least squares projection, which is the version Hyndman, Ahmed, Athanasopoulos and Shang published in Computational Statistics and Data Analysis in 2011 as an optimal combination. It solves for the bottom-level vector that minimises the squared distance to all seven base forecasts at once, and on these numbers it returns 12,062 and 9,562 for region A's shows, and 14,995 and 7,595 for region B's.
Look at what moved. Both region A shows rose by 62 registrations. Both region B shows fell by 205. Region A's reconciled total is 21,624, which is 276 below its own base forecast. Region B's is 22,590, ten below its base. The portfolio total comes out at 44,214, which sits between the 44,000 the total model produced and the 44,500 the shows summed to.
That last number is the point of the whole exercise. Neither original candidate survived. The reconciled total used the evidence from the total's own model, from both regional models and from all four show models, and landed 214 above one and 286 below the other. The 2011 paper describes the resulting forecasts as adding up appropriately across the hierarchy, unbiased, and of minimum variance among all combination forecasts under some simple assumptions.
Notice also that the adjustment was not shared out evenly. Region B absorbed 410 registrations of correction against region A's 124, because region B's internal disagreement was larger. The projection allocates by how much each part of the structure disagrees with the rest.
What minimum trace adds to the projection
The least squares version treats every base forecast as equally reliable. It is not, and that assumption costs accuracy in exactly the case a show portfolio produces.
Wickramasuriya, Athanasopoulos and Hyndman took this up in the Journal of the American Statistical Association in 2019. Their reconciliation weights levels by the covariance of the base forecast errors, giving the reconciliation matrix as the inverse of S transpose times W inverse times S, all times S transpose times W inverse, where W is that covariance. Minimising the trace of the reconciled error covariance is what gives the method its name.
Their paper also disposes of the obvious earlier idea. Reconciling by the covariance of the reconciliation errors themselves, they note, is unusable because "this matrix is impossible to estimate in practice due to identifiability conditions". Using the base forecast error covariance instead is what makes the thing computable.
For an event portfolio the consequence is direct. The portfolio total forecast typically has a much smaller relative error than any individual show forecast, because the aggregate is smoother. Under minimum trace the total is therefore treated as more reliable and moved less, and the shows are moved more. Under least squares, as in the worked example, the total moved 214 registrations, which on a reliable aggregate is more than it should.
You need error variances per series to run it. Five editions gives five one-step errors per series, which is thin but usable if you shrink the covariance toward a diagonal, and a diagonal weighting by each series' own error variance is the practical version most teams will get to first.
Does reconciliation make the individual show forecasts worse?
For some shows, yes, and this needs saying before the method is presented to show directors.
In the worked example, show three's forecast fell from 15,200 to 14,995. If the show three model was the best model in the building and the region B total model was weak, that 205 registration reduction is an error being imported from elsewhere. The guarantee is about the variance of the whole reconciled set, and an individual series can be moved in the wrong direction.
What makes this acceptable in practice is that the alternative is worse. Without reconciliation, somebody in the review meeting adjusts a number by hand to close the gap, the adjustment is undocumented, and the same 500 registrations get allocated by whoever speaks last. A projection with a stated weighting is at least a rule that behaves the same way every quarter.
The defensible presentation is both numbers side by side: base forecast and reconciled forecast, per show, with the size of the adjustment. A show director seeing minus 205 next to their number will ask why, and the answer, that region B's own model disagreed with the sum of its shows by 400, is a real answer.
Where this stops
Reconciliation fixes coherence. It does not fix a hierarchy that describes the business badly.
If your shows are grouped by region but the thing that actually moves them together is vertical, then the regional level carries no signal and reconciling through it adds noise. Grouped structures, where a show belongs to a region and a vertical at the same time, are supported by the same algebra, and they are usually a better description of a portfolio than a strict tree. They also multiply the number of series whose error variances you have to estimate.
The second limit is data. Minimum trace wants a covariance matrix estimated from held-out forecast errors, and a portfolio of eight shows with five editions each yields very few of them. With a diagonal approximation you are estimating eight variances from five errors apiece, which is the same short-series constraint that runs through everything else in this cluster, and it is why the least squares version is often the honest choice.
Reconciliation across levels of a portfolio is one axis. The same algebra applies down the time axis, reconciling a weekly pacing path with a quarterly and an annual forecast of the same show, and temporal hierarchy forecasting covers that in O19. Choosing to derive every level from one instead, either by summing up or by splitting down, is bottom up versus top down in O18. Estimating the shared portfolio movement inside a single regression, rather than reconciling separate models afterwards, is pooled panel regression in O16.
Start by measuring your own incoherence. Take the last forecast pack, add up the show numbers, and subtract the portfolio number somebody produced separately. If the gap is under half a per cent, reconciliation will change little and you can leave it. If it is 500 registrations on 44,000, somebody is currently closing that gap by hand every quarter, and it is worth knowing who. The wider question of which of these methods a portfolio system should run sits with the forecasting methods behind it.
Questions people ask about hierarchical forecast reconciliation
- What is the summing matrix in hierarchical forecasting?
- It is the matrix that encodes how the bottom level aggregates. For a portfolio of four shows in two regions it has seven rows, one per series, and four columns, one per show. The top row is all ones, the region rows have ones against their own shows, and the show rows form an identity block.
- Which forecasts change when you reconcile?
- All of them. Reconciliation is a projection, so the portfolio total, the region totals and every individual show forecast move. On a four show example with a 500 registration discrepancy, two shows rose by 62 each, two fell by 205 each, and the total settled at 44,214 between the two original candidates.
- Why is minimum trace better than reconciling with equal weights?
- Because base forecasts at different levels have different error variances, and equal weights ignore that. Wickramasuriya, Athanasopoulos and Hyndman derived a weighting from the base forecast error covariance that minimises the total variance of the reconciled set, and it uses information from every level rather than one.
Related reading
- Pooled panel regression turns eight shows and five years into forty rows
- Bottom up versus top down forecasting for a portfolio of eight shows
- Temporal hierarchy forecasting reconciles the weekly pace with the annual number