Skip to content

Temporal hierarchy forecasting reconciles the weekly pace with the annual number

Forecasting methodsUpdated 2026-08-237 min read

In short

Temporal hierarchy forecasting aggregates a single series to weekly, monthly, quarterly and annual levels, forecasts each level separately, then reconciles so the levels agree. For a show it forces the pacing model and the board number into one answer instead of two that never quite match.

The pacing report says the show lands at 13,890. The annual forecast in the budget says 14,600. Both were built by competent people from the same registration file, and the 710 registration gap between them gets resolved every Monday by whoever is presenting.

Temporal hierarchy forecasting is the method that stops that happening. The same series is aggregated to several frequencies, each frequency is forecast on its own, and the results are reconciled so that the weekly path and the annual number are the same statement.

The two numbers for the same show that never agree

The gap has a structural cause. A weekly pacing model is fitted to weekly increments, which are dominated by campaign mechanics: an email send, a price tier closing, an exhibitor pushing guest passes. An annual model is fitted to five final numbers and sees only the trend.

Neither model is looking at the other's evidence. The weekly model has no view about whether this edition should be bigger than last, because year-on-year information barely appears in weekly increments. The annual model has no view about whether the campaign is currently four weeks behind, because it has never seen a week.

So the two numbers disagree, and the disagreement is information rather than a defect. Both models know something the other does not.

What does temporal aggregation do to the signal?

It changes which components of the series you can see.

Kourentzes, Petropoulos and Trapero built a forecasting method on exactly this property in the International Journal of Forecasting in 2014. Multiple series are constructed from the original by temporal aggregation, and in their words "these derivative series highlight different aspects of the original data, as temporal aggregation helps in strengthening or attenuating the signals of different time series components". They fit an exponential smoothing model at each aggregation level, forecast the components separately, and combine.

For a registration campaign the effect is easy to see. At weekly resolution the series is mostly noise with sharp spikes at deadlines. Sum to four-weekly and the spikes flatten into a smooth acceleration toward doors. Sum again to twelve-weekly and what remains is the shape of the campaign in three or four numbers. Sum to the edition and you have a single point on a five-point trend.

Their reported result is that this approach delivers improvements in forecasting accuracy, and that the improvement is largest for long-term forecasts. Long-term is where a show forecast lives, since the number that matters is the one at doors and the question is being asked in week ten.

Building the hierarchy for one campaign

A 48 week campaign aggregates cleanly. Forty eight weekly observations, twelve four-weekly, four twelve-weekly, one annual figure. Each divides the one below it exactly, which is what non-overlapping aggregation requires.

Athanasopoulos, Hyndman, Kourentzes and Petropoulos set out the framework in the European Journal of Operational Research in 2017. Their opening statement is the general one: "A temporal hierarchy can be constructed for any time series by means of non-overlapping temporal aggregation. Predictions constructed at all aggregation levels are combined with the proposed framework to result in temporally reconciled, accurate and robust forecasts."

The algebra is the same algebra used to reconcile shows against a portfolio total, with the aggregation running down the time axis instead of across the business. That is worth knowing because it means one implementation covers both, and hierarchical forecast reconciliation in O17 sets out the summing matrix and the projection that both share.

The 2017 paper makes a second point that matters operationally. The framework is independent of the forecasting models used at each level, and it can take in managerial forecasts alongside statistical ones. A show director's own number for the final total is a legitimate input at the annual level, and reconciliation will move it and everything else toward agreement instead of leaving it as a separate opinion in a separate document.

The reconciliation, worked on the last four weeks

Four weeks to doors, 9,400 registrations already in the file. Three models are running.

The weekly model forecasts new registrations of 620, 780, 1,150 and 1,940 across the remaining weeks, summing to 4,490.

A two-weekly model, fitted to two-week blocks across previous editions, forecasts 1,500 for the first block and 3,250 for the second, summing to 4,750.

A four-week model, fitted to the final four-week block of each previous edition, forecasts 5,200 for the whole remaining period.

Three answers: 4,490, 4,750 and 5,200. The least squares projection onto the coherent space returns weekly figures of 730, 890, 1,280 and 2,070.

Check what happened. The two-week blocks now come out at 1,620 and 3,350, both above their own base forecasts of 1,500 and 3,250. The four-week total is 4,970, below its base of 5,200 and well above the weekly sum of 4,490. Every one of the four weekly figures rose, the first two by 110 registrations each and the last two by 130.

The final forecast is 9,400 plus 4,970, which is 14,370. That number is not the pacing report's 13,890 and not the budget's 14,600. It uses both, and it comes with a weekly path that operations can staff against, because the reconciled weekly figures still sum to it.

Why does combining frequencies improve accuracy?

Because it spreads the risk of choosing the wrong model, and on short show histories the model choice is the largest single source of error.

If you forecast at one frequency you have made one bet on one specification. If you forecast at four frequencies and reconcile, a badly specified model at one level is diluted by three others fitted to differently aggregated versions of the same data, which respond to different failure modes. Kourentzes, Petropoulos and Trapero framed their own method the same way, as an algorithm that aims to mitigate the importance of model selection while increasing accuracy.

The 2017 framework reports its largest gains where modelling uncertainty is highest, illustrated with a case study on emergency department arrivals. A trade show with five editions is squarely in the high-uncertainty case. Nobody knows whether the annual series is a trend or a random walk with drift, and five observations cannot settle it, which is O11's subject. Reconciling across frequencies means you do not have to settle it before Friday.

What it costs to run

More models, and one awkward data requirement.

Four levels means four fitted models per show, and eight shows means thirty two. That is a scheduling job rather than a hard problem, and each model is small.

The awkward part is the aggregate levels on a short history. The annual level of the hierarchy has five observations, so any model fitted to it is subject to the same constraints as every other annual model. The twelve-weekly level has four blocks per edition times five editions, which is twenty observations, and the weekly level has 240. The hierarchy is deep in data at the bottom and thin at the top, which is the reverse of the portfolio case where the aggregate is the well-measured series.

That has a direct consequence for the weighting. If you weight levels by the inverse of their forecast error variance, the annual level will get a small weight because its errors are large and estimated from five numbers. Some teams find that uncomfortable, because the annual number is the one in the budget. The uncomfortable answer is that the budget number was always the least evidenced figure in the pack.

Where this stops

Non-overlapping aggregation needs the levels to divide exactly. A 48 week campaign divides into four-weekly and twelve-weekly blocks cleanly. A 51 week campaign does not, and the choice is either to drop weeks or to use awkward block sizes, both of which introduce edge effects at the start of the campaign where the data is thinnest anyway.

A registration campaign is also not a stationary series in the way the method's usual applications are. It has a hard start and a hard end, so the weekly series is a curve rather than a process, and aggregating a curve to annual gives you one number per edition rather than a longer series. Temporal hierarchies work best on the increments and the pacing behaviour inside the campaign, and less well as a way of extending your annual history, which they do not do.

The last limit is that reconciliation makes the numbers agree without making them right. If all four levels are fitted to editions that preceded a venue move, all four are wrong and the reconciled answer is a carefully weighted average of four wrong answers, delivered with more apparent authority than any of them had alone. Detecting that break is O31's problem and it belongs upstream of everything here.

Choosing to derive one level from another instead of reconciling, either by summing the weekly path or by splitting the annual figure, is the same decision made on the portfolio axis, and bottom up versus top down covers it in O18.

Take one show's current campaign and produce the remaining-weeks total three ways: from the weekly pacing model, from four-week blocks, and from the annual model minus what is already in the file. Write the three numbers next to each other. The spread between them is the amount of disagreement your reporting currently resolves by hand every week, and it is usually larger than anyone expects. What a system needs to hold to run all three at once sits with the forecasting methods behind it.

Questions people ask about temporal hierarchy forecasting

What is a temporal hierarchy?
It is one series aggregated to several frequencies by non-overlapping sums. A registration campaign becomes a weekly series, a four-weekly series, a twelve-weekly series and a single annual total. Each level is forecast independently, then the forecasts are reconciled so the weekly path adds up to the annual figure.
Why forecast the same data at several frequencies?
Because aggregation changes what is visible. Kourentzes, Petropoulos and Trapero found that temporal aggregation strengthens or attenuates the signals of different components, so short intervals show the campaign dynamics and long ones show the trend. Combining the levels reduces how much rests on any single model choice.
Does temporal reconciliation change the weekly forecast as well as the total?
Yes. It is a projection, so every level moves. On a four week example the weekly figures rose from 620, 780, 1,150 and 1,940 to 730, 890, 1,280 and 2,070, the two-week blocks rose to 1,620 and 3,350, and the four-week total settled at 4,970 between the disagreeing base forecasts.

Related reading

All forecasting methods articles