Adaptive conformal inference after a venue move breaks exchangeability
Adaptive conformal inference treats the miscoverage level as a parameter that moves. After an interval misses, the level tightens; after it contains the outcome, the level loosens slightly. Gibbs and Candes (2021) prove the long-run coverage frequency converges to the target under arbitrary distribution shift.
The show is leaving the hall it has used for eleven years. New site, other side of the city, better rail link, 18 per cent more gross square metres, and a car park that holds half as many cars. The 2026 edition is the first one there.
Every residual in your calibration set came from the old building. The band you would compute from them has a coverage guarantee that is, as of the move, a statement about a world that no longer exists. Adaptive conformal inference is the method for exactly this position: it stops treating the coverage level as fixed and starts treating it as something you steer.
Gibbs and Candes (2021) set the approach out in "Adaptive Conformal Inference Under Distribution Shift", and their claim is stronger than it sounds. They provably achieve the desired coverage frequency over long-time intervals irrespective of the true data generating process. No exchangeability, no stationarity, no assumption about how the shift arrives.
What exactly does a venue move break?
Split conformal, covered in the rank arithmetic of a calibration set in O21, rests on one assumption. The calibration points and the point you are predicting have to be exchangeable, meaning you could shuffle them into any order and the joint distribution would be unchanged.
A venue move violates that in several directions at once. The catchment changes, so the day-visitor share moves. The hall geometry changes, so the exhibitor mix and the aisle counts change. Parking capacity changes, so a whole segment of your audience faces a different travel cost. Whatever your model was getting wrong in the old venue, it is now getting wrong differently, and the size of the error is drawn from a different distribution.
The failure mode is specific and unpleasant. A fixed conformal quantile computed from pre-move errors will undercover, and it will undercover most in the first edition after the move, which is exactly the edition where somebody is making a hall-hire commitment against your number. The guarantee does not bend. It stops applying.
The update rule, worked one miss at a time
The mechanism is a single scalar. Instead of using the nominal miscoverage level alpha for every round, you carry a working level alpha_t that changes after each outcome.
Gibbs and Candes (2021) write the recursion as alpha_t+1 equals alpha_t plus gamma times (alpha minus err_t), where err_t is 1 if the realised value fell outside the interval and 0 if it fell inside, and gamma is a positive step size.
Put numbers on it. Target alpha is 0.1, so nominal coverage is 90 per cent. Set gamma at 0.05 and start at alpha_1 equal to 0.1.
Round one misses. err_1 is 1, so alpha_2 equals 0.1 plus 0.05 times (0.1 minus 1), which is 0.1 minus 0.045, giving 0.055. The next interval is built at a working level of 0.055, so it is a 94.5 per cent interval and it is wider.
Round two contains the outcome. err_2 is 0, so alpha_3 equals 0.055 plus 0.05 times 0.1, which is 0.060. The interval narrows, by a small amount.
The asymmetry is the whole design. A miss moves the level by gamma times (1 minus alpha), which is 0.045. A hit moves it by gamma times alpha, which is 0.005. The ratio is nine to one, which is exactly (1 minus alpha) over alpha at a 90 per cent target. The method punishes misses hard and relaxes slowly, and in the long run the two forces balance when the miss frequency equals alpha.
Run the miss case three times in a row and something useful happens. Starting at 0.1, the sequence is 0.055, then 0.010, then minus 0.035. Gibbs and Candes prove alpha_t stays inside the range from minus gamma to 1 plus gamma, and once it goes below zero the interval is infinite. That is the correct behaviour. Three consecutive misses at these settings is the method telling you that it has no idea how wide the band should be, and an infinite interval is a more honest output than a confident wrong one.
How long is the long run on an annual series?
This is where event data makes trouble, and it deserves the arithmetic rather than a warning.
Gibbs and Candes give a finite-time bound: the absolute difference between the realised average miscoverage over T rounds and the target alpha is at most the quantity (max of alpha_1 and 1 minus alpha_1, plus gamma) divided by (T times gamma).
With alpha_1 at 0.1 and gamma at 0.05, the numerator is 0.9 plus 0.05, which is 0.95. Now vary T.
At T equal to 20, which is twenty consecutive annual editions, the bound is 0.95 divided by 1.0, which is 0.95. The guarantee is that your realised miscoverage is within 95 percentage points of 10 per cent. That is no guarantee at all.
At T equal to 168, which is eight shows each producing 21 weekly checkpoints inside a 150-day booking window over one year, the bound is 0.95 divided by 8.4, which is 0.113.
At T equal to 1,200, which is eight shows checked daily across the same window, the bound is 0.95 divided by 60, which is 0.016.
The conclusion is not subtle. Adaptive conformal inference run on the annual final-attendance figure will never accumulate enough rounds to mean anything within the working life of a show director. Run it inside the booking window, on the forecast of the final figure made at each checkpoint, and the round count becomes large enough for the bound to bite. The unit of adaptation is a forecast-and-outcome pair, and you have to manufacture those at a cadence faster than once a year.
There is a cost to that, which is that consecutive weekly checkpoints on the same edition are heavily correlated, so 21 checkpoints are worth rather less than 21 independent observations. The Gibbs and Candes bound holds regardless, because it makes no independence assumption, but the practical rate of learning is slower than the count suggests.
Ensemble intervals when you cannot spare a calibration set
Splitting off a calibration set is expensive when the whole series is five editions long. Xu and Xie (2021), in "Conformal prediction interval for dynamic time-series", published in the Proceedings of the 38th International Conference on Machine Learning, take a different route.
Their method, EnbPI, wraps a bootstrap ensemble. Train B models on bootstrap resamples, and for each training point form an aggregate prediction using only the ensemble members that did not see that point. The resulting leave-one-out residuals act as calibration scores, so every observation contributes to both fitting and calibration. They describe it as requiring neither data-splitting nor training multiple ensemble estimators, and the assumption it replaces exchangeability with is that the error process is stationary and strongly mixing.
For a portfolio where each show has four or five usable editions, that trade is often the right one. You keep every observation in the fit, and the coverage you get is approximate and asymptotic instead of exact and finite-sample. Given that the exact guarantee was going to be void after the venue move anyway, an approximate guarantee that uses all your data is the better deal.
The two ideas compose. Use an ensemble to produce the scores without splitting, and run the adaptive level update on top of the intervals it produces.
Choosing gamma
Gamma controls how fast the level moves and therefore how jumpy the interval widths are.
At gamma equal to 0.01, a miss moves the level by 0.009, and it takes eleven consecutive misses to drive the working level to zero. The intervals are stable and the recovery after a genuine break is slow.
At gamma equal to 0.1, a miss moves the level by 0.09, and two consecutive misses take the level from 0.1 to negative. Recovery is immediate and the width series looks like a sawtooth.
The finite-time bound argues for larger gamma, since it sits in the denominator. Interval stability argues for smaller. My own preference on event data is to set gamma so that roughly three consecutive misses exhausts the level, which at a 90 per cent target puts gamma near 0.04, and then to accept that the first two editions after a structural change will produce wide bands. Detecting that a change happened at all is a separate exercise, and it belongs to structural break detection in O31.
Where this stops
The guarantee is about a long-run frequency, and a long-run frequency is not a statement about any particular edition.
If the 2026 edition in the new venue is the one you need a defensible band for, adaptive conformal inference does not give you one. It gives you a procedure whose average miscoverage across many rounds tends to the target. On the specific round after the specific break, the working level is still carrying information from the old venue, and the interval will be wrong in whatever direction the shift went until enough rounds have passed to unlearn it.
There is also nothing in the method that decides how much pre-break history to throw away. The level adapts; the underlying model does not. If your point forecast is fitted on eleven editions in the old hall, adaptive conformal widens the band around a biased centre, and a wide band around a wrong number is a poor substitute for refitting. Choosing the estimation window after a break is its own decision and belongs to O32.
The first step is cheap and does not need any of the machinery. Take your last two seasons of weekly pacing forecasts, mark each one as inside or outside whatever band you published at the time, and run the alpha recursion by hand in a spreadsheet with gamma at 0.05. You will see immediately whether your published bands have been chronically too narrow, and the drifting level is a better diagnostic of that than any single coverage percentage. The wider set of forecasting methods can wait until you know the answer.
Questions people ask about adaptive conformal inference
- Why does a venue move break a conformal prediction interval?
- Split conformal needs the calibration errors and the new case to be exchangeable, meaning any ordering of them is equally likely. Errors recorded in an old venue with its own catchment, hall layout and travel cost come from a different data-generating process, so the finite-sample coverage guarantee no longer applies to the new edition.
- How does the adaptive conformal update rule work?
- The target miscoverage level for the next round equals the current level plus a step size times the difference between the nominal level and an indicator of whether the last interval missed. A miss pushes the level down and widens the next interval. A hit pushes it up by a much smaller amount and narrows the next one.
- How do you choose the step size gamma?
- Larger values adapt faster and make interval widths jumpier; smaller values are stable but take longer to recover after a shift. The finite-time bound in Gibbs and Candes (2021) scales as one over the number of rounds times gamma, so the useful question is how many update rounds you will accumulate before anyone needs the answer.