Our solution to the ADIA Lab Structural Break Challenge: Real-Time Edition

Hishaam Abdul Razik, Simreen Siraj Shehzadi, Muhammed Hayyan · October 2026

We took part in the ADIA Lab Structural Break Challenge: Real-Time Edition on CrunchDAO and finished with a public leaderboard TS-AUC of 64.33, up from about 52 for the organizers’ quickstart baseline. In this post we describe the ideas behind our model and, perhaps more usefully, how we arrived at them.

Problem setting. Each series consists of a history of 1,000–5,000 points, known to contain no break, followed by an online part of 10–999 points that arrive one at a time. In about half of the series the generating process changes at some point τ of the online part. After every new point, the model must output a score for “has the break already happened?”. The label is 0 before τ and 1 from τ onwards. For full details, see the challenge page.

HISTORYno break (last 300 points)ONLINE PARTarrives one point at a timethe breaklabel 0 · no break yetlabel 1 · broken (and it stays 1)labelAfter every point: a score for “has the break happened yet?”
Figure 1. One training series (#23). Its break comes 254 steps into the online part.

The metric, TS-AUC, computes the AUC across series at every online step t and averages these, weighting each step by the number of (broken, not broken) pairs at that step:

TS-AUC = Σt wt · AUCt / Σt wt,   wt = npost · nnegt

Only the ranking of series at the same step matters. Early steps carry little weight because few series have broken yet, and very late steps because few series are still running. Most of the weight lies between steps 50 and 400.

Three things make the problem hard. First, the series are very different from each other: some behave like white noise, some zig-zag, some wander, and many have calm and volatile spells. A move that is ordinary for one series is alarming for another, yet the metric compares them directly. Second, many breaks are subtle. They change the volatility or the autocorrelation rather than the level, and some cannot be seen by eye (Figure 2). Third, the model runs live: it may use only the points seen so far, one series at a time, and about 10% of series are rerun and must reproduce the same scores to within 10−8.

The volatility jumps (series #23)The moves are 2.4× bigger after the break.breakThe rhythm changes (series #7696)Same size of moves (1.02×), but the oscillation pattern shifts (lag-4 autocorrelation −0.72 → −0.42).breakA wanderer turns jittery (series #1206)Step-to-step moves become 3× bigger and start to reverse each other.break
last 200 points of the historyonline part
Figure 2. Three breaks from the training set. Only the first is easy to see.

To put it simply, at every online step we compute 443 causal features, most of which measure how unusual the recent points are compared with the series’ own history, and score them with a weighted blend of 45 LightGBM and CatBoost models. An overview of the approach is shown below.

ONCE PER SERIES, FROM THE HISTORYHistory1,000–5,000 points① Learn what “normal” looks like for this seriesrhythm (AR fits)volatility speedhistory windowsscalestarting stateEVERY ONLINE STEPNew pointxₜ② Update five blocks of streaming featuresbase 274A 36B 52D 33K 21Feature row443 numbers③ 45 decision-tree models (5 groups × 9 copies)lgb_AKlgb_ADKlgb_ADcat_ABcat_ABD④ Blendweighted average⑤ Offsetper-series baselinescore between 0 and 1: how sure we are that the break has already happened
Figure 3. Overview. The top row runs once per series; the rest runs at every online step.

Below we first summarize the five ideas that made the biggest difference, then tell the story of how we got there in two phases, and finally describe the final model, the ideas that did not work, and a few lessons.

Key ideas

1. Judge every series only against its own history

We checked whether a series’ history says anything about whether a break will come. It does not: a model given only history statistics scores an AUC of 0.505. What the history does tell us is what “normal” looks like for that series. So almost every feature is a comparison with the series’ own past: window means against history means, innovations converted to ranks against their own historical distribution, a volatility filter whose speed is fitted to each history by quasi-maximum likelihood (QML), and the latest window z-scored against every window of the same length in the history. The QML fit matters more than it might seem. Some series have long calm and volatile spells and need a volatility filter that adapts quickly; others are homogeneous and need a slow one. Fitting the filter’s speed to each history means every new point is judged against the right baseline for that series, so a burst that is routine for a volatile series is not mistaken for a break, while the same burst in a calm series stands out. This puts calm and wild series on one scale, which is exactly what a cross-sectional metric needs. It also explains a surprising result: hiding the raw history descriptors from the model improved the score by about a point in our internal tests, because they let the trees take shortcuts. Only two kinds were worth adding back, residual tail heaviness (+0.38) and volatility clustering (+0.20), and only as context for reading the other features.

2. Exact Bayesian change-point evidence, updated live

Windowed statistics need a guess of how long ago the break happened. Bayesian online change-point detection (Adams and MacKay) avoids the guess by averaging over every possible start. We whiten the series with an AR(8) model fitted on its history and compute, at every step, the evidence that the residuals changed at some unknown point k:

Bt = log Σk ≤ t π(k) · p(ek, …, et | change at k) / p(ek, …, et | no change)

With a Normal–Inverse-Gamma prior on the post-change mean and variance, each term has a closed form, and the sum is exact and causal. Further variants mix over a grid of change ages for mean shifts, scale changes and AR(1) changes separately. Together these features added about +0.2 in the first phase.

3. Detect fast, never forget

Adaptive detectors such as an EWMA volatility forecast react quickly to a break, but they also adapt to the new regime and their evidence fades within about 100 steps. The label never fades, and the metric punishes a forgotten break at every remaining step. We therefore pair adaptive statistics with frozen references calibrated on the history, which keep the evidence for as long as the series stays different (Figure 7). This is specific to the real-time setting, and in our experience it is the easiest mistake to make.

In the final model, our feature blocks cover the whole range between the two extremes:

4. More data, not more features

A three-point learning curve showed that our model was short of data (Figure 8). We created label-preserving copies of each training series by re-cutting it into a new history and online part with the break inside the online part (Figure 9). This added more than half a point without a single new feature.

5. Train the way you are graded

Each model sees a random sample of 64–128 steps per series, and each row is weighted by the number of (broken, not broken) pairs it takes part in at its step, so that training counts errors the way TS-AUC does. At inference, a small causal per-series offset subtracts part of each series’ own early score, so that series which look alarming from the very first step do not crowd out real breaks.

1. Our journey

Our work had two phases. In the first (10–18 September) we built a streaming feature set and a LightGBM + CatBoost + ExtraTrees model that reached 62.69. In the second we rebuilt the training pipeline on top of those features and reached 64.33. Figure 4 shows every official submission.

phase 1 submissionphase 2 submissionbest so far
phase 1: the earlier packagephase 26162636461.0562.6964.33finalsubmissions in order →
Data table
#PhaseOfficialChange
1161.05first scored package
2161.29evidence-only features (history descriptors masked)
3161.30three-forest blend
4161.39LightGBM
5161.59updated LightGBM
6161.80+ Bayesian change evidence (NIG)
7162.08+ component evidence
8162.32+ residual-tail context
9161.94reference-context variant
10162.21linear-tree blend
11162.38+ CatBoost (half blend)
12162.43CatBoost 320 trees
13162.51+ Bayesian cell-width mixtures
14162.50+ history-gated ExtraTrees
15162.42learned-onset CatBoost
16162.69+ volatility-clustering descriptor (end of phase 1)
17162.64linear-tree variant
18162.64without ExtraTrees
19263.16v1: block A + pair-weighted LightGBM
20264.05v5: blocks B, D + re-windowed copies
21264.33v10: block K, bagging, score offset (final)
Figure 4. Every official submission, in order. Hover over a point to see what changed.

Throughout, we evaluated mostly with 5-fold cross-validation split by series and changed one thing at a time. In the first phase, several internal gains did not carry over to the leaderboard (for example 63.05 internal against 61.94 official), which taught us to be strict: a change had to win on most folds, and we preferred a few strong features over many weak ones.

Phase 1: From the baseline to 62.69

The quickstart baseline, a streaming EWMA z-score, scores about 52 locally. Our first step was supervised learning on causal features: every series is standardized by its history and an AR(8) model is fitted on it, giving a raw and a residual channel. For each channel we track windowed means and variances of several transforms (value, square, absolute value, lagged products) over windows of 16, 64, 256 points and the whole online part, each divided by the standard error the history implies for that window size, plus CUSUMs and decaying maxima. Rank features convert the AR innovations to midranks against their own history, and multi-scale scans take the maximum contrast over windows from 8 to 512 points, a crude stand-in for not knowing when the break happened. With pair-balanced sample weights, a gradient-boosted model on these features scored about 61 on the leaderboard.

Two habits from this phase shaped everything after it. First, we screened every new feature on synthetic series with known mean, volatility and autocorrelation breaks, and on break-free series to check for false alarms, before spending real cross-validation on it. The screen caught bad ideas cheaply (lag-16 features failed the false-alarm check), but it was not enough on its own: features based on the ranks of magnitudes gained +0.27 AUC on synthetic data and then lost 0.56 TS-AUC on the real data. Second, we set the minimum leaf size to 512 while sampling at most 64 rows per series, so every leaf of every tree contains at least 8 different series. The trees cannot memorize individual series, which matters when neighbouring rows of one series are almost identical.

The next gains came from three directions:

The phase ended with a two-column volatility-clustering descriptor of the history (+0.20), giving 62.69 on the leaderboard with 301 features.

Phase 2: From 62.69 to 64.33

We kept the 301 features as the base block and rebuilt everything around them. On our new folds the phase-1 LightGBM + CatBoost pair scores 62.82 CV (we dropped the ExtraTrees specialist, whose gain was within noise). From there, the official score stayed about 0.36 below CV for every version we submitted, so gains carried over one to one.

our cross-validation scoreofficial leaderboard score
startstage 1stage 2stage 362.563.063.564.064.565.0startv1v2v3v4v5v6v7v8v9v1062.6963.1664.05CV 64.70official 64.33
Data table
VersionChangeCVOfficial
startPhase-1 features, LightGBM + CatBoost (official 62.69 was with the ExtraTrees specialist)62.8262.69
v1Detector A (volatility surprise); LightGBM with pair weights63.4963.16
v2Detector B; CatBoost group; two-group blend63.75–
v3Detector D (history-calibrated windows); four groups63.90–
v4Three re-cut copies of every training series64.27–
v5Six copies for LightGBM; bigger trees64.4064.05
v6128 sampled steps per series; more regularization64.45–
v7Six model copies per group (stability only)64.45–
v8Detector K (Kalman rhythm drift); five groups64.56–
v9Nine model copies per group; per-series score offset64.64–
v10Twelve training copies for lgb_ADK (final)64.7064.33
Figure 5. Phase 2: cross-validation and official scores of our ten versions. Hover over a version to see what changed.

Stage 1: Better features (v1–v3, CV 62.82 → 63.90)

We first made whitening adaptive. An EWMA of past squared AR errors gives the expected size of the next error:

σ̂t+12 = λ σ̂t2 + (1 − λ) et2

How fast this forecast should adapt differs a lot between series: a series with long calm and volatile spells needs a fast filter, a homogeneous one a slow filter. So we choose the decay λ separately for every series by Gaussian quasi-maximum likelihood (QML) on its own history. For each λ on a grid from 0.90 to 0.999 we run the filter through the history residuals and keep the λ that minimizes

QML(λ) = meant [ log σ̂t2(λ) + et2 / σ̂t2(λ) ],

the Gaussian negative log-likelihood. It is “quasi” because the residuals are not Gaussian, but the criterion still picks a sensible speed. The selected λ and the QML gain over a constant variance are themselves features: they tell the trees how much volatility clustering the series normally has. Block A also keeps four filters with fixed decays (0.97 to 0.998) next to the QML one. The ratio of the squared error to the forecast is a surprise statistic that is about 1 when nothing has changed:

ut = et2 / σ̂t2

A. The raw seriesB. Prediction error eₜ of the AR(8) model fitted on the historyC. EWMA volatility forecast (band = ±2σ̂ₜ)D. Surprise uₜ = eₜ² / σ̂ₜ² (16-point average)1 = as expectedsurprise jumps……then fades as the forecast adaptshistoryonline, before the breakonline, after the break
Figure 6. Whitening series #23. After the break the errors grow (B) and the surprise jumps (D). The volatility forecast then adapts (C) and the surprise fades back towards 1.

Block A averages ut over windows of 16 points up to the whole online part, adds CUSUM statistics, and repeats both against the volatility frozen at the end of the history. With nothing else changed, it raised a LightGBM model from 62.80 to 63.37 CV, the largest single gain of the phase. At the same time we switched to sampling steps per series and weighting rows by their pair count (Key idea 5), worth +0.26 for LightGBM.

Block B compares the last m points (m from 4 to 512) with the forecast made just before those points began, and ranks each comparison against the same comparison on stretches of the history. We added it together with a CatBoost model group (+0.26 in total).

Looking at the errors of these models, we found the forgetting problem (Key idea 3). Block D measures the volatility, autocorrelation and tail shape of the latest window (32 to 512 points) and z-scores each value against the same measurement on every window of the same length in the history. It also asks a very direct question: is the current window more volatile than any window this series has ever had, and by how much? (+0.15)

Block A: surprise of the last 64 points against the adaptive forecast (log scale; 0 = as expected)-10+1+2sees the break, then forgets itBlock D: the last 128 points against every 128-point window of the history (z-score)-4-20+2+4+6keeps the evidence0200400600

x-axis: online step of series #23 (break at step 254). Both lines are real feature values from the package code.

Figure 7. The forgetting problem on series #23. Block A sees the break, and its evidence fades within about 100 steps. Block D keeps it.

Stage 2: More data (v4–v5, CV 63.90 → 64.40)

We trained the same model on 50%, 75% and 100% of the training series, and the score was still climbing steeply at 100% (Figure 8).

61.562.062.563.063.564.064.5original training seriescopies added →← fewer series0.5×50% of series0.75×75% of series1×all series2×+1 copy each4×+3 copies each64.0161.9
Data table
Training dataCross-validation
0.5× (50% of series)61.90
0.75× (75% of series)63.10
1× (all series)63.60
2× (+1 copy each)63.78
4× (+3 copies each)64.01
Figure 8. CV score of one model against the amount of training data. The augmented copies continue the curve.

Outside data was not allowed, so we created new series from the existing ones. For each training series we join the history and the online part and cut out a new window: the new history length, online length and break position are drawn from the same distributions as in the competition data, and the break always falls inside the new online part, so the new history is break-free and the labels stay correct. A copy always goes into the same fold as its source series, and only original series are scored.

Original serieshistory 3000 · online 800Copy 1history 2100 · online 650Copy 2history 1200 · online 500Copy 3history 3200 · online 300the same break01,0002,0003,000position in the joined series (history + online)
history (always break-free)online part (always contains the break)
Figure 9. One series and three copies cut from it. Each copy has a break-free history and the break inside its online part.

Three copies per series gave +0.37, and with six copies for the LightGBM groups and larger trees the CV reached 64.40.

Stage 3: Polish (v6–v10, CV 64.40 → 64.70)

2. The final model

Features

Every feature is computed by streaming code that keeps a small state per series (running sums, buffers, filter values), so each step costs the same regardless of position, and the same code produced the training features and runs in the submission. The 443 columns come from five blocks:

BlockWhat it measuresColumnsPhase
baseWindow contrasts against history means, history-midranked AR innovations, multi-scale scans, CUSUMs and decaying maxima2891
baseAR displacement (lasso and ridge corrections to the history AR(8) fit) and conditional-variance displacement61
baseBayesian change evidence: single-change NIG Bayes factor and mixtures over change age for mean, scale and AR(1) changes41
baseA few fixed history descriptors: residual tail heaviness, raw shape, volatility clustering (the other 27 are masked)6 of 331
ASurprise of the recent points relative to EWMA volatility forecasts (decay chosen per series by QML, plus four fixed decays), over several window lengths, plus CUSUMs362
BThe last m points against the forecast made just before they began, ranked against the history; AR-coefficient change; Bayesian change-point evidence522
DVolatility, autocorrelation and tail shape of the latest window, z-scored against every window of the same length in the history332
KDrift of AR(4) coefficients tracked by a Kalman filter, relative to their drift during the history212

Models

Five model groups, each with nine bags, use different learners or feature subsets so that their errors differ. Their scores are averaged with fixed weights, and the blend beats every single group.

GroupLearnerFeature blocksAugmented copiesBlend weightCV alone
lgb_AKLightGBM, 31 leaves, 400 treesbase + A + K6 (weight 1.0)0.2064.26
lgb_ADKLightGBM, 63 leaves, 400 treesbase + A + D + K12 (weight 0.35)0.3064.55
lgb_ADLightGBM, 63 leaves, 300 treesbase + A + D6 (weight 0.35)0.1564.39
cat_ABCatBoost, depth 6, 320 treesbase + A + B30.1563.79
cat_ABDCatBoost, depth 6, 320 treesbase + A + B + D30.2063.92

Finally, with Lt the blended score on the log-odds scale, the submitted score is

scoret = σ( Lt − 0.15 · mean(L1, …, L20) ),

where the mean is over the series’ first 20 online steps (or all steps so far, if fewer).

Engineering for real time

In this competition a model that scores well offline but behaves differently live is worthless, and reruns must match to 10−8. Early on, a one-ulp float32 difference between our batch and streaming feature code showed us how easily the two drift apart. From then on, all training features were produced by the exact online update function, one point at a time, and every release was tested for causality: changing future points must not change past scores, and resetting or reordering series must not change anything. At inference, each of the 16 workers runs single-threaded, CatBoost trees are evaluated by our own short NumPy code, and in phase 1 the ExtraTrees splits, learned on ranks, were compiled into plain float thresholds, so reruns reproduce the scores exactly.

Results

The final model scores 64.70 in cross-validation and 64.33 on the public leaderboard. Breaks younger than 10 steps are nearly invisible (AUC 53), but by 100–200 steps after the break the AUC reaches about 67, and 70 by 300–500 steps. Figure 10 shows the model on three held-out series.

Series #23 · the volatility jump
historyonline partbreak00.51model score →

The score rises within a few steps of the break and stays high.

Series #7696 · the hidden rhythm change
historyonline partbreak00.51model score →

Invisible to the eye; the rhythm features push the score up about 20 steps after the break.

Series #9883 · a break that is detected, then partly forgotten
historyonline partbreak00.51model score →

The moves grow 2.5× bigger; the score jumps, then drifts down as the new volatility starts to look normal. Some forgetting remains in the final model.

Figure 10. The final model’s out-of-fold score after every point, for three training series. Hover to read the values.

3. What did not work

Over both phases we tried far more ideas than we kept. A pattern stands out: on this problem, careful statistics read by gradient-boosted trees beat every neural and pre-trained model we tried. Gains are in TS-AUC points against the model of the time:

What we triedResult
Neural networks on raw windows or on our features: TCN, GRU, ResNet, MLPall well below the trees (e.g. TCN 54.6 vs 60.8 on one fold; GRU blend −2.3)
Pre-trained models: Chronos-Bolt forecast surprise, TabICLv2−0.03 and −0.69
Other learners: EBM, XGBoost pairwise, QDA, kNN, TabNet, NODE−0.7 to −1.3
A per-step ranking objective (RankNet) instead of log-loss−1.01
Wavelet denoising, FFT and Haar-energy featuresno gain
Rank tests (Kolmogorov–Smirnov, Mann–Whitney, runs) as an extra blockno gain
Bayesian AR(1) mixture, learned change-onset models−0.10; internal gain did not transfer (official 62.42)
Age-stratified sampling, up-weighting rows after the break−0.28 and no gain
Smoothing or stacking the score pathno gain
Extra copies cut from the break-free stretch before a break−0.36
Re-standardizing each copy like the original historiesno gain
Deeper CatBoost, more CatBoost copies, 256 sampled steps per seriesno gain
Blend weights optimized on out-of-fold predictionsdid not hold on held-out folds

4. Final thoughts

Looking back, the biggest gains came from thinking about the structure of the problem rather than from bigger models. The metric compares series with each other, so every piece of evidence has to be measured against the series’ own history. The label never switches back, so a detector that adapts must be paired with one that does not. And the change point is unknown, so it pays to average over it exactly where we can, as the Bayesian features do.

We also learned to treat the cross-validation score with care. In the first phase, internal gains often failed to transfer, so we kept only changes that won on most folds. In the second, splitting folds by series, keeping augmented copies with their source and training with the metric’s pair weights gave a CV score that moved together with the leaderboard, and a simple learning curve showed us that the next gain would come from data rather than features.

We thank ADIA Lab and CrunchDAO for organizing the challenge.