Our solution to the ADIA Lab Structural Break Challenge: Real-Time Edition
Hishaam Abdul Razik, Simreen Siraj Shehzadi, Muhammed Hayyan · October 2026
We took part in the ADIA Lab Structural Break Challenge: Real-Time Edition on CrunchDAO and finished with a public leaderboard TS-AUC of 64.33, up from about 52 for the organizers’ quickstart baseline. In this post we describe the ideas behind our model and, perhaps more usefully, how we arrived at them.
Problem setting. Each series consists of a history of 1,000–5,000 points, known to contain no break, followed by an online part of 10–999 points that arrive one at a time. In about half of the series the generating process changes at some point τ of the online part. After every new point, the model must output a score for “has the break already happened?”. The label is 0 before τ and 1 from τ onwards. For full details, see the challenge page.
The metric, TS-AUC, computes the AUC across series at every online step t and averages these, weighting each step by the number of (broken, not broken) pairs at that step:
TS-AUC = Σt wt · AUCt / Σt wt, wt = npost · nnegt
Only the ranking of series at the same step matters. Early steps carry little weight because few series have broken yet, and very late steps because few series are still running. Most of the weight lies between steps 50 and 400.
Three things make the problem hard. First, the series are very different from each other: some behave like white noise, some zig-zag, some wander, and many have calm and volatile spells. A move that is ordinary for one series is alarming for another, yet the metric compares them directly. Second, many breaks are subtle. They change the volatility or the autocorrelation rather than the level, and some cannot be seen by eye (Figure 2). Third, the model runs live: it may use only the points seen so far, one series at a time, and about 10% of series are rerun and must reproduce the same scores to within 10−8.
To put it simply, at every online step we compute 443 causal features, most of which measure how unusual the recent points are compared with the series’ own history, and score them with a weighted blend of 45 LightGBM and CatBoost models. An overview of the approach is shown below.
Below we first summarize the five ideas that made the biggest difference, then tell the story of how we got there in two phases, and finally describe the final model, the ideas that did not work, and a few lessons.
Key ideas
1. Judge every series only against its own history
We checked whether a series’ history says anything about whether a break will come. It does not: a model given only history statistics scores an AUC of 0.505. What the history does tell us is what “normal” looks like for that series. So almost every feature is a comparison with the series’ own past: window means against history means, innovations converted to ranks against their own historical distribution, a volatility filter whose speed is fitted to each history by quasi-maximum likelihood (QML), and the latest window z-scored against every window of the same length in the history. The QML fit matters more than it might seem. Some series have long calm and volatile spells and need a volatility filter that adapts quickly; others are homogeneous and need a slow one. Fitting the filter’s speed to each history means every new point is judged against the right baseline for that series, so a burst that is routine for a volatile series is not mistaken for a break, while the same burst in a calm series stands out. This puts calm and wild series on one scale, which is exactly what a cross-sectional metric needs. It also explains a surprising result: hiding the raw history descriptors from the model improved the score by about a point in our internal tests, because they let the trees take shortcuts. Only two kinds were worth adding back, residual tail heaviness (+0.38) and volatility clustering (+0.20), and only as context for reading the other features.
2. Exact Bayesian change-point evidence, updated live
Windowed statistics need a guess of how long ago the break happened. Bayesian online change-point detection (Adams and MacKay) avoids the guess by averaging over every possible start. We whiten the series with an AR(8) model fitted on its history and compute, at every step, the evidence that the residuals changed at some unknown point k:
Bt = log Σk ≤ t π(k) · p(ek, …, et | change at k) / p(ek, …, et | no change)
With a Normal–Inverse-Gamma prior on the post-change mean and variance, each term has a closed form, and the sum is exact and causal. Further variants mix over a grid of change ages for mean shifts, scale changes and AR(1) changes separately. Together these features added about +0.2 in the first phase.
3. Detect fast, never forget
Adaptive detectors such as an EWMA volatility forecast react quickly to a break, but they also adapt to the new regime and their evidence fades within about 100 steps. The label never fades, and the metric punishes a forgotten break at every remaining step. We therefore pair adaptive statistics with frozen references calibrated on the history, which keep the evidence for as long as the series stays different (Figure 7). This is specific to the real-time setting, and in our experience it is the easiest mistake to make.
In the final model, our feature blocks cover the whole range between the two extremes:
- Fully adaptive (block A): surprise against an EWMA forecast that keeps learning. Fastest to react, first to forget.
- Delayed (block B): the last m points against the forecast made just before they began, so the reference has not yet seen the change it is testing.
- Adaptive against frozen (block K): the log-likelihood ratio of a Kalman filter that keeps re-estimating the AR coefficients against the AR model frozen at the end of the history. The longer the series stays in its new regime, the more the adaptive model wins.
- Fully frozen (block D): the latest window against every window of the same length in the history, including how far it lies beyond the most volatile (or calmest) window the series has ever had.
4. More data, not more features
A three-point learning curve showed that our model was short of data (Figure 8). We created label-preserving copies of each training series by re-cutting it into a new history and online part with the break inside the online part (Figure 9). This added more than half a point without a single new feature.
5. Train the way you are graded
Each model sees a random sample of 64–128 steps per series, and each row is weighted by the number of (broken, not broken) pairs it takes part in at its step, so that training counts errors the way TS-AUC does. At inference, a small causal per-series offset subtracts part of each series’ own early score, so that series which look alarming from the very first step do not crowd out real breaks.
1. Our journey
Our work had two phases. In the first (10–18 September) we built a streaming feature set and a LightGBM + CatBoost + ExtraTrees model that reached 62.69. In the second we rebuilt the training pipeline on top of those features and reached 64.33. Figure 4 shows every official submission.
Data table
| # | Phase | Official | Change |
|---|---|---|---|
| 1 | 1 | 61.05 | first scored package |
| 2 | 1 | 61.29 | evidence-only features (history descriptors masked) |
| 3 | 1 | 61.30 | three-forest blend |
| 4 | 1 | 61.39 | LightGBM |
| 5 | 1 | 61.59 | updated LightGBM |
| 6 | 1 | 61.80 | + Bayesian change evidence (NIG) |
| 7 | 1 | 62.08 | + component evidence |
| 8 | 1 | 62.32 | + residual-tail context |
| 9 | 1 | 61.94 | reference-context variant |
| 10 | 1 | 62.21 | linear-tree blend |
| 11 | 1 | 62.38 | + CatBoost (half blend) |
| 12 | 1 | 62.43 | CatBoost 320 trees |
| 13 | 1 | 62.51 | + Bayesian cell-width mixtures |
| 14 | 1 | 62.50 | + history-gated ExtraTrees |
| 15 | 1 | 62.42 | learned-onset CatBoost |
| 16 | 1 | 62.69 | + volatility-clustering descriptor (end of phase 1) |
| 17 | 1 | 62.64 | linear-tree variant |
| 18 | 1 | 62.64 | without ExtraTrees |
| 19 | 2 | 63.16 | v1: block A + pair-weighted LightGBM |
| 20 | 2 | 64.05 | v5: blocks B, D + re-windowed copies |
| 21 | 2 | 64.33 | v10: block K, bagging, score offset (final) |
Throughout, we evaluated mostly with 5-fold cross-validation split by series and changed one thing at a time. In the first phase, several internal gains did not carry over to the leaderboard (for example 63.05 internal against 61.94 official), which taught us to be strict: a change had to win on most folds, and we preferred a few strong features over many weak ones.
Phase 1: From the baseline to 62.69
The quickstart baseline, a streaming EWMA z-score, scores about 52 locally. Our first step was supervised learning on causal features: every series is standardized by its history and an AR(8) model is fitted on it, giving a raw and a residual channel. For each channel we track windowed means and variances of several transforms (value, square, absolute value, lagged products) over windows of 16, 64, 256 points and the whole online part, each divided by the standard error the history implies for that window size, plus CUSUMs and decaying maxima. Rank features convert the AR innovations to midranks against their own history, and multi-scale scans take the maximum contrast over windows from 8 to 512 points, a crude stand-in for not knowing when the break happened. With pair-balanced sample weights, a gradient-boosted model on these features scored about 61 on the leaderboard.
Two habits from this phase shaped everything after it. First, we screened every new feature on synthetic series with known mean, volatility and autocorrelation breaks, and on break-free series to check for false alarms, before spending real cross-validation on it. The screen caught bad ideas cheaply (lag-16 features failed the false-alarm check), but it was not enough on its own: features based on the ranks of magnitudes gained +0.27 AUC on synthetic data and then lost 0.56 TS-AUC on the real data. Second, we set the minimum leaf size to 512 while sampling at most 64 rows per series, so every leaf of every tree contains at least 8 different series. The trees cannot memorize individual series, which matters when neighbouring rows of one series are almost identical.
The next gains came from three directions:
- Less history, more evidence. Masking the 32 history descriptors helped (Key idea 1). Restoring the residual tail heaviness was one of the largest steps of the phase (+0.38).
- AR and Bayesian evidence. Lasso and ridge corrections to the history AR(8) coefficients over recent windows measure how much the dynamics have changed; the NIG Bayes factor and the cell-width mixtures add exact change-point evidence (Key idea 2).
- Model diversity. Averaging LightGBM with CatBoost added about +0.2. For the half of series with light-tailed histories, decided before any online point arrives, an ExtraTrees model on rank-transformed features gets half of the weight.
The phase ended with a two-column volatility-clustering descriptor of the history (+0.20), giving 62.69 on the leaderboard with 301 features.
Phase 2: From 62.69 to 64.33
We kept the 301 features as the base block and rebuilt everything around them. On our new folds the phase-1 LightGBM + CatBoost pair scores 62.82 CV (we dropped the ExtraTrees specialist, whose gain was within noise). From there, the official score stayed about 0.36 below CV for every version we submitted, so gains carried over one to one.
Data table
| Version | Change | CV | Official |
|---|---|---|---|
| start | Phase-1 features, LightGBM + CatBoost (official 62.69 was with the ExtraTrees specialist) | 62.82 | 62.69 |
| v1 | Detector A (volatility surprise); LightGBM with pair weights | 63.49 | 63.16 |
| v2 | Detector B; CatBoost group; two-group blend | 63.75 | – |
| v3 | Detector D (history-calibrated windows); four groups | 63.90 | – |
| v4 | Three re-cut copies of every training series | 64.27 | – |
| v5 | Six copies for LightGBM; bigger trees | 64.40 | 64.05 |
| v6 | 128 sampled steps per series; more regularization | 64.45 | – |
| v7 | Six model copies per group (stability only) | 64.45 | – |
| v8 | Detector K (Kalman rhythm drift); five groups | 64.56 | – |
| v9 | Nine model copies per group; per-series score offset | 64.64 | – |
| v10 | Twelve training copies for lgb_ADK (final) | 64.70 | 64.33 |
Stage 1: Better features (v1–v3, CV 62.82 → 63.90)
We first made whitening adaptive. An EWMA of past squared AR errors gives the expected size of the next error:
σ̂t+12 = λ σ̂t2 + (1 − λ) et2
How fast this forecast should adapt differs a lot between series: a series with long calm and volatile spells needs a fast filter, a homogeneous one a slow filter. So we choose the decay λ separately for every series by Gaussian quasi-maximum likelihood (QML) on its own history. For each λ on a grid from 0.90 to 0.999 we run the filter through the history residuals and keep the λ that minimizes
QML(λ) = meant [ log σ̂t2(λ) + et2 / σ̂t2(λ) ],
the Gaussian negative log-likelihood. It is “quasi” because the residuals are not Gaussian, but the criterion still picks a sensible speed. The selected λ and the QML gain over a constant variance are themselves features: they tell the trees how much volatility clustering the series normally has. Block A also keeps four filters with fixed decays (0.97 to 0.998) next to the QML one. The ratio of the squared error to the forecast is a surprise statistic that is about 1 when nothing has changed:
ut = et2 / σ̂t2
Block A averages ut over windows of 16 points up to the whole online part, adds CUSUM statistics, and repeats both against the volatility frozen at the end of the history. With nothing else changed, it raised a LightGBM model from 62.80 to 63.37 CV, the largest single gain of the phase. At the same time we switched to sampling steps per series and weighting rows by their pair count (Key idea 5), worth +0.26 for LightGBM.
Block B compares the last m points (m from 4 to 512) with the forecast made just before those points began, and ranks each comparison against the same comparison on stretches of the history. We added it together with a CatBoost model group (+0.26 in total).
Looking at the errors of these models, we found the forgetting problem (Key idea 3). Block D measures the volatility, autocorrelation and tail shape of the latest window (32 to 512 points) and z-scores each value against the same measurement on every window of the same length in the history. It also asks a very direct question: is the current window more volatile than any window this series has ever had, and by how much? (+0.15)
x-axis: online step of series #23 (break at step 254). Both lines are real feature values from the package code.
Stage 2: More data (v4–v5, CV 63.90 → 64.40)
We trained the same model on 50%, 75% and 100% of the training series, and the score was still climbing steeply at 100% (Figure 8).
Data table
| Training data | Cross-validation |
|---|---|
| 0.5× (50% of series) | 61.90 |
| 0.75× (75% of series) | 63.10 |
| 1× (all series) | 63.60 |
| 2× (+1 copy each) | 63.78 |
| 4× (+3 copies each) | 64.01 |
Outside data was not allowed, so we created new series from the existing ones. For each training series we join the history and the online part and cut out a new window: the new history length, online length and break position are drawn from the same distributions as in the competition data, and the break always falls inside the new online part, so the new history is break-free and the labels stay correct. A copy always goes into the same fold as its source series, and only original series are scored.
Three copies per series gave +0.37, and with six copies for the LightGBM groups and larger trees the CV reached 64.40.
Stage 3: Polish (v6–v10, CV 64.40 → 64.70)
- Regularization. 128 sampled steps per series, larger minimum leaf sizes and feature subsampling (+0.05).
- Block K. Some breaks change only the rhythm of a series (the second example in Figure 2). Block K tracks the AR(4) coefficients with a Kalman filter and measures how far they drift compared with their drift during the history (+0.11).
- Bagging. Nine copies of every model group, trained on different sampled rows, for stability.
- Score offset (Key idea 5). We chose its two parameters on three folds and checked them on the other two (+0.08 overall).
- Twelve copies for the strongest LightGBM group (+0.06).
2. The final model
Features
Every feature is computed by streaming code that keeps a small state per series (running sums, buffers, filter values), so each step costs the same regardless of position, and the same code produced the training features and runs in the submission. The 443 columns come from five blocks:
| Block | What it measures | Columns | Phase |
|---|---|---|---|
| base | Window contrasts against history means, history-midranked AR innovations, multi-scale scans, CUSUMs and decaying maxima | 289 | 1 |
| base | AR displacement (lasso and ridge corrections to the history AR(8) fit) and conditional-variance displacement | 6 | 1 |
| base | Bayesian change evidence: single-change NIG Bayes factor and mixtures over change age for mean, scale and AR(1) changes | 4 | 1 |
| base | A few fixed history descriptors: residual tail heaviness, raw shape, volatility clustering (the other 27 are masked) | 6 of 33 | 1 |
| A | Surprise of the recent points relative to EWMA volatility forecasts (decay chosen per series by QML, plus four fixed decays), over several window lengths, plus CUSUMs | 36 | 2 |
| B | The last m points against the forecast made just before they began, ranked against the history; AR-coefficient change; Bayesian change-point evidence | 52 | 2 |
| D | Volatility, autocorrelation and tail shape of the latest window, z-scored against every window of the same length in the history | 33 | 2 |
| K | Drift of AR(4) coefficients tracked by a Kalman filter, relative to their drift during the history | 21 | 2 |
Models
Five model groups, each with nine bags, use different learners or feature subsets so that their errors differ. Their scores are averaged with fixed weights, and the blend beats every single group.
| Group | Learner | Feature blocks | Augmented copies | Blend weight | CV alone |
|---|---|---|---|---|---|
| lgb_AK | LightGBM, 31 leaves, 400 trees | base + A + K | 6 (weight 1.0) | 0.20 | 64.26 |
| lgb_ADK | LightGBM, 63 leaves, 400 trees | base + A + D + K | 12 (weight 0.35) | 0.30 | 64.55 |
| lgb_AD | LightGBM, 63 leaves, 300 trees | base + A + D | 6 (weight 0.35) | 0.15 | 64.39 |
| cat_AB | CatBoost, depth 6, 320 trees | base + A + B | 3 | 0.15 | 63.79 |
| cat_ABD | CatBoost, depth 6, 320 trees | base + A + B + D | 3 | 0.20 | 63.92 |
Finally, with Lt the blended score on the log-odds scale, the submitted score is
scoret = σ( Lt − 0.15 · mean(L1, …, L20) ),
where the mean is over the series’ first 20 online steps (or all steps so far, if fewer).
Engineering for real time
In this competition a model that scores well offline but behaves differently live is worthless, and reruns must match to 10−8. Early on, a one-ulp float32 difference between our batch and streaming feature code showed us how easily the two drift apart. From then on, all training features were produced by the exact online update function, one point at a time, and every release was tested for causality: changing future points must not change past scores, and resetting or reordering series must not change anything. At inference, each of the 16 workers runs single-threaded, CatBoost trees are evaluated by our own short NumPy code, and in phase 1 the ExtraTrees splits, learned on ranks, were compiled into plain float thresholds, so reruns reproduce the scores exactly.
Results
The final model scores 64.70 in cross-validation and 64.33 on the public leaderboard. Breaks younger than 10 steps are nearly invisible (AUC 53), but by 100–200 steps after the break the AUC reaches about 67, and 70 by 300–500 steps. Figure 10 shows the model on three held-out series.
The score rises within a few steps of the break and stays high.
Invisible to the eye; the rhythm features push the score up about 20 steps after the break.
The moves grow 2.5× bigger; the score jumps, then drifts down as the new volatility starts to look normal. Some forgetting remains in the final model.
3. What did not work
Over both phases we tried far more ideas than we kept. A pattern stands out: on this problem, careful statistics read by gradient-boosted trees beat every neural and pre-trained model we tried. Gains are in TS-AUC points against the model of the time:
| What we tried | Result |
|---|---|
| Neural networks on raw windows or on our features: TCN, GRU, ResNet, MLP | all well below the trees (e.g. TCN 54.6 vs 60.8 on one fold; GRU blend −2.3) |
| Pre-trained models: Chronos-Bolt forecast surprise, TabICLv2 | −0.03 and −0.69 |
| Other learners: EBM, XGBoost pairwise, QDA, kNN, TabNet, NODE | −0.7 to −1.3 |
| A per-step ranking objective (RankNet) instead of log-loss | −1.01 |
| Wavelet denoising, FFT and Haar-energy features | no gain |
| Rank tests (Kolmogorov–Smirnov, Mann–Whitney, runs) as an extra block | no gain |
| Bayesian AR(1) mixture, learned change-onset models | −0.10; internal gain did not transfer (official 62.42) |
| Age-stratified sampling, up-weighting rows after the break | −0.28 and no gain |
| Smoothing or stacking the score path | no gain |
| Extra copies cut from the break-free stretch before a break | −0.36 |
| Re-standardizing each copy like the original histories | no gain |
| Deeper CatBoost, more CatBoost copies, 256 sampled steps per series | no gain |
| Blend weights optimized on out-of-fold predictions | did not hold on held-out folds |
4. Final thoughts
Looking back, the biggest gains came from thinking about the structure of the problem rather than from bigger models. The metric compares series with each other, so every piece of evidence has to be measured against the series’ own history. The label never switches back, so a detector that adapts must be paired with one that does not. And the change point is unknown, so it pays to average over it exactly where we can, as the Bayesian features do.
We also learned to treat the cross-validation score with care. In the first phase, internal gains often failed to transfer, so we kept only changes that won on most folds. In the second, splitting folds by series, keeping augmented copies with their source and training with the metric’s pair weights gave a CV score that moved together with the leaderboard, and a simple learning curve showed us that the next gain would come from data rather than features.
We thank ADIA Lab and CrunchDAO for organizing the challenge.