Every model, with its record.
Intended use, training data, held-out results, calibration, leakage checks, drift, challengers and known limits, for each model that prices a game or values a shot. Studies that have not run are marked absent, never blank.
- Models
- 5
- Register entries
- 38
- Studies present
- 7 of 7
- Adopted / rejected
- 10 / 17
Expected goals
Value of an unblocked shot from its location, type and context; the base of Drive, Setup, Finish, GSAx, RAPM and the Elo family. Shooter-agnostic by design so Finish means something.
What it learned from
- seasons
- 2023-24, 2024-25
- shots
- 227076
- features
- 34
How it did on seasons it had not seen
- testedOn
- 2025-26
- logLoss
- 0.229
- baselineLogLoss
- 0.259
- auc
- 0.756
- clockRepair
- Repaired 2026-09-29 (hub card 149362): from 2023-24 the league stamps a goal about a second later than other shots, so the rebound (3 s) and rush (4 s) windows
- logLossBeforeRepair
- 0.228
- aucBeforeRepair
- 0.757
- logLossAfterRepair
- 0.228
- aucAfterRepair
- 0.756
- deciles
- (expected: 69.1, actual: 75); (expected: 143.4, actual: 126); (expected: 243.5, actual: 236); (expected: 373.4, actual: 332); (expected: 535.7, actual: 485); (expected: 718.8, actual: 724); (expected: 936.3, actual: 942); (expected: 1204.5, actual: 1206); (expected: 1576.7, actual: 1604); (expected: 2734.9, actual: 2356)
- seasonShift
- -0.064
| Decile | Expected goals | Actual goals | Actual / expected |
|---|---|---|---|
| 1 | 69.1 | 75 | 1.085 |
| 2 | 143.4 | 126 | 0.879 |
| 3 | 243.5 | 236 | 0.969 |
| 4 | 373.4 | 332 | 0.889 |
| 5 | 535.7 | 485 | 0.905 |
| 6 | 718.8 | 724 | 1.007 |
| 7 | 936.3 | 942 | 1.006 |
| 8 | 1204.5 | 1206 | 1.001 |
| 9 | 1576.7 | 1604 | 1.017 |
| 10 | 2734.9 | 2356 | 0.861 |
Do 60% calls win 60% of the time?
- available
- true
- shots
- 108568
- games
- 1271
- totalXg
- 8273.9
- totalGoals
- 7823
- ratioGoalsOverXg
- 0.946
- ece
- 0.005
- logLossPerShot
- 0.228
- brierPerShot
- 0.062
- deciles
- (decile: 1, xgLo: 0, xgHi: 0.009, shots: 10736, expectedGoals: 64.2, actualGoals: 71, meanXg: 0.006, goalRate: 0.007, ratioActualOverExpected: 1.106, z: 0.85); (decile: 2, xgLo: 0.009, xgHi: 0.017, shots: 10883, expectedGoals: 137, actualGoals: 122, meanXg: 0.013, goalRate: 0.011, ratioActualOverExpected: 0.891, z: -1.29); (decile: 3, xgLo: 0.017, xgHi: 0.027, shots: 10905, expectedGoals: 236.5, actualGoals: 227, meanXg: 0.022, goalRate: 0.021, ratioActualOverExpected: 0.96, z: -0.62); (decile: 4, xgLo: 0.027, xgHi: 0.04, shots: 10884, expectedGoals: 362.8, actualGoals: 315, meanXg: 0.033, goalRate: 0.029, ratioActualOverExpected: 0.868, z: -2.55); (decile: 5, xgLo: 0.04, xgHi: 0.055, shots: 10836, expectedGoals: 517.5, actualGoals: 484, meanXg: 0.048, goalRate: 0.045, ratioActualOverExpected: 0.935, z: -1.51); (decile: 6, xgLo: 0.056, xgHi: 0.073, shots: 10861, expectedGoals: 695.2, actualGoals: 709, meanXg: 0.064, goalRate: 0.065, ratioActualOverExpected: 1.02, z: 0.54); (decile: 7, xgLo: 0.073, xgHi: 0.094, shots: 10888, expectedGoals: 907.7, actualGoals: 876, meanXg: 0.083, goalRate: 0.081, ratioActualOverExpected: 0.965, z: -1.1); (decile: 8, xgLo: 0.095, xgHi: 0.121, shots: 10856, expectedGoals: 1163.6, actualGoals: 1158, meanXg: 0.107, goalRate: 0.107, ratioActualOverExpected: 0.995, z: -0.17); (decile: 9, xgLo: 0.121, xgHi: 0.164, shots: 10845, expectedGoals: 1523.7, actualGoals: 1547, meanXg: 0.141, goalRate: 0.143, ratioActualOverExpected: 1.015, z: 0.64); (decile: 10, xgLo: 0.164, xgHi: 0.908, shots: 10874, expectedGoals: 2665.6, actualGoals: 2314, meanXg: 0.245, goalRate: 0.213, ratioActualOverExpected: 0.868, z: -8.19)
- crossFitIsotonic
- split: games permuted with seed 11, halves of 635/636 games, logLossBefore: 0.228, logLossAfter: 0.228, brierBefore: 0.062, brierAfter: 0.062, eceAfter: 0.001, diffAfterMinusBefore: (mean: 0, range: 0, 0, excludesZero: no, B: 5000, clusters: games), mapHalfAAtGrid: (0.01: 0.007, 0.02: 0.012, 0.05: 0.046, 0.1: 0.101, 0.15: 0.154, 0.2: 0.172, 0.3: 0.213, 0.4: 0.369, 0.5: 0.369), mapHalfBAtGrid: (0.01: 0.01, 0.02: 0.021, 0.05: 0.042, 0.1: 0.109, 0.15: 0.154, 0.2: 0.173, 0.3: 0.218, 0.4: 0.348, 0.5: 0.514), plattSecondary: (note: cross-fitted logit(q) = a + b*logit(xg), same game halves; secondary, not the registered isotonic test, halfA: (a: -0.21, b: 0.933), halfB: (a: -0.189, b: 0.942), logLossAfter: 0.228, diffAfterMinusBefore: (mean: 0, range: 2 items, excludesZero: yes, B: 5000, clusters: games))
Pre-registered challengers, judged on an unseen season
Seven pre-registered challengers to the expected-goals model have been judged on the unseen 2025-26 season; three passed the registered rule. The desk withholds all three and puts none forward, so the champion stays.
Trained on 2023-24 and 2024-25, judged on 2025-26. Champion log loss 0.22830 on 1,312 games (0.22861 on the 1,232 with an events file). Rule: adopt only if the bootstrap 95% range of (challenger - champion) log loss lies below zero; otherwise reject; adoption itself is a hub decision card, never automatic.
| Challenger | Result | Log loss against champion | 95% range | Details |
|---|---|---|---|---|
| A rink-adjusted distance (+ angle by the same geometry) | fails the rule | +0.000104 | −0.000033 to +0.000236 | |
| B strength-state models (four fits) | passes, withheld | −0.000491 | −0.000713 to −0.000248 | |
| C flurry features added | passes, withheld | −0.001587 | −0.001970 to −0.001228 | |
| D rink-adjusted + strength-state fits + flurry | passes, withheld | −0.001965 | −0.002436 to −0.001503 | |
| F champion features, rebound and rush on the corrected clock | fails the rule | +0.000225 | +0.000107 to +0.000351 | |
| G strength-state fits with the corrected flags | fails the rule | −0.000261 | −0.000543 to +0.000026 | |
| H G plus flurry features on the corrected clock | fails the rule | −0.000264 | −0.000548 to +0.000018 | |
| B strength-state fits, flags as in the shot file (re-scored on these rows) | passes, withheld | −0.000539 | −0.000788 to −0.000298 |
The desk: no candidate passes the registered rule, is attributable and shows no placebo gain, so the desk puts none forward. B passes the rule but its flags are built on the raw clock. The clock finding goes to whoever owns tools/pull-pbp.mjs.
Goals are stamped later than other shots
- Goals are stamped about one second later than other shots in 2023-24, 2024-25 and 2025-26; the model's rebound and rush inputs are short windows on that clock.
- With goal times corrected the champion scores 0.22883 against 0.22861 on the same 105,288 shots (+0.000225 per shot, 95% range +0.000107 to +0.000351): part of its published accuracy comes from the clock.
- 70% of shots flagged as rush in 2025-26 follow the club's own blocked attempt and are not rushes.
| Placebo input | Clock | Change in log loss | 95% range | Lowers log loss |
|---|---|---|---|---|
| four events | raw | −0.000661 | −0.000881 to −0.000466 | yes |
| four events | goal times moved back 1 s | +0.000008 | −0.000002 to +0.000016 | no |
| hits only | raw | −0.000028 | −0.000143 to +0.000097 | no |
| hits only | goal times moved back 1 s | +0.000010 | −0.000007 to +0.000029 | no |
Where shots are recorded differently, 2025-26
- 8 of 32 buildings recorded shot distance differently from the same clubs' road shots in 2025-26 (25% at p below 0.01, against 3.5% by chance).
- Season to season the effects correlate at r 0.519 (registered threshold 0.5); the spread between buildings is 1.00 ft against a noise floor of about 0.56 ft.
- A distance correction repaired 22 arena-seasons on unseen games and broke 20 (exact p 0.88), so the model carries no rink adjustment.
| Building | Recorded distance against road shots (ft) | About (95%) | Details |
|---|---|---|---|
| STL Enterprise Center | +1.89 | +0.87 to +2.92 | |
| UTA Delta Center | +1.67 | +0.66 to +2.67 | |
| ANA Honda Center | −1.63 | −2.54 to −0.71 | |
| CAR Lenovo Center | −1.61 | −2.69 to −0.52 | |
| BUF KeyBank Center | +1.47 | +0.49 to +2.46 | |
| EDM Rogers Place | −1.33 | −2.27 to −0.39 | |
| TBL Benchmark International Arena | +1.31 | +0.25 to +2.38 | |
| VAN Rogers Arena | −1.26 | −2.22 to −0.31 |
Where it stops
- Top decile of chances over-predicted by about 15% and the bottom decile under-predicted (decile table); an isotonic recalibration is on the register.
- No rink (scorer) adjustment of shot locations: the correction tried did not help on unseen games (see Arena scorer effects); goalies who play half their games in one building inherit its recording habits.
- Repaired 2026-09-29 (hub card 149362): from 2023-24 the league stamps a goal about a second later than other shots, so the rebound (3 s) and rush (4 s) windows missed goals and the model learned to mark those chances down. Goal times are now moved back 1 s before the windows are applied (2022-23 and earlier unchanged): 592 rebound and 258 rush flags gained, all of them goals. Published accuracy fell, log loss 0.2283 to 0.2285 and AUC 0.757 to 0.756 on 112,091 2025-26 shots, because a leak of the outcome into the inputs was removed, not because the model got worse at hockey.
- Clock repair not complete in the test season: 80 of 1,312 2025-26 games (6,803 shots, 485 goals) have no events file on disk, so their flags stay on the raw clock; test season only, training rows fully corrected.
- One model with power-play and shorthanded flags rather than four strength-state models.
- Goalie quality leaks into the label: a shot on a weak goalie is a goal more often, so xG absorbs some goaltending.
Game win model (champion)
Pre-game probability that the home club wins; the price on every game page, the daily email, the public ledger and the house picker.
What it learned from
- trainedOn
- 2018-19 to 2025-26
- games
- 9187
- features
- b2b, eG, eX, form, h2h, lEvA, lEvF, lFin, lPK, lPP, lSv, luMiss … +4 more
How it did on seasons it had not seen
- tournament
- folds: 2019, 2020; 2020, 2021; 2021, 2022; 2022, 2023; 2023, 2024; 2024, 2025, leader: (name: logistic: everything minus trap/schedule spot, avgLogLoss: 0.657, avgAccuracy: 0.604); (name: logistic: everything, travel = time zones + 3-in-4 only, avgLogLoss: 0.657, avgAccuracy: 0.604); (name: logistic: everything minus travel (ridge 50), avgLogLoss: 0.657, avgAccuracy: 0.604); (name: logistic: everything minus goalie rotation, avgLogLoss: 0.657, avgAccuracy: 0.606); (name: logistic: everything (ridge 50), avgLogLoss: 0.657, avgAccuracy: 0.604); (name: logistic: + lineups/goalie, avgLogLoss: 0.658, avgAccuracy: 0.605)
- tiers
- (tier: coin flip, games: 1760, share: 0.251, predicted: 0.525, right: 0.517, rightTight: 0.487, rightLoose: 0.547); (tier: lean, games: 1588, share: 0.227, predicted: 0.574, right: 0.567, rightTight: 0.547, rightLoose: 0.587); (tier: solid, games: 1331, share: 0.19, predicted: 0.624, right: 0.606, rightTight: 0.586, rightLoose: 0.626); (tier: strong, games: 991, share: 0.141, predicted: 0.674, right: 0.679, rightTight: 0.653, rightLoose: 0.705); (tier: heavy, games: 1336, share: 0.191, predicted: 0.756, right: 0.748, rightTight: 0.749, rightLoose: 0.747)
- blockSearch
- versions: 49152, championProof: 0.664, bestFinalist: (families: special teams, finishing+goalie, lineup, rivalry+race, l2: 50, selection: 0.648, proof: 0.664, vsChampion: -0.001, 0, better: no)
- search
- versions: 3000, luckGap: 0.027, newChampion: —
- gameModelV2
- validation: (games: 1235, logLoss: 0.659, brier: 0.233, accuracy: 0.602), settings: (kG: 8, kX: 22, carry: 0.85, a: 0.05), testedOn: 2025-26
| Model | Avg log loss | Avg accuracy |
|---|---|---|
| logistic: everything minus trap/schedule spot | 0.6567 | 0.604 |
| logistic: everything, travel = time zones + 3-in-4 only | 0.657 | 0.604 |
| logistic: everything minus travel (ridge 50) | 0.6571 | 0.604 |
| logistic: everything minus goalie rotation | 0.6572 | 0.606 |
| logistic: everything (ridge 50) | 0.6574 | 0.604 |
| logistic: + lineups/goalie | 0.6578 | 0.605 |
| Confidence tier | Games | Predicted right | Actually right |
|---|---|---|---|
| coin flip | 1760 | 53% | 52% |
| lean | 1588 | 57% | 57% |
| solid | 1331 | 62% | 61% |
| strong | 991 | 67% | 68% |
| heavy | 1336 | 76% | 75% |
Do 60% calls win 60% of the time?
- perSeason
- 2020: (n: 803, logLoss: 0.648, brier: 0.228, baseRate: 0.538, meanPredicted: 0.535, reliabilityWidth: (bins: 10 items, ece: 0.035, mode: width), reliabilityFreq: (bins: 10 items, ece: 0.039, mode: freq), eceWidth: 0.035, eceFreq: 0.039, decomposition: (bins: 10, brier: 5 fields, logLoss: 5 fields)), 2021: (n: 1210, logLoss: 0.626, brier: 0.218, baseRate: 0.539, meanPredicted: 0.539, reliabilityWidth: (bins: 10 items, ece: 0.024, mode: width), reliabilityFreq: (bins: 10 items, ece: 0.038, mode: freq), eceWidth: 0.024, eceFreq: 0.038, decomposition: (bins: 10, brier: 5 fields, logLoss: 5 fields)), 2022: (n: 1217, logLoss: 0.646, brier: 0.227, baseRate: 0.532, meanPredicted: 0.541, reliabilityWidth: (bins: 10 items, ece: 0.031, mode: width), reliabilityFreq: (bins: 10 items, ece: 0.028, mode: freq), eceWidth: 0.031, eceFreq: 0.028, decomposition: (bins: 10, brier: 5 fields, logLoss: 5 fields)), 2023: (n: 1230, logLoss: 0.654, brier: 0.231, baseRate: 0.538, meanPredicted: 0.535, reliabilityWidth: (bins: 10 items, ece: 0.03, mode: width), reliabilityFreq: (bins: 10 items, ece: 0.025, mode: freq), eceWidth: 0.03, eceFreq: 0.025, decomposition: (bins: 10, brier: 5 fields, logLoss: 5 fields)), 2024: (n: 1235, logLoss: 0.653, brier: 0.231, baseRate: 0.564, meanPredicted: 0.54, reliabilityWidth: (bins: 10 items, ece: 0.039, mode: width), reliabilityFreq: (bins: 10 items, ece: 0.035, mode: freq), eceWidth: 0.039, eceFreq: 0.035, decomposition: (bins: 10, brier: 5 fields, logLoss: 5 fields)), 2025: (n: 1311, logLoss: 0.681, brier: 0.244, baseRate: 0.522, meanPredicted: 0.541, reliabilityWidth: (bins: 10 items, ece: 0.049, mode: width), reliabilityFreq: (bins: 10 items, ece: 0.051, mode: freq), eceWidth: 0.049, eceFreq: 0.051, decomposition: (bins: 10, brier: 5 fields, logLoss: 5 fields))
- pooled
- n: 7006, logLoss: 0.652, brier: 0.23, baseRate: 0.539, meanPredicted: 0.539, reliabilityWidth: (bins: (6 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields), ece: 0.009, mode: width), reliabilityFreq: (bins: (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields), ece: 0.014, mode: freq), eceWidth: 0.009, eceFreq: 0.014, decomposition: (bins: 10, brier: (score: 0.23, reliability: 0, resolution: 0.018, uncertainty: 0.248, residualWithinBin: -0.001), logLoss: (score: 0.652, reliability: 0, resolution: 0.037, uncertainty: 0.69, residualWithinBin: -0.002))
Could the model have seen the future?
- generatedAt
- 2026-09-29T02:58:55Z
- source
- data/models/features.csv
- games
- 9781
- gamesWithOutcome
- 9698
- rowsWithoutHomeWin
- 83
- featureColumns
- 98
- method
- adversarial: LightGBM binary (300 rounds, lr 0.05, 31 leaves, min 50 per leaf, feature/bagging fraction 0.8), 5-fold out-of-fold predictions, rank AUC; gain share averaged over folds, univariate: rank AUC of each column against homeWin, reported as max(AUC, 1-AUC); flag above 0.75, canary: leaked columns appended and re-tested by the same rules; multivariate detector = LightGBM 5-fold homeWin AUC and gain rank, seed: 11
- adversarial
- target: 2025 vs 2018-2024, result: (season: 2025, nCurrent: 1312, nPrior: 8469, auc: 1, topGain: (2 fields); (2 fields); (2 fields); (2 fields); (2 fields); (2 fields); (2 fields); (2 fields); (2 fields); (2 fields); (2 fields); (2 fields) … +3 more, ablations: (3 fields); (3 fields); (3 fields)), contextBySeason: (season: 2020, nCurrent: 868, nPrior: 2353, auc: 0.993, topGain: 15 items); (season: 2021, nCurrent: 1312, nPrior: 3221, auc: 0.993, topGain: 15 items); (season: 2022, nCurrent: 1312, nPrior: 4533, auc: 0.998, topGain: 15 items); (season: 2023, nCurrent: 1312, nPrior: 5845, auc: 0.997, topGain: 15 items); (season: 2024, nCurrent: 1312, nPrior: 7157, auc: 0.999, topGain: 15 items), reading: AUC near 0.5 means the 2025 rows are indistinguishable from earlier seasons; a high AUC driven by a few features (and collapsing after their removal) indicates those inputs shifted level between seasons (league scoring rate, lineup availability), not a leak of outcomes.
- univariate
- flagThreshold: 0.75, flagged: none, raw: (feature: h_lEvF, auc: 0.546, directionalAuc: 0.546, flag: no); (feature: h_lEvA, auc: 0.545, directionalAuc: 0.545, flag: no); (feature: h_lPP, auc: 0.547, directionalAuc: 0.547, flag: no); (feature: h_lPK, auc: 0.522, directionalAuc: 0.522, flag: no); (feature: h_lFin, auc: 0.52, directionalAuc: 0.52, flag: no); (feature: h_lSv, auc: 0.538, directionalAuc: 0.538, flag: no); (feature: h_home, auc: 0.5, directionalAuc: 0.5, flag: no); (feature: h_b2b, auc: 0.49, directionalAuc: 0.51, flag: no); (feature: h_oppB2b, auc: 0.518, directionalAuc: 0.518, flag: no); (feature: h_travelK, auc: 0.501, directionalAuc: 0.501, flag: no); (feature: h_tzCross, auc: 0.502, directionalAuc: 0.502, flag: no); (feature: h_threeIn4, auc: 0.493, directionalAuc: 0.507, flag: no) … +86 more, differences: (feature: diff_altitude, auc: 0.507, directionalAuc: 0.507, flag: no); (feature: diff_b2b, auc: 0.474, directionalAuc: 0.526, flag: no); (feature: diff_eG, auc: 0.616, directionalAuc: 0.616, flag: no); (feature: diff_eX, auc: 0.616, directionalAuc: 0.616, flag: no); (feature: diff_eliminated, auc: 0.471, directionalAuc: 0.529, flag: no); (feature: diff_foShare, auc: 0.539, directionalAuc: 0.539, flag: no); (feature: diff_form, auc: 0.576, directionalAuc: 0.576, flag: no); (feature: diff_h2h, auc: 0.559, directionalAuc: 0.559, flag: no); (feature: diff_hits, auc: 0.498, directionalAuc: 0.502, flag: no); (feature: diff_home, auc: 0.5, directionalAuc: 0.5, flag: no); (feature: diff_lEvA, auc: 0.59, directionalAuc: 0.59, flag: no); (feature: diff_lEvF, auc: 0.589, directionalAuc: 0.589, flag: no) … +37 more, strongestRaw: (feature: h_luSk, auc: 0.618, directionalAuc: 0.618, flag: no); (feature: a_luSk, auc: 0.382, directionalAuc: 0.618, flag: no); (feature: h_eG, auc: 0.616, directionalAuc: 0.616, flag: no); (feature: a_eG, auc: 0.384, directionalAuc: 0.616, flag: no); (feature: h_eX, auc: 0.616, directionalAuc: 0.616, flag: no); (feature: a_eX, auc: 0.385, directionalAuc: 0.616, flag: no); (feature: h_form, auc: 0.576, directionalAuc: 0.576, flag: no); (feature: a_form, auc: 0.424, directionalAuc: 0.576, flag: no); (feature: h_h2h, auc: 0.559, directionalAuc: 0.559, flag: no); (feature: a_h2h, auc: 0.441, directionalAuc: 0.559, flag: no), strongestDiff: (feature: diff_luSk, auc: 0.618, directionalAuc: 0.618, flag: no); (feature: diff_eG, auc: 0.616, directionalAuc: 0.616, flag: no); (feature: diff_eX, auc: 0.616, directionalAuc: 0.616, flag: no); (feature: diff_lEvA, auc: 0.59, directionalAuc: 0.59, flag: no); (feature: diff_lEvF, auc: 0.589, directionalAuc: 0.589, flag: no); (feature: diff_lPP, auc: 0.578, directionalAuc: 0.578, flag: no); (feature: diff_form, auc: 0.576, directionalAuc: 0.576, flag: no); (feature: diff_h2h, auc: 0.559, directionalAuc: 0.559, flag: no); (feature: diff_trap, auc: 0.556, directionalAuc: 0.556, flag: no); (feature: diff_lSv, auc: 0.551, directionalAuc: 0.551, flag: no)
- multivariateBaseline
- auc: 0.608, topGain: (feature: a_hits, gainShare: 0.028); (feature: h_prevOppStr, gainShare: 0.027); (feature: h_luSk, gainShare: 0.026); (feature: a_lPP, gainShare: 0.025); (feature: a_lPK, gainShare: 0.025); (feature: h_lPP, gainShare: 0.025); (feature: a_nextOppStr, gainShare: 0.025); (feature: h_refRate, gainShare: 0.025); (feature: h_eG, gainShare: 0.025); (feature: h_eX, gainShare: 0.024)
- canary
- rows: (feature: canary_final_margin, auc: 0.977, directionalAuc: 0.977, flag: yes, expectedAuc: —); (feature: canary_homeWin_plus_noise_sd0.5, auc: 0.925, directionalAuc: 0.925, flag: yes, expectedAuc: 0.921); (feature: canary_homeWin_plus_noise_sd1.5_weak, auc: 0.681, directionalAuc: 0.681, flag: no, expectedAuc: 0.681), multivariateAucWithCanary: 0.923, canaryGainRank: 1, canaryGainShare: 0.539, topGainWithCanary: (feature: canary_homeWin_plus_noise_sd0.5, gainShare: 0.539); (feature: a_oppGkShare, gainShare: 0.013); (feature: h_lPP, gainShare: 0.013); (feature: h_luSk, gainShare: 0.013); (feature: a_hits, gainShare: 0.012), detectorProven: yes, weakCanaryFlagged: no, note: the sd1.5 canary (theoretical AUC 0.68) is BELOW the 0.75 rule on purpose: it documents the smallest leak the univariate rule can miss
- runtimeSeconds
- 328.9
Is a challenger really different?
- generatedAt
- 2026-09-29T02:58:55Z
- seed
- 11
- method
- dieboldMariano: stat = mean(d)/sqrt(LRV/n), d = loss_A - loss_B per game; LRV with Bartlett lags h-1 = 0 (one-step forecasts) plus a Newey-West automatic-lag variant; pNormal from N(0,1); pHLN = Harvey-Leybourne-Newbold correction against t(n-1), bootstrap: paired percentile bootstrap of mean(d), 5000 resamples of games with replacement, minEvaluable: 100
- live
- a: champion, b: ratings-only, n: 14, meanLossA: 0.691, meanLossB: 0.718, meanDiff: -0.027, dieboldMariano: (n: 14, meanDiff: -0.027, lags: 0, stat: -0.55, pNormal: 0.582, statHLN: -0.53, pHLN: 0.605, neweyWest: (lags: 2, stat: -0.613, pNormal: 0.54)), pairedBootstrap: (n: 14, mean: -0.027, range: -0.119, 0.071, shareBelowZero: 0.71, excludesZero: no, B: 5000), verdict: not evaluable: only 14 games (needs 100+); reported for the record, context: (source: data/ledger/live-games.csv, gameTypes: 1, note: preseason games; lineups known), accuracy: (champion: 0.5, ratingsOnly: 0.5)
- engines2025
- win: (n: 1311, engines: (closed: 2 fields, markov: 2 fields, mc: 2 fields), pairs: (9 fields); (9 fields); (9 fields)), finish: (n: 1311, engines: (closed: 2 fields, markov: 2 fields, mc: 2 fields), pairs: (9 fields); (9 fields); (9 fields)), score: (n: 1311, engines: (closed: 2 fields, markov: 2 fields, mc: 2 fields), pairs: (9 fields); (9 fields); (9 fields))
- championVsClosed2025
- a: champion, b: closed engine, n: 1311, meanLossA: 0.681, meanLossB: 0.676, meanDiff: 0.005, dieboldMariano: (n: 1311, meanDiff: 0.005, lags: 0, stat: 1.532, pNormal: 0.126, statHLN: 1.531, pHLN: 0.126, neweyWest: (lags: 7, stat: 1.544, pNormal: 0.123)), pairedBootstrap: (n: 1311, mean: 0.005, range: -0.001, 0.011, shareBelowZero: 0.062, excludesZero: no, B: 5000), verdict: not distinguishable, context: (season: 2025, matchedGames: 1311, columnSemantics: preds-2025.csv *_win = probability assigned to the actual winner (not P(home win)); champion converted to -log P(actual))
- columnSemanticsNote
- preds-2025.csv columns *_win/*_finish/*_score are the probabilities each engine assigned to what actually happened (winner, finish type, exact score); the log s
- realityCheck
- —
- realityCheckReason
- White (2000) reality check / Hansen SPA over the 49,152 block-search versions needs the per-game loss series of every version; data/models/block-search.json sto
- runtimeSeconds
- 0.6
Has the world moved?
- generatedAt
- 2026-09-29T15:09:32Z
- method
- Drift: feature distribution shift in features.csv and change-points in the champion's per-game residuals. Features: for every h_/a_ column, the population stab
- features
- columns: 98, summaryTable: (season: 2018, vs: pooled2018_2022, reference: 2018-2022 (includes itself), nCurrent: 1271, nReference: 5845, psiOver0.10: 26, psiOver0.25: 8, ksPBelow0.01: 51, medianPsi: 0.026, top10: 10 items); (season: 2019, vs: previous, reference: 2018, nCurrent: 1082, nReference: 1271, psiOver0.10: 27, psiOver0.25: 21, ksPBelow0.01: 40, medianPsi: 0.02, top10: 10 items); (season: 2019, vs: pooled2018_2022, reference: 2018-2022 (includes itself), nCurrent: 1082, nReference: 5845, psiOver0.10: 13, psiOver0.25: 4, ksPBelow0.01: 33, medianPsi: 0.013, top10: 10 items); (season: 2020, vs: previous, reference: 2019, nCurrent: 868, nReference: 1082, psiOver0.10: 25, psiOver0.25: 4, ksPBelow0.01: 37, medianPsi: 0.039, top10: 10 items); (season: 2020, vs: pooled2018_2022, reference: 2018-2022 (includes itself), nCurrent: 868, nReference: 5845, psiOver0.10: 16, psiOver0.25: 7, ksPBelow0.01: 38, medianPsi: 0.03, top10: 10 items); (season: 2021, vs: previous, reference: 2020, nCurrent: 1312, nReference: 868, psiOver0.10: 24, psiOver0.25: 9, ksPBelow0.01: 47, medianPsi: 0.03, top10: 10 items); (season: 2021, vs: pooled2018_2022, reference: 2018-2022 (includes itself), nCurrent: 1312, nReference: 5845, psiOver0.10: 2, psiOver0.25: 2, ksPBelow0.01: 33, medianPsi: 0.011, top10: 10 items); (season: 2022, vs: previous, reference: 2021, nCurrent: 1312, nReference: 1312, psiOver0.10: 17, psiOver0.25: 6, ksPBelow0.01: 32, medianPsi: 0.013, top10: 10 items); (season: 2022, vs: pooled2018_2022, reference: 2018-2022 (includes itself), nCurrent: 1312, nReference: 5845, psiOver0.10: 17, psiOver0.25: 6, ksPBelow0.01: 41, medianPsi: 0.013, top10: 10 items); (season: 2023, vs: previous, reference: 2022, nCurrent: 1312, nReference: 1312, psiOver0.10: 10, psiOver0.25: 6, ksPBelow0.01: 20, medianPsi: 0.01, top10: 10 items); (season: 2023, vs: pooled2018_2022, reference: 2018-2022, nCurrent: 1312, nReference: 5845, psiOver0.10: 23, psiOver0.25: 8, ksPBelow0.01: 37, medianPsi: 0.02, top10: 10 items); (season: 2024, vs: previous, reference: 2023, nCurrent: 1312, nReference: 1312, psiOver0.10: 14, psiOver0.25: 6, ksPBelow0.01: 28, medianPsi: 0.01, top10: 10 items) … +3 more, pairs: (2018_vs_pooled2018_2022: (summary: 10 fields, features: 98 items), 2019_vs_previous: (summary: 10 fields, features: 98 items), 2019_vs_pooled2018_2022: (summary: 10 fields, features: 98 items), 2020_vs_previous: (summary: 10 fields, features: 98 items), 2020_vs_pooled2018_2022: (summary: 10 fields, features: 98 items), 2021_vs_previous: (summary: 10 fields, features: 98 items), 2021_vs_pooled2018_2022: (summary: 10 fields, features: 98 items), 2022_vs_previous: (summary: 10 fields, features: 98 items), 2022_vs_pooled2018_2022: (summary: 10 fields, features: 98 items), 2023_vs_previous: (summary: 10 fields, features: 98 items), 2023_vs_pooled2018_2022: (summary: 10 fields, features: 98 items), 2024_vs_previous: (summary: 10 fields, features: 98 items) … +3 fields), reading: PSI: <0.10 stable, 0.10-0.25 moderate, >0.25 major. Season-level inputs (league scoring rate, referee rates, takeaway/giveaway tracking) shift every season by construction; the champion uses home-minus-away differences, which cancel league-level shifts.
- championCusum
- ordering: schedule date, then game id, surprise: (label: champion log-loss surprise (L - H(p)), n: 7006, meanResidual: 0.004, sd: 0.284, maxDeviation: 41.51, maxDeviationScaled: 1.748, changePointIndex: 1977, changePointDate: 2022-04-26, meanBefore: -0.017, meanAfter: 0.012, permutationP: 0.009, permutations: 2000 … +2 fields), outcomeResidual: (label: champion outcome residual (y - p), n: 7006, meanResidual: 0, sd: 0.48, maxDeviation: 29.253, maxDeviationScaled: 0.728, changePointIndex: 5655, changePointDate: 2025-04-12, meanBefore: 0.005, meanAfter: -0.022, permutationP: 0.65, permutations: 2000 … +2 fields), perSeason: (2020: (n: 803, meanSurprise: -0.013, meanOutcomeResidual: 0.003, logLoss: 0.648), 2021: (n: 1210, meanSurprise: -0.017, meanOutcomeResidual: 0, logLoss: 0.626), 2022: (n: 1217, meanSurprise: 0.014, meanOutcomeResidual: -0.009, logLoss: 0.646), 2023: (n: 1230, meanSurprise: 0.012, meanOutcomeResidual: 0.003, logLoss: 0.654), 2024: (n: 1235, meanSurprise: -0.001, meanOutcomeResidual: 0.024, logLoss: 0.653), 2025: (n: 1311, meanSurprise: 0.021, meanOutcomeResidual: -0.019, logLoss: 0.681))
- liveLedgerCusum
- n: 14, insufficient: yes, minGames: 50, undated: 0, reason: only 14 games; change-point tests need at least 50, meanSurprise: 0.036, meanOutcomeResidual: -0.021, cusumPathSurprise: -0.149, -0.385, -0.457, -0.406, -0.109, 0.161, 0.276, 0.286, 0.009, -0.04, -0.288, 0.561 … +2 more, dates: 2026-09-26, 2026-09-26, 2026-09-26, 2026-09-26, 2026-09-26, 2026-09-26, 2026-09-26, 2026-09-26, 2026-09-26, 2026-09-26, 2026-09-26, 2026-09-26 … +2 more, forTheRecord: (surprise: (label: live surprise, n: 14, meanResidual: 0.036, sd: 0.298, maxDeviation: 0.682, maxDeviationScaled: 0.611, changePointIndex: 10, changePointDate: 2026-09-26, meanBefore: -0.026, meanAfter: 0.263, permutationP: 0.705, permutations: 500 … +2 fields), outcome: (label: live outcome residual, n: 14, meanResidual: -0.021, sd: 0.517, maxDeviation: 0.977, maxDeviationScaled: 0.505, changePointIndex: 11, changePointDate: 2026-09-26, meanBefore: -0.102, meanAfter: 0.468, permutationP: 0.876, permutations: 500 … +2 fields))
- runtimeSeconds
- 1.3
Pre-registered alternatives
stateSpace
- generatedAt
- 2026-09-29T03:06:57Z
- registerIds
- state-space-kalman-v1, state-space-kalman-v2
- finalVersion
- state-space-kalman-v2
- decision
- display rating only; remains a challenger (pooled proof does not beat the champion with the range excluding zero)
- method
- State-space (Kalman) team-strength model: a pre-registered challenger to the champion. Model (linear-Gaussian on regulation goals, documented choice): state
- data
- regularSeasonGames: 9781, seasons: 2018-2025 (start years), scheduleGamesWithoutScore: 0, preseason2026Games: 32, clubAlias: (ARI: UTA), observation: regulation goals incl. empty-net
- tuning
- objective: home-win log loss on 2020-2022 (walk-forward, 2018-2019 burn-in), versions: (state-space-kalman-v1: (grid: 3 fields, best: 5 fields, tunedOnGridEdge: 3 fields, table: 60 items), state-space-kalman-v2: (grid: 3 fields, best: 5 fields, tunedOnGridEdge: 3 fields, table: 100 items)), best: (q: 0, sigma2: 3.5, seasonInflation: 0.02, logLoss: 0.657, games: 3491), finalVersionRule: the version with the lower selection (2020-2022) log loss; never chosen on the proof seasons
- proof
- state-space-kalman-v1: (2023: (n: 1230, logLossStateSpace: 0.66, logLossChampion: 0.654, brierStateSpace: 0.234, brierChampion: 0.231, accuracyStateSpace: 0.608, accuracyChampion: 0.616, diffStateSpaceMinusChampion: 6 fields, dieboldMariano: 8 fields, blend5050LogLoss: 0.655, corrProbabilities: 0.893, winnerMismatchesVsChampionFile: 0), 2024: (n: 1235, logLossStateSpace: 0.664, logLossChampion: 0.653, brierStateSpace: 0.236, brierChampion: 0.231, accuracyStateSpace: 0.598, accuracyChampion: 0.602, diffStateSpaceMinusChampion: 6 fields, dieboldMariano: 8 fields, blend5050LogLoss: 0.656, corrProbabilities: 0.84, winnerMismatchesVsChampionFile: 0), 2025: (n: 1311, logLossStateSpace: 0.695, logLossChampion: 0.681, brierStateSpace: 0.25, brierChampion: 0.244, accuracyStateSpace: 0.526, accuracyChampion: 0.556, diffStateSpaceMinusChampion: 6 fields, dieboldMariano: 8 fields, blend5050LogLoss: 0.685, corrProbabilities: 0.797, winnerMismatchesVsChampionFile: 0), pooled: (n: 3776, logLossStateSpace: 0.673, logLossChampion: 0.663, brierStateSpace: 0.24, brierChampion: 0.236, accuracyStateSpace: 0.576, accuracyChampion: 0.591, diffStateSpaceMinusChampion: 6 fields, dieboldMariano: 8 fields, blend5050LogLoss: 0.666, corrProbabilities: 0.85)), state-space-kalman-v2: (2023: (n: 1230, logLossStateSpace: 0.659, logLossChampion: 0.654, brierStateSpace: 0.234, brierChampion: 0.231, accuracyStateSpace: 0.611, accuracyChampion: 0.616, diffStateSpaceMinusChampion: 6 fields, dieboldMariano: 8 fields, blend5050LogLoss: 0.654, corrProbabilities: 0.894, winnerMismatchesVsChampionFile: 0), 2024: (n: 1235, logLossStateSpace: 0.664, logLossChampion: 0.653, brierStateSpace: 0.236, brierChampion: 0.231, accuracyStateSpace: 0.594, accuracyChampion: 0.602, diffStateSpaceMinusChampion: 6 fields, dieboldMariano: 8 fields, blend5050LogLoss: 0.656, corrProbabilities: 0.844, winnerMismatchesVsChampionFile: 0), 2025: (n: 1311, logLossStateSpace: 0.695, logLossChampion: 0.681, brierStateSpace: 0.251, brierChampion: 0.244, accuracyStateSpace: 0.529, accuracyChampion: 0.556, diffStateSpaceMinusChampion: 6 fields, dieboldMariano: 8 fields, blend5050LogLoss: 0.685, corrProbabilities: 0.801, winnerMismatchesVsChampionFile: 0), pooled: (n: 3776, logLossStateSpace: 0.673, logLossChampion: 0.663, brierStateSpace: 0.24, brierChampion: 0.236, accuracyStateSpace: 0.577, accuracyChampion: 0.591, diffStateSpaceMinusChampion: 6 fields, dieboldMariano: 8 fields, blend5050LogLoss: 0.665, corrProbabilities: 0.853))
- clubs
- endOf2025_26: (homeAdvantage: 0.207, homeAdvantageSD: 0.089, level: 2.929, levelSD: 0.073, clubs: (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields) … +20 more, reading: attack: goals per game above league level scored; defence: goals per game below league level conceded; net = attack + defence, asOf: 2026-04-16), latest: (homeAdvantage: 0.209, homeAdvantageSD: 0.09, level: 2.916, levelSD: 0.074, clubs: (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields); (8 fields) … +20 more, reading: attack: goals per game above league level scored; defence: goals per game below league level conceded; net = attack + defence, asOf: 2026-09-22, includes: 32 2026 preseason games from goals-pre2026.json (observation variance x3.0), after the 2026-27 season-start inflation)
- perGamePredictionsFile
- data/models/state-space-rates.json
- runtimeSeconds
- 79.4
scoreDistributions
- generatedAt
- 2026-09-29T03:13:03Z
- registerId
- score-dist-challengers-v1
- decision
- keep independent Poisson (no family beats it with the range excluding zero and no worse 6+ share)
- method
- Score-distribution challengers for 2025-26, all fed by the SAME walk-forward goal rates. Rates: the Kalman state-space filtered means (data/models/state-space-
- data
- ratesFile: data/models/state-space-rates.json, ratesParams: (q: 0, sigma2: 3.5, seasonInflation: 0.02, logLoss: 0.657, games: 3491), ratesVersion: state-space-kalman-v2, trainSeasons: 2023, 2024, testSeason: 2025, nTrain: 2624, nTest: 1312, gridMaxGoals: 12, rateCalibration: (home: 1.002, away: 0.986), leagueAverage: (home: 3.102, away: 2.844), scoreDefinition: regulation score (periods 1-3) from the goals files; the score at the end of regulation, not the final
- fits
- dixonColes: (rho: -0.116, grid: 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items … +29 more), bivariatePoisson: (lambda3: 0, grid: 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items … +13 more), negativeBinomial: (r: 522.36, grid: 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items; 2 items … +13 more, impliedVarianceInflationAtLambda3: 1.006)
- test2025
- poissonLeagueAverage: (n: 1312, scoreLogScore: 3.851, crpsTotalGoals: 1.307, sixPlus: (predicted: 0.546, actual: 0.573), sevenPlus: (predicted: 0.385, actual: 0.379), meanTotalPredicted: 5.946, meanTotalActual: 6.005, varianceTotalPredicted: 5.943, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.102, tieShare: (predicted: 0.167, actual: 0.248), homeWinLogLoss: 0.693 … +3 fields), poissonIndependent: (n: 1312, scoreLogScore: 3.844, crpsTotalGoals: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), sevenPlus: (predicted: 0.382, actual: 0.379), meanTotalPredicted: 5.924, meanTotalActual: 6.005, varianceTotalPredicted: 5.919, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.097, tieShare: (predicted: 0.163, actual: 0.248), homeWinLogLoss: 0.695 … +2 fields), dixonColes: (n: 1312, scoreLogScore: 3.844, crpsTotalGoals: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), sevenPlus: (predicted: 0.382, actual: 0.379), meanTotalPredicted: 5.924, meanTotalActual: 6.005, varianceTotalPredicted: 5.924, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.094, tieShare: (predicted: 0.168, actual: 0.248), homeWinLogLoss: 0.695 … +3 fields), bivariatePoisson: (n: 1312, scoreLogScore: 3.844, crpsTotalGoals: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), sevenPlus: (predicted: 0.382, actual: 0.379), meanTotalPredicted: 5.924, meanTotalActual: 6.005, varianceTotalPredicted: 5.919, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.097, tieShare: (predicted: 0.163, actual: 0.248), homeWinLogLoss: 0.695 … +3 fields), negativeBinomial: (n: 1312, scoreLogScore: 3.845, crpsTotalGoals: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), sevenPlus: (predicted: 0.382, actual: 0.379), meanTotalPredicted: 5.923, meanTotalActual: 6.005, varianceTotalPredicted: 5.952, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.097, tieShare: (predicted: 0.162, actual: 0.248), homeWinLogLoss: 0.695 … +3 fields)
- notComparableWith
- preds-2025.csv *_score (final score incl. OT/SO, different engines)
- dispersionNote
- The actual variance of total regulation goals is BELOW every family here (under-dispersion, in line with the refractory finding in momentum.json). Negative bino
- runtimeSeconds
- 1.2
Floor and ceiling
- naive
- home club wins; last season's points; ratings only (logged in the ledger as a challenger)
- market
- closing lines not yet licensed (The Odds API key pending); the honest ceiling is unknown until then
- bookSimulation
- bookLogLoss: 0.671
Where it stops
- Saturated on public box-score data: 49,152 versions found nothing better on proof seasons.
- 60 to 70% tier was about four points over-confident in the v2 validation.
- 2025-26 was the hardest season for every model (~0.68).
- Starting goalies are not named by the league feed before puck drop; a starters feed is connected but keyless.
Score engine
Expected goals per side, finish-type probabilities and most-likely scores on game pages and in the email.
What it learned from
- trainedOn
- 2023-24, 2024-25
- testedOn
- 2025-26
How it did on seasons it had not seen
- results
- games: 1311, winLogLoss: 0.677, finishLogLoss: 1.234, scoreLogLoss: 3.644, totalGoalsMAE: 1.838, mostLikelyScoreHit: 0.076, sixPlusGoals: (predicted: 0.524, actual: 0.574)
- calibration
- (bucket: 10-20%, games: 1, predicted: 0.182, actual: 0); (bucket: 20-30%, games: 10, predicted: 0.265, actual: 0); (bucket: 30-40%, games: 104, predicted: 0.362, actual: 0.404); (bucket: 40-50%, games: 353, predicted: 0.456, actual: 0.473); (bucket: 50-60%, games: 507, predicted: 0.548, actual: 0.503); (bucket: 60-70%, games: 277, predicted: 0.642, actual: 0.632); (bucket: 70-80%, games: 53, predicted: 0.735, actual: 0.792); (bucket: 80-90%, games: 6, predicted: 0.81, actual: 0.5)
- tieInflation
- 1.46
Do 60% calls win 60% of the time?
Not on file.
Pre-registered alternatives
present
Nothing to show.
generatedAt
Nothing to show.
registerId
Nothing to show.
decision
Nothing to show.
method
Nothing to show.
data
- ratesFile
- data/models/state-space-rates.json
- ratesParams
- q: 0, sigma2: 3.5, seasonInflation: 0.02, logLoss: 0.657, games: 3491
- ratesVersion
- state-space-kalman-v2
- trainSeasons
- 2023, 2024
- testSeason
- 2025
- nTrain
- 2624
- nTest
- 1312
- gridMaxGoals
- 12
- rateCalibration
- home: 1.002, away: 0.986
- leagueAverage
- home: 3.102, away: 2.844
- scoreDefinition
- regulation score (periods 1-3) from the goals files; the score at the end of regulation, not the final
fits
- dixonColes
- rho: -0.116, grid: -0.2, -3.864; -0.19, -3.863; -0.18, -3.863; -0.17, -3.863; -0.16, -3.863; -0.15, -3.862; -0.14, -3.862; -0.13, -3.862; -0.12, -3.862; -0.11, -3.862; -0.1, -3.862; -0.09, -3.862 … +29 more
- bivariatePoisson
- lambda3: 0, grid: 0, -3.864; 0.025, -3.865; 0.05, -3.866; 0.075, -3.866; 0.1, -3.867; 0.125, -3.868; 0.15, -3.869; 0.175, -3.87; 0.2, -3.872; 0.225, -3.873; 0.25, -3.874; 0.275, -3.876 … +13 more
- negativeBinomial
- r: 522.36, grid: 2, -4.137; 2.67, -4.057; 3.56, -3.997; 4.74, -3.952; 6.32, -3.921; 8.43, -3.9; 11.25, -3.886; 15, -3.877; 20, -3.872; 26.67, -3.869; 35.57, -3.867; 47.43, -3.866 … +13 more, impliedVarianceInflationAtLambda3: 1.006
test2025
- poissonLeagueAverage
- n: 1312, scoreLogScore: 3.851, crpsTotalGoals: 1.307, sixPlus: (predicted: 0.546, actual: 0.573), sevenPlus: (predicted: 0.385, actual: 0.379), meanTotalPredicted: 5.946, meanTotalActual: 6.005, varianceTotalPredicted: 5.943, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.102, tieShare: (predicted: 0.167, actual: 0.248), homeWinLogLoss: 0.693 … +3 fields
- poissonIndependent
- n: 1312, scoreLogScore: 3.844, crpsTotalGoals: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), sevenPlus: (predicted: 0.382, actual: 0.379), meanTotalPredicted: 5.924, meanTotalActual: 6.005, varianceTotalPredicted: 5.919, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.097, tieShare: (predicted: 0.163, actual: 0.248), homeWinLogLoss: 0.695 … +2 fields
- dixonColes
- n: 1312, scoreLogScore: 3.844, crpsTotalGoals: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), sevenPlus: (predicted: 0.382, actual: 0.379), meanTotalPredicted: 5.924, meanTotalActual: 6.005, varianceTotalPredicted: 5.924, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.094, tieShare: (predicted: 0.168, actual: 0.248), homeWinLogLoss: 0.695 … +3 fields
- bivariatePoisson
- n: 1312, scoreLogScore: 3.844, crpsTotalGoals: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), sevenPlus: (predicted: 0.382, actual: 0.379), meanTotalPredicted: 5.924, meanTotalActual: 6.005, varianceTotalPredicted: 5.919, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.097, tieShare: (predicted: 0.163, actual: 0.248), homeWinLogLoss: 0.695 … +3 fields
- negativeBinomial
- n: 1312, scoreLogScore: 3.845, crpsTotalGoals: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), sevenPlus: (predicted: 0.382, actual: 0.379), meanTotalPredicted: 5.923, meanTotalActual: 6.005, varianceTotalPredicted: 5.952, varianceTotalActual: 5.463, regulationOutcomeLogScore3way: 1.097, tieShare: (predicted: 0.162, actual: 0.248), homeWinLogLoss: 0.695 … +3 fields
notComparableWith
Nothing to show.
dispersionNote
Nothing to show.
runtimeSeconds
Nothing to show.
Do goals cluster?
- generatedAt
- 2026-09-29T03:14:33Z
- registerId
- momentum-hawkes-v1
- verdict
- no momentum: goals are LESS clustered than the time-varying, score-dependent null (a refractory window after a goal; holds in core time before 55:00 as well), a
- method
- Do goals cluster in time more than a Poisson process with a time-varying rate predicts? ("momentum") Data: goals files 2021-2025 (regular season, start years 2
- data
- seasons: 2021, 2022, 2023, 2024, 2025, games: 6549, regulationGoals: 39414, intervals: 32886
- nullModel
- replicates: 200, cells: 60 minutes x |score diff| {0,1,2,3+}, pseudoCountSeconds: 5, rateGoalsPer60ByDiff: 5.498, 6.368, 6.541, 6.553, rateLastMinuteByDiffGoalsPer60: 4.027, 31.768, 25.425, 4.663, homeShareBySignedDiff: (0: 0.513, 1: 0.528, 2: 0.563, 3: 0.536, -3: 0.489, -2: 0.49, -1: 0.503), goalsByMinute: 405, 540, 508, 576, 676, 549, 612, 605, 570, 568, 589, 583 … +48 more
- intervals
- data: (count: 32886, mean: 497.365, q10: 68, q25: 161, median: 366, q75: 695, q90: 1126, under30: 0.032, under60: 0.085, under120: 0.186, replyGoals60: 2797, replyGoals120: 6125), dataVsNull: (mean: (data: 497.365, nullMean: 489.147, nullSD: 2.562, null2.5: 484.291, null97.5: 493.925, z: 3.21, mcPTwoSided: 0.005), q10: (data: 68, nullMean: 50.894, nullSD: 0.899, null2.5: 49, null97.5: 52.025, z: 19.03, mcPTwoSided: 0.005), q25: (data: 161, nullMean: 142.331, nullSD: 1.623, null2.5: 139, null97.5: 145, z: 11.5, mcPTwoSided: 0.005), median: (data: 366, nullMean: 349.048, nullSD: 2.958, null2.5: 344, null97.5: 355, z: 5.73, mcPTwoSided: 0.005), q75: (data: 695, nullMean: 693.928, nullSD: 4.739, null2.5: 685, null97.5: 703, z: 0.23, mcPTwoSided: 0.935), q90: (data: 1126, nullMean: 1126.455, nullSD: 7.803, null2.5: 1110.975, null97.5: 1141, z: -0.06, mcPTwoSided: 0.935), under30: (data: 0.032, nullMean: 0.059, nullSD: 0.001, null2.5: 0.057, null97.5: 0.062, z: -20.57, mcPTwoSided: 0.005), under60: (data: 0.085, nullMean: 0.116, nullSD: 0.002, null2.5: 0.113, null97.5: 0.119, z: -17.9, mcPTwoSided: 0.005), under120: (data: 0.186, nullMean: 0.215, nullSD: 0.002, null2.5: 0.212, null97.5: 0.221, z: -12.73, mcPTwoSided: 0.005), replyGoals60: (data: 2797, nullMean: 3808.26, nullSD: 65, null2.5: 3701, null97.5: 3942.1, z: -15.56, mcPTwoSided: 0.005), replyGoals120: (data: 6125, nullMean: 7091.535, nullSD: 95.555, null2.5: 6910, null97.5: 7277, z: -10.11, mcPTwoSided: 0.005), count: (data: 32886, nullMean: 32901.155, nullSD: 187.357, null2.5: 32525.85, null97.5: 33292.025, z: -0.08, mcPTwoSided: 0.955) … +1 fields), excessReplyGoalsWithin60s: -1011.3, coreTime: (definition: intervals whose second goal is before 55:00, intervalsData: 27789, intervalsNullMean: 27799.5, under60: (data: 0.084, nullMean: 0.11, null2.5: 0.106, null97.5: 0.114, z: -13.99), windows: (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields)), direction: anti-clustering (refractory)
- hawkes
- data: (alpha: 0, beta: 0.005, meanDelaySeconds: 212.158, muPerSecond: 0.001, 0.002, 0.002, muPerPeriodGoals: 1.77, 2.096, 2.152, logLik: -291280.265, iterations: 98, goals: 39414, games: 6549, logLikPoissonByPeriod: -291280.234), bootstrap: (B: 200, alpha95: 0, 0, beta95: 0.004, 0.005, meanDelaySeconds95: 198.8, 224.6), null: (replicatesFitted: 40, alphaMean: 0.006, alpha95: 0.001, 0.014, meanDelayMean: 43.5, alphaMcP: 1, reading: alpha the misspecified Hawkes (period-constant background) attributes to self-excitation on data with NONE), byPeriodData: (period1: (alpha: 0, meanDelaySeconds: 121.1, goals: 11590), period2: (alpha: 0, meanDelaySeconds: 127, goals: 13728), period3: (alpha: 0, meanDelaySeconds: 102.7, goals: 14096)), branchingRatioReading: alpha = expected number of goals triggered by one goal under the Hawkes reading; compare with null alpha before believing it, betaNote: when alpha is (numerically) zero the kernel carries no mass and beta / the mean delay are not identified; they are the EM values at convergence and should not be read as a timescale
- runtimeSeconds
- 73.9
Where it stops
- Totals under-predicted: 52.9% of games predicted at six or more goals against 57.3% observed.
- Poisson independence; Dixon-Coles, bivariate Poisson and negative binomial are on the register as challengers.
Player props (anytime goal, anytime point)
Per-dressed-skater probabilities recorded in the ledger at the lineup stage.
What it learned from
- goalLayer
- 2019-20 to 2024-25
- pointLayer
- 2022-23 to 2024-25
- constants
- DECAY: 0.97, K_SHOT: 8, K_FIN: 120
How it did on seasons it had not seen
- goal
- season: 2025-26, logLoss: 0.381, baseline: 0.386
- point
- season: 2025-26, logLoss: 0.592, baseline: 0.597
Do 60% calls win 60% of the time?
Not on file.
Where it stops
- Stars (40%+) about 3.5 points over-confident.
- Needs the confirmed starter; the league feed lists both goalies.
Lineup strength
Lineup-aware repricing when rosters appear 80 to 160 minutes before puck drop.
What it learned from
Not on file.
How it did on seasons it had not seen
- note
- Adds 0.6777 to 0.6772 win log loss on 2025-26 (register entry hist-2026-09-26-lineups).
Do 60% calls win 60% of the time?
Not on file.
Where it stops
- Relative on-ice xG, not isolated: RAPM values replace these ratings when the measurement desk's run completes.
Pre-registration register
Hypothesis, metric, proof seasons and decision rule, written before the run; the outcome, whatever it was.
| Entry | Registered | Status | Hypothesis | Outcome | Details |
|---|---|---|---|---|---|
| prospect-growth-age-v1 | 2026-09-29 | rejected | P1. The site's hand-set age line (1.00 at 18, minus 0.06 a year, floor 0.58) is flatter than measured same-league growth in points per game (whole classes: steeper through 22). Replacing it with the measured curve, selection-corrected (same-league pairs reweighted so their production mix matches every line at that age in that league, because stayers are the unpromoted), ranks prospects better. Within one age the change only reorders by months of age, so the expected per-age gain is small (under 0.01); the effect is across ages. | proofMeanRho 0.473 · championMeanRho 0.471 · gain mean: 0.002, range95: -0.002, 0.006, shareAboveZero: 0.807, players: 644, B: 2000, seed: 5 · logLossVsChampion -0.003 · devGain mean: 0.003, range95: 0, 0.005, shareAboveZero: 0.969, players: 1236, B: 2000, seed: 5 · decision REJECT: the 95% range of the gain does not lie above zero | |
| prospect-pick-blend-v1 | 2026-09-29 | adopted | P2. Draft position ranks NHL games through 23 better than the formula at 18 and 19 (whole classes 0.605 vs 0.448 at 18). A blend of the two percentiles with one weight w, chosen on the development classes, beats the formula alone. Expected: a large gain at 18, shrinking by 20. | proofMeanRho 0.636 · championMeanRho 0.471 · gain mean: 0.165, range95: 0.127, 0.203, shareAboveZero: 1, players: 644, B: 2000, seed: 5 · logLossVsChampion -0.088 · devGain mean: 0.136, range95: 0.105, 0.163, shareAboveZero: 1, players: 1236, B: 2000, seed: 5 · decision PASS: gain range above zero and graduation log loss not worse | |
| prospect-corrected-ladder-v1 | 2026-09-29 | rejected | P3. The league ladder is measured on players who moved to the NHL, who are the league's best scorers, so it understates what a typical line is worth. Re-measuring each league's translation with movers reweighted to the league's production deciles (inverse probability of moving), on 2010-2015 classes only, and using it in place of the site's ladder ranks prospects better. | proofMeanRho 0.476 · championMeanRho 0.471 · gain mean: 0.005, range95: -0.004, 0.015, shareAboveZero: 0.855, players: 644, B: 2000, seed: 5 · logLossVsChampion -0.002 · devGain mean: -0.005, range95: -0.015, 0.004, shareAboveZero: 0.138, players: 1236, B: 2000, seed: 5 · decision REJECT: the 95% range of the gain does not lie above zero | |
| prospect-fitted-benchmark-v1 | 2026-09-29 | benchmark | P4. A fitted benchmark (L2-regularised logistic regression on graduation, numpy only because lightgbm is not installed) on the formula's parts, age, position, league family, height, draft-year production and draft position shows how much ranking the formula leaves on the table. It is a ceiling, not a site candidate: it uses draft position and a fitted black box the page could not print. | proofMeanRho 0.63 · championMeanRho 0.471 · gain mean: 0.159, range95: 0.122, 0.197, shareAboveZero: 1, players: 644, B: 2000, seed: 5 · logLossVsChampion -0.095 · devGain mean: 0.146, range95: 0.12, 0.17, shareAboveZero: 1, players: 1236, B: 2000, seed: 5 · decision PASS: gain range above zero and graduation log loss not worse | |
| cba-retrieval-r1-vocab-repair-v1 | 2026-09-29 | rejected | The contract-term rule of the question expansion fires on 'long-term', 'how long' and 'length' in questions that are not about contract length, and several rules append words that are common across the agreement. Repairing the triggers and dropping the common words raises the rank of the passage that answers a fan-worded question. Expected: a gain concentrated on long-term-injury and duration questions and small overall (the rule fired on 24 of 110 development questions); 60 held-out questions may be too few to show it. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params dropAboveShare: 0.05 · development leaveOneOut: (mrr: 0.674, recallAt5: 0.8, picks: (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.573, recallAt5: 0.8, recallAt8: 0.836, mrr: 0.674), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: 0.036, lo: 0.009, hi: 0.073, better: 4, worse: 0), mrr: (difference: 0.043, lo: 0.014, hi: 0.075, better: 17, worse: 3)) · heldOut n: 60, candidate: (recallAt1: 0.25, recallAt5: 0.517, recallAt8: 0.583, mrr: 0.366), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: -0.017, lo: -0.05, hi: 0, better: 0, worse: 1), mrr: (difference: -0.002, lo: -0.058, hi: 0.052, better: 9, worse: 12) | |
| cba-retrieval-r2-weighted-expansion-v1 | 2026-09-29 | rejected | Expansion words at full weight can outweigh the fan's own words. Entering them at a reduced weight keeps the bridge between a fan's vocabulary and the agreement's while letting the question lead. Expected: a small gain in MRR; possibly none if R1 has already removed the expansions that did harm. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params weight: 0.25 · development leaveOneOut: (mrr: 0.652, recallAt5: 0.782, picks: (2 fields); (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.564, recallAt5: 0.782, recallAt8: 0.855, mrr: 0.664), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: 0.018, lo: -0.018, hi: 0.054, better: 3, worse: 1), mrr: (difference: 0.033, lo: -0.003, hi: 0.072, better: 17, worse: 3)) · heldOut n: 60, candidate: (recallAt1: 0.25, recallAt5: 0.483, recallAt8: 0.55, mrr: 0.354), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: -0.05, lo: -0.133, hi: 0.017, better: 1, worse: 4), mrr: (difference: -0.015, lo: -0.088, hi: 0.057, better: 9, worse: 17) | |
| cba-retrieval-r3-ocr-word-split-v1 | 2026-09-29 | rejected | In the scanned 2025 MOU, words run together by the OCR are single tokens that no question contains. Splitting them, for the index only, into words of the clean documents (2013 CBA, 2020 MOU) makes those passages findable; the passage text shown and quoted is unchanged. Expected: gains only where the gold sits in the 2025 MOU (12 of the 60 held-out questions), so the overall effect may be too small to pass the rule. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params — · development leaveOneOut: (mrr: 0.627, recallAt5: 0.754, picks: (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.518, recallAt5: 0.754, recallAt8: 0.818, mrr: 0.627), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: -0.009, lo: -0.027, hi: 0, better: 0, worse: 1), mrr: (difference: -0.004, lo: -0.02, hi: 0.011, better: 7, worse: 13)) · heldOut n: 60, candidate: (recallAt1: 0.233, recallAt5: 0.517, recallAt8: 0.55, mrr: 0.364), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: -0.017, lo: -0.05, hi: 0, better: 0, worse: 1), mrr: (difference: -0.005, lo: -0.024, hi: 0.008, better: 6, worse: 12) | |
| cba-retrieval-r3b-ocr-letter-repair-v1 | 2026-09-29 | rejected | The same scan reads the letter m as 'in', 'im', 'iin' or 'rn' ('terin' for term, 'aimount' for amount, 'liimit' for limit, 'gaine' for game). Replacing, for the index only, a token that is absent from the clean vocabulary by the one clean-vocabulary word those confusions lead to makes the agreement's commonest words match in the 2025 MOU. Expected: larger than R3 on 2025 MOU questions, because the damaged words are the ones fans type; still limited to the 12 held-out questions whose gold sits there. Added by the desk; not in the coordinator's list. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params — · development leaveOneOut: (mrr: 0.628, recallAt5: 0.764, picks: (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.518, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.628), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: 0, lo: 0, hi: 0, better: 0, worse: 0), mrr: (difference: -0.003, lo: -0.013, hi: 0.005, better: 5, worse: 10)) · heldOut n: 60, candidate: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.55, mrr: 0.368), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: 0, lo: 0, hi: 0, better: 0, worse: 0), mrr: (difference: -0.001, lo: -0.001, hi: 0, better: 1, worse: 6) | |
| cba-retrieval-r4-field-weights-v1 | 2026-09-29 | rejected | The title is counted twice and the score is raised 8% per document rank; both were set by hand. A title weight or document boost chosen on development questions ranks the gold passage higher. Expected: small. On the development questions the boost was worth 2 points of recall@5, inside the noise, and the 2020 and 2025 MOU chunks all carry the same title, so the title can only help in the 2013 CBA. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params titleWeight: 1, rankBoost: 0.08 · development leaveOneOut: (mrr: 0.609, recallAt5: 0.746, picks: (2 fields); (2 fields); (2 fields); (2 fields); (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.527, recallAt5: 0.754, recallAt8: 0.818, mrr: 0.634), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: -0.009, lo: -0.027, hi: 0, better: 0, worse: 1), mrr: (difference: 0.003, lo: -0.011, hi: 0.018, better: 12, worse: 14)) · heldOut n: 60, candidate: (recallAt1: 0.233, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.36), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: 0, lo: 0, hi: 0, better: 0, worse: 0), mrr: (difference: -0.009, lo: -0.04, hi: 0.02, better: 9, worse: 15) | |
| cba-retrieval-r5-feedback-pass-v1 | 2026-09-29 | rejected | A second ranking pass that appends the most distinctive terms of the first pass's top three passages, at low weight, finds passages the fan's words miss. Expected: uncertain; feedback helps when the first pass is mostly right and hurts when it is wrong, and a fan's first pass is often wrong. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params terms: 5, weight: 0.1 · development leaveOneOut: (mrr: 0.609, recallAt5: 0.736, picks: (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.491, recallAt5: 0.736, recallAt8: 0.818, mrr: 0.609), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: -0.027, lo: -0.064, hi: 0, better: 0, worse: 3), mrr: (difference: -0.022, lo: -0.046, hi: -0.001, better: 13, worse: 22)) · heldOut n: 60, candidate: (recallAt1: 0.233, recallAt5: 0.517, recallAt8: 0.567, mrr: 0.363), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: -0.017, lo: -0.067, hi: 0.033, better: 1, worse: 2), mrr: (difference: -0.005, lo: -0.037, hi: 0.025, better: 16, worse: 17) | |
| cba-fallback-refusal-v1 | 2026-09-29 | adopted | While the model is unavailable the passages-only fallback answers every question, including the 30 of 30 development questions the agreement cannot answer. Retrieval evidence alone (top score, margin to the second, share of the question's words found in the top passages, a named player or club) separates enough of them to refuse or warn without often refusing a question the agreement does answer. Expected: the name rule catches the named-player and club questions; the evidence rule catches about half of the rest at 5% wrong refusals (the top score alone caught 50% at that point on development). | testedAt 2026-09-29T14:02:40.456Z · configuration B0 · rule mean: 27.339, 3.971, 0.694, 0.597, scale: 11.952, 4.368, 0.165, 0.183, weights: -2.448, 0.626, -0.833, -0.156, intercept: -2.515, weak: 0.289, none: 0.557 · development answerable: 110, unanswerable: 30, method: leave-one-out, levelC: (wrongRefusals: 5, wrongRate: 0.045, caught: 15, tpr: 0.5, tprInterval: 0.332, 0.668), levelBorC: (wronglyFlagged: 16, wrongRate: 0.145, caught: 23, tpr: 0.767, tprInterval: 0.591, 0.882), auc: 0.906 · heldOut withNameRule: (levelC: (caught: 15, tpr: 0.75, tprInterval: 2 items, wrongRefusals: 13, wrongRate: 0.217, wrongInterval: 2 items), levelBorC: (caught: 18, tpr: 0.9, wronglyFlagged: 25, wrongRate: 0.417), byNameRule: (unanswerable: 9, answerable: 3, answerableIds: 3 items)), evidenceRuleAlone: (levelC: (caught: 14, tpr: 0.7, tprInterval: 2 items, wrongRefusals: 11, wrongRate: 0.183, wrongInterval: 2 items), levelBorC: (caught: 18, tpr: 0.9, wronglyFlagged: 23, wrongRate: 0.383)), guard: (answerableNamingNoOne: 57, wronglyRefusedAtC: 10, rate: 0.175, limit: 0.1, holds: no, consequence: level (c) switched off for evidence; levels (a) and (b) only; the name rule stays) · shipped levels (a) confident and (b) weak from evidence; level (c) from evidence switched OFF by the registered guard (held-out answerable naming no one wrongly refused | |
| cba-number-gate-exact-v1 | 2026-09-29 | adopted | The check that every number in an answer appears in the passage it cites is a substring test, so a claimed '5' passes against a passage that says '50' or '2025'. An exact test on number tokens, compared as canonical decimals, closes that hole without failing answers that are right. | selfTest tools/assay/eval/selftest.mjs 20/20 incl. the 5-vs-fifty (50) false accept; vitest src/server/__tests__/grounded-core.test.ts: 8.50==8.5, 5!=50, 5!=15/2025, 50! · developmentAnswers no model answers exist (model mode not run); instead every gold answer and every gold number of the 160-item set was put through both gates as a sentence citing · sentences 258 · passOld 258 · passNew 258 · passOldFailNew 0 | |
| cba-clause-split-fix-v1 | 2026-09-29 | adopted | The fallback is meant to quote clauses of semicolon lists, but its split pattern has a literal 's' where whitespace was meant, so it never splits at '; ' and quotes whole sentences. Splitting on whitespace after a semicolon gives shorter quotes that are nearer the question. | split /(?<=;)s+/ · quotesBefore 657 · quotesAfter 657 · medianQuoteLengthBefore 362 · medianQuoteLengthAfter 329 · everyQuoteVerbatim true | |
| xg-clock-corrected-flags-v1 | 2026-09-29 | adopted | With the champion's rebound and rush flags recomputed on a clock that does not depend on the outcome (goal times moved back 1 s, the offset measured on 2023-24 + 2024-25, before the 3 s and 4 s windows are applied) and everything else unchanged, held-out log loss does NOT improve and is expected to get slightly worse, by about 0.0002 to 0.0003 per shot, because the raw flags carry information about the outcome through the clock. A worse log loss here means part of the champion's published accuracy was the clock. | offsetSeconds 1 · rows trainRows: 227076, testRows: 105288, testGames: 1232, testRowsInShotFile: 112091, testGamesInShotFile: 1312, trainGoals: 15987, testGoals: 7601 · trainLogLoss 0.221 · testLogLoss 0.229 · testAuc 0.756 · championReplicaTestLogLoss 0.229 | |
| xg-clock-corrected-split-v1 | 2026-09-29 | run | F's corrected flags inside the four strength-state fits of challenger B. Expected against the champion: the strength split's gain on the proof season (about 0.0005 with two seasons of training) less the cost of correcting the clock (about 0.0002 to 0.0003), a net gain near 0.0002 that may not clear the rule. The split's gain is not robust to the amount of training data: trained on one season it was worse than a single fit. | offsetSeconds 1 · rows trainRows: 227076, testRows: 105288, testGames: 1232, testRowsInShotFile: 112091, testGamesInShotFile: 1312, trainGoals: 15987, testGoals: 7601 · trainLogLoss 0.22 · testLogLoss 0.228 · testAuc 0.757 · championReplicaTestLogLoss 0.229 | |
| xg-clock-corrected-flurry-v1 | 2026-09-29 | run | G plus the flurry features of challenger C (seconds since the previous unblocked shot by the same club in the same period, capped at 30 s, and an indicator for under 3 s) computed on the corrected clock. Expected: little or nothing beyond G. The gain C showed with goal times moved back is expected to have been mostly a repair of the raw flags, which G already has. | offsetSeconds 1 · rows trainRows: 227076, testRows: 105288, testGames: 1232, testRowsInShotFile: 112091, testGamesInShotFile: 1312, trainGoals: 15987, testGoals: 7601 · trainLogLoss 0.22 · testLogLoss 0.228 · testAuc 0.757 · championReplicaTestLogLoss 0.229 | |
| xg-rink-adjusted-v1 | 2026-09-29 | run | Replacing the recorded shot distance with its league-equivalent (the arena quantile map estimated on 2023-24 + 2024-25 only, applied to the training rows and carried unchanged to 2025-26; angle recomputed from the same geometry where well conditioned) removes scorer error from the most important xG input and lowers held-out log loss. Expected effect: very small and possibly negative - the arena study found the scorer effect has shrunk to about 1 ft standard deviation against about 0.55 ft of noise, and that an unshrunk map carried into the next season did not improve agreement. | trainLogLoss 0.221 · testLogLoss 0.228 · testAuc 0.757 · championReplicaTestLogLoss 0.228 · championReplicaTestAuc 0.757 · diffVsChampion 0 | |
| xg-strength-split-v1 | 2026-09-29 | run | Four separate fits by strength state (5v5 = strengthDiff 0 and net occupied, power play, shorthanded, empty net) with the champion's features minus the strength flags let distance, angle and shot-type effects differ by state and lower the combined held-out log loss. Expected effect: a small gain concentrated on the power play; the shorthanded and empty-net fits have few rows and may give some of it back. | trainLogLoss 0.22 · testLogLoss 0.228 · testAuc 0.758 · championReplicaTestLogLoss 0.228 · championReplicaTestAuc 0.757 · diffVsChampion 0 | |
| xg-flurry-feature-v1 | 2026-09-29 | run | Adding the time since the previous unblocked shot by the same club in the same period (capped at 30 s) and an indicator for under 3 s to the champion's 34 features lowers held-out log loss, because shots inside a flurry are taken against a goalie out of position beyond what the rebound flag already records. | trainLogLoss 0.219 · testLogLoss 0.227 · testAuc 0.761 · championReplicaTestLogLoss 0.228 · championReplicaTestAuc 0.757 · diffVsChampion -0.002 | |
| xg-combined-abc-v1 | 2026-09-29 | run | Rink-adjusted distance, strength-state fits and the flurry features together lower held-out log loss by more than any of them alone. | trainLogLoss 0.218 · testLogLoss 0.226 · testAuc 0.762 · championReplicaTestLogLoss 0.228 · championReplicaTestAuc 0.757 · diffVsChampion -0.002 | |
| arena-scorer-profiles-v1 | 2026-09-29 | run | Some NHL arenas record unblocked shots at systematically different distances than the same clubs' shots are recorded elsewhere (a scorer effect in the recorded location, visible in |x| as well as in distance), the effect persists from one season to the next, and subjective counts (hits, giveaways, takeaways, blocked shots) differ by building beyond what the clubs involved explain. Expected size in the 2021+ feed: a standard deviation across arenas near 1 ft, far below the 3-5 ft effects published for the 2007-2013 feed. | meanConsecutiveR_2021on 0.519 · meanConsecutiveR_allSeasons 0.591 · consecutiveR 2018->2019: 0.803, 2019->2020: 0.763, 2020->2021: 0.494, 2021->2022: 0.57, 2022->2023: 0.561, 2023->2024: 0.64, 2024->2025: 0.306 · shareKsPBelow0.01_2025 0.25 · placeboShareKsPBelow0.01 0.035 · sdOfArenaMeanDiff_ft 2018: 1.698, 2019: 1.67, 2020: 1.814, 2021: 1.713, 2022: 1.174, 2023: 0.801, 2024: 0.832, 2025: 0.996 | |
| score-dist-challengers-v1 | 2026-09-29 | rejected | Given the same walk-forward Kalman goal rates, a low-score dependence (Dixon-Coles), a shared component (bivariate Poisson) or per-side over-dispersion (negative binomial) improves the regulation-score log score over independent Poisson on 2025-26; NHL regulation scores are close to independent Poisson so any gain is expected to be small (under 0.01 nats per game) and the negative binomial may fit worse than Poisson. | fits dixonColes: (rho: -0.116), bivariatePoisson: (lambda3: 0), negativeBinomial: (r: 522.36, impliedVarianceInflationAtLambda3: 1.006) · test2025 poissonLeagueAverage: (scoreLogScore: 3.851, crps: 1.307, sixPlus: (predicted: 0.546, actual: 0.573), vsPoissonRange95: -0.009, 0.022), poissonIndependent: (scoreLogScore: 3.844, crps: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), vsPoissonRange95: —), dixonColes: (scoreLogScore: 3.844, crps: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), vsPoissonRange95: -0.003, 0.002), bivariatePoisson: (scoreLogScore: 3.844, crps: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), vsPoissonRange95: 0, 0), negativeBinomial: (scoreLogScore: 3.845, crps: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), vsPoissonRange95: 0, 0) · decision keep independent Poisson (no family beats it with the range excluding zero and no worse 6+ share) · evaluatedAt 2026-09-29T03:13:03Z | |
| momentum-hawkes-v1 | 2026-09-29 | run | Within NHL regulation play, goals cluster in time beyond what a Poisson process with a rate varying by game minute and by score state predicts (a "reply goal" / momentum effect); if it exists, the Hawkes branching ratio fitted on data exceeds the ratio the same fit produces on simulated null games, and short inter-goal intervals (under 60 s) are more common than in the null. | alphaData 0 · alphaBootstrap95 0, 0 · alphaNull95 0.001, 0.014 · alphaMcP 1 · under60Data 0.085 · under60Null95 0.113, 0.119 | |
| state-space-kalman-v2 | 2026-09-29 | run | A Kalman random-walk attack/defence model on regulation goals (linear-Gaussian observation, daily process noise, season-start inflation, home advantage and scoring level as slowly drifting states) predicts home wins walk-forward about as well as the champion; it is expected NOT to beat the champion on log loss because it sees only goals, but it may add information in a blend and gives a transparent rating. | tuned q: 0, sigma2: 3.5, seasonInflation: 0.02, logLoss: 0.657, games: 3491 · tunedOnGridEdge q: no, sigma2: yes, seasonInflation: no · proof 2023: (n: 1230, logLossStateSpace: 0.659, logLossChampion: 0.654, blend5050LogLoss: 0.654, diffRange95: -0.003, 0.013), 2024: (n: 1235, logLossStateSpace: 0.664, logLossChampion: 0.653, blend5050LogLoss: 0.656, diffRange95: 0.002, 0.019), 2025: (n: 1311, logLossStateSpace: 0.695, logLossChampion: 0.681, blend5050LogLoss: 0.685, diffRange95: 0.006, 0.022), pooled: (n: 3776, logLossStateSpace: 0.673, logLossChampion: 0.663, blend5050LogLoss: 0.665, diffRange95: 0.005, 0.015) · decision display rating only; remains a challenger (pooled proof does not beat the champion with the range excluding zero) · evaluatedAt 2026-09-29T03:06:56Z | |
| state-space-kalman-v1 | 2026-09-29 | run | A Kalman random-walk attack/defence model on regulation goals (linear-Gaussian observation, daily process noise, season-start inflation, home advantage and scoring level as slowly drifting states) predicts home wins walk-forward about as well as the champion; it is expected NOT to beat the champion on log loss because it sees only goals, but it may add information in a blend and gives a transparent rating. | tuned q: 0.001, sigma2: 4.5, seasonInflation: 0.02, logLoss: 0.657, games: 3491 · tunedOnGridEdge q: yes, sigma2: yes, seasonInflation: yes · proof 2023: (n: 1230, logLossStateSpace: 0.66, logLossChampion: 0.654, blend5050LogLoss: 0.655, diffRange95: -0.002, 0.014), 2024: (n: 1235, logLossStateSpace: 0.664, logLossChampion: 0.653, blend5050LogLoss: 0.656, diffRange95: 0.003, 0.019), 2025: (n: 1311, logLossStateSpace: 0.695, logLossChampion: 0.681, blend5050LogLoss: 0.685, diffRange95: 0.006, 0.022), pooled: (n: 3776, logLossStateSpace: 0.673, logLossChampion: 0.663, blend5050LogLoss: 0.666, diffRange95: 0.006, 0.015) · decision display rating only; remains a challenger (pooled proof does not beat the champion with the range excluding zero) · evaluatedAt 2026-09-29T03:06:09Z | |
| isotonic-calibrator-v1 | 2026-09-29 | rejected | The champion home-win probabilities carry a miscalibration that a monotone (isotonic) map fitted on earlier seasons corrects out of sample; because the walk-forward reliability is already close to the diagonal the expected gain is small (under 0.003 log loss) and may be zero or negative. | fit2020to2024 logLossBefore: 0.681, logLossAfter: 0.683, diffRange95: 0, 0.005 · fit2023to2024 logLossBefore: 0.681, logLossAfter: 0.683, diffRange95: -0.001, 0.006 · decision rejected · evaluatedAt 2026-09-29T02:49:22Z | |
| hist-2026-09-26-game-model-v2 | 2026-09-26 | adopted | Even-strength xG Elo + goals Elo + PP/PK EWMA + rest and back-to-back lowers held-out log loss versus ratings only. | ratingsOnly 0.675 · plusRest 0.672 · plusSpecialTeams no gain · note 60-70% bucket ~4 points over-confident | |
| hist-2026-09-26-starting-goalie-feature | 2026-09-26 | rejected | Adding the starting goalie's GSAx as a pre-game feature lowers log loss. | before 0.678 · after 0.679 · note worse; dropped | |
| hist-2026-09-26-tournament | 2026-09-26 | adopted | Boosted trees or added feature families beat a ridge logistic on ratings + rest. | winner logistic, everything, ridge 50 · avgLogLoss 0.651 · accuracy 0.614 · ratingsPlusRest 0.654 · trees worse on every fold · note numbers before shootout results were added; later comparable figures ~0.656-0.667 | |
| hist-2026-09-26-search-3000 | 2026-09-26 | rejected | One of 3,000 feature-set versions beats the champion on a proof season it was not selected on. | top10SelectionLL 0.658 · proofLL 0.682-0.687 vs champion 0.6806 · luckGap 0.027 · note every finalist lost on proof; champion holds | |
| hist-2026-09-26-block-search-49k | 2026-09-26 | rejected | Some combination of 14 input families (special teams, finishing+goalie, lineup, rivalry, travel, goalie rotation, schedule spot, GM context, discipline+referees, puck play, grudge scrums, league scoring) beats the champion. | versions 49152 · championProof 0.664 · bestFinalist 0.664 · range -0.001, 0 · note not better; model saturated on public box data. Top-200 family shares: special teams 100%, lineup 84%, finishing+goalie 62% | |
| hist-2026-09-26-travel | 2026-09-26 | rejected | Travel distance, time zones crossed, three-in-four and road-trip length carry information beyond rest and back-to-backs. | all 0.656 · minusTravel 0.656 · note no gain | |
| hist-2026-09-26-goalie-rotation-trap | 2026-09-26 | rejected | Goalie rotation (backup/tired) and schedule-spot (trap game) features lower log loss. | all 0.657 · minusTrap 0.657 · minusRotation 0.657 · note neither adopted | |
| hist-2026-09-26-score-engines | 2026-09-26 | rejected | Monte Carlo minute-level or Markov score engines beat the closed-form Poisson on score and finish log loss. | win indistinguishable · score indistinguishable (3.648 vs 3.650) · finish closed form clearly best; MC and Markov under-predict ties · note late-tied caution measured: last 5 minutes tied x0.77 | |
| hist-2026-09-26-lineups | 2026-09-26 | adopted | Walk-forward skater relative-xG ratings plus starter GSAx and missing-vs-usual lineup lower log loss. | before 0.678 · after 0.677 · note small gain; confident (>=70%) games ~78% right on ~70 of 1,301 | |
| hist-2026-09-26-props | 2026-09-26 | adopted | A learned logistic layer over hand features (shot rate, finishing, opponent shots allowed, venue, ice time, individual xG rate, opposing starter form) beats the base rate for anytime-goal and anytime-point. | goal logistic: 0.381, trees: 0.381, baseline: 0.386 · point blend: 0.59, baselinePtsPerGame: 0.597 · note stars (>=40%) ~3.5 points over-confident; training on 2019-24 was WORSE for points than 2022-24 (scoring drift) | |
| hist-2026-09-26-confidence | 2026-09-26 | rejected | Bootstrap spread across 100 refits marks the games the model gets right more often. | medianSpread 0.024 · note loose estimates were right MORE often; shrink-by-spread gave no gain. Confidence = lean tier, calibrated within ~2 points | |
| hist-2026-09-22-xg-v1 | 2026-09-22 | adopted | A logistic expected-goals model on distance, angle, shot type, strength, rebound, rush and empty-net flags beats the league-average goal rate on an unseen season. | logLoss 0.228 · baseline 0.259 · auc 0.757 · note top decile over-predicted by ~15%; bottom decile under-predicted |
calibration
Run by the model desk; results are on the relevant card.
leakage
Run by the model desk; results are on the relevant card.
comparisons
Run by the model desk; results are on the relevant card.
drift
Run by the model desk; results are on the relevant card.
stateSpace
Run by the model desk; results are on the relevant card.
scoreDist
Run by the model desk; results are on the relevant card.
momentum
Run by the model desk; results are on the relevant card.
Reading a card
Training data names seasons, not just counts. Evaluation is always on seasons the model did not see. Calibration is the share that actually won inside each probability bucket. Leakage checks include a deliberately leaked canary the detector must catch. Challengers are pre-registered alternatives with their proof results, adopted only through the hub.