Model cards · built 2026-09-29

Every model, with its record.

Intended use, training data, held-out results, calibration, leakage checks, drift, challengers and known limits, for each model that prices a game or values a shot. Studies that have not run are marked absent, never blank.

Models
5
Register entries
38
Studies present
7 of 7
Adopted / rejected
10 / 17
logistic regression (ridge, fitted by Adam)

Expected goals

Value of an unblocked shot from its location, type and context; the base of Drive, Setup, Finish, GSAx, RAPM and the Elo family. Shooter-agnostic by design so Finish means something.

Usersevery measure page; the game model's xG Elo
Ownermodel desk
Version2026-09-29
Public ceilingabout 0.78 AUC from location and type alone in the public literature
Training data

What it learned from

seasons
2023-24, 2024-25
shots
227076
features
34
Evaluation

How it did on seasons it had not seen

testedOn
2025-26
logLoss
0.229
baselineLogLoss
0.259
auc
0.756
clockRepair
Repaired 2026-09-29 (hub card 149362): from 2023-24 the league stamps a goal about a second later than other shots, so the rebound (3 s) and rush (4 s) windows
logLossBeforeRepair
0.228
aucBeforeRepair
0.757
logLossAfterRepair
0.228
aucAfterRepair
0.756
deciles
(expected: 69.1, actual: 75); (expected: 143.4, actual: 126); (expected: 243.5, actual: 236); (expected: 373.4, actual: 332); (expected: 535.7, actual: 485); (expected: 718.8, actual: 724); (expected: 936.3, actual: 942); (expected: 1204.5, actual: 1206); (expected: 1576.7, actual: 1604); (expected: 2734.9, actual: 2356)
seasonShift
-0.064
All 4 columns shown
Columns
DecileExpected goalsActual goalsActual / expected
169.1751.085
2143.41260.879
3243.52360.969
4373.43320.889
5535.74850.905
6718.87241.007
7936.39421.006
81204.512061.001
91576.716041.017
102734.923560.861
Calibration

Do 60% calls win 60% of the time?

available
true
shots
108568
games
1271
totalXg
8273.9
totalGoals
7823
ratioGoalsOverXg
0.946
ece
0.005
logLossPerShot
0.228
brierPerShot
0.062
deciles
(decile: 1, xgLo: 0, xgHi: 0.009, shots: 10736, expectedGoals: 64.2, actualGoals: 71, meanXg: 0.006, goalRate: 0.007, ratioActualOverExpected: 1.106, z: 0.85); (decile: 2, xgLo: 0.009, xgHi: 0.017, shots: 10883, expectedGoals: 137, actualGoals: 122, meanXg: 0.013, goalRate: 0.011, ratioActualOverExpected: 0.891, z: -1.29); (decile: 3, xgLo: 0.017, xgHi: 0.027, shots: 10905, expectedGoals: 236.5, actualGoals: 227, meanXg: 0.022, goalRate: 0.021, ratioActualOverExpected: 0.96, z: -0.62); (decile: 4, xgLo: 0.027, xgHi: 0.04, shots: 10884, expectedGoals: 362.8, actualGoals: 315, meanXg: 0.033, goalRate: 0.029, ratioActualOverExpected: 0.868, z: -2.55); (decile: 5, xgLo: 0.04, xgHi: 0.055, shots: 10836, expectedGoals: 517.5, actualGoals: 484, meanXg: 0.048, goalRate: 0.045, ratioActualOverExpected: 0.935, z: -1.51); (decile: 6, xgLo: 0.056, xgHi: 0.073, shots: 10861, expectedGoals: 695.2, actualGoals: 709, meanXg: 0.064, goalRate: 0.065, ratioActualOverExpected: 1.02, z: 0.54); (decile: 7, xgLo: 0.073, xgHi: 0.094, shots: 10888, expectedGoals: 907.7, actualGoals: 876, meanXg: 0.083, goalRate: 0.081, ratioActualOverExpected: 0.965, z: -1.1); (decile: 8, xgLo: 0.095, xgHi: 0.121, shots: 10856, expectedGoals: 1163.6, actualGoals: 1158, meanXg: 0.107, goalRate: 0.107, ratioActualOverExpected: 0.995, z: -0.17); (decile: 9, xgLo: 0.121, xgHi: 0.164, shots: 10845, expectedGoals: 1523.7, actualGoals: 1547, meanXg: 0.141, goalRate: 0.143, ratioActualOverExpected: 1.015, z: 0.64); (decile: 10, xgLo: 0.164, xgHi: 0.908, shots: 10874, expectedGoals: 2665.6, actualGoals: 2314, meanXg: 0.245, goalRate: 0.213, ratioActualOverExpected: 0.868, z: -8.19)
crossFitIsotonic
split: games permuted with seed 11, halves of 635/636 games, logLossBefore: 0.228, logLossAfter: 0.228, brierBefore: 0.062, brierAfter: 0.062, eceAfter: 0.001, diffAfterMinusBefore: (mean: 0, range: 0, 0, excludesZero: no, B: 5000, clusters: games), mapHalfAAtGrid: (0.01: 0.007, 0.02: 0.012, 0.05: 0.046, 0.1: 0.101, 0.15: 0.154, 0.2: 0.172, 0.3: 0.213, 0.4: 0.369, 0.5: 0.369), mapHalfBAtGrid: (0.01: 0.01, 0.02: 0.021, 0.05: 0.042, 0.1: 0.109, 0.15: 0.154, 0.2: 0.173, 0.3: 0.218, 0.4: 0.348, 0.5: 0.514), plattSecondary: (note: cross-fitted logit(q) = a + b*logit(xg), same game halves; secondary, not the registered isotonic test, halfA: (a: -0.21, b: 0.933), halfB: (a: -0.189, b: 0.942), logLossAfter: 0.228, diffAfterMinusBefore: (mean: 0, range: 2 items, excludesZero: yes, B: 5000, clusters: games))
Challengers

Pre-registered challengers, judged on an unseen season

Seven pre-registered challengers to the expected-goals model have been judged on the unseen 2025-26 season; three passed the registered rule. The desk withholds all three and puts none forward, so the champion stays.

Trained on 2023-24 and 2024-25, judged on 2025-26. Champion log loss 0.22830 on 1,312 games (0.22861 on the 1,232 with an events file). Rule: adopt only if the bootstrap 95% range of (challenger - champion) log loss lies below zero; otherwise reject; adoption itself is a hub decision card, never automatic.

4 of 5 columns shown · open a row for the other 1
Columns
ChallengerResultLog loss against champion95% rangeDetails

The desk: no candidate passes the registered rule, is attributable and shows no placebo gain, so the desk puts none forward. B passes the rule but its flags are built on the raw clock. The clock finding goes to whoever owns tools/pull-pbp.mjs.

Full study on /integrity.

Known limitation: event clock

Goals are stamped later than other shots

  • Goals are stamped about one second later than other shots in 2023-24, 2024-25 and 2025-26; the model's rebound and rush inputs are short windows on that clock.
  • With goal times corrected the champion scores 0.22883 against 0.22861 on the same 105,288 shots (+0.000225 per shot, 95% range +0.000107 to +0.000351): part of its published accuracy comes from the clock.
  • 70% of shots flagged as rush in 2025-26 follow the club's own blocked attempt and are not rushes.
All 5 columns shown
Columns
Placebo inputClockChange in log loss95% rangeLowers log loss
four eventsraw−0.000661−0.000881 to −0.000466yes
four eventsgoal times moved back 1 s+0.000008−0.000002 to +0.000016no
hits onlyraw−0.000028−0.000143 to +0.000097no
hits onlygoal times moved back 1 s+0.000010−0.000007 to +0.000029no

The clock finding in full.

Arena scorer effects

Where shots are recorded differently, 2025-26

  • 8 of 32 buildings recorded shot distance differently from the same clubs' road shots in 2025-26 (25% at p below 0.01, against 3.5% by chance).
  • Season to season the effects correlate at r 0.519 (registered threshold 0.5); the spread between buildings is 1.00 ft against a noise floor of about 0.56 ft.
  • A distance correction repaired 22 arena-seasons on unseen games and broke 20 (exact p 0.88), so the model carries no rink adjustment.
3 of 4 columns shown · open a row for the other 1
Columns
BuildingRecorded distance against road shots (ft)About (95%)Details

Every building on /integrity.

Known limits

Where it stops

  • Top decile of chances over-predicted by about 15% and the bottom decile under-predicted (decile table); an isotonic recalibration is on the register.
  • No rink (scorer) adjustment of shot locations: the correction tried did not help on unseen games (see Arena scorer effects); goalies who play half their games in one building inherit its recording habits.
  • Repaired 2026-09-29 (hub card 149362): from 2023-24 the league stamps a goal about a second later than other shots, so the rebound (3 s) and rush (4 s) windows missed goals and the model learned to mark those chances down. Goal times are now moved back 1 s before the windows are applied (2022-23 and earlier unchanged): 592 rebound and 258 rush flags gained, all of them goals. Published accuracy fell, log loss 0.2283 to 0.2285 and AUC 0.757 to 0.756 on 112,091 2025-26 shots, because a leak of the outcome into the inputs was removed, not because the model got worse at hockey.
  • Clock repair not complete in the test season: 80 of 1,312 2025-26 games (6,803 shots, 485 goals) have no events file on disk, so their flags stay on the raw clock; test season only, training rows fully corrected.
  • One model with power-play and shorthanded flags rather than four strength-state models.
  • Goalie quality leaks into the label: a shot on a weak goalie is a goal more often, so xG absorbs some goaltending.
present

calibration

Run by the model desk; results are on the relevant card.

present

leakage

Run by the model desk; results are on the relevant card.

present

comparisons

Run by the model desk; results are on the relevant card.

present

drift

Run by the model desk; results are on the relevant card.

present

stateSpace

Run by the model desk; results are on the relevant card.

present

scoreDist

Run by the model desk; results are on the relevant card.

present

momentum

Run by the model desk; results are on the relevant card.

Reading a card

Training data names seasons, not just counts. Evaluation is always on seasons the model did not see. Calibration is the share that actually won inside each probability bucket. Leakage checks include a deliberately leaked canary the detector must catch. Challengers are pre-registered alternatives with their proof results, adopted only through the hub.