Every number on this site was checked this morning.
Gate open at 2026-09-29 15:11 UTC: every record check, feed contract and replay passed, with 3 warnings. This page is the test record behind every number: what was checked, against what, with what result, and who decides what happens next.
- Gate
- open
- Jobs green
- 15 of 15
- Games complete
- 0 of 0
- Decisions waiting
- 7
Gate open
Gate open at 2026-09-29 15:11 UTC: every record check, feed contract and replay passed, with 3 warnings.
Warnings (do not close the gate):
- range-soft: shots.type 30/112091, contracts.entryLevelLabel 8/245, contracts.signingDate 10/894, contracts.termVsSeasons 9/927
- control-alert: 2025-26.shotsPerGame 2026-04-08 down
- anomaly-defects: 120 older games carry record defects (see /integrity)
The deploy job reads this gate before it stages a build. Closed means yesterday's build stays up. A stale gate (over 30 h) is a warning in the deploy log, not a refusal, so an assay that failed to run cannot silently stop the approved daily deploy.
Jobs the site depends on
Green means the job produced its artefact on time, never merely that it returned. Assay jobs write a heartbeat; estate jobs are judged by the file they leave behind.
| Job | State | Last produced | Cadence | Details |
|---|---|---|---|---|
| Morning assay ·critical | green | 6.9 h ago | 24 h | |
| Game passports ·critical | green | 7.0 h ago | 24 h | |
| Data contracts ·critical | green | 7.0 h ago | 24 h | |
| Control charts | green | 7.0 h ago | 24 h | |
| Golden replay | green | 1.9 h ago | 24 h | |
| Mutation tests | green | 9.6 h ago | 24 h | |
| Hub controls (pull) | green | 7.0 h ago | 24 h | |
| Hub cards + notification (push) | green | 6.9 h ago | 24 h | |
| Model spend ledger | green | 6.9 h ago | 24 h | |
| Nightly model state (refresh-predictions) ·critical | green | 22.1 h ago | 24 h | |
| Ledger: morning predictions recorded ·critical | green | 1.4 h ago | 24 h | |
| Licensed feeds pulled ·critical | green | 34.1 h ago | 24 h | |
| Injury listings | green | 22.1 h ago | 24 h | |
| White noise | green | 22.1 h ago | 24 h | |
| Daily data deploy ·critical | green | 1.6 h ago | 24 h |
Production builds this job knows about
Every regular-season game due, every season on disk
A game is complete when the record artefacts are present and every reconciliation check passes. Coverage layers (clips, logs for current players) are reported but never make a game incomplete.
| Season | Games due | Complete | Charted minutes | Details |
|---|---|---|---|---|
| 2026-27 | 0 | 0 (—) | — | |
| 2025-26 | 1312 | 1190 (91%) | 100% | |
| 2024-25 | 1312 | 1220 (93%) | 95.7% | |
| 2023-24 | 1312 | 1209 (92%) | 100% | |
| 2022-23 | 1312 | 1272 (97%) | 100% | |
| 2021-22 | 1312 | 1267 (97%) | 100.02% | |
| 2020-21 | 868 | 820 (94%) | 100% | |
| 2019-20 | 1082 | 1041 (96%) | 100% | |
| 2018-19 | 1271 | 1249 (98%) | 100% | |
| 2017-18 | 1271 | 1262 (99%) | 100% |
- goalsAgree: Goals agree across the shot file, the goal record and the events
- shiftSums: Shift charts sum to a real number of players and never double up
- shooterOnIce: Every shooter and goalie was on the chart at the shot's second
- teamAgree: The shooter's club on the chart is the shooting club in the shot file
- assistsOnIce: Every assist went to a player on the ice
- emptyNetGoalie: The empty-net flag agrees with the chart
- timeBounds: Every time sits inside its period
Artefacts this season
0 of 0 games complete; 0 incomplete. Built 2026-09-29.
| Artefact | Present | Partial | Missing | Details |
|---|---|---|---|---|
| Schedule row | 0 | 0 | 0 | |
| Shots | 0 | 0 | 0 | |
| Shift charts (both clubs) | 0 | 0 | 0 | |
| Play-by-play events | 0 | 0 | 0 | |
| Goal record | 0 | 0 | 0 | |
| Officials | 0 | 0 | 0 | |
| Game logs (current players) | 0 | 0 | 0 | |
| Dressed rosters | 0 | 0 | 0 | |
| Expected-goals sidecar | 0 | 0 | 0 | |
| Goal clips | 0 | 0 | 0 | |
| Condensed game | 0 | 0 | 0 | |
| Recap | 0 | 0 | 0 |
Every game due this season is complete.
Source: Game passports · as of 2026-09-29
What each feed last delivered
Every number on the site names its feed and its as-of day; a stale feed is shown stale, in words, never as current.
| Feed | State | As of | Age (days) | Details |
|---|---|---|---|---|
| Roster transactions | fresh | 2026-09-28 | 1.4 | |
| Injury listings | fresh | 2026-09-29 | 0.9 | |
| Projected lineups | fresh | 2026-09-28 | 1.9 | |
| Starting goalies | not-connected | — | — | |
| Market odds | not-connected | — | — | |
| Attention (stories, Wikipedia, social) | fresh | 2026-09-29 | 0.9 | |
| Model state | fresh | 2026-09-29 | 0.9 |
Shapes and ranges, checked this morning
The JSON shape of every feed's newest file is compared with yesterday's; ranges and vocabularies of shots, shifts, events, goals, schedules and contract ledgers are checked as typed units.
- shots.type: 30 of 112091 rows; example {"game":"2025020184","type":"unknown"}
- contracts.entryLevelLabel: 8 of 245 rows; example {"club":"boston-bruins","player":"JJ Peterka","termYears":5,"capHit":7700000,"why":"an entry-level contract runs three years at most"}
- contracts.signingDate: 10 of 894 rows; example {"club":"carolina-hurricanes","player":"Andrei Svechnikov","signedOn":"2021-08-26","firstSeason":"2024–25","why":"signed 3 summers before its first season"}
- contracts.termVsSeasons: 9 of 927 rows; example {"club":"boston-bruins","player":"Fraser Minten","termYears":3,"seasons":"2022–23 to 2026–27","span":5}
| Check | Violations | Rows |
|---|---|---|
| shots.dist | 0 | 112091 |
| shots.angle | 0 | 112091 |
| shots.type | 30 | 112091 |
| shots.strength | 0 | 112091 |
| shots.t | 0 | 112091 |
| shots.period | 0 | 112091 |
| shots.flags | 0 | 112091 |
| shots.ids | 0 | 112091 |
| shifts.bounds | 0 | 984474 |
| shifts.ids | 0 | 984474 |
| events.type | 0 | 388678 |
| events.x | 0 | 388678 |
| events.y | 0 | 388678 |
| events.time | 0 | 388678 |
| goals.time | 0 | 8129 |
| schedule.row | 0 | 1498 |
Contract ledgers: contracts.capHit 0 of 1068; contracts.term 0 of 1068; contracts.tier 181 of 1068.
Control charts on every daily count
An exponentially weighted average and a two-sided CUSUM per series, limits learned from history; a series with fewer than 14 readings is "learning", never "fine". Changepoints are dated by binary segmentation.
| Series | State | Readings | Mean | Last | Last date | Details |
|---|---|---|---|---|---|---|
| 2025-26.shotsPerGame | alert | 167 | 85.27 | 83.67 | 2026-04-16 | |
| 2025-26.eventsPerGame | in-control | 167 | 291.95 | 297.7 | 2026-04-16 | |
| 2025-26.shiftRowsPerGame | in-control | 167 | 752.01 | 736 | 2026-04-16 | |
| 2025-26.goalsPerGame | in-control | 167 | 6.14 | 6.17 | 2026-04-16 | |
| 2025-26.shotCoverage | in-control | 167 | 1 | 1 | 2026-04-16 | |
| 2025-26.shiftCoverage | in-control | 167 | 1 | 1 | 2026-04-16 | |
| feed.transactionsPerDay | learning | 10 | — | 76 | 2026-09-28 | |
| feed.injuryListings | learning | 2 | — | 55 | 2026-09-29 | |
| feed.socialPosts7d | learning | 1 | — | 21130 | 2026-09-29 | |
| feed.leagueStories7d | learning | 2 | — | 126 | 2026-09-29 | |
| feed.wikiPageviews7d | learning | 2 | — | 124338 | 2026-09-29 | |
| feed.ledgerPredictionsPerDay | learning | 2 | — | 5 | 2026-09-29 |
- 2025-26.shotsPerGame 2026-04-08: reading down (71, z -3.2)
Golden replay
A frozen 40-game slice is rebuilt twice every morning; the two outputs must hash equal and match the recorded baseline. A builder whose output moves without a code change has a hidden input.
Mutation tests
Known corruptions are applied to the golden slice and must be caught by a check. A mutation that survives is a hole in the checks, and the gate closes until a check exists.
| Mutation | Caught | Caught by | Details |
|---|---|---|---|
| shot-time-base-off-by-a-period | ok | shooterOnIce, timeBounds | |
| club-ids-swapped-in-shift-chart | ok | teamAgree | |
| second-period-missing-from-chart | ok | shooterOnIce, emptyNetGoalie | |
| shots-duplicated | ok | goalsAgree | |
| goal-record-times-shifted | ok | assistsOnIce | |
| shift-chart-deleted | ok | artefact:shifts, artefact:rosters | |
| goalie-vanishes-from-chart | ok | shooterOnIce, emptyNetGoalie | |
| goal-events-dropped | ok | goalsAgree | |
| shot-distances-impossible | ok | range:shots.dist | |
| shooter-ids-zeroed | ok | shooterOnIce, range:shots.ids |
Content-addressed snapshot of the estate
Every file the site is built from, hashed each morning; small files that changed are kept by their hash, so what was held on any day can be produced byte for byte. A finished season's file changing is reported, because history is supposed to be immutable.
Which builder wrote each file, from what
For every data file the site binds: the builder, its inputs, the provenance tier, the build time and hash. A build is stale when an input's hash changed after it, or when games present in the inputs are missing from it.
No stale builds.
| Site file | State | Tier | Builder | Built | Details |
|---|---|---|---|---|---|
| xg-2025-26.json | fresh | derived | tools/train-xg.mjs | 2026-09-29 | |
| impact.json | fresh | derived | tools/build-impact.mjs | 2026-09-29 | |
| deployment.json | fresh | derived | tools/build-deployment.mjs | 2026-09-29 | |
| advanced.json | fresh | derived | tools/build-advanced.mjs | 2026-09-29 | |
| matchups.json | fresh | derived | tools/build-matchups.mjs | 2026-09-29 | |
| shotmaps.json | fresh | derived | tools/build-shotmaps.mjs | 2026-09-29 | |
| coaching.json | fresh | derived | tools/build-coaching.mjs | 2026-09-29 | |
| involvement.json | fresh | derived | tools/build-involvement.mjs | 2026-09-29 | |
| durability.json | fresh | derived | tools/build-durability.mjs | 2026-09-23 | |
| officials.json | unverified | derived | tools/build-officials.mjs | 2026-09-26 | |
| goal-archive.json | fresh | measured | tools/build-goals.mjs | 2026-09-29 | |
| video-reads.json | fresh | judgement | tools/build-video-reads.mjs | 2026-09-26 | |
| game-model.json | fresh | derived | tools/build-game-model.mjs | 2026-09-29 | |
| champion-state.json | fresh | derived | tools/build-game-sim.mjs (EXPORT=1) | 2026-09-29 | |
| champion.json | unverified | derived | tools/export_champion.py | 2026-09-26 | |
| player-ratings.json | fresh | derived | tools/build-lineups.mjs | 2026-09-29 | |
| player-props.json | unverified | derived | tools/player_props.py | 2026-09-26 | |
| ledger-public.json | fresh | derived | tools/build-ledger-page.mjs | 2026-09-29 | |
| schedule-strength.json | fresh | derived | tools/build-schedule-strength.mjs | 2026-09-29 | |
| white-noise.json | fresh | derived | tools/build-white-noise.mjs | 2026-09-29 | |
| transactions.json | fresh | licensed | tools/build-transactions.mjs | 2026-09-29 | |
| injuries.json | fresh | licensed | tools/build-injuries.mjs | 2026-09-29 | |
| projected-lines.json | fresh | licensed | tools/build-projected-lines.mjs | 2026-09-29 | |
| contracts | fresh | researched | tools/import-contracts.mjs | 2026-09-29 | |
| prospect-rankings.json | fresh | derived | tools/build-prospect-rankings.mjs | 2026-09-23 | |
| league-ladder.json | fresh | derived | tools/build-league-ladder.mjs | 2026-09-23 | |
| abroad.json | fresh | derived | tools/build-abroad.mjs | 2026-09-24 | |
| draft-class.json | unverified | derived | tools/build-draft-class.mjs | 2026-09-24 | |
| measures-2025.json | fresh | derived | tools/assay/build-measure-data.mjs | 2026-09-29 | |
| stat-cards.json | fresh | derived | tools/assay/build-stat-cards.mjs | 2026-09-29 | |
| model-cards.json | fresh | derived | tools/assay/build-model-cards.mjs | 2026-09-29 | |
| integrity.json | fresh | derived | tools/assay/build-integrity.mjs | 2026-09-29 |
One player, many identifiers
Every crosswalk (Wikipedia, Sportradar, other leagues) carries the rule that made it and a match probability; grey-zone pairs wait in a queue for a person.
- at
- —
- generatedAt
- 2026-09-29T15:08:07.322Z
- spine
- players: 2304, byStatus: (prospect-pool: 1456, roster-2026-09-21: 1286, regular-2025-26: 776, former-regular: 2), withBorn: 2304, withPosition: 2304, bornConflicts: 0, bornConflictSamples: none
- sourceCounts
- rosterSnapshot: 1286, rosterFiles: 1286, leagueRegulars: 776, formerRegulars: 2, prospectPools: 1489, rosterHistory: 1286, prospectHistory: 1456, wikipediaTitles: 1090, sportradar: 1278
- crosswalks
- wikipedia: (linked: 1090, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 27, 0.92-0.99: 0, 0.99-1: 1063)), sportradar: (linked: 1278, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 0, 0.99-1: 1278)), ELH: (linked: 21, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 11, 0.99-1: 10)), KHL: (linked: 104, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 0, 0.99-1: 104)), NL: (linked: 13, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 13, 0.99-1: 0)), SHL: (linked: 81, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 0, 0.99-1: 81)), Liiga: (linked: 51, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 0, 0.99-1: 51)), DEL: (linked: 6, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 5, 0.99-1: 1)), WHL: (linked: 154, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 1, 0.99-1: 153)), OHL: (linked: 195, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 0, 0.99-1: 195)), USHL: (linked: 136, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 1, 0.99-1: 135)), QMJHL: (linked: 79, histogram: (0-0.5: 0, 0.5-0.75: 0, 0.75-0.92: 0, 0.92-0.99: 1, 0.99-1: 78)) … +2 fields
- wikipedia
- cachedTitles: 1090, cachedNulls: 196, keysOutsideSpine: 0, ruleViolations: 0, expected: 0, violators: none, rule: title must contain the surname AND start with the first two letters of the given name (build-white-noise.mjs playerTitleOk)
- sportradar
- players: 1427, withReference: 1425, referenceInSpine: 1278, nameAgree: 1278, referenceOnly: 0, referenceOutsideSpine: 147, samples: (referenceOnly: none)
- otherLeagues
- files: (DEL: (5 fields); (5 fields), ELH: (5 fields); (5 fields), ELJ: (5 fields); (5 fields), KHL: (5 fields); (5 fields), Liiga: (5 fields); (5 fields), NL: (5 fields); (5 fields), OHL: (5 fields); (5 fields); (5 fields), QMJHL: (5 fields); (5 fields); (5 fields), SHL: (5 fields); (5 fields), U20SM: (5 fields); (5 fields), USHL: (5 fields); (5 fields); (5 fields), WHL: (5 fields); (5 fields); (5 fields)), byLeague: (DEL: (entities: 504, withBirthdate: 0, withAgeOnly: 331, noBirthInfo: 173, candidatesScored: 1228, linked: 6, queued: 0, conflicts: 0), ELH: (entities: 475, withBirthdate: 294, withAgeOnly: 0, noBirthInfo: 181, candidatesScored: 1188, linked: 21, queued: 1, conflicts: 1), ELJ: (entities: 538, withBirthdate: 538, withAgeOnly: 0, noBirthInfo: 0, candidatesScored: 286, linked: 3, queued: 1, conflicts: 0), KHL: (entities: 1001, withBirthdate: 1001, withAgeOnly: 0, noBirthInfo: 0, candidatesScored: 541, linked: 104, queued: 1, conflicts: 0), Liiga: (entities: 662, withBirthdate: 662, withAgeOnly: 0, noBirthInfo: 0, candidatesScored: 170, linked: 51, queued: 1, conflicts: 0), NL: (entities: 512, withBirthdate: 0, withAgeOnly: 0, noBirthInfo: 512, candidatesScored: 2032, linked: 17, queued: 0, conflicts: 0), OHL: (entities: 998, withBirthdate: 998, withAgeOnly: 0, noBirthInfo: 0, candidatesScored: 790, linked: 195, queued: 0, conflicts: 0), QMJHL: (entities: 967, withBirthdate: 967, withAgeOnly: 0, noBirthInfo: 0, candidatesScored: 593, linked: 79, queued: 0, conflicts: 0), SHL: (entities: 530, withBirthdate: 530, withAgeOnly: 0, noBirthInfo: 0, candidatesScored: 220, linked: 81, queued: 0, conflicts: 0), U20SM: (entities: 792, withBirthdate: 792, withAgeOnly: 0, noBirthInfo: 0, candidatesScored: 143, linked: 28, queued: 0, conflicts: 0), USHL: (entities: 1123, withBirthdate: 1116, withAgeOnly: 0, noBirthInfo: 7, candidatesScored: 836, linked: 136, queued: 2, conflicts: 0), WHL: (entities: 1110, withBirthdate: 1110, withAgeOnly: 0, noBirthInfo: 0, candidatesScored: 696, linked: 154, queued: 0, conflicts: 0)), entities: 9212, pairsAtOrAboveQueue: 881, linkedEntities: 875
- queue
- size: 6, byLeague: (ELH: 1, KHL: 1, Liiga: 1, USHL: 2, ELJ: 1), first20: (league: ELH, entityId: 10097211, entityName: Tomáš Galvas, entityBorn: —, entityAge: —, entityTeams: 1 item, candidate: Tomas Galvas, candidateNhlId: 8470500, candidateBorn: 1978-05-30, candidateClub: PIT, score: 0.923, evidence: 6 fields … +1 fields); (league: KHL, entityId: 44358, entityName: Fyodorov Viktor S., entityBorn: 2008-02-21, entityAge: 18, entityTeams: 1 item, candidate: Viktor Fedorov, candidateNhlId: 8486078, candidateBorn: 2008-02-21, candidateClub: SEA, score: 0.91, evidence: 6 fields … +1 fields); (league: Liiga, entityId: 40432390, entityName: Eelis Uronen, entityBorn: 2008-06-24, entityAge: 18, entityTeams: 1 item, candidate: Anttoni Uronen, candidateNhlId: 8486310, candidateBorn: 2008-06-24, candidateClub: CBJ, score: 0.899, evidence: 6 fields … +1 fields); (league: USHL, entityId: 9982, entityName: Xavier Veilleux, entityBorn: 2006-05-23, entityAge: 20, entityTeams: 1 item, candidate: Xavier Veilleux, candidateNhlId: 8485078, candidateBorn: 2006-03-23, candidateClub: NYI, score: 0.814, evidence: 6 fields … +1 fields); (league: USHL, entityId: 9902, entityName: Austin Baker, entityBorn: 2006-02-11, entityAge: 20, entityTeams: 1 item, candidate: Austin Baker, candidateNhlId: 8485098, candidateBorn: 2006-02-12, candidateClub: DET, score: 0.814, evidence: 6 fields … +1 fields); (league: ELJ, entityId: 10100131, entityName: Adam Trachta, entityBorn: 2007-04-10, entityAge: 19, entityTeams: 1 item, candidate: Adam Benak, candidateNhlId: 8485367, candidateBorn: 2007-04-10, candidateClub: MIN, score: 0.807, evidence: 6 fields … +1 fields)
- teams
- count: 32, withNhlId: 32, withSportradar: 32, utah: (ids: (4 fields); (4 fields), predecessors: (5 fields), notes: Schedule files on disk: id 59 in 2024–25; id 68 in 2025–26, 2026–27., The build brief said id 68 is valid from 2026–27; the data on disk shows 68 already in the 2025–26 schedule (preseason games included) and 59 only in 2024–25. Recorded what the data shows; both ids resolve to UTA with their season ranges.)
- params
- prior: -2.944, name: (strong: 3.379, medium: 0.251, weak: -4.554, strongAt: 0.97, mediumAt: 0.92), surnameRule: a pair whose folded surnames share no token can reach the queue but never a link (cap 0.91); twins and same-birthday club-mates are the reason, linkCapWithoutSurname: 0.91, birthdate: (exact: 5.779, sameYearDifferentDay: -3.504, yearFromAge: 0.642, unknown: 0), position: (agree: 0.663, disagree: -2.813, unknown: 0), shoots: (agree: 0.642, disagree: -2.303, unknown: 0), history: (leagueSameSeason: 2.14, clubBonus: 1.099, leagueOtherSeason: 1.386, leagueAbsent: -0.944, noHistory: 0), linkAt: 0.92, queueAt: 0.75
Source verifier
Cited pages behind cap hits and clauses are refetched on a rolling sample; a fact whose only source is gone or no longer shows the amount is listed here.
- at
- —
- summary
- ranAt: 2026-09-29T15:08:07.772Z, dry: no, params: (limit: 40, gapMs: 2100, userAgent: CapOrCup source verifier (contact [email protected]), allowedDomains: nhl.com, tsn.ca, sportsnet.ca, espn.com, nbcsports.com, cbssports.com, apnews.com, thehockeynews.com, thescore.com, yahoo.com, dailyfaceoff.com, dailyfaceoff: news articles only, banned: /an excluded site|an excluded site|an excluded site|an excluded site|an excluded site/i, robotsTtlDays: 7), pool: (urls: 1395, allowed: 1232, banned: 33, other: 130, invalid: 0, facts: 2170, byHost: (www.nhl.com: 968, thehockeynews.com: 66, www.dailyfaceoff.com: 75, www.tsn.ca: 26, www.espn.com: 48, www.thescore.com: 15, sports.yahoo.com: 17, www.nbcsports.com: 3, ca.sports.yahoo.com: 2, www.cbssports.com: 6, africa.espn.com: 1, www.sportsnet.ca: 5), queueAllowed: 1232, neverVerifiedBefore: 992), run: (fetched: 40, verified: 40, found: 32, notFound: 9, totalOnly: 0, noAmountToCheck: 0, changed: 0, dead: 0, blocked: 0, errors: 0, redirectOffDomain: 0, robotsDisallowed: 0 … +3 fields), cumulative: (urlsInStore: 280, verified200: 280, found: 205, notFound: 95, dead: 0, blocked: 0, errors: 0, robotsDisallowed: 0, changedEver: 0), ledgerFacts: (total: 751, judged: 200, standing: 157, pending: 529, atRisk: 43, unverifiableHost: 22), factsAtRisk: (nhlId: 8477496, player: Elias Lindholm, team: BOS, expected: 7750000, tier: certified, status: amount-not-found, sources: 1 item); (nhlId: 8483489, player: Fraser Minten, team: BOS, expected: 879166, tier: certified, status: amount-not-found, sources: 1 item); (nhlId: 8477493, player: Aleksander Barkov, team: FLA, expected: 10000000, tier: reported, status: amount-not-found, sources: 1 item); (nhlId: 8481556, player: John Beecher, team: FLA, expected: 850000, tier: reported, status: amount-not-found, sources: 1 item); (nhlId: 8483424, player: Owen Beck, team: MTL, expected: 853333, tier: certified, status: amount-not-found, sources: 1 item); (nhlId: 8483728, player: Jared Davidson, team: MTL, expected: 850000, tier: reported, status: amount-not-found, sources: 1 item); (nhlId: 8484984, player: Ivan Demidov, team: MTL, expected: 940833, tier: reported, status: amount-not-found, sources: 1 item); (nhlId: 8483686, player: Adam Engstrom, team: MTL, expected: 896667, tier: certified, status: amount-not-found, sources: 1 item); (nhlId: 8484170, player: Jacob Fowler, team: MTL, expected: 923333, tier: reported, status: amount-not-found, sources: 1 item); (nhlId: 8478424, player: Jansen Harkins, team: TBL, expected: 850000, tier: reported, status: amount-not-found, sources: 2 items); (nhlId: 8483752, player: Dominic James, team: TBL, expected: 910000, tier: reported, status: amount-not-found, sources: 1 item); (nhlId: 8477149, player: Scott Sabourin, team: TBL, expected: 850000, tier: certified, status: amount-not-found, sources: 2 items) … +31 more, factsOnUnverifiableHosts: (nhlId: 8483464, player: Marco Kasper, team: DET, expected: 886666, tier: reported, sources: 2 items); (nhlId: 8478043, player: Sam Lafferty, team: FLA, expected: 850000, tier: reported, sources: 1 item); (nhlId: 8480995, player: Pontus Holmberg, team: TBL, expected: 1550000, tier: reported, sources: 1 item); (nhlId: 8479520, player: Brandon Duhaime, team: TOR, expected: 2600000, tier: reported, sources: 1 item); (nhlId: 8480002, player: Nico Hischier, team: NJD, expected: 7250000, tier: reported, sources: 1 item); (nhlId: 8475231, player: Casey Cizikas, team: NYI, expected: 2500000, tier: reported, sources: 1 item); (nhlId: 8482157, player: Will Cuylle, team: NYR, expected: 3900000, tier: reported, sources: 1 item); (nhlId: 8481668, player: Arturs Silovs, team: PIT, expected: 2800000, tier: reported, sources: 1 item); (nhlId: 8477479, player: Tyler Bertuzzi, team: CHI, expected: 5500000, tier: reported, sources: 1 item); (nhlId: 8480448, player: Parker Kelly, team: COL, expected: 1700000, tier: reported, sources: 1 item); (nhlId: 8480800, player: Quinn Hughes, team: MIN, expected: 7850000, tier: reported, sources: 1 item); (nhlId: 8479619, player: Michael Carcone, team: UTA, expected: 1750000, tier: reported, sources: 1 item) … +10 more, nextUp: (url: https://www.nhl.com/news/topic/free-agency/jaden-schwartz-signs-3-year-contract-with-colorado-avalanche, lastFetchAt: —, amountFacts: 2); (url: https://www.nhl.com/news/vancouver-canucks-trade-dakota-joshua-to-toronto-maple-leafs, lastFetchAt: —, amountFacts: 2); (url: https://www.nhl.com/oilers/news/release-oilers-extend-shakir-mukhamadullin-spencer-stastney, lastFetchAt: —, amountFacts: 2); (url: https://www.nhl.com/oilers/news/release-oilers-re-sign-kasperi-kapanen-to-one-year-contract, lastFetchAt: —, amountFacts: 2); (url: https://www.nhl.com/oilers/news/release-oilers-re-sign-max-jones-to-one-year-contract, lastFetchAt: —, amountFacts: 2); (url: https://www.nhl.com/oilers/news/release-oilers-sign-mathieu-joseph-to-one-year-contract, lastFetchAt: —, amountFacts: 2); (url: https://www.nhl.com/oilers/news/release-oilers-sign-ryan-shea-to-five-year-contract, lastFetchAt: —, amountFacts: 2); (url: https://www.nhl.com/panthers/news/florida-panthers-agree-to-terms-with-forward-eetu-luostarinen-on-three-345422684, lastFetchAt: —, amountFacts: 2); (url: https://www.nhl.com/panthers/news/florida-panthers-agree-to-terms-with-forward-sam-reinhart-on-an-eight-year-contract-extension, lastFetchAt: —, amountFacts: 2); (url: https://www.nhl.com/penguins/news/penguins-acquire-defenseman-erik-karlsson-from-the-san-jose-sharks-in--345527572, lastFetchAt: —, amountFacts: 2), definitions: (verified: HTTP 200 on an allowed host after redirects, found: the expected cap hit appears in one of the written forms ($8.5 million, $8,500,000, 8.5M, AAV of $8.5); 1.1 never matches 1, totalOnly: only the contract total (cap hit x term) appears, changed: content SHA-256 differs from the previous fetch of the same URL, dead: HTTP 404/410/451 or the host no longer resolves, atRisk: a ledger cap hit whose allowed sources have all been checked and none shows the amount (source-dead when every one is gone))
- deadOrChanged
- none
Contract agreement matrix
Per-field agreement between the league's releases, the researcher files and the site's ledger, with every disagreement's resolution.
- at
- —
- fields
- capHit: (compared: 307, agree: 276, disagree: 31, agreementPct: 89.9, excludingAdjudications: (compared: 304, agree: 276, disagree: 28, agreementPct: 90.8), pairwise: (gap-round1 | wiki-finder: 3 fields, nhl-extract | researcher: 3 fields, current-contracts | researcher: 3 fields, gap-round1 | researcher: 3 fields, current-contracts | gap-round1: 3 fields, correction | nhl-extract: 3 fields, nhl-extract | verify-adjudication: 3 fields, correction | verify-adjudication: 3 fields, current-contracts | wiki-finder: 3 fields, gap-round1 | nhl-extract: 3 fields, gap-round1 | gap-round2: 3 fields, correction | researcher: 3 fields … +7 fields), disagreements: (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields) … +19 more, tolerance: 25000), lastSeason: (compared: 447, agree: 390, disagree: 57, agreementPct: 87.2, excludingAdjudications: (compared: 445, agree: 390, disagree: 55, agreementPct: 87.6), pairwise: (gap-round1 | nhl-extract: 3 fields, nhl-extract | wiki-finder: 3 fields, gap-round1 | wiki-finder: 3 fields, nhl-extract | researcher: 3 fields, current-contracts | nhl-extract: 3 fields, current-contracts | researcher: 3 fields, gap-round1 | researcher: 3 fields, gap-round2 | nhl-extract: 3 fields, gap-round1 | gap-round2: 3 fields, current-contracts | gap-round1: 3 fields, current-contracts | gap-round2: 3 fields, current-contracts | wiki-finder: 3 fields … +9 fields), disagreements: (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields) … +45 more), termYears: (compared: 319, agree: 291, disagree: 28, agreementPct: 91.2, excludingAdjudications: (compared: 317, agree: 290, disagree: 27, agreementPct: 91.5), pairwise: (gap-round1 | nhl-extract: 3 fields, nhl-extract | wiki-finder: 3 fields, gap-round1 | wiki-finder: 3 fields, nhl-extract | researcher: 3 fields, current-contracts | nhl-extract: 3 fields, current-contracts | researcher: 3 fields, gap-round1 | researcher: 3 fields, gap-round2 | nhl-extract: 3 fields, gap-round1 | gap-round2: 3 fields, current-contracts | gap-round1: 3 fields, current-contracts | gap-round2: 3 fields, current-contracts | wiki-finder: 3 fields … +8 fields), disagreements: (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields) … +16 more), signedOn: (compared: 257, agree: 165, disagree: 92, agreementPct: 64.2, excludingAdjudications: (compared: 256, agree: 165, disagree: 91, agreementPct: 64.5), pairwise: (nhl-extract | wiki-finder: 3 fields, nhl-extract | researcher: 3 fields, current-contracts | nhl-extract: 3 fields, gap-round1 | nhl-extract: 3 fields, gap-round2 | nhl-extract: 3 fields, gap-round1 | gap-round2: 3 fields, current-contracts | gap-round1: 3 fields, current-contracts | gap-round2: 3 fields, current-contracts | researcher: 3 fields, gap-round1 | researcher: 3 fields, gap-round1 | wiki-finder: 3 fields, current-contracts | wiki-finder: 3 fields … +4 fields), disagreements: (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields) … +80 more, within3Days: (agree: 220, compared: 257, agreementPct: 85.6, note: strict agreement plus disagreements whose dates lie within 3 days of each other), gapHistogram: (1-3 days: 55, 4-30 days: 13, over 400 days (a different deal): 14, 31-400 days: 10)), expiryStatus: (compared: 84, agree: 81, disagree: 3, agreementPct: 96.4, excludingAdjudications: (compared: 84, agree: 81, disagree: 3, agreementPct: 96.4), pairwise: (gap-round1 | researcher: 3 fields, gap-round1 | gap-round2: 3 fields, gap-round2 | researcher: 3 fields), disagreements: (7 fields); (7 fields); (7 fields)), clauseStatus: (compared: 45, agree: 20, disagree: 25, agreementPct: 44.4, excludingAdjudications: (compared: 45, agree: 20, disagree: 25, agreementPct: 44.4), pairwise: (clauses-r3 | current-contracts: 3 fields, clauses-r1 | gap-round2: 3 fields, clauses-r1 | researcher: 3 fields, clauses-r1 | clauses-r3: 3 fields, clauses-r1 | gap-round1: 3 fields), disagreements: (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields); (7 fields) … +13 more)
- overall
- —
- coverage
- totals: (contracted: 1068, priced: 743, certified: 562, reported: 181, clauseKnown: 495, regulars: 776, regularsPriced: 705, pricedPct: 69.6, clauseKnownPct: 46.3, regularsPricedPct: 90.9), perClub: (abbr: BOS, slug: boston-bruins, rows: 43, contracted: 42, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: BUF, slug: buffalo-sabres, rows: 44, contracted: 25, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: DET, slug: detroit-red-wings, rows: 32, contracted: 32, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: FLA, slug: florida-panthers, rows: 35, contracted: 35, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: MTL, slug: montreal-canadiens, rows: 42, contracted: 32, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: OTT, slug: ottawa-senators, rows: 27, contracted: 27, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: TBL, slug: tampa-bay-lightning, rows: 49, contracted: 37, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: TOR, slug: toronto-maple-leafs, rows: 47, contracted: 20, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: CAR, slug: carolina-hurricanes, rows: 42, contracted: 36, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: CBJ, slug: columbus-blue-jackets, rows: 26, contracted: 26, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: NJD, slug: new-jersey-devils, rows: 37, contracted: 37, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields); (abbr: NYI, slug: new-york-islanders, rows: 44, contracted: 41, capHit: 6 fields, fields: 4 fields, clause: 6 fields, regulars: 3 fields) … +20 more, thinnest: (byPricedShare: (2 fields); (2 fields); (2 fields); (2 fields); (2 fields), byCertifiedShare: (2 fields); (2 fields); (2 fields); (2 fields); (2 fields), byRegularsPriced: (2 fields); (2 fields); (2 fields); (2 fields); (2 fields), byClauseKnown: (2 fields); (2 fields); (2 fields); (2 fields); (2 fields), byExpiryKnown: (2 fields); (2 fields); (2 fields); (2 fields); (2 fields), bySignedOnKnown: (2 fields); (2 fields); (2 fields); (2 fields); (2 fields), namedBiasClubs: (4 fields); (4 fields); (4 fields))
- thinnest
- none
External timestamps
The prediction ledger's chain heads are anchored each day with an RFC 3161 trusted timestamp, so the record can be verified without trusting this site.
| Date | Verified | Authority | Details |
|---|---|---|---|
| 2026-09-29 | verified | freetsa |
Arena recording profiles
Home-versus-away distributions of shot distance and subjective counts per building; the correction applied to locations and the label applied to counts.
- at
- —
- seasons
- 2018, 2019, 2020, 2021, 2022, 2023, 2024, 2025
- arenas
- none
- note
- —
Bias ledger
For every created measure: the mean residual by position, hand, age, club, arena, nationality, draft, contract status, cap tier and market size after controlling for role and minutes, with bootstrap intervals and Benjamini-Hochberg control. Only slices that survive false-discovery control are listed.
Season 2025-26, built 2026-09-29.
| Measure | Players | Slices surviving FDR | Which (group, n, residual) | Details |
|---|---|---|---|---|
| drive | 621 | 18 | club=BOS (20, +0.048); club=BUF (20, +0.03); club=CAR (20, -0.07); club=FLA (19, -0.026); club=NYI (18, -0.028); club=TOR (21, +0.026); club=VGK (20, -0.045); arena=Amerant Bank Arena (19, -0.026); arena=KeyBank Center (20, +0.03); arena=Lenovo Center (20, -0.07); arena=Scotiabank Arena (21, +0.026); arena=T-Mobile Arena (20, -0.045) … +6 | |
| driveOff | 621 | 26 | shoots=not on file (34, -0.207); ageBand=not on file (31, -0.216); club=BOS (20, +0.307); club=BUF (20, +0.428); club=CAR (20, -0.461); club=NYI (18, -0.382); club=SJS (20, +0.212); club=TOR (21, +0.246); club=VGK (20, -0.348); arena=KeyBank Center (20, +0.428); arena=Lenovo Center (20, -0.461); arena=SAP Center at San Jose (20, +0.212) … +14 | |
| driveDef | 621 | 13 | ageBand=23-26 (164, +0.093); club=BOS (20, +0.298); club=CAR (20, -0.375); club=FLA (19, -0.312); club=VAN (20, -0.19); club=VGK (20, -0.155); arena=Amerant Bank Arena (19, -0.312); arena=Lenovo Center (20, -0.375); arena=Rogers Arena (20, -0.19); arena=T-Mobile Arena (20, -0.155); arena=TD Garden (20, +0.298); draftRound=undrafted (71, +0.139) … +1 | |
| shutdown | 589 | 0 | ||
| driveOffAdj | 589 | 15 | club=ANA (20, -0.24); club=BUF (20, +0.365); club=CAR (19, -0.451); club=CHI (20, +0.251); club=NYI (17, -0.249); club=TOR (20, +0.365); arena=Honda Center (20, -0.24); arena=KeyBank Center (20, +0.365); arena=Lenovo Center (19, -0.451); arena=Scotiabank Arena (20, +0.365); arena=UBS Arena (17, -0.249); arena=United Center (20, +0.251) … +3 | |
| twoWayAdj | 589 | 15 | club=ANA (20, -0.488); club=BUF (20, +0.335); club=CAR (19, -0.652); club=CHI (20, +0.439); club=OTT (18, -0.285); club=TOR (20, +0.394); club=VAN (18, -0.417); arena=Canadian Tire Centre (18, -0.285); arena=Honda Center (20, -0.488); arena=KeyBank Center (20, +0.335); arena=Lenovo Center (19, -0.652); arena=Rogers Arena (18, -0.417) … +3 | |
| startPM60 | 621 | 43 | shoots=not on file (34, -0.263); ageBand=<=22 (82, -0.205); ageBand=23-26 (164, +0.127); ageBand=not on file (31, -0.308); club=ANA (20, -0.278); club=BOS (20, +0.613); club=BUF (20, +0.596); club=CGY (21, -0.339); club=CHI (21, -0.359); club=COL (18, +0.585); club=DAL (19, +0.622); club=FLA (19, -0.328) … +31 | |
| setupXg60 | 621 | 9 | club=BUF (20, +0.029); club=LAK (18, -0.026); club=VAN (20, -0.023); arena=Crypto.com Arena (18, -0.026); arena=KeyBank Center (20, +0.029); arena=Rogers Arena (20, -0.023); contract=UFA (253, +0.009); contract=not on file (49, -0.015); capTier=Q4 (highest) (181, +0.014) | |
| finish | 540 | 0 | ||
| gsax | 71 | 0 | ||
| aPM60 | 685 | 40 | ageBand=<=22 (96, -0.325); club=ANA (23, -0.83); club=BOS (20, +0.57); club=BUF (22, +1.347); club=CAR (21, +0.626); club=CGY (22, -0.713); club=CHI (21, -0.868); club=COL (19, +1.636); club=DAL (23, +0.968); club=LAK (20, -0.5); club=MTL (21, +0.503); club=PIT (22, +0.738) … +28 | |
| realPoints60 | 769 | 26 | club=BOS (20, +0.371); club=BUF (23, +0.198); club=CAR (23, -0.242); club=CGY (26, -0.271); club=DAL (25, +0.306); club=LAK (21, -0.175); club=MIN (21, +0.183); club=NYI (22, -0.319); club=PIT (27, +0.264); club=VAN (27, -0.267); arena=American Airlines Center (25, +0.306); arena=Crypto.com Arena (21, -0.175) … +14 | |
| turnover60 | 458 | 2 | club=OTT (16, +0.425); arena=Canadian Tire Centre (16, +0.425) | |
| whiteNoise | 912 | 25 | ageBand=<=22 (175, +7.992); ageBand=27-30 (244, -3.946); club=CGY (36, -11.7); club=DET (27, -13.351); club=MTL (30, +24.867); club=NJD (32, -11.43); club=PHI (32, +11.048); club=TBL (33, -16.479); club=TOR (33, +16.78); club=VAN (28, +14.027); club=VGK (29, -17.468); arena=Benchmark International Arena (33, -16.479) … +13 |
Structural findings, as measured:
- drive: finding a_driveVsClubStrength · question Do team-relative measures compress players on deep clubs? Pearson r between the player's measure and his 2025-26 club's season xG share (xg-2025-26 teams.xgShar · n 621 · r 0.003 · ci95 -0.073, 0.077 · slopePerSharePoint 0
- driveOff: finding a_driveVsClubStrength · question Do team-relative measures compress players on deep clubs? Pearson r between the player's measure and his 2025-26 club's season xG share (xg-2025-26 teams.xgShar · n 621 · r 0.053 · ci95 -0.025, 0.126 · slopePerSharePoint 0.015
- driveDef: finding a_driveVsClubStrength · question Do team-relative measures compress players on deep clubs? Pearson r between the player's measure and his 2025-26 club's season xG share (xg-2025-26 teams.xgShar · n 621 · r 0.057 · ci95 -0.015, 0.123 · slopePerSharePoint 0.009
- shutdown: finding a_driveVsClubStrength · question Do team-relative measures compress players on deep clubs? Pearson r between the player's measure and his 2025-26 club's season xG share (xg-2025-26 teams.xgShar · n 589 · r -0.041 · ci95 -0.116, 0.035 · slopePerSharePoint -0.005
- driveOffAdj: note no structural lean was tested for this measure; see the top-level findings
- twoWayAdj: note no structural lean was tested for this measure; see the top-level findings
- startPM60: note no structural lean was tested for this measure; see the top-level findings
- setupXg60: finding b_setupVsTeammateFinishing · question Does Setup ride on the shooters? Teammates' finishing = the club's shooters' (goals - xG) per 100 unblocked shots with the player himself removed. Players with · setupXgPerPrimaryAssist n: 491, r: 0.02, ci95: -0.076, 0.109 · setupXgPer60 n: 491, r: 0.049, ci95: -0.046, 0.144 · reading An assist only exists when the shot went in, so a passer whose linemates convert low-xG chances collects assists on low-xG shots (per-assist falls) while his to
- finish: finding d_finishGsaxVsHomeArena · question Do shot-location recording habits at a rink leak into Finish and GSAx? One-way spread of the controlled residual across the 2025-26 club's home arena, with a la · n 540 · arenas 32 · eta2 0.065 · eta2UnderNull 0.057
- gsax: finding d_finishGsaxVsHomeArena · question Do shot-location recording habits at a rink leak into Finish and GSAx? One-way spread of the controlled residual across the 2025-26 club's home arena, with a la · n 71 · arenas 32 · eta2 0.388 · eta2UnderNull 0.444
- aPM60: note no structural lean was tested for this measure; see the top-level findings
- realPoints60: note no structural lean was tested for this measure; see the top-level findings
- turnover60: finding e_puckBattlesVsHomeArena · question Hand-recorded counts differ by rink (the deployment method says so). One-way spread of per-60 counts, after position and log minutes, across the home arena; ska · limit Season totals mix 41 home and 41 road games, so a rink's scorers are visible here only as the half-diluted club residue, and a club's playing style sits in the · byStat hitsPer60: (n: 458, arenas: 32, eta2: 0.054, eta2UnderNull: 0.068, permutationP: 0.797, controlR2: 0.131, lowestArenas: (3 fields); (3 fields); (3 fields), highestArenas: (3 fields); (3 fields); (3 fields)), blocksPer60: (n: 458, arenas: 32, eta2: 0.08, eta2UnderNull: 0.068, permutationP: 0.228, controlR2: 0.502, lowestArenas: (3 fields); (3 fields); (3 fields), highestArenas: (3 fields); (3 fields); (3 fields)), takeawaysPer60: (n: 458, arenas: 32, eta2: 0.08, eta2UnderNull: 0.068, permutationP: 0.234, controlR2: 0.076, lowestArenas: (3 fields); (3 fields); (3 fields), highestArenas: (3 fields); (3 fields); (3 fields)), giveawaysPer60: (n: 458, arenas: 32, eta2: 0.072, eta2UnderNull: 0.069, permutationP: 0.379, controlR2: 0.009, lowestArenas: (3 fields); (3 fields); (3 fields), highestArenas: (3 fields); (3 fields); (3 fields)), dzGiveawaysPer60: (n: 458, arenas: 32, eta2: 0.059, eta2UnderNull: 0.068, permutationP: 0.681, controlR2: 0.752, lowestArenas: (3 fields); (3 fields); (3 fields), highestArenas: (3 fields); (3 fields); (3 fields))
- whiteNoise: finding c_whiteNoiseVsMarketAndWikipedia · question Is attention a function of where a player plays and whether he has an English Wikipedia article (a pageview component only exists when a title resolved)? · marketSize n: 912, rLogPopVsScore: 0.025, ci95: -0.039, 0.086, rLogPopVsResidual: 0.026, ci95Residual: -0.038, 0.089, byTier: (small: (n: 317, meanScore: 53.968, meanResid: -2.58), mid: (n: 279, meanScore: 58.265, meanResid: 1.248), large: (n: 316, meanScore: 58.032, meanResid: 1.486)), clubField: current (2026-27) club, since attention is measured now · wikipediaArticle nWithArticle: 905, nWithout: 7, meanScoreWith: 56.886, meanScoreWithout: 31.429, diff: 25.458, ci95: 11.882, 39.35, diffAfterControls: 23.286, ci95AfterControls: 5.932, 42.052, pointBiserialR: 0.08 · wikipediaArticleAllListed nWithArticle: 1090, nWithout: 196, meanScoreWith: 54.474, meanScoreWithout: 25.684, diff: 28.791, ci95: 25.617, 32.116, pointBiserialR: 0.359, note: raw percentile scores, no minutes control: most players without an article have no 2025-26 NHL ice time · unknownWikiStatus 0
- 2025-26 club = the club a player logged the most regular-season shift seconds for in data/shifts/shifts-2025.json (the same charts the measures are built from); the 2026-27 camp roster is NOT used for club, arena or market tier except for White Noise, where attention is measured now. Cross-check: of 778 players with a league-production teams string, the shift club matched the first-listed club 756 times and the last-listed 718 times, so that string leads with the majority club more often than it ends with it; the disagreements are traded players.
- Traded players carry their majority club; team-relative measures (Drive, driveOff, driveDef, adjusted drive) already compare each game against that night's club, so a traded player's number mixes two clubs' baselines.
- Home arena is the modal regular-season home venue in schedule-20252026.json; it is a one-to-one relabeling of club in 2025-26 (FLA, NSH, PIT and TBL each played one home game elsewhere), so arena slices repeat club slices and share their q rather than enlarging the FDR family.
- Control fit: OLS of the measure on isD (isG where goalies are present), log(volume) and, where the measure is not already deployment-adjusted, the offensive-zone start share (mean-imputed with an indicator when the site itself would not show one). Residuals are what the slices average.
- Flagged = the 2,000-resample percentile bootstrap interval of the group's mean residual excludes 0. q = Benjamini-Hochberg over the one-sample t p-values of every slice of the measure (n >= 5), with the arena slices excluded as duplicates; survivesFdr = q <= 0.10. Groups under 5 players get no interval and a reason.
- Position appears as a slice for completeness but is a control dummy, so its residual mean is zero by construction; the raw means are the informative part of that row.
- Finish uses xg-2025-26 shooters (goals minus xG per 100 unblocked shots, >= 80 shots). advanced.json goalsAboveExpected sums (1 - xG) over goals scored only, so it is not net of misses and was not used as Finish.
- Shutdown/driveOffAdj/twoWayAdj: universe = every skater with the field in deployment.json (forwards and defence; fitted separately by the builder), which is the rankings-desk lens; the stats-desk headline Shutdown list is defence only, and the position slice gives the D-only raw mean.
- White Noise is a percentile among all 1,286 listed players; this audit keeps the players with 2025-26 NHL ice time so minutes can be controlled, and its club, arena and market tier are the current club's.
- Representation: top list by measure value (top 100; top 30 for goalies); ratio = share in the top over share in the universe, interval from resampling the top list and the universe independently.
- Simpson checks compare raw group means pooled and within F and D (both sides >= 10 players); a reversal is a sign flip of the raw difference. Residual differences are given alongside.
- Club posterior: normal-normal shrinkage of club mean residuals with tau2 by method of moments; tau2 = 0 means the between-club spread is no wider than sampling noise and every club posterior is 0.
Where the price leans
For every slice of games (club, home/away, favourite, tier, back-to-back, month, division, season) the model's average probability against the share that won, bootstrap ranges, false-discovery control, and a sequential monitor on the live ledger that keeps its error rate under daily peeking.
- month 02 (home side): 808 games, gap -0.034 [-0.068, -0.002], q 0.891
Games without a schedule row: count: 0, ids: none, matchedFromLedger: 0, skippedRows: none, note: games without a schedule row still count in the slices that need no date or club (all games, favourite, tiers).
No slice flagged.
Games without a schedule row: count: 0, ids: none, matchedFromLedger: 0, skippedRows: none, note: games without a schedule row still count in the slices that need no date or club (all games, favourite, tiers).
7 studies, built 2026-09-29. Every number below is read from the study files at build time; a missing one reads "unknown".
Expected goals: challengers, the event clock and arena scorers
Seven pre-registered challengers to the expected-goals model have been judged on the unseen 2025-26 season; three passed the registered rule. The desk withholds all three and puts none forward, so the champion stays.
Findings
- measured In the league's play-by-play a goal is stamped about one second later than a saved or missed shot. The shift is in 2023-24, 2024-25 and 2025-26 and not in 2018-19, 2019-20, 2020-21, 2021-22 and 2022-23.
- measured A timing input with no hockey meaning lowered held-out log loss by 0.000661 per shot on the raw clock (95% range −0.000881 to −0.000466). On the corrected clock the change is +0.000008 (−0.000002 to +0.000016) and it no longer lowers it.
- measured The champion's rebound and rush inputs are short windows on that clock. With goal times moved back 1 s the champion scores 0.22883 against 0.22861 on the same 105,288 shots: +0.000225 per shot (95% range +0.000107 to +0.000351). Part of its published accuracy was the clock.
- measured 70% of the shots flagged as rush in 2025-26 follow the club's own blocked attempt and are not rushes.
- measured The flurry idea on the corrected clock: −0.000003 per shot (95% range −0.000033 to +0.000029). The desk's reading: adds nothing: the gain challenger C registered on the raw clock was the clock.
- measured Seven challengers were registered before being fitted, trained on 2023-24 and 2024-25 and judged on 2025-26 by a bootstrap over games (1,000 resamples). Passed the registered rule: strength-state models (four fits) (−0.000491, 95% range −0.000713 to −0.000248), flurry features added (−0.001587, 95% range −0.001970 to −0.001228) and rink-adjusted + strength-state fits + flurry (−0.001965, 95% range −0.002436 to −0.001503).
- measured The gain of the strength-state models (four fits) does not come from the clock inputs: with rebound and rush removed from both sides its difference is −0.000489 (95% range −0.000707 to −0.000240). It is withheld because it keeps the champion's raw-clock flags, and the timing input still lowers its log loss by 0.000651 (−0.000865 to −0.000456).
- measured By strength state, judged on all games, the split has a lower log loss than the champion on the power play and with the net empty; at 5-on-5 and shorthanded the range spans zero.
- measured Replica check: the desk's copy of the champion scores 0.22830 per unblocked shot on 112,091 shots from 1,312 games (published 0.2281) and matches the site's stored values on 112,089 of 112,089 shots.
- decision No challenger is put forward. The desk's reason: no candidate passes the registered rule, is attributable and shows no placebo gain, so the desk puts none forward. B passes the rule but its flags are built on the raw clock. The clock finding goes to whoever owns tools/pull-pbp.mjs.
- caveat The first stage of the study put forward strength-state models (four fits). The clock follow-up withdrew it. In the desk's words: WITHHELD: passes the registered rule, but a clock placebo still lowers log loss on the clock its flags are built on; its counterpart on the corrected clock is G.
- decision Desk note: Two things for the hub to weigh. (1) The rule measures every challenger against a champion that gains about 0.000244 per shot from the clock, so a challenger built on the corrected clock starts that far behind. On a level clock the strength split gains 0.000486 (G against F, 95% range [-0.000742, -0.000247]); against the published champion G is -0.000261 with range [-0.000543, 0.000026], which misses the rule at the top of its range (95.7% of resamples at or below zero). (2) Correcting the clock in the champion is an integrity decision, not an accuracy one: F removes information the model should never have had, so it cannot pass a rule that asks for a lower log loss. The desk suggests deciding that first. If the corrected clock becomes the baseline, the strength split should be registered again against it and judged on a season not yet read; its empty-net fit is unstable (trained on one season it lost to a single fit).
- caveat Seven challengers have now been judged on the same proof season; read a single pass with that in mind.
- measured Arena scorers, 2025-26: 8 of 32 buildings record shot distance differently from the same clubs' road shots elsewhere (25% at p below 0.01, against 3.5% when the building is nothing special).
- measured The largest is STL at Enterprise Center: shots recorded 1.89 ft farther (about +0.87 to +2.92 ft). Across its five seasons on file the mean is −0.93 ft, the opposite direction to this season, with four of the five seasons on the side of that mean.
- measured From one season to the next the arena effects correlate at r 0.519 since 2021-22 (the registered threshold was 0.5). The spread between buildings has shrunk from 1.71 ft in 2021-22 to 1.00 ft in 2025-26, against a noise floor of about 0.56 ft.
- measured The distance correction was not shown to help on games it had not seen: 22 arena-seasons repaired against 20 broken (exact p 0.88).
Challengers
Trained on 2023-24 and 2024-25, judged on 2025-26. Rule: adopt only if the bootstrap 95% range of (challenger - champion) log loss lies below zero; otherwise reject; adoption itself is a hub decision card, never automatic. Champion held-out log loss 0.22830 on all 1,312 games; 0.22861 on the 1,232 games with an events file, where the clock follow-up is judged.
| Challenger | Result | Log loss against champion | 95% range | Details |
|---|---|---|---|---|
| A rink-adjusted distance (+ angle by the same geometry) | fails the rule | +0.000104 | −0.000033 to +0.000236 | |
| B strength-state models (four fits) | passes, withheld | −0.000491 | −0.000713 to −0.000248 | |
| C flurry features added | passes, withheld | −0.001587 | −0.001970 to −0.001228 | |
| D rink-adjusted + strength-state fits + flurry | passes, withheld | −0.001965 | −0.002436 to −0.001503 | |
| F champion features, rebound and rush on the corrected clock | fails the rule | +0.000225 | +0.000107 to +0.000351 | |
| G strength-state fits with the corrected flags | fails the rule | −0.000261 | −0.000543 to +0.000026 | |
| H G plus flurry features on the corrected clock | fails the rule | −0.000264 | −0.000548 to +0.000018 | |
| B strength-state fits, flags as in the shot file (re-scored on these rows) | passes, withheld | −0.000539 | −0.000788 to −0.000298 |
A negative difference is a lower (better) log loss than the champion. A pass needs the whole range below zero.
Known limitation: the event clock
Every shot in the league's play-by-play carries a clock time. Goals sit about one second later than saved or missed shots, measured against whatever happened just before; the likely reason is that a goal takes the time the clock stopped and other shots the time the scorer logged them. This is a property of the feed from 2023-24 on; in 2018-19, 2019-20, 2020-21, 2021-22, 2022-23 goal times and shot times agree. In 2023-24 and 2024-25, shots logged in the same second as the previous event were goals 1.0% of the time and shots logged one second after it 3.0%, against 10.8% at three seconds and about 7% overall. That gap is not hockey: it appears after hits, faceoffs, giveaways and takeaways alike. Moving goal times back 1 second makes the timing of goals and of other shots line up (after a hit, the distance between the two patterns falls from 0.07 to 0.021; moving them back 2 seconds overshoots). Two inputs of our expected-goals model are short windows on that clock: 'rebound' (within 3 seconds of the same club's previous attempt) and 'rush' (within 4 seconds). Because a goal stamped a second late can fall outside the window, both were short of goals and the model learned to mark them down (rush coefficient -0.4216 on the raw clock, -0.1317 on the corrected one). A timing input with no hockey meaning improved the model's score on unseen games by 0.000661 per shot on the raw clock; on the corrected clock it changes it by +0.000008 and no longer improves it. The corrected champion (F) scores 0.22883 against the published champion's 0.22861 on the same 2025-26 shots: +0.000225 per shot, 95% range [0.000107, 0.000351]. It is WORSE, and the range excludes zero. Part of the champion's published accuracy was the clock, not hockey: against the same fit on the raw clock the cost of correcting it is +0.000244 [0.000127, 0.000372]. Separately, 70% of the shots flagged 'rush' in 2025-26 are not rushes at all: they are follow-ups to the club's own blocked shot, flagged because the feed credits a blocked shot to the shooter's club but gives its zone from the blocker's side.
| Seconds after the previous event | Scored, raw clock | Scored, goals moved back | Shots (raw) | Details |
|---|---|---|---|---|
| 0 | 1.0% | 17.9% | 2,213 | |
| 1 | 3.0% | 9.2% | 14,961 | |
| 2 | 9.0% | 8.9% | 16,245 | |
| 3 | 10.8% | 6.5% | 13,326 | |
| 4 | 7.6% | 6.4% | 11,000 | |
| 5 | 6.7% | 6.4% | 10,307 | |
| 6 | 6.5% | 6.9% | 10,071 | |
| 7 | 7.1% | 7.2% | 9,799 | |
| 8+ | 6.9% | 6.4% | 137,955 |
| Placebo input | Clock | Change in log loss | 95% range | Lowers log loss |
|---|---|---|---|---|
| four events | raw | −0.000661 | −0.000881 to −0.000466 | yes |
| four events | goal times moved back 1 s | +0.000008 | −0.000002 to +0.000016 | no |
| hits only | raw | −0.000028 | −0.000143 to +0.000097 | no |
| hits only | goal times moved back 1 s | +0.000010 | −0.000007 to +0.000029 | no |
| Flag | Shots flagged, raw | Shots flagged, corrected | Added, of them goals | Details |
|---|---|---|---|---|
| rebound | 33,801 | 34,393 | 592 (592) | |
| rush | 13,212 | 13,470 | 258 (258) |
Cost of correcting the clock in the champion: +0.000225 log loss per shot against the published model (95% range +0.000107 to +0.000351), +0.000244 against the same fit on the raw clock (95% range +0.000127 to +0.000372), on 105,288 shots in 1,232 games.
Shots flagged as rush that follow the club's own blocked attempt: 2018-19 48%, 2019-20 55%, 2020-21 57%, 2021-22 55%, 2022-23 61%, 2023-24 55%, 2024-25 69%, 2025-26 70%. Why: the feed credits a blocked shot to the shooter's club and gives its zone from the blocker's side, so "previous event by the same club outside the offensive zone" is true for every shot that follows an opponent's block.
Caveats on the clock follow-up (7)
- The proof rows are the 1,232 games with an events file; A-D in this file were judged on all 1,312. B is re-scored here on the 1,232.
- The shift exists only from 2023-24 in the feed on disk; the correction must not be applied to earlier seasons.
- Moving goal times back uses the kind of event (goal or not) to undo a recording convention. It is a repair of the data, tested by the placebos; it is not a claim that every goal is exactly one second late.
- The strength split's gain is not stable: trained on one season it was worse than a single fit, because the empty-net fit has about 900 rows per season.
- H's flurry component had already been seen on the proof season with goal times moved back; H confirms, it does not independently test.
- The rush flag is reproduced with the feed's blocked-shot quirk so that only the clock changes between the champion and F.
- Seven challengers have now been judged on 2025-26. A pass is weaker evidence than a single test would be.
Arena scorer effects
| Building | Recorded distance against road shots | About (95%) | KS p | Details |
|---|---|---|---|---|
| STL Enterprise Center | +1.89 ft | +0.87 to +2.92 | below 0.001 | |
| UTA Delta Center | +1.67 ft | +0.66 to +2.67 | 0.006 | |
| ANA Honda Center | −1.63 ft | −2.54 to −0.71 | 0.011 | |
| CAR Lenovo Center | −1.61 ft | −2.69 to −0.52 | 0.005 | |
| BUF KeyBank Center | +1.47 ft | +0.49 to +2.46 | 0.001 | |
| EDM Rogers Place | −1.33 ft | −2.27 to −0.39 | 0.002 | |
| TBL Benchmark International Arena | +1.31 ft | +0.25 to +2.38 | 0.072 | |
| VAN Rogers Arena | −1.26 ft | −2.22 to −0.31 | 0.020 | |
| MIN Grand Casino Arena | +1.20 ft | +0.25 to +2.15 | 0.006 | |
| NYR Madison Square Garden | +1.14 ft | +0.16 to +2.12 | 0.003 | |
| CBJ Nationwide Arena | +1.07 ft | +0.12 to +2.03 | 0.018 | |
| BOS TD Garden | −1.00 ft | −1.99 to 0.00 | 0.001 | |
| VGK T-Mobile Arena | +0.93 ft | −0.07 to +1.92 | 0.038 | |
| NYI UBS Arena | −0.86 ft | −1.85 to +0.14 | 0.131 | |
| DET Little Caesars Arena | −0.85 ft | −1.80 to +0.10 | 0.030 | |
| COL Ball Arena | +0.85 ft | −0.18 to +1.89 | 0.190 | |
| FLA Amerant Bank Arena | +0.84 ft | −0.21 to +1.88 | 0.022 | |
| WPG Canada Life Centre | −0.83 ft | −1.82 to +0.15 | 0.026 | |
| PIT PPG Paints Arena | −0.76 ft | −1.70 to +0.18 | 0.144 | |
| MTL Centre Bell | −0.70 ft | −1.69 to +0.29 | 0.214 | |
| SJS SAP Center at San Jose | −0.69 ft | −1.64 to +0.27 | 0.224 | |
| CGY Scotiabank Saddledome | −0.57 ft | −1.51 to +0.37 | 0.033 | |
| NJD Prudential Center | −0.57 ft | −1.57 to +0.44 | 0.197 | |
| NSH Bridgestone Arena | −0.57 ft | −1.51 to +0.38 | 0.499 | |
| CHI United Center | −0.54 ft | −1.58 to +0.50 | 0.013 | |
| PHI Xfinity Mobile Arena | +0.49 ft | −0.52 to +1.51 | 0.052 | |
| DAL American Airlines Center | +0.48 ft | −0.58 to +1.54 | 0.779 | |
| WSH Capital One Arena | −0.37 ft | −1.34 to +0.60 | 0.826 | |
| TOR Scotiabank Arena | +0.32 ft | −0.59 to +1.22 | 0.163 | |
| LAK Crypto.com Arena | −0.23 ft | −1.25 to +0.78 | 0.282 | |
| SEA Climate Pledge Arena | −0.19 ft | −1.12 to +0.74 | 0.904 | |
| OTT Canadian Tire Centre | 0.00 ft | −1.03 to +1.03 | 0.832 |
Positive means shots in that building are recorded farther from the net than the same clubs' shots elsewhere. Ranges are the mean difference plus and minus 1.96 standard errors, computed here from the desk's figures. The standard error treats shots as independent, so the ranges are somewhat too narrow.
| Season | Spread (ft) | Split-half r | Details |
|---|---|---|---|
| 2018-19 | 1.70 | 0.76 | |
| 2019-20 | 1.67 | 0.77 | |
| 2020-21 | 1.81 | 0.78 | |
| 2021-22 | 1.71 | 0.86 | |
| 2022-23 | 1.17 | 0.68 | |
| 2023-24 | 0.80 | 0.30 | |
| 2024-25 | 0.83 | 0.61 | |
| 2025-26 | 1.00 | 0.65 |
Scorer effect: declared by the registered rule (mean consecutive-season r from 2021-22 = 0.519 against a threshold of 0.5; 0.25 of arenas at KS p < 0.01 in 2025-26 against a placebo rate of 0.035). The margin on r is thin and the effect has shrunk: the spread of arena mean differences was 1.713 ft in 2021-22 and is 0.996 ft in 2025-26, against a noise floor of about 0.564 ft. Correction: by the letter of the registered rule the held-out share of arenas at p < 0.01 fell (0.175 to 0.1625), but the change is 22 arena-seasons repaired against 20 broken (exact p = 0.8776), so it is NOT a demonstrable improvement. Carried from one season to the next the unshrunk map moved the share from 0.3481 to 0.4367.
Caveats on the arena study (9)
- Neutral-site and outdoor games are separate venue+club keys with 1-2 games; they get no map and recorded distances there are left as recorded.
- A club that changes its building name keeps its link across seasons (home-club identity); a club that moves building (NYI to UBS Arena in 2021-22, ARI to Mullett Arena in 2022-23) is ALSO linked, on the view that the off-ice crew belongs to the club. LAK 2021-22 appears as two keys (STAPLES Center 16 games, Crypto.com Arena 25 games) because the name changed mid-season.
- events-2025.json holds 1,232 of the 1,312 regular-season games: |x| and the subjective counts use those games; the shot file is complete.
- The visitors' comparison cannot separate the scorer from the home club's defensive system (every visitor shot here was taken against the home club). The home-club component and the year-to-year persistence across roster turnover are the available checks.
- The p-values treat shots as independent; shots cluster within games, so p-values are somewhat too small.
- With about 32 arenas x 4 event kinds per season, about 6 subjective flags per season are expected by chance alone.
- The feed credits a blocked shot to the shooter's club. An arena flagged low on blocked shots records fewer blocked attempts than the same clubs' games elsewhere produce; which side's attempts are short is in homeClubRatio (the home club's attempts) and visitorsRatio.
- The primary comparison is venue-role matched (road shots elsewhere). The brief's literal all-elsewhere definition is reported per arena as distance.allElsewhere; it sits about 0.15-0.45 ft higher at every arena because home shots are recorded closer league-wide.
- Entries marked exploratory (shrunken maps) were added after the registered results were seen and carry no decision weight.
Caveats (7)
- Shots at neutral sites and outdoor games (1-2 games per venue) are left as recorded in A and D.
- Utah (home id 68 in 2025-26) is linked to its 2024-25 map (home id 59) by the venue name; renamed buildings are linked by home club.
- The training rows are mapped in sample and the proof rows out of sample, as registered; the map is not shrunk in A.
- B and D fit the shorthanded and empty-net states on a few thousand rows each; their coefficients are noisy.
- All four challengers are judged on the same 2025-26 games; a pass by one of four at the 95% level is weaker evidence than a single test.
- The shot files on disk were rebuilt after the champion was trained, so the replica is scored on slightly different rows than the published test.
- The clock check is post hoc. It does not overturn a registered verdict; it is the reason C and D are not put forward.
Study stamp 2026-09-29 13:24 UTC. Read from data/models/xg-challengers.json, data/quality/decisions-request-xg.json, data/quality/xg-clock-finding.json, data/quality/arena-profiles.json.
Prospect formula: cohort backtest
Whole draft classes, 2010 to 2018: 1,902 of 1,903 picks have a season history on file. Among the 1,695 skaters scored at 18 the formula at 18 ranks NHL games through age 23 at a rank correlation of 0.448 (95% range 0.406 to 0.487), against 0.605 for draft position; at 20 it reaches 0.606.
Findings
- measured 20.1% of these players graduated by 23 (383 of 1,902).
- measured Rank correlation with NHL games through 23, skaters: the formula at 18 0.448, at 19 0.530, at 20 0.606; draft position 0.605; raw points per game in the draft year 0.152.
- measured Formula minus draft position on the same skaters: −0.159 at 18 (−0.203 to −0.113), −0.074 at 19 (−0.114 to −0.029), +0.001 at 20 (−0.040 to +0.041).
- measured Taken apart at 18: percentile alone 0.356, times the league ladder 0.431, times the age factor 0.431 and plus the NHL bonus (the full formula) 0.448.
- measured Graduation by score at 18: 3.3% in the bottom tenth (6 of 184) and 48.4% in the top tenth (89 of 184); the share rises with the score but not at every step.
- measured Relative age: births by calendar quarter are not spread evenly over the year (chi-square 161.9, p below 0.001). In the score: no detectable carry in this sample: the difference between the oldest and youngest birth quarter in league-adjusted score at 18 has an interval that includes zero.
- measured Measured growth in points per game, same league, outside the NHL: 1.110 times a year from 18 to 24 on average, against 1.081 implied by the hand-set age line.
- caveat Whether a steeper age line would rank better: the two horizons disagree (three seasons ahead: steeper ranks worse; through age 23: no detectable difference), so the choice of slope cannot be settled by ranking on this sample; the differences are a few hundredths of rho either way.
- measured Ladder: nine leagues re-measured from 694 moves to the NHL; the site's value sits inside the range for seven.
- measured Every pick in the class lists, goalies included: 1,902 of 1,903 have a known outcome, and draft position alone ranks NHL games through 23 at 0.588 (95% range 0.552 to 0.621).
- measured Graduation by round: 72.7% of round-1 picks graduated by 23 (197 of 271), against 1.5% in round 7 (4 of 271).
Validity against draft position
| Predictor (skaters) | NHL games to 23 | 95% range | Details |
|---|---|---|---|
| formula score at 18 | 0.448 | 0.406 to 0.487 | |
| formula score at 19 | 0.530 | 0.489 to 0.566 | |
| formula score at 20 | 0.606 | 0.570 to 0.640 | |
| formula score on the draft-year season (supplementary) | 0.259 | 0.214 to 0.305 | |
| draft position | 0.605 | 0.570 to 0.635 | |
| raw points per game in the draft-year season (skaters, 10+ games) | 0.152 | 0.104 to 0.198 | |
| league only: ladder value of the main league at 18 | 0.196 | 0.149 to 0.241 | |
| league only: ladder value of the main league at 19 | 0.318 | 0.269 to 0.365 | |
| league only: ladder value of the main league at 20 | 0.536 | 0.500 to 0.576 |
| Formula at | Difference in rank correlation | 95% range | Reading | Details |
|---|---|---|---|---|
| 18 | −0.159 | −0.203 to −0.113 | draft position ranks better | |
| 19 | −0.074 | −0.114 to −0.029 | draft position ranks better | |
| 20 | +0.001 | −0.040 to +0.041 | no detectable difference | |
| draftYear | −0.344 | −0.395 to −0.295 | draft position ranks better |
Whole classes: graduation by round
1,902 of 1,903 picks, goalies included. Draft position alone ranks NHL games through 23 at 0.588 (95% range 0.552 to 0.621).
Outcome: NHL regular-season games and points summed over every season up to and including the age-23 season (age = whole years on 15 September of the season's start year). Graduated = 82 NHL games within any two consecutive seasons, both no later than the age-23 season. gpAfter / ptsAfter at an age count only the seasons after that age's season.
| Round | Graduated by 23 | Played an NHL game | Mean NHL games to 23 | Details |
|---|---|---|---|---|
| 1 | 72.7% (197 of 271) | 98.2% | 210.8 | |
| 2 | 32.4% (89 of 275) | 70.2% | 74.7 | |
| 3 | 11.8% (32 of 270) | 47.8% | 27.3 | |
| 4 | 10.3% (28 of 272) | 36.4% | 25.0 | |
| 5 | 7.0% (19 of 271) | 31.4% | 16.1 | |
| 6 | 5.1% (14 of 272) | 22.1% | 12.0 | |
| 7 | 1.5% (4 of 271) | 17.0% | 5.1 |
Caveats on the whole-class count (2)
- a pick with no history and no birthdate takes draft year + 5 as his age-23 season, which is one or two seasons generous to an overager
- a name match can miss a player whose name is spelled differently in the class list and in the season files (Alex / Alexander, accents are folded); such a player would be counted as never having played
Graduation by score at 18
| Tenth of the score | Graduated by 23 | Scores |
|---|---|---|
| 1 | 3.3% (6 of 184) | 0.0 to 3.8 |
| 2 | 5.4% (10 of 184) | 3.9 to 8.2 |
| 3 | 10.3% (19 of 184) | 8.2 to 11.1 |
| 4 | 8.7% (16 of 184) | 11.1 to 14.1 |
| 5 | 15.2% (28 of 184) | 14.1 to 17.0 |
| 6 | 14.7% (27 of 184) | 17.0 to 19.6 |
| 7 | 20.1% (37 of 184) | 19.6 to 22.4 |
| 8 | 35.3% (65 of 184) | 22.4 to 25.5 |
| 9 | 45.6% (84 of 184) | 25.5 to 33.7 |
| 10 | 48.4% (89 of 184) | 33.7 to 154.7 |
The league ladder, re-measured
| League | Measured translation | 95% range | Site value | Site inside range | Details |
|---|---|---|---|---|---|
| AHL | 0.447 | 0.408 to 0.476 | 0.455 | yes | |
| NCAA | 0.369 | 0.301 to 0.414 | 0.383 | yes | |
| OHL | 0.273 | 0.238 to 0.302 | 0.283 | yes | |
| WHL | 0.250 | 0.214 to 0.322 | 0.280 | yes | |
| SHL | 0.480 | 0.365 to 0.648 | 0.570 | yes | |
| QMJHL | 0.229 | 0.198 to 0.300 | 0.267 | yes | |
| KHL | 0.657 | 0.504 to 0.765 | 0.772 | no | |
| Liiga | 0.500 | 0.351 to 0.579 | 0.491 | yes | |
| HockeyAllsvenskan | 0.718 | 0.229 to 0.989 | 0.200 | no |
Caveats (10)
- Reference set for percentiles: the site ranks a prospect against the prospects in its own pools; the backtest ranks a player against every player whose history is on disk (the cohort plus other draft years and undrafted players) whose main line was in the same league that season and the same position group.
- Overlap: a score at 19 or 20 can contain NHL games (the NHL bonus, and NHL as the main league) that are also inside the games-through-23 outcome. Games and points counted only after the scoring season remove the overlap and are the fairer reading at 19 and 20.
- Percentile ties use average rank; the site's builder breaks ties by sort order. Clubs within one league-season are summed before the main line is chosen; the site picks the single club line with the most games.
- League labels are canonised with the site ladder builder's alias table before the ladder is applied (older seasons spell SHL as "Sweden", Liiga as "Finland", KHL as "Russia", NCAA as its conferences). The site's prospect builder does not canonise; on 2025-26 labels that makes no difference.
- The ladder applied to 2010-2021 seasons is the one measured in September 2026 (league-ladder.json), and its translations were estimated from histories that include these players: the ladder has seen this cohort's future. A clean test would re-measure the ladder on movers before 2010 only.
- Overagers: a player drafted at 19 or 20 has an age-18 season BEFORE his draft. His score at 18 is what the formula would have said, not what any page could have shown. The table for first-year-eligible skaters repeats the pooled table without them.
- Goalies are scored on save percentage and have no NHL bonus (as on the site). Many goalie lines outside North America carry no save percentage in the feed and those goalies have no score. The skaters tables are the ones to read for points.
- 64 scored lines carry the feed's label "U-18" or "U-17", which is the USNTDP club team. The site's builder has no ladder entry those labels contain, so they take the 0.2 fallback and not NTDP's 0.30; the backtest does the same.
- 3 scored lines at 18-20 have an international tournament as the line with the most games that season; they take the 0.2 fallback ladder value, as the site's builder would give them.
- The age curve is measured on points per game, while the formula applies its age factor to a percentile. A 20% rise in points per game is not a 20% rise in percentile; the comparison says whether the hand-set line has the right shape and rough size, not its exact value.
Study stamp 2026-09-29 13:57 UTC. Read from data/assay/prospect-backtest.json.
Contracts: what the market pays for, and how well a cap hit can be predicted
On 666 signings the market pays for production, games played, term and position; age, the UFA flag and a goalie's save performance add nothing once those are held (the fit explains 82% of the variation in cap share). Tested on 462 later deals, the fit's median error was 17.5% of the cap hit on terms it had seen, against 21.5% for the better of two simple baselines, and 51.5% on 1- and 2-year deals, which it had never seen.
Findings
- measured What the market pays for, other things held. Each extra 0.1 points per game in the season before: +12% (95% range +10% to +14%). Each extra tenth of the schedule played in the season before: +10% (95% range +8% to +12%). Each extra year of term: +17% (95% range +15% to +19%). A defenceman against a forward with the same points: +22% (95% range +13% to +31%).
- measured Not distinguishable from zero once those are held: a goalie's goals saved above average (p 0.98), age (p 0.31), age squared (p 0.39), the free-agent flag (p 0.84) and the second-contract flag (p 0.12).
- measured For goalies the market prices games played and not save performance: in four goalie-only fits on 67 deals the performance term was never distinguishable from zero (p 0.32 to 0.86).
- measured Backtest: fitted on 204 deals starting 2019-20 to 2024-25, tested on 462 deals starting later. On 3- to 8-year deals (201) the mean error was $1.24M, against $1.49M for production-quantile matching and $2.69M for the position median.
- caveat The training deals held no 1- and 2-year contracts, and on those the fit prices too high: 40% on 1-year deals (77) and 33% on 2-year deals (184).
- caveat The backtest prices a deal of known length. Without the term the mean error over all 462 test deals is $1.49M, against $1.06M with it.
- measured Bands: the 80% band held 58.0% of test deals (268 of 462): 80.6% on terms seen in training and 40.6% on terms not seen.
- measured Where the residuals lean: 55 groups tested, five survive false-discovery control: CAR (club) −16.6% against what the fit prices (22 deals), round-1 picks +7.0% against what the fit prices (247 deals), cap hits of the reported tier −8.6% against what the fit prices (154 deals), CHI (club) +44.6% against what the fit prices (15 deals) and undrafted players −11.5% against what the fit prices (89 deals).
- measured Nothing survives by nationality, market size or position.
- measured The site's constant, "+20% a point, 2024→2026": Reproduced (+19.6%). A rise in dollars is supported: 7 of 10 readings exclude zero (the site's own reading does not). A premium is not: no reading separates the rise from the cap's own +18.2%, and once terms are matched the open market rose as much. '+20% a point' describes what the salary cap did, not a second-contract premium.
- measured In numbers: the site's arithmetic gives +19.6% (95% range −3.3% to +39.8%) while the salary cap's upper limit grew +18.2%; measured against the cap the change is +1.2% (−18.2% to +18.3%).
- measured The ledger behind it: 1,205 rows, 449 without a cap hit, 76 labelled entry-level, 10 set aside for a signing date that cannot belong to their seasons, 666 fitted.
- caveat 14 rows labelled entry-level carry a term an entry-level contract cannot have, so the site's market panel skips second contracts it means to count.
Backtest against simple baselines
Fitted on 204 deals starting 2019-20 to 2024-25 (terms 3, 4, 5, 6, 7, 8 years only). Median error is the middle absolute error as a share of the real cap hit.
| Test deals | Deals | Fit: median error | Production-quantile match | Position median | Details |
|---|---|---|---|---|---|
| Both seasons, all test deals | 462 | 28.9% | 39.3% | 137.3% | |
| Both seasons, terms seen in training | 201 | 17.5% | 21.5% | 47.1% | |
| Both seasons, terms not seen | 261 | 51.5% | 86.0% | 469.6% | |
| 2025-26, all test deals | 205 | 22.4% | 33.2% | 93.5% | |
| 2025-26, terms seen in training | 103 | 15.5% | 19.4% | 46.3% | |
| 2025-26, terms not seen | 102 | 42.2% | 65.6% | 337.1% | |
| 2026-27, all test deals | 257 | 33.8% | 50.4% | 191.0% | |
| 2026-27, terms seen in training | 98 | 22.6% | 23.2% | 49.8% | |
| 2026-27, terms not seen | 159 | 53.4% | 100.5% | 580.7% |
| Term (years) | Deals | Seen in training | Fit prices too high by | Median error |
|---|---|---|---|---|
| 1 | 77 | no | +40% | 38.2% |
| 2 | 184 | no | +33% | 39.3% |
| 3 | 59 | yes | −10% | 23.7% |
| 4 | 35 | yes | −1% | 18.5% |
| 5 | 34 | yes | −5% | 18.6% |
| 6 | 26 | yes | +2% | 16.8% |
| 7 | 14 | yes | −3% | 9.5% |
| 8 | 33 | yes | +12% | 12.7% |
Band coverage
| First season | Inside the band | Terms seen | Terms not seen | Details |
|---|---|---|---|---|
| 2025-26 | 62.4% (128 of 205) | 81.6% of 103 | 43.1% of 102 | |
| 2026-27 | 54.5% (140 of 257) | 79.6% of 98 | 39.0% of 159 | |
| Both seasons | 58.0% (268 of 462) | 80.6% of 201 | 40.6% of 261 |
What the market pays for
| Other things held | Effect on cap hit | 95% range | p |
|---|---|---|---|
| each extra 0.1 points per game in the season before | +12% | +10% to +14% | below 0.001 |
| each extra tenth of the schedule played in the season before | +10% | +8% to +12% | below 0.001 |
| no NHL games in the season before | +32% | +15% to +52% | below 0.001 |
| each 10 goals saved above average, goalies | 0% | −8% to +8% | 0.979 |
| each extra tenth of the schedule played, goalies | +23% | +17% to +29% | below 0.001 |
| each year of age, at 27 | +1% | −1% to +3% | 0.311 |
| curvature in age (per squared year from 27) | 0% | 0% to 0% | 0.394 |
| each extra year of term | +17% | +15% to +19% | below 0.001 |
| a defenceman against a forward with the same points | +22% | +13% to +31% | below 0.001 |
| a goalie against a forward | +55% | +20% to +100% | below 0.001 |
| unrestricted against restricted free agent | −1% | −9% to +8% | 0.841 |
| a second contract (24 or under) | +11% | −3% to +27% | 0.121 |
Where the residuals lean
| Group | Paid against the fit | 95% range | Deals | Details |
|---|---|---|---|---|
| CAR (club) | −16.6% | −24.5% to −7.9% | 22 | |
| round-1 picks | +7.0% | +2.6% to +11.4% | 247 | |
| cap hits of the reported tier | −8.6% | −13.2% to −3.8% | 154 | |
| CHI (club) | +44.6% | +19.1% to +77.6% | 15 | |
| undrafted players | −11.5% | −18.9% to −4.0% | 89 |
Caveats (19)
- The ledgers are a snapshot of contracts on the books, not a history of signings: a deal from an earlier season is on file only if its term carried it to 2026-27. The shortest term on file is 8 years for 2019-20, 7 years for 2020-21, 6 years for 2021-22, 5 years for 2022-23, 4 years for 2023-24, 2 years for 2024-25, 2 years for 2025-26 and 1 year for 2026-27. Older cohorts are long, expensive deals by construction.
- So the training deals (first seasons to 2024-25, n = 204) carry terms 3, 4, 5, 6, 7 and 8 only and 4 of them are under $1M, while 261 of the 462 test deals have a term the fit never saw. Errors on those deals measure extrapolation as much as pricing; every backtest figure is given for both groups.
- Both naive baselines learn from the same long, expensive training deals, which makes them easy to beat on the whole test set. The margin inside the terms seen in training is the honest one.
- Bought-out, terminated and retired contracts have left the ledgers, so what remains from earlier summers leans toward the deals that worked.
- The backtest takes the deal's own term as an input. It prices a deal of known length; it does not forecast the length.
- Term enters as a straight line, as specified. The residuals show it is not one: three- and four-year deals sit above the line, one-, two-, seven- and eight-year deals below. Every tested group in the bias ledger is rechecked with term as categories.
- 256 of the 666 fitted deals were signed a full season before they start. For those the 'season before the first season' had not been played when the price was agreed; the specified definition is kept for the main fit and the dated-at-signing fit is reported beside it.
- 10 rows carry a signing date that cannot belong to the seasons on the row; they are set aside, not corrected. Other rows may be off by one season in ways this desk cannot detect without the source pages.
- 14 rows labelled entry-level carry a term an entry-level contract cannot have.
- The UFA flag is the ledger's expiry status where present (372 deals) and an age proxy elsewhere (294). expiry status is the player's status when the deal ends, not when he signed it, and it disagrees with the age-27 rule on 107 deals. The coefficient on the flag should not be read as the price of free agency.
- Production is points per game and games played for skaters, goals saved above average and games for goalies. Nothing here measures defence, role, injury history, signing bonuses, trade protection or tax; all of that is in the residual.
- The league minimum is a floor: 108 fitted deals sit under $1M and the log-linear form prices them only roughly. Every tested group in the bias ledger is rechecked without them.
- Nationality is on file for 479 of 666 deals (careers.json covers regulars) and draft round for 644. Players without a record are listed and not tested.
- Club is the club holding the contract in September 2026, not always the club that signed it. A club lean can be a trading pattern as much as a signing pattern.
- Residuals in the bias ledger are in-sample. Intervals are percentile bootstraps of the mean (2,000 resamples), p-values are one-sample t, q-values are Benjamini-Hochberg over every tested group at once. Standard errors are robust to unequal variance, not to clustering by club.
- Market tiers come from market-size.json, whose populations were typed from general knowledge and carry their own citation caveat.
- Cap hits for two acquired contracts (Kadri, Karlsson) are restored to the signed average from the ledger's own notes; the ledger stores the acquiring club's post-retention figure.
- The 2026-27 upper limit is the announced figure and 2027-28 is a joint league and union projection; cap shares for extensions use the projection.
- Predicted cap hit is exp(predicted log share) x upper limit, the conditional median. It is not a mean: summed across a roster it would understate payroll.
Study stamp 2026-09-29 12:51 UTC. Read from data/assay/pricing-backtest.json.
Ask the CBA: evaluation of the grounded assistant
Retrieval only: the right passage is in the top five for 81.1% of 90 answerable questions (95% range 71.8% to 87.9%) and first for 57.8%; asked in a fan's words the top-five figure is 30% (10 pairs). On questions written after the ranking was frozen the top-five figure falls to 53.3% (60 questions). The model's answers have not been scored (not run: no credit and the hub does not allow model spend); while the model is dark the fallback answered 19 of 30 unanswerable questions with quoted passages instead of refusing.
Findings
- measured 160 questions were written against the corpus of 1,269 passages: 90 answerable, 30 unanswerable, 20 adversarial and 20 paraphrases in 10 pairs.
- measured Recall at 1, 3, 5 and 10: 57.8%, 76.7%, 81.1% and 86.7%. Mean reciprocal rank 0.683 (95% range 0.599 to 0.760). A passage that states the answer reaches the model for 92.2% of questions.
- measured Weakest topics at five: long-term injury 50% (3 of 6), minimum salary 50% (3 of 6), performance bonuses 60% (3 of 5) and waivers 67% (4 of 6). Each rests on a handful of questions: they say where to look.
- measured Paraphrase: on the same 10 facts the top five holds the right passage for 80% of questions in the agreement's words and 30% in a fan's words; the two wordings share a first passage on 10% of pairs.
- measured Refusing by retrieval score alone: the best threshold refuses 26 of 30 unanswerable questions (86.7%, range 70.3% to 94.7%) and wrongly refuses 16 of 90 answerable ones (17.8%, range 11.3% to 26.9%). 25 of the unanswerable questions score above the lowest-scoring answerable one, so no threshold separates them.
- measured Mirror check: on 15 questions sent to the live route the top five passages matched on 100%, all passages handed over on 100% and the fallback text on 100%; 21 of 21 mirrored constants match the app's source.
- caveat The fallback does not refuse: it answered 19 of 30 unanswerable questions with quoted passages. The quotes are real; they are not answers.
- measured One defect in the app's own code was found while measuring and is listed below; the study changed nothing in the app.
- measured Against the earlier run: recall at five was 81.1% and is 81.1%.
- measured Held-out questions (60 answerable, written after the ranking was frozen): the right passage is in the top five for 53.3% and first for 25.0%, against 81.1% on the questions the ranking was developed on.
- decision Six changes to the ranking were tried against the held-out questions; none was adopted, so the site keeps its ranking.
- decision Model and judge: not run: no credit and the hub does not allow model spend.
Retrieval
| Within the top | Right passage found | Questions | 95% range |
|---|---|---|---|
| 1 | 57.8% | 52 of 90 | 47.5% to 67.5% |
| 3 | 76.7% | 69 of 90 | 67.0% to 84.2% |
| 5 | 81.1% | 73 of 90 | 71.8% to 87.9% |
| 8 | 85.6% | 77 of 90 | 76.8% to 91.4% |
| 10 | 86.7% | 78 of 90 | 78.1% to 92.2% |
Mean reciprocal rank 0.683 (95% range 0.599 to 0.760). A passage that states the answer reaches the model for 92.2% of questions. The range on recall at 5 is the desk's. The others are Wilson 95% ranges computed here from the desk's counts; the same arithmetic reproduces the desk's range.
| Topic | Top five | First | Questions |
|---|---|---|---|
| long-term injury | 50% | 33% | 6 |
| minimum salary | 50% | 50% | 6 |
| performance bonuses | 60% | 40% | 5 |
| waivers | 67% | 17% | 6 |
| no-move and no-trade clauses | 80% | 60% | 5 |
| buy-outs | 83% | 50% | 6 |
| term limits | 83% | 50% | 6 |
| free-agency groups | 83% | 50% | 6 |
| arbitration | 83% | 50% | 6 |
| qualifying offers | 83% | 67% | 6 |
| entry-level contracts | 86% | 71% | 7 |
| retained salary | 100% | 60% | 5 |
| 2025 MOU changes | 100% | 67% | 6 |
| expansion | 100% | 67% | 3 |
| escrow | 100% | 100% | 4 |
| payroll range | 100% | 100% | 6 |
| offer sheets | 100% | 100% | 1 |
Questions written after the ranking was frozen
| Change | Decision | Why |
|---|---|---|
| R1 | rejected | recall@5 difference -0.0167 (95% -0.05 to 0), not above zero; MRR difference -0.0023 (95% -0.0581 to 0.0522), not above zero; original 90 recall@5 84.4% against a floor of 79.1% |
| R2 | rejected | recall@5 difference -0.05 (95% -0.1333 to 0.0167), not above zero; MRR difference -0.0145 (95% -0.0883 to 0.0573), not above zero; original 90 recall@5 83.3% against a floor of 79.1% |
| R3 | rejected | recall@5 difference -0.0167 (95% -0.05 to 0), not above zero; MRR difference -0.005 (95% -0.0242 to 0.008), not above zero; original 90 recall@5 80.0% against a floor of 79.1% |
| R3B | rejected | recall@5 difference 0 (95% 0 to 0), not above zero; MRR difference -0.0005 (95% -0.0012 to 0), not above zero; original 90 recall@5 81.1% against a floor of 79.1% |
| R4 | rejected | recall@5 difference 0 (95% 0 to 0), not above zero; MRR difference -0.0086 (95% -0.0401 to 0.0201), not above zero; original 90 recall@5 80.0% against a floor of 79.1% |
| R5 | rejected | recall@5 difference -0.0167 (95% -0.0667 to 0.0333), not above zero; MRR difference -0.0051 (95% -0.0366 to 0.0251), not above zero; original 90 recall@5 78.9% against a floor of 79.1%: BELOW THE FLOOR |
Refusing by retrieval score
| Threshold chosen for | Unanswerable refused | Answerable refused | Details |
|---|---|---|---|
| best threshold | 26 of 30 | 16 of 90 | |
| wrong refusals held to 5% | 15 of 30 | 2 of 90 | |
| wrong refusals held to 10% | 19 of 30 | 9 of 90 | |
| refusing 90% of unanswerable | 27 of 30 | 20 of 90 | |
| refusing all unanswerable | 30 of 30 | 49 of 90 |
Defects found in the app
- The contract-term rule, /\bextension\b|\bterm\b|how long|length|how many years|.../, appends "SPC term years extension greater than six seven years". It fired on 24 of the 110 answerable and paraphrase questions, triggered by "how long" (12), "term" (7), "length" (1), "how many years" (2), "extension" (2). "\bterm\b" matches inside "long-term", so 6 long-term-injury questions are searched as if they were about contract length, and "how long" does the same to "How long is the waiver claim period?". With the rule the gold chunk ranks lower on 13 of those 24 and higher on 2 (sign test p = 0.0074). Removing this rule alone brings 3 into the top 5 (ans-006, pp-09-a, pp-09-b) and drops 0.
Caveats (14)
- Model path not exercised: the spend ledger (2026-09-29T12:35:00.563Z) reads "exhausted" ($-0.12 of $55 left), and all 15 live calls to /api/ask came back "extractive" (14), "not-in-sources" (1). model and judge stay null until a run with credit.
- A hit means one of the item's 1-3 gold chunk ids is in the top k. Gold lists were checked against what the ranking returned for every missed item, and a chunk that states the answer was added (31 items have more than one gold chunk).
- Reachable is the ceiling for the model: a passage stating the answer is among the 8 corpus passages or the 3 mechanics passages the route hands it. 92.2% of answerable items reach it (85.6% through the corpus passages alone). For 7 answerable questions no passage on the item's gold list is handed over (ans-009, ans-012, ans-033, ans-056, ans-057, ans-061, ans-066); the 11 passages handed over for each were read, and none is the passage that settles the question, so a correct answer cited to the right section is not available to the model today. Section 9.4 (cba-59, and mou2025-1053 after it) gives the same $35,000 floor for the entry-level default. It is a different rule from 11.12(b), so it is not gold; an answer that cites it has the right figure from the wrong section. One of the seven shows why chunking matters: section 50.10(d)(iii) is cut mid-sentence between cba-730 and cba-731, and only the first half is retrieved.
- By the document the gold sits in: mou2025 recall@5 79.2%, recall@1 58.3% (n=24); cba recall@5 81.4%, recall@1 54.2% (n=59); mou2020 recall@5 85.7%, recall@1 85.7% (n=7). The 2025 MOU is a scan whose OCR runs words together ("NHLMiniinunSalary", "totaloffive(5)regularRecalls"); a run-together word is one token that no question contains. The intervals overlap, so this sample does not show that the documents differ.
- VOCAB expansion, measured on the same 90 questions: recall@5 81.1% with it, 86.7% without; MRR 0.6831 with, 0.7118 without. It brought 0 question(s) into the top 5 and pushed 5 out (sign test p = 0.0625); by rank, the gold chunk sits lower with the expansion on 12 questions and higher on 3 (p = 0.0352). Pushed out: ans-006, ans-016, ans-056, ans-061, ans-064. Caveat: these questions were written from the agreement's text, so they already use its words, and the expansion exists for people who do not. On the ten fan-worded paraphrases recall@5 is 30.0% with the expansion and 20.0% without (a rule fired on 5 of the ten): ten questions cannot settle it.
- Document-rank multiplier, same 90 questions: recall@5 81.1% with it, 78.9% without; MRR 0.6831 with, 0.6691 without (sign test p = 0.5). Top-1 passage by document: mou2025 23, mou2020 11, cba 56; gold by document: mou2025 24, cba 59, mou2020 7.
- Planted text and retrieval, each adversarial question against the same question asked plainly: planted-instruction recall@8 80.0% against 70.0% clean (gold rank worse on 4 of 10, better on 4, sign test p = 1); false-premise recall@8 100.0% against 90.0% clean (worse on 1, better on 3, p = 0.625). The planted words are query terms like any others.
- Paraphrase: wording "a" stays close to the agreement's own words, wording "b" is how a fan would ask. Recall@5 80.0% for "a" and 30.0% for "b" on the same ten facts. BM25 matches words, not meaning, and the VOCAB list covers only the 19 phrasings someone thought of.
- Refusal by score: the threshold 21.4329 was chosen on the same 120 items it is scored on (TPR 86.7%, FPR 17.8%). Leave-one-out gives TPR 86.7% and FPR 17.8% (the threshold did not move in any fold). AUC 0.8985 needs no threshold. The 30 unanswerable questions were written for this set and most are plainly off the subject; a user who asks in the agreement's own vocabulary about something it does not settle is the hard case, and a score cannot tell "the corpus discusses this" from "the corpus answers this". A question about a named player's no-move clause already scores higher than many questions the corpus does answer.
- The passages-only fallback now grades its evidence (register cba-fallback-refusal-v1): a question naming a current player or club is answered "not-in-sources" with a pointer to that club's desk; weak evidence keeps the quotes behind one warning sentence; the evidence level for "not-in-sources" is off because it failed its registered held-out guard. Unanswerable questions answered anyway: 19 of 30 (before: 30 of 30). Answerable questions refused: 0 of 90. A weak-evidence answer still counts as answered here; the reader sees the warning.
- Clause split: now /(?<=;)\s+/ (was /(?<=;)s+/, a literal "s"), so the fallback quotes clauses of semicolon lists.
- Number gate: now an exact test on number tokens compared as canonical decimals (8.50 = 8.5; 5 is not supported by 50 or 2025; a season label is one token). Of the 7 adversarial items whose planted number the substring gate let through (adv-fp-01, adv-fp-03, adv-fp-07, adv-fp-08, adv-pi-01, adv-pi-07, adv-pi-08), 0 still pass the exact gate.
- Small samples throughout. Intervals are Wilson 95%. Topic figures rest on 1 to 7 questions each: they say where to look, they do not measure a topic. The sign tests were run after looking at the data and are not adjusted for how many were run; a p-value near 0.05 here is a lead to follow, not a finding.
- The questions were written by reading the corpus, so they lean on the agreement's own words and on facts that sit in one place. Recall on questions real users type will be lower than recall here; the fan-worded paraphrases (30.0% at 5) are the closer guide.
Study stamp 2026-09-29 14:08 UTC. Read from data/assay/assistant-eval.json, data/assay/assistant-eval-mirror.json, data/assay/assistant-eval-items.json, data/assay/assistant-eval-heldout.json, data/assay/assistant-eval-heldout-results.json.
Start +/-: how much the sequence window matters
The window matters: top-30 rank-biased overlap is 0.596 between 15 s and 35 s and 0.8 between 25 s and 35 s (1 = identical lists), and the number of goals that get a start moves from 5127 to 7082. Start +/- is therefore shown with its window named, and the card treats 25 s as a convention, not a finding.
Findings
- measured Goals that get a start: 5,127 at 15 s, 6,369 at 25 s and 7,082 at 35 s, over 695 players in 2025-26.
- measured Top-thirty overlap between windows (1 means identical lists): 15 s against 25 s, 0.571, 15 s against 35 s, 0.596 and 25 s against 35 s, 0.800.
- caveat One entry in the top-ten lists is a player id with no name on file.
| Window | Goals with a start | Players | Top three |
|---|---|---|---|
| 15 s | 5,127 | 695 | Jason Robertson 53, Jack Eichel 48, Nathan MacKinnon 46 |
| 25 s | 6,369 | 695 | Nathan MacKinnon 67, Martin Necas 64, Jack Eichel 61 |
| 35 s | 7,082 | 695 | Nathan MacKinnon 77, Martin Necas 73, Jason Robertson 64 |
| Player | Narrowest window | Widest window | Moved |
|---|---|---|---|
| Andrei Vasilevskiy | 14 | 48 | 34 |
| Marcus Pettersson | -30 | -62 | -32 |
| Nathan MacKinnon | 46 | 77 | 31 |
| Martin Necas | 45 | 73 | 28 |
| Nikita Kucherov | 29 | 56 | 27 |
| Scott Wedgewood | 28 | 54 | 26 |
| Kirill Kaprizov | 26 | 51 | 25 |
| Filip Forsberg | 23 | 47 | 24 |
Study stamp 2026-09-29 12:11 UTC. Read from data/assay/sensitivity-start.json.
Video reads: today's blind review sample
Today's blind sample: 511 read windows drawn from 1,566 in 32 games, up to 40 per event type, listed without the model's events so the reviewer is not led. No human verdicts are on file yet, so the precision of the video reads is unmeasured.
Findings
- measured 17 event types; 10 have the full 40 windows in the sample, the rest have fewer windows than that on file.
- measured The sample touches 32 of the 32 games read. It is the same every time it is drawn on the same day (seed 20260929).
- caveat About 138 reviewed windows per event type bound precision within ±5 points at 80% confidence when precision is near 0.7; two reviewers on the same windows give Cohen's kappa for the task ceiling.
| Event type | In the sample | Windows read | Details |
|---|---|---|---|
| shot | 40 | 980 | |
| save | 40 | 894 | |
| board battle won | 40 | 871 | |
| zone entry carried | 40 | 651 | |
| faceoff won | 40 | 597 | |
| other | 40 | 471 | |
| (no events reported) | 40 | 460 | |
| goal | 40 | 312 | |
| zone exit with possession | 40 | 273 | |
| clearance | 40 | 134 | |
| penalty | 36 | 36 | |
| takeaway | 24 | 24 | |
| zone entry dumped | 23 | 23 | |
| giveaway | 20 | 20 | |
| zone entry denied at the blue line | 6 | 6 | |
| entry denied at the blue line | 1 | 1 | |
| blocked shot | 1 | 1 |
How to review: Open /video/<gameId> and jump to the window's start second. Record every entry (carried/dumped/denied), exit, takeaway, giveaway, board battle, shot and save you see, with jersey numbers, BEFORE revealing the model's events. Save the flags; the morning build applies them (wrong → drop, edited → retype, correct → confidence 1, missed → add) and reports precision and recall per type on /integrity.
Study stamp 2026-09-29 12:20 UTC. Read from data/condensed/review-sample-2026-09-29.json.
Career arcs: aging curves, draft projections and career labels
Corrected for players who leave the league, forwards peak at 24 and defencemen at 26 in points per 82 games; goalies at 24 in save % above league, on far fewer seasons. Shipped for skaters: flash in the pan, sustained breakout and late bloomer, and exceeded ceiling at the draft. Not shipped: early peak, fell short, projections at 20 and 22, and every goalie label.
Findings
- measured Forwards: scoring rate (points per 60) peaks at 24 (95% range 23-24) and then changes by -0.052 a year over the next 5 years (-2.5% of the peak level a year); ignoring the players who drop out would put the peak at 24 with a decline of -0.041 a year.
- measured Forwards: ice-time share peaks at 26 (95% range 25-27) and then changes by -0.00529 a year over the next 5 years (-2.0% of the peak level a year); ignoring the players who drop out would put the peak at 26 with a decline of -0.00323 a year.
- measured Forwards: Drive peaks at 24 (95% range 19-25) and then changes by -0.00428 a year over the next 5 years (-27.1% of the peak level a year); ignoring the players who drop out would put the peak at 24 with a decline of -0.00292 a year.
- measured Defencemen: scoring rate (points per 60) peaks at 23 (95% range 19-26) and then changes by -0.00854 a year over the next 5 years (-0.8% of the peak level a year); ignoring the players who drop out would put the peak at 23 with a decline of -0.00391 a year.
- measured Defencemen: ice-time share peaks at 26 (95% range 25-27) and then changes by -0.00326 a year over the next 5 years (-1.0% of the peak level a year); ignoring the players who drop out would put the peak at 27 with a decline of -0.00349 a year.
- measured Defencemen: Drive peaks at 24 (95% range 20-28) and then changes by -0.00223 a year over the next 5 years (-41.1% of the peak level a year); ignoring the players who drop out would put the peak at 24 with a decline of -0.00228 a year.
- measured Goalies: save % above league peaks at 24 (95% range 24-25) and then changes by -0.237 a year over the next 5 years (-23.3% of the peak level a year); ignoring the players who drop out would put the peak at 24 with a decline of -0.151 a year.
- measured A season 2 or more standard deviations above what a player's earlier seasons and age predicted came back further than ordinary regression to the mean expects: in 120 such seasons from 2018-19 on, the next season landed 0.28 standard deviations below the regression-to-the-mean forecast (95% range 0.07-0.49 below). 'Flash in the pan' is shipped.
- measured Sustained breakout: 42 label-holders tested on seasons from 2019-20 on; the season after the label sat +1.52 SD from the old expectation (95% range 1.32-1.72) and -0.14 SD from a regression-to-the-mean forecast that already knew the new level. Shipped.
- measured Late bloomer: 34 label-holders tested on seasons from 2019-20 on; the season after the label sat +1.24 SD from the old expectation (95% range 0.99-1.48) and -0.12 SD from a regression-to-the-mean forecast that already knew the new level. Shipped.
- measured Early peak: 12 label-holders tested on seasons from 2019-20 on; the season after the label sat -1.12 SD from the old expectation (95% range -1.38--0.86) and +0.36 SD from a regression-to-the-mean forecast that already knew the new level. Not shipped (only 12 proof label-holders (< 20)).
- measured Projection at the draft (NHL ice time through age 25), held-out draft classes 2015-2018: 10% landed above the 90th percentile and 14% of 124 players with a non-zero floor landed below the 10th. Shipped.
- measured Projection at 20 (NHL ice time through age 25), held-out draft classes 2015-2018: 14% landed above the 90th percentile and 21% of 292 players with a non-zero floor landed below the 10th. Not shipped: the floor is set too high.
- measured Projection at 22 (NHL ice time through age 25), held-out draft classes 2015-2018: 11% landed above the 90th percentile and 20% of 341 players with a non-zero floor landed below the 10th. Not shipped: the floor is set too high.
- measured Goalies, tested on their own (registered careers-goalie-labels-v1): Flash in the pan: not shipped: the regression-to-the-mean forecaster's one-step error SD on goalie proof seasons is 1.16 (274 seasons; must lie in 0.85-1.15). Sustained breakout: not shipped: the regression-to-the-mean forecaster's one-step error SD on goalie proof seasons is 1.16 (274 seasons; must lie in 0.85-1.15). Late bloomer: not shipped: the regression-to-the-mean forecaster's one-step error SD on goalie proof seasons is 1.16 (274 seasons; must lie in 0.85-1.15). Early peak: not shipped (not shipped for any position). Exceeded ceiling: not shipped: 18% of 92 held-out goalies (draft classes 2015-2018) landed above the draft projection's 90th percentile, against 10% nominal (allowed 5-15%); goalie draft projections are not shown.
- measured Bias checks: 251 tests across nationality, draft round, birth quarter, league of origin and market, 27 survive false-discovery control at 10% (at the draft checkpoint: 2).
| Label | Position | Shipped | Decision |
|---|---|---|---|
| flash in the pan | Forwards | yes | SHIP: spikes regress more than regression to the mean predicts |
| sustained breakout | Forwards | yes | proof mean z range above 0 |
| late bloomer | Forwards | yes | proof mean z range above 0 |
| early peak | Forwards | no | only 12 proof label-holders (< 20) |
| flash in the pan | Defencemen | yes | SHIP: spikes regress more than regression to the mean predicts |
| sustained breakout | Defencemen | yes | proof mean z range above 0 |
| late bloomer | Defencemen | yes | proof mean z range above 0 |
| early peak | Defencemen | no | only 12 proof label-holders (< 20) |
| flash in the pan | Goalies | no | not shipped: the regression-to-the-mean forecaster's one-step error SD on goalie proof seasons is 1.16 (274 seasons; must lie in 0.85-1.15). |
| sustained breakout | Goalies | no | not shipped: the regression-to-the-mean forecaster's one-step error SD on goalie proof seasons is 1.16 (274 seasons; must lie in 0.85-1.15). |
| late bloomer | Goalies | no | not shipped: the regression-to-the-mean forecaster's one-step error SD on goalie proof seasons is 1.16 (274 seasons; must lie in 0.85-1.15). |
| early peak | Goalies | no | not shipped (not shipped for any position). |
| exceeded ceiling (draft) | Forwards | yes | unknown |
| exceeded ceiling (draft) | Defencemen | yes | unknown |
| exceeded ceiling (draft) | Goalies | no | not shipped: 18% of 92 held-out goalies (draft classes 2015-2018) landed above the draft projection's 90th percentile, against 10% nominal (allowed 5-15%); goalie draft projections are not shown. |
| Checkpoint | Landed above the 90th percentile | Landed below the 10th (non-zero floor) | Shipped |
|---|---|---|---|
| At the draft | 10.4% | 13.7% | yes |
| At 20 | 13.7% | 20.9% | no |
| At 22 | 10.9% | 20.2% | no |
The full desk, with every chart and list, is on /careers.
Caveats (1)
- Birth dates cover 32.4% of goalie shots against 89.8% of skater ice time, so goalie curves rest on a few hundred seasons.
Study stamp 2026-09-29 22:04 UTC. Read from data/assay/careers.json.
Pre-registration register
Every model change is written down before it runs: hypothesis, metric, proof seasons and decision rule. Outcomes are recorded whatever they were. Adoption of a challenger is a decision card in the hub.
| Entry | Registered | Status | Hypothesis | Outcome | Details |
|---|---|---|---|---|---|
| careers-goalie-labels-v1 | 2026-09-29 | not-shipped | G, goalie-specific arc labels and ceiling label. REGISTERED AFTER the pooled label and ceiling results (all positions together) were seen; the goalie-only proof numbers below had NOT been computed. Diagnosis that prompted it (development data and code only): (1) the labels were validated with K built on the development goalie aging curve, which is empty (no age step with >= 30 pairs before 2018, so zero drift), but applied on the site with the all-season curve (pattern-mixture, peak 24, -0.24 save-% points a year, from 477 qualifying goalie-seasons with a known age; the mixed model on the same rows says -0.02 a year). That drift pulled a 31-year-old's expectation down ~1.5 points, so 'expected' in a goalie sentence was on a different scale from the one validated (Hellebuyck 2024-25: expected -0.07 on the site, +0.77 under the validated K). (2) The maximum-likelihood tau^2 sat on the lower edge of the search grid. (3) Goalie label-holders were a small share of the pooled validation, so no goalie label was ever tested on goalies. (4) Goalie draft-checkpoint upper-tail misses run 18-20% against a nominal 10%, and the position check only asked the Wilson interval to overlap the window. K_G: zero aging drift for goalies everywhere (fit, labels, site bands); tau^2, omega^2 by maximum likelihood on seasons <= 2018 over an extended grid (tau^2 0.002-1.5 x level variance, omega^2 0.0005-0.3 x); observation variance binomial by shots faced (shot-volume-aware shrinkage, unchanged). Label rules and thresholds unchanged (z 2.0 spike, 1.0 flash, 1.5 sustained / late bloomer). | calibration proofSeasons: 274, residualSd: 1.162, window: 0.85, 1.15, ok: no, residualMean: -0.156 · flash beta: 0.602, devSpikePairs: 2, devSpikeMeanResidual: 0.602, devRange95: 0.602, 0.602, proofPairs: 274, proofSpikePairs: 2, proofSpikeMeanResidual: 0.173, proofSpikeRange95: 0.007, 0.338, proofNonSpikeMeanResidual: -0.159, proofMseImprovement: -0.001, proofMseImprovementRange95: -0.004, 0, mseA: 1.375 … +3 fields · sustained dev: (n: 0), proof: (n: 0), ship: no, decision: only 0 proof label-holders (< 20) · lateBloomer dev: (n: 0), proof: (n: 3, players: 3, meanZvsPreWindow: 0.745, range95: -0.439, 1.607, shareAbove0: 0.962, meanZvsFullK: 0.276, fullKRange95: -0.9, 1.39), ship: no, decision: only 3 proof label-holders (< 20) · ship flashInThePan: no, sustainedBreakout: no, lateBloomer: no, earlyPeak: no · kalman measure: savePctAboveLeague, tau2: 0.007, omega2: 0.023, levelVar: 1.23, atGridEdge: no, drift: none (careers-goalie-labels-v1), bloomCut75: 0.277, curveEstimator: pm, players: 153 | |
| careers-checkpoint-projection-v5 | 2026-09-29 | registered | P, next blind test (registered 2026-09-29, NOT run: its proof data do not exist yet). After v2-v4 all failed the lower tail at ages 20 and 22 on the 2015-2018 classes with the median too high (PIT mean 0.37-0.41), while the draft checkpoint passed: those classes played ages ~21-25 through the shortened 2019-20 (~70 games) and 2020-21 (56 games) schedules, which a model fitted on 2010-2014 cannot know. v5 = v4 (increment target, Mondrian CQR) with every NHL season's TOI, games and points scaled to an 82-game schedule (x 82 / club games that season) in both features and outcomes, fitted on classes 2010-2018. | not run yet | |
| careers-checkpoint-projection-v4 | 2026-09-29 | shipped | P, structural repair. REGISTERED AFTER the v2 and v3 proof results were seen (v3: lower-tail misses 16.5% at 20 and 16.8% at 22). A development-only diagnosis (leave-one-class-out, 2010-2014) shows the cause: the models predict cumulative totals and are then floored at the minutes already banked; for about half the rows the raw 10th percentile is below what is banked, so the floor lifts the median (P(y < q50) 0.48 -> 0.56 at 20, 0.51 -> 0.58 at 22) and raw quantile boosting under-covers the lower tail (35-38% misses). v4 models the INCREMENT after the checkpoint (>= 0) and adds the banked amount back; Mondrian CQR as v3. Models, features and model choice otherwise unchanged. | draft proof: (players: 855, upperMiss: 0.095, lowerMiss: 0.089, lowerEligible: 124, coverage80: 0.892, pitMean: 0.488, pitHist10: 75, 101, 94, 98, 59, 118, 64, 109, 56, 81, pinball: 5.81, pinballQ0: 5.81, gainOverQ0: (mean: 0, range95: 2 items, shareAbove0: 0), spearmanMedianVsOutcome: 0.479), dev: (players: 1047, upperMiss: 0.098, lowerMiss: 0.104, lowerEligible: 154, coverage80: 0.886, pitMean: 0.485, pitHist10: 99, 108, 140, 98, 87, 147, 64, 123, 78, 103, pinball: 6.36, pinballQ0: 6.36, gainOverQ0: (mean: 0, range95: 2 items, shareAbove0: 0), spearmanMedianVsOutcome: 0.518), ship: yes, decision: PROVISIONAL SHIP (dev and non-blind proof tails) · age20 proof: (players: 854, upperMiss: 0.126, lowerMiss: 0.191, lowerEligible: 278, coverage80: 0.811, pitMean: 0.404, pitHist10: 194, 154, 63, 76, 35, 72, 47, 60, 45, 108, pinball: 4.842, pinballQ0: 5.408, gainOverQ0: (mean: 0.567, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.58), dev: (players: 1046, upperMiss: 0.126, lowerMiss: 0.228, lowerEligible: 351, coverage80: 0.797, pitMean: 0.402, pitHist10: 294, 171, 74, 50, 36, 68, 70, 74, 76, 133, pinball: 5.01, pinballQ0: 5.923, gainOverQ0: (mean: 0.913, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.608), ship: no, decision: NOT SHIPPED (dev and non-blind proof tails) · age22 proof: (players: 855, upperMiss: 0.101, lowerMiss: 0.171, lowerEligible: 334, coverage80: 0.833, pitMean: 0.37, pitHist10: 232, 154, 65, 69, 47, 59, 49, 46, 48, 86, pinball: 2.587, pinballQ0: 4.067, gainOverQ0: (mean: 1.48, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.752), dev: (players: 1047, upperMiss: 0.124, lowerMiss: 0.163, lowerEligible: 424, coverage80: 0.81, pitMean: 0.381, pitHist10: 280, 191, 85, 73, 54, 67, 46, 60, 61, 130, pinball: 2.803, pinballQ0: 4.326, gainOverQ0: (mean: 1.523, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.774), ship: no, decision: NOT SHIPPED (dev and non-blind proof tails) · shippedFrom draft: (ship: yes, mode: v2, provisional: no), age20: (ship: no, mode: v2, provisional: no), age22: (ship: no, mode: v2, provisional: no) · evaluatedAt 2026-09-29T22:04:29Z | |
| careers-checkpoint-projection-v3 | 2026-09-29 | not-shipped | P, calibration repair. REGISTERED AFTER v2's proof result was seen: v2 failed the lower-tail test at age 20 (21% misses) and age 22 (20%); the development folds show the same (~25%), so the cause is structural: the single pooled conformal correction is set by the mass of players at 0 NHL minutes and never moves the lower bound of players already on an NHL path. v3 = v2 with Mondrian (group-conditional) CQR: separate corrections per checkpoint x target x target age x group, group = (q10 > 0 or not) for the lower bound and tercile of q90 for the upper, each from development leave-one-class-out scores only (groups under 30 rows fall back to the pooled correction). Models, features and choice unchanged. | draft proof: (players: 855, upperMiss: 0.095, lowerMiss: 0.089, lowerEligible: 124, coverage80: 0.892, pitMean: 0.488, pitHist10: 75, 101, 94, 98, 59, 118, 64, 109, 56, 81, pinball: 5.81, pinballQ0: 5.81, gainOverQ0: (mean: 0, range95: 2 items, shareAbove0: 0), spearmanMedianVsOutcome: 0.479), dev: (players: 1047, upperMiss: 0.098, lowerMiss: 0.104, lowerEligible: 154, coverage80: 0.886, pitMean: 0.485, pitHist10: 99, 108, 140, 98, 87, 147, 64, 123, 78, 103, pinball: 6.36, pinballQ0: 6.36, gainOverQ0: (mean: 0, range95: 2 items, shareAbove0: 0), spearmanMedianVsOutcome: 0.518), ship: yes, decision: SHIP: upper miss 0.095 (in [0.05,0.15]), lower miss 0.089 on 124 eligible (ok) · age20 proof: (players: 854, upperMiss: 0.122, lowerMiss: 0.165, lowerEligible: 254, coverage80: 0.829, pitMean: 0.406, pitHist10: 204, 147, 59, 72, 30, 81, 49, 56, 52, 104, pinball: 4.884, pinballQ0: 5.782, gainOverQ0: (mean: 0.898, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.581), dev: (players: 1046, upperMiss: 0.125, lowerMiss: 0.224, lowerEligible: 331, coverage80: 0.804, pitMean: 0.403, pitHist10: 298, 163, 75, 53, 33, 66, 66, 85, 76, 131, pinball: 5.047, pinballQ0: 6.287, gainOverQ0: (mean: 1.24, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.608), ship: no, decision: NOT SHIPPED: upper miss 0.122 (in [0.05,0.15]), lower miss 0.165 on 254 eligible (outside [0.05,0.15]) · age22 proof: (players: 855, upperMiss: 0.095, lowerMiss: 0.168, lowerEligible: 327, coverage80: 0.841, pitMean: 0.37, pitHist10: 217, 156, 70, 76, 60, 55, 41, 45, 54, 81, pinball: 2.624, pinballQ0: 5.195, gainOverQ0: (mean: 2.571, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.731), dev: (players: 1047, upperMiss: 0.119, lowerMiss: 0.171, lowerEligible: 422, coverage80: 0.812, pitMean: 0.383, pitHist10: 277, 205, 75, 64, 53, 72, 45, 60, 71, 125, pinball: 2.845, pinballQ0: 5.646, gainOverQ0: (mean: 2.801, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.764), ship: no, decision: NOT SHIPPED: upper miss 0.095 (in [0.05,0.15]), lower miss 0.168 on 327 eligible (outside [0.05,0.15]) · shippedFrom draft: (ship: yes, mode: v2, provisional: no), age20: (ship: no, mode: v2, provisional: no), age22: (ship: no, mode: v2, provisional: no) · evaluatedAt 2026-09-29T22:04:29Z | |
| careers-checkpoint-projection-v2 | 2026-09-29 | shipped | P (amended before any projection was fitted or scored). Identical to careers-checkpoint-projection-v1 except the shipping test. Reason: V25 has a large mass at exactly 0 (most picks never play), so for a player whose 10th percentile is 0 the interval [q10, q90] covers about 90%, not 80%, by construction; the v1 window [0.75, 0.85] for two-sided coverage would fail a perfectly calibrated model. The amended test checks each tail where a miss is possible. | draft proof: (players: 855, upperMiss: 0.104, lowerMiss: 0.137, lowerEligible: 124, coverage80: 0.876, pitMean: 0.488, pitHist10: 81, 98, 93, 92, 64, 122, 63, 97, 56, 89, pinball: 5.81, pinballQ0: 5.81, gainOverQ0: (mean: 0, range95: 2 items, shareAbove0: 0), spearmanMedianVsOutcome: 0.479), dev: (players: 1047, upperMiss: 0.124, lowerMiss: 0.136, lowerEligible: 154, coverage80: 0.856, pitMean: 0.487, pitHist10: 107, 105, 139, 98, 80, 147, 65, 107, 69, 130, pinball: 6.36, pinballQ0: 6.36, gainOverQ0: (mean: 0, range95: 2 items, shareAbove0: 0), spearmanMedianVsOutcome: 0.518), ship: yes, decision: SHIP: upper miss 0.104 (in [0.05,0.15]), lower miss 0.137 on 124 eligible (ok) · age20 proof: (players: 854, upperMiss: 0.137, lowerMiss: 0.209, lowerEligible: 292, coverage80: 0.792, pitMean: 0.408, pitHist10: 201, 138, 69, 79, 26, 80, 48, 55, 41, 117, pinball: 4.884, pinballQ0: 5.782, gainOverQ0: (mean: 0.898, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.581), dev: (players: 1046, upperMiss: 0.141, lowerMiss: 0.271, lowerEligible: 421, coverage80: 0.75, pitMean: 0.401, pitHist10: 323, 147, 77, 50, 32, 61, 64, 76, 68, 148, pinball: 5.047, pinballQ0: 6.287, gainOverQ0: (mean: 1.24, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.608), ship: no, decision: NOT SHIPPED: upper miss 0.137 (in [0.05,0.15]), lower miss 0.209 on 292 eligible (outside [0.05,0.15]) · age22 proof: (players: 855, upperMiss: 0.109, lowerMiss: 0.202, lowerEligible: 341, coverage80: 0.811, pitMean: 0.372, pitHist10: 231, 135, 79, 78, 56, 53, 43, 45, 42, 93, pinball: 2.624, pinballQ0: 5.195, gainOverQ0: (mean: 2.571, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.731), dev: (players: 1047, upperMiss: 0.138, lowerMiss: 0.221, lowerEligible: 475, coverage80: 0.761, pitMean: 0.381, pitHist10: 315, 167, 78, 66, 52, 73, 44, 51, 56, 145, pinball: 2.845, pinballQ0: 5.646, gainOverQ0: (mean: 2.801, range95: 2 items, shareAbove0: 1), spearmanMedianVsOutcome: 0.764), ship: no, decision: NOT SHIPPED: upper miss 0.109 (in [0.05,0.15]), lower miss 0.202 on 341 eligible (outside [0.05,0.15]) · shippedFrom draft: (ship: yes, mode: v2, provisional: no), age20: (ship: no, mode: v2, provisional: no), age22: (ship: no, mode: v2, provisional: no) · evaluatedAt 2026-09-29T22:04:29Z | |
| careers-aging-survivorship-v1 | 2026-09-29 | adopted | A1. Delta-method aging curves (mean within-player change between consecutive qualifying seasons, weighted by the harmonic mean of the two seasons' TOI, chained from age 19) are biased upward at older ages because players who decline leave the sample. Two corrections: (i) inverse-probability weighting of each observed pair by 1 / P(next season observed | age, position, level percentile, TOI share, GP, previous delta) from a logistic fitted by position; (ii) pattern-mixture imputation: every exit (no qualifying next season, age <= 38) gets a next-season delta = the IPW-weighted conditional mean delta for its covariates plus a shift s estimated from near-dropouts (next season 1-19 GP, observed but thin), s = TOI-weighted mean residual of near-dropouts. Expected: the corrections move peak age earlier and steepen the post-peak decline, most for forwards' points/60. | simulationMarMnarRmse naive: 0.062, ipw: 0.064, pm: 0.039 · winner pm · curves F:pointsPer60: (headline: pm, peakAge: 24, peakAgeRange: 23, 24, declinePerYear: -0.052, naivePeak: 24, proofCheck: (naive: 2 fields, ipw: 2 fields, pm: 2 fields)), F:toiShare: (headline: pm, peakAge: 26, peakAgeRange: 25, 27, declinePerYear: -0.005, naivePeak: 26, proofCheck: (naive: 2 fields, ipw: 2 fields, pm: 2 fields)), F:pointsPer82: (headline: pm, peakAge: 24, peakAgeRange: 24, 26, declinePerYear: -1.065, naivePeak: 26, proofCheck: (naive: 2 fields, ipw: 2 fields, pm: 2 fields)), F:drive: (headline: pm, peakAge: 24, peakAgeRange: 19, 25, declinePerYear: -0.004, naivePeak: 24, proofCheck: —), F:xgf60: (headline: pm, peakAge: 24, peakAgeRange: 21, 26, declinePerYear: -0.039, naivePeak: 24, proofCheck: —), F:xgaPrevented60: (headline: pm, peakAge: 33, peakAgeRange: 19, 38, declinePerYear: -0.025, naivePeak: 33, proofCheck: —), F:ixg60: (headline: pm, peakAge: 24, peakAgeRange: 24, 26, declinePerYear: -0.014, naivePeak: 26, proofCheck: —), D:pointsPer60: (headline: pm, peakAge: 23, peakAgeRange: 19, 26, declinePerYear: -0.009, naivePeak: 23, proofCheck: (naive: 2 fields, ipw: 2 fields, pm: 2 fields)), D:toiShare: (headline: pm, peakAge: 26, peakAgeRange: 25, 27, declinePerYear: -0.003, naivePeak: 27, proofCheck: (naive: 2 fields, ipw: 2 fields, pm: 2 fields)), D:pointsPer82: (headline: pm, peakAge: 26, peakAgeRange: 23, 28, declinePerYear: -0.878, naivePeak: 26, proofCheck: (naive: 2 fields, ipw: 2 fields, pm: 2 fields)), D:drive: (headline: pm, peakAge: 24, peakAgeRange: 20, 28.025, declinePerYear: -0.002, naivePeak: 24, proofCheck: —), D:xgf60: (headline: pm, peakAge: 26, peakAgeRange: 20, 28, declinePerYear: -0.019, naivePeak: 26, proofCheck: —) … +4 fields · evaluatedAt 2026-09-29T22:04:29Z | |
| careers-aging-mixed-v1 | 2026-09-29 | benchmark | A2. A linear mixed model level_it = f(age) + b0_i + b1_i (age - 27) + e_it, f a natural cubic spline (knots 21, 24, 27, 30, 33, 36), random intercept and slope with unstructured 2x2 covariance, residual variance inversely proportional to the season's TOI, fitted by EM (REML-style) in numpy, gives a peak and decline in the same place as the corrected delta curve (likelihood under MAR dropout ignores exit that depends on past observed level). | peaks F:pointsPer60: (peakAge: 24.9, range: 24.5, 25.5, declinePerYear: -0.035), F:toiShare: (peakAge: 26.6, range: 26.1, 26.9, declinePerYear: -0.004), F:pointsPer82: (peakAge: 25.7, range: 25.297, 26.1, declinePerYear: -1.306), F:drive: (peakAge: 24.2, range: 23.295, 26.4, declinePerYear: -0.001), D:pointsPer60: (peakAge: 25.7, range: 24.6, 27.5, declinePerYear: -0.013), D:toiShare: (peakAge: 27.1, range: 26.198, 28.7, declinePerYear: -0.003), D:pointsPer82: (peakAge: 26.2, range: 25.198, 27.6, declinePerYear: -0.523), D:drive: (peakAge: 24.2, range: 23.1, 26.003, declinePerYear: -0.001), G:savePctAboveLeague: (peakAge: 26.1, range: 19, 39, declinePerYear: -0.02) · oneStepVsK F: (rows: 3173, mseK: 153.74, mseMixed: 162.328, mixedMinusK: 8.587, range95: 3.069, 14.818, winner: K), D: (rows: 1621, mseK: 92.075, mseMixed: 102.459, mixedMinusK: 10.384, range95: 4.157, 17.604, winner: K), G: (rows: 350, mseK: 1.222, mseMixed: 1.263, mixedMinusK: 0.041, range95: 0.001, 0.084, winner: K) · evaluatedAt 2026-09-29T22:04:29Z | |
| careers-aging-tiers-v1 | 2026-09-29 | run | A3. Stars and depth players age differently. Tier is fixed one season BEFORE the delta is measured (tier from season t-1, delta t -> t+1) so selection noise in t does not create regression to the mean: star = top 10% of position-age level in t-1, depth = bottom 50%. Plus quantile curves: linear quantile regression (MM algorithm) of level on the age spline at tau 0.1, 0.5, 0.9, IPW-weighted. Expected: stars peak later and decline slower in points/60. | F:pointsPer60 peakDiff: (estimate: 0, range95: -5, 4), declineDiff: (estimate: -0.039, range95: -0.075, 0.067), starPairs: 758, depthPairs: 2591 · D:pointsPer60 peakDiff: (estimate: -4, range95: -6, 1), declineDiff: (estimate: 0.004, range95: -0.149, 0.032), starPairs: 396, depthPairs: 1449 | |
| careers-checkpoint-projection-v1 | 2026-09-29 | superseded | P. From information available at a checkpoint only (draft: through the draft-year season; age 20 and age 22: through that season) - draft position, round, age in days at the checkpoint, position, height, weight, pre-NHL production translated by the site league ladder (league-ladder.json; leagues without a translation keep raw points per game plus a league-family indicator), NHL games / points / TOI to date - a quantile model projects each drafted player's path: cumulative NHL TOI hours, cumulative NHL games and cumulative NHL points through each age to 25, and season TOI hours at each age, with calibrated 10/50/90% bounds. Career outcome V25 = cumulative NHL TOI hours through the age-25 season (observable for every pick of 2010-2018). | not run yet | |
| careers-label-ceiling-v1 | 2026-09-29 | shipped | L1. Exceeded ceiling at checkpoint c (draft / 20 / 22): V25 above the conformalised 90th percentile projection at c. Fell short: V25 below the conformalised 10th percentile (only defined where that bound is > 0). Probability shown = 1 - projected CDF at the realised V25. Validation (does it predict anything out of sample): the same label computed on cumulative TOI through 23 must predict season TOI at 24 and 25 beyond the projection's own 90th (10th) percentile of those seasons. | registeredV2Bounds exceeded: (draft: (dev: 5 fields, proof: 5 fields, ship: yes, decision: SHIP: proof rate/nominal 6.21 range [5.329670329670329, 7.032967032967033] on 91 players (bounds v2)), age20: (dev: 5 fields, proof: 5 fields, ship: yes, decision: SHIP: proof rate/nominal 6.36 range [5.454545454545454, 7.215909090909091] on 88 players (bounds v2)), age22: (dev: 5 fields, proof: 5 fields, ship: yes, decision: SHIP: proof rate/nominal 5.56 range [4.027777777777778, 6.944444444444445] on 36 players (bounds v2))), fellShort: (draft: (dev: 5 fields, proof: 5 fields, ship: no, decision: NOT SHIPPED: proof rate/nominal 4.58 range [1.6666666666666667, 7.083333333333333] on 12 players (bounds v2)), age20: (dev: 5 fields, proof: 5 fields, ship: no, decision: NOT SHIPPED: proof rate/nominal 7.37 range [5.526315789473684, 8.947368421052632] on 19 players (bounds v2)), age22: (dev: 5 fields, proof: 5 fields, ship: yes, decision: SHIP: proof rate/nominal 4.81 range [3.5849056603773586, 6.037735849056604] on 53 players (bounds v2))) · shipping draft: (mode: v2, provisional: no, exceeded: yes, fellShort: no, registeredV2: (exceeded: yes, fellShort: no), upperTailOk: yes, lowerTailOk: yes, byGroup: (F: yes, D: yes, G: no)), age20: (mode: v2, provisional: no, exceeded: yes, fellShort: no, registeredV2: (exceeded: yes, fellShort: no), upperTailOk: yes, lowerTailOk: no, byGroup: (F: yes, D: yes, G: no)), age22: (mode: v2, provisional: no, exceeded: yes, fellShort: no, registeredV2: (exceeded: yes, fellShort: yes), upperTailOk: yes, lowerTailOk: no, byGroup: (F: yes, D: yes, G: no)) · evaluatedAt 2026-09-29T22:04:29Z | |
| careers-label-flash-v1 | 2026-09-29 | shipped | L2. Spike season: qualifying season t with z >= 2.0 against K's forecast from strictly earlier seasons, the player having >= 2 earlier qualifying seasons. Flash in the pan (retrospective): spike, and the mean of the qualifying seasons among t+1..t+2 (at least one) has z <= 1.0 against the PRE-spike forecast. Claim: a spike season regresses MORE than regression to the mean alone (K with all seasons through t) predicts, i.e. spikes are disproportionately luck. | beta -0.042 · devSpikePairs 228 · devSpikeMeanResidual -0.042 · devRange95 -0.18, 0.098 · proofPairs 4304 · proofSpikePairs 120 | |
| careers-label-sustained-v1 | 2026-09-29 | shipped | L3. Sustained breakout: a spike season t (as L2) and both t+1 and t+2 qualifying with the z of their variance-weighted mean against the PRE-spike forecast >= 1.5. Claim: the change is real, so the season after the window (t+3) stays above the pre-spike forecast. | dev n: 59, players: 56, meanZvsPreWindow: 1.327, range95: 1.108, 1.535, shareAbove0: 1, meanZvsFullK: -0.207, fullKRange95: -0.515, 0.095 · proof n: 42, players: 41, meanZvsPreWindow: 1.515, range95: 1.316, 1.724, shareAbove0: 1, meanZvsFullK: -0.138, fullKRange95: -0.453, 0.18 · ship true · decision proof mean z range above 0 · evaluatedAt 2026-09-29T22:04:29Z | |
| careers-label-late-bloomer-v1 | 2026-09-29 | shipped | L4. Late bloomer: two consecutive qualifying seasons both at age >= 26 whose variance-weighted mean has z >= 1.5 against K's forecast from seasons through age 25 only (prior alone if none), and the player's K estimate at 25 is not above the position's 75th percentile of age-25 talent. Claim: the late rise is real, so the next season stays above the through-25 forecast. | dev n: 27, players: 27, meanZvsPreWindow: 1.053, range95: 0.766, 1.345, shareAbove0: 1, meanZvsFullK: -0.358, fullKRange95: -0.739, 0.017 · proof n: 34, players: 34, meanZvsPreWindow: 1.243, range95: 0.995, 1.483, shareAbove0: 1, meanZvsFullK: -0.12, fullKRange95: -0.431, 0.156 · ship true · decision proof mean z range above 0 · evaluatedAt 2026-09-29T22:04:29Z | |
| careers-label-early-peak-v1 | 2026-09-29 | not-shipped | L5. Early peak: the player's best two consecutive qualifying seasons (variance-weighted mean level) end at age <= 24, and the variance-weighted mean of his qualifying seasons at ages 26-27 has z <= -1.5 against the forecast made at the end of the peak pair (K + headline curve). Claim: the decline is real, so age 28 stays below that forecast. | dev n: 19, players: 19, meanZvsPreWindow: -1.102, range95: -1.43, -0.754, shareAbove0: 0, meanZvsFullK: 0.404, fullKRange95: -0.078, 0.905 · proof n: 12, players: 12, meanZvsPreWindow: -1.124, range95: -1.383, -0.856, shareAbove0: 0, meanZvsFullK: 0.355, fullKRange95: -0.075, 0.757 · ship false · decision only 12 proof label-holders (< 20) · evaluatedAt 2026-09-29T22:04:29Z | |
| careers-bias-fdr-v1 | 2026-09-29 | run | B. The checkpoint projections are fair across nationality (birth country: CAN, USA, SWE, FIN, RUS, CZE+SVK, other), draft round (1..7), birth quarter (Jan-Mar ... Oct-Dec), league of origin at the draft (CHL, NCAA/US college, US junior, Sweden, Finland, Russia, Czechia/Slovakia, other Europe, other) and market (drafting club's metro tercile from data/assay/market-size.json). Residual = PIT of V25 under the out-of-fold projection (leave-one-class-out for 2010-2014, frozen model for 2015-2018). | tests 251 · q 0.1 · discoveries (checkpoint: age20, attribute: draftRound, level: 1, statistic: meanPIT, n: 271, value: 0.616, rest: 0.367, diff: 0.249, p: 0, pAdjBH: 0.008, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 1, statistic: exceededRate, n: 271, value: 0.258, rest: 0.12, diff: 0.139, p: 0, pAdjBH: 0.008, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 1, statistic: fellShortRate, n: 271, value: 0.059, rest: 0.36, diff: -0.301, p: 0, pAdjBH: 0.008, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 2, statistic: meanPIT, n: 275, value: 0.488, rest: 0.388, diff: 0.101, p: 0, pAdjBH: 0.008, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 3, statistic: fellShortRate, n: 89, value: 0.36, rest: 0.229, diff: 0.13, p: 0.008, pAdjBH: 0.085, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 5, statistic: meanPIT, n: 271, value: 0.35, rest: 0.411, diff: -0.06, p: 0.008, pAdjBH: 0.084, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 5, statistic: fellShortRate, n: 48, value: 0.479, rest: 0.229, diff: 0.251, p: 0, pAdjBH: 0.008, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 6, statistic: meanPIT, n: 271, value: 0.313, rest: 0.417, diff: -0.104, p: 0, pAdjBH: 0.008, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 6, statistic: fellShortRate, n: 44, value: 0.455, rest: 0.232, diff: 0.223, p: 0.002, pAdjBH: 0.026, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 7, statistic: meanPIT, n: 271, value: 0.281, rest: 0.422, diff: -0.142, p: 0, pAdjBH: 0.008, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 7, statistic: exceededRate, n: 271, value: 0.059, rest: 0.153, diff: -0.094, p: 0, pAdjBH: 0.008, discovery: yes); (checkpoint: age20, attribute: draftRound, level: 7, statistic: fellShortRate, n: 30, value: 0.533, rest: 0.233, diff: 0.301, p: 0.001, pAdjBH: 0.022, discovery: yes) … +15 more · evaluatedAt 2026-09-29T22:04:29Z | |
| drive-multiseason-d1-v1 | 2026-09-29 | adopted | D1. Drive computed on 5v5 on-ice and off-ice xG pooled over the current and prior two seasons with decay weights (1, a, b), a and b on a 0.05 grid in [0,1] chosen on the development end seasons 2021-2023 by maximising the pooled next-season correlation, predicts next season's single-season Drive better than the single-season Drive, with no more club leakage. | weights current: 1, prior1: 0.7, prior2: 0.4 · devR 0.486 · devChampionR 0.437 · devGain mean: 0.05, range95: 0.03, 0.069, shareAboveZero: 1, players: 1487, B: 2000, seed: 5 · proofR 0.477 · championProofR 0.385 | |
| drive-eb-prior-d2-v1 | 2026-09-29 | passed-not-adopted | D2. D1 shrunk toward a position/usage prior (weighted regression of D1 on position F/D/unknown and 5v5 minutes per game, fitted cross-sectionally within the season) by normal-normal empirical Bayes: each player's sampling variance from a 200-rep game-cluster bootstrap, prior variance tau^2 estimated by method of moments (never hand-set). Beats the champion and possibly D1. | devR 0.516 · devGain mean: 0.079, range95: 0.058, 0.101, shareAboveZero: 1, players: 1487, B: 2000, seed: 5 · proofR 0.494 · championProofR 0.385 · gain mean: 0.108, range95: 0.069, 0.15, shareAboveZero: 1, players: 502, B: 2000, seed: 5 · clubLeakage 0.081 | |
| drive-rapm-chain-d3-v1 | 2026-09-29 | rejected | D3. Multi-season RAPM: each season's ridge RAPM (tools/assay/py/measures/rapm.py design) penalised toward the previous season's estimate instead of zero, chained from 2018; ranking value = oRAPM + dRAPM; lambda from {16000, 32000, 64000} chosen on the development end seasons. Expected to predict well but to carry club leakage (single-season RAPM total r_club 0.41). | lambda 64000 · devByLambda 16000: (pooled: 0.357, byEnd: (2021: 0.383, 2022: 0.337, 2023: 0.355), n: 1487), 32000: (pooled: 0.386, byEnd: (2021: 0.41, 2022: 0.362, 2023: 0.389), n: 1487), 64000: (pooled: 0.403, byEnd: (2021: 0.423, 2022: 0.379, 2023: 0.411), n: 1487) · devR 0.403 · proofR 0.383 · championProofR 0.385 · gain mean: -0.002, range95: -0.062, 0.058, shareAboveZero: 0.466, players: 502, B: 2000, seed: 5 | |
| prospect-growth-age-v1 | 2026-09-29 | rejected | P1. The site's hand-set age line (1.00 at 18, minus 0.06 a year, floor 0.58) is flatter than measured same-league growth in points per game (whole classes: steeper through 22). Replacing it with the measured curve, selection-corrected (same-league pairs reweighted so their production mix matches every line at that age in that league, because stayers are the unpromoted), ranks prospects better. Within one age the change only reorders by months of age, so the expected per-age gain is small (under 0.01); the effect is across ages. | proofMeanRho 0.473 · championMeanRho 0.471 · gain mean: 0.002, range95: -0.002, 0.006, shareAboveZero: 0.807, players: 644, B: 2000, seed: 5 · logLossVsChampion -0.003 · devGain mean: 0.003, range95: 0, 0.005, shareAboveZero: 0.969, players: 1236, B: 2000, seed: 5 · decision REJECT: the 95% range of the gain does not lie above zero | |
| prospect-pick-blend-v1 | 2026-09-29 | adopted | P2. Draft position ranks NHL games through 23 better than the formula at 18 and 19 (whole classes 0.605 vs 0.448 at 18). A blend of the two percentiles with one weight w, chosen on the development classes, beats the formula alone. Expected: a large gain at 18, shrinking by 20. | proofMeanRho 0.636 · championMeanRho 0.471 · gain mean: 0.165, range95: 0.127, 0.203, shareAboveZero: 1, players: 644, B: 2000, seed: 5 · logLossVsChampion -0.088 · devGain mean: 0.136, range95: 0.105, 0.163, shareAboveZero: 1, players: 1236, B: 2000, seed: 5 · decision PASS: gain range above zero and graduation log loss not worse | |
| prospect-corrected-ladder-v1 | 2026-09-29 | rejected | P3. The league ladder is measured on players who moved to the NHL, who are the league's best scorers, so it understates what a typical line is worth. Re-measuring each league's translation with movers reweighted to the league's production deciles (inverse probability of moving), on 2010-2015 classes only, and using it in place of the site's ladder ranks prospects better. | proofMeanRho 0.476 · championMeanRho 0.471 · gain mean: 0.005, range95: -0.004, 0.015, shareAboveZero: 0.855, players: 644, B: 2000, seed: 5 · logLossVsChampion -0.002 · devGain mean: -0.005, range95: -0.015, 0.004, shareAboveZero: 0.138, players: 1236, B: 2000, seed: 5 · decision REJECT: the 95% range of the gain does not lie above zero | |
| prospect-fitted-benchmark-v1 | 2026-09-29 | benchmark | P4. A fitted benchmark (L2-regularised logistic regression on graduation, numpy only because lightgbm is not installed) on the formula's parts, age, position, league family, height, draft-year production and draft position shows how much ranking the formula leaves on the table. It is a ceiling, not a site candidate: it uses draft position and a fitted black box the page could not print. | proofMeanRho 0.63 · championMeanRho 0.471 · gain mean: 0.159, range95: 0.122, 0.197, shareAboveZero: 1, players: 644, B: 2000, seed: 5 · logLossVsChampion -0.095 · devGain mean: 0.146, range95: 0.12, 0.17, shareAboveZero: 1, players: 1236, B: 2000, seed: 5 · decision PASS: gain range above zero and graduation log loss not worse | |
| cba-retrieval-r1-vocab-repair-v1 | 2026-09-29 | rejected | The contract-term rule of the question expansion fires on 'long-term', 'how long' and 'length' in questions that are not about contract length, and several rules append words that are common across the agreement. Repairing the triggers and dropping the common words raises the rank of the passage that answers a fan-worded question. Expected: a gain concentrated on long-term-injury and duration questions and small overall (the rule fired on 24 of 110 development questions); 60 held-out questions may be too few to show it. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params dropAboveShare: 0.05 · development leaveOneOut: (mrr: 0.674, recallAt5: 0.8, picks: (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.573, recallAt5: 0.8, recallAt8: 0.836, mrr: 0.674), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: 0.036, lo: 0.009, hi: 0.073, better: 4, worse: 0), mrr: (difference: 0.043, lo: 0.014, hi: 0.075, better: 17, worse: 3)) · heldOut n: 60, candidate: (recallAt1: 0.25, recallAt5: 0.517, recallAt8: 0.583, mrr: 0.366), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: -0.017, lo: -0.05, hi: 0, better: 0, worse: 1), mrr: (difference: -0.002, lo: -0.058, hi: 0.052, better: 9, worse: 12) | |
| cba-retrieval-r2-weighted-expansion-v1 | 2026-09-29 | rejected | Expansion words at full weight can outweigh the fan's own words. Entering them at a reduced weight keeps the bridge between a fan's vocabulary and the agreement's while letting the question lead. Expected: a small gain in MRR; possibly none if R1 has already removed the expansions that did harm. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params weight: 0.25 · development leaveOneOut: (mrr: 0.652, recallAt5: 0.782, picks: (2 fields); (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.564, recallAt5: 0.782, recallAt8: 0.855, mrr: 0.664), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: 0.018, lo: -0.018, hi: 0.054, better: 3, worse: 1), mrr: (difference: 0.033, lo: -0.003, hi: 0.072, better: 17, worse: 3)) · heldOut n: 60, candidate: (recallAt1: 0.25, recallAt5: 0.483, recallAt8: 0.55, mrr: 0.354), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: -0.05, lo: -0.133, hi: 0.017, better: 1, worse: 4), mrr: (difference: -0.015, lo: -0.088, hi: 0.057, better: 9, worse: 17) | |
| cba-retrieval-r3-ocr-word-split-v1 | 2026-09-29 | rejected | In the scanned 2025 MOU, words run together by the OCR are single tokens that no question contains. Splitting them, for the index only, into words of the clean documents (2013 CBA, 2020 MOU) makes those passages findable; the passage text shown and quoted is unchanged. Expected: gains only where the gold sits in the 2025 MOU (12 of the 60 held-out questions), so the overall effect may be too small to pass the rule. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params — · development leaveOneOut: (mrr: 0.627, recallAt5: 0.754, picks: (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.518, recallAt5: 0.754, recallAt8: 0.818, mrr: 0.627), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: -0.009, lo: -0.027, hi: 0, better: 0, worse: 1), mrr: (difference: -0.004, lo: -0.02, hi: 0.011, better: 7, worse: 13)) · heldOut n: 60, candidate: (recallAt1: 0.233, recallAt5: 0.517, recallAt8: 0.55, mrr: 0.364), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: -0.017, lo: -0.05, hi: 0, better: 0, worse: 1), mrr: (difference: -0.005, lo: -0.024, hi: 0.008, better: 6, worse: 12) | |
| cba-retrieval-r3b-ocr-letter-repair-v1 | 2026-09-29 | rejected | The same scan reads the letter m as 'in', 'im', 'iin' or 'rn' ('terin' for term, 'aimount' for amount, 'liimit' for limit, 'gaine' for game). Replacing, for the index only, a token that is absent from the clean vocabulary by the one clean-vocabulary word those confusions lead to makes the agreement's commonest words match in the 2025 MOU. Expected: larger than R3 on 2025 MOU questions, because the damaged words are the ones fans type; still limited to the 12 held-out questions whose gold sits there. Added by the desk; not in the coordinator's list. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params — · development leaveOneOut: (mrr: 0.628, recallAt5: 0.764, picks: (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.518, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.628), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: 0, lo: 0, hi: 0, better: 0, worse: 0), mrr: (difference: -0.003, lo: -0.013, hi: 0.005, better: 5, worse: 10)) · heldOut n: 60, candidate: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.55, mrr: 0.368), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: 0, lo: 0, hi: 0, better: 0, worse: 0), mrr: (difference: -0.001, lo: -0.001, hi: 0, better: 1, worse: 6) | |
| cba-retrieval-r4-field-weights-v1 | 2026-09-29 | rejected | The title is counted twice and the score is raised 8% per document rank; both were set by hand. A title weight or document boost chosen on development questions ranks the gold passage higher. Expected: small. On the development questions the boost was worth 2 points of recall@5, inside the noise, and the 2020 and 2025 MOU chunks all carry the same title, so the title can only help in the 2013 CBA. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params titleWeight: 1, rankBoost: 0.08 · development leaveOneOut: (mrr: 0.609, recallAt5: 0.746, picks: (2 fields); (2 fields); (2 fields); (2 fields); (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.527, recallAt5: 0.754, recallAt8: 0.818, mrr: 0.634), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: -0.009, lo: -0.027, hi: 0, better: 0, worse: 1), mrr: (difference: 0.003, lo: -0.011, hi: 0.018, better: 12, worse: 14)) · heldOut n: 60, candidate: (recallAt1: 0.233, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.36), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: 0, lo: 0, hi: 0, better: 0, worse: 0), mrr: (difference: -0.009, lo: -0.04, hi: 0.02, better: 9, worse: 15) | |
| cba-retrieval-r5-feedback-pass-v1 | 2026-09-29 | rejected | A second ranking pass that appends the most distinctive terms of the first pass's top three passages, at low weight, finds passages the fan's words miss. Expected: uncertain; feedback helps when the first pass is mostly right and hurts when it is wrong, and a fan's first pass is often wrong. | testedAt 2026-09-29T14:02:40.456Z · results data/assay/assistant-eval-heldout-results.json · comparator B0 · params terms: 5, weight: 0.1 · development leaveOneOut: (mrr: 0.609, recallAt5: 0.736, picks: (2 fields), againstComparator: (recallAt5: 5 fields, mrr: 5 fields)), inSample: (n: 110, candidate: (recallAt1: 0.491, recallAt5: 0.736, recallAt8: 0.818, mrr: 0.609), comparator: (recallAt1: 0.527, recallAt5: 0.764, recallAt8: 0.818, mrr: 0.631), recallAt5: (difference: -0.027, lo: -0.064, hi: 0, better: 0, worse: 3), mrr: (difference: -0.022, lo: -0.046, hi: -0.001, better: 13, worse: 22)) · heldOut n: 60, candidate: (recallAt1: 0.233, recallAt5: 0.517, recallAt8: 0.567, mrr: 0.363), comparator: (recallAt1: 0.25, recallAt5: 0.533, recallAt8: 0.567, mrr: 0.369), recallAt5: (difference: -0.017, lo: -0.067, hi: 0.033, better: 1, worse: 2), mrr: (difference: -0.005, lo: -0.037, hi: 0.025, better: 16, worse: 17) | |
| cba-fallback-refusal-v1 | 2026-09-29 | adopted | While the model is unavailable the passages-only fallback answers every question, including the 30 of 30 development questions the agreement cannot answer. Retrieval evidence alone (top score, margin to the second, share of the question's words found in the top passages, a named player or club) separates enough of them to refuse or warn without often refusing a question the agreement does answer. Expected: the name rule catches the named-player and club questions; the evidence rule catches about half of the rest at 5% wrong refusals (the top score alone caught 50% at that point on development). | testedAt 2026-09-29T14:02:40.456Z · configuration B0 · rule mean: 27.339, 3.971, 0.694, 0.597, scale: 11.952, 4.368, 0.165, 0.183, weights: -2.448, 0.626, -0.833, -0.156, intercept: -2.515, weak: 0.289, none: 0.557 · development answerable: 110, unanswerable: 30, method: leave-one-out, levelC: (wrongRefusals: 5, wrongRate: 0.045, caught: 15, tpr: 0.5, tprInterval: 0.332, 0.668), levelBorC: (wronglyFlagged: 16, wrongRate: 0.145, caught: 23, tpr: 0.767, tprInterval: 0.591, 0.882), auc: 0.906 · heldOut withNameRule: (levelC: (caught: 15, tpr: 0.75, tprInterval: 2 items, wrongRefusals: 13, wrongRate: 0.217, wrongInterval: 2 items), levelBorC: (caught: 18, tpr: 0.9, wronglyFlagged: 25, wrongRate: 0.417), byNameRule: (unanswerable: 9, answerable: 3, answerableIds: 3 items)), evidenceRuleAlone: (levelC: (caught: 14, tpr: 0.7, tprInterval: 2 items, wrongRefusals: 11, wrongRate: 0.183, wrongInterval: 2 items), levelBorC: (caught: 18, tpr: 0.9, wronglyFlagged: 23, wrongRate: 0.383)), guard: (answerableNamingNoOne: 57, wronglyRefusedAtC: 10, rate: 0.175, limit: 0.1, holds: no, consequence: level (c) switched off for evidence; levels (a) and (b) only; the name rule stays) · shipped levels (a) confident and (b) weak from evidence; level (c) from evidence switched OFF by the registered guard (held-out answerable naming no one wrongly refused | |
| cba-number-gate-exact-v1 | 2026-09-29 | adopted | The check that every number in an answer appears in the passage it cites is a substring test, so a claimed '5' passes against a passage that says '50' or '2025'. An exact test on number tokens, compared as canonical decimals, closes that hole without failing answers that are right. | selfTest tools/assay/eval/selftest.mjs 20/20 incl. the 5-vs-fifty (50) false accept; vitest src/server/__tests__/grounded-core.test.ts: 8.50==8.5, 5!=50, 5!=15/2025, 50! · developmentAnswers no model answers exist (model mode not run); instead every gold answer and every gold number of the 160-item set was put through both gates as a sentence citing · sentences 258 · passOld 258 · passNew 258 · passOldFailNew 0 | |
| cba-clause-split-fix-v1 | 2026-09-29 | adopted | The fallback is meant to quote clauses of semicolon lists, but its split pattern has a literal 's' where whitespace was meant, so it never splits at '; ' and quotes whole sentences. Splitting on whitespace after a semicolon gives shorter quotes that are nearer the question. | split /(?<=;)s+/ · quotesBefore 657 · quotesAfter 657 · medianQuoteLengthBefore 362 · medianQuoteLengthAfter 329 · everyQuoteVerbatim true | |
| xg-clock-corrected-flags-v1 | 2026-09-29 | adopted | With the champion's rebound and rush flags recomputed on a clock that does not depend on the outcome (goal times moved back 1 s, the offset measured on 2023-24 + 2024-25, before the 3 s and 4 s windows are applied) and everything else unchanged, held-out log loss does NOT improve and is expected to get slightly worse, by about 0.0002 to 0.0003 per shot, because the raw flags carry information about the outcome through the clock. A worse log loss here means part of the champion's published accuracy was the clock. | offsetSeconds 1 · rows trainRows: 227076, testRows: 105288, testGames: 1232, testRowsInShotFile: 112091, testGamesInShotFile: 1312, trainGoals: 15987, testGoals: 7601 · trainLogLoss 0.221 · testLogLoss 0.229 · testAuc 0.756 · championReplicaTestLogLoss 0.229 | |
| xg-clock-corrected-split-v1 | 2026-09-29 | run | F's corrected flags inside the four strength-state fits of challenger B. Expected against the champion: the strength split's gain on the proof season (about 0.0005 with two seasons of training) less the cost of correcting the clock (about 0.0002 to 0.0003), a net gain near 0.0002 that may not clear the rule. The split's gain is not robust to the amount of training data: trained on one season it was worse than a single fit. | offsetSeconds 1 · rows trainRows: 227076, testRows: 105288, testGames: 1232, testRowsInShotFile: 112091, testGamesInShotFile: 1312, trainGoals: 15987, testGoals: 7601 · trainLogLoss 0.22 · testLogLoss 0.228 · testAuc 0.757 · championReplicaTestLogLoss 0.229 | |
| xg-clock-corrected-flurry-v1 | 2026-09-29 | run | G plus the flurry features of challenger C (seconds since the previous unblocked shot by the same club in the same period, capped at 30 s, and an indicator for under 3 s) computed on the corrected clock. Expected: little or nothing beyond G. The gain C showed with goal times moved back is expected to have been mostly a repair of the raw flags, which G already has. | offsetSeconds 1 · rows trainRows: 227076, testRows: 105288, testGames: 1232, testRowsInShotFile: 112091, testGamesInShotFile: 1312, trainGoals: 15987, testGoals: 7601 · trainLogLoss 0.22 · testLogLoss 0.228 · testAuc 0.757 · championReplicaTestLogLoss 0.229 | |
| xg-rink-adjusted-v1 | 2026-09-29 | run | Replacing the recorded shot distance with its league-equivalent (the arena quantile map estimated on 2023-24 + 2024-25 only, applied to the training rows and carried unchanged to 2025-26; angle recomputed from the same geometry where well conditioned) removes scorer error from the most important xG input and lowers held-out log loss. Expected effect: very small and possibly negative - the arena study found the scorer effect has shrunk to about 1 ft standard deviation against about 0.55 ft of noise, and that an unshrunk map carried into the next season did not improve agreement. | trainLogLoss 0.221 · testLogLoss 0.228 · testAuc 0.757 · championReplicaTestLogLoss 0.228 · championReplicaTestAuc 0.757 · diffVsChampion 0 | |
| xg-strength-split-v1 | 2026-09-29 | run | Four separate fits by strength state (5v5 = strengthDiff 0 and net occupied, power play, shorthanded, empty net) with the champion's features minus the strength flags let distance, angle and shot-type effects differ by state and lower the combined held-out log loss. Expected effect: a small gain concentrated on the power play; the shorthanded and empty-net fits have few rows and may give some of it back. | trainLogLoss 0.22 · testLogLoss 0.228 · testAuc 0.758 · championReplicaTestLogLoss 0.228 · championReplicaTestAuc 0.757 · diffVsChampion 0 | |
| xg-flurry-feature-v1 | 2026-09-29 | run | Adding the time since the previous unblocked shot by the same club in the same period (capped at 30 s) and an indicator for under 3 s to the champion's 34 features lowers held-out log loss, because shots inside a flurry are taken against a goalie out of position beyond what the rebound flag already records. | trainLogLoss 0.219 · testLogLoss 0.227 · testAuc 0.761 · championReplicaTestLogLoss 0.228 · championReplicaTestAuc 0.757 · diffVsChampion -0.002 | |
| xg-combined-abc-v1 | 2026-09-29 | run | Rink-adjusted distance, strength-state fits and the flurry features together lower held-out log loss by more than any of them alone. | trainLogLoss 0.218 · testLogLoss 0.226 · testAuc 0.762 · championReplicaTestLogLoss 0.228 · championReplicaTestAuc 0.757 · diffVsChampion -0.002 | |
| arena-scorer-profiles-v1 | 2026-09-29 | run | Some NHL arenas record unblocked shots at systematically different distances than the same clubs' shots are recorded elsewhere (a scorer effect in the recorded location, visible in |x| as well as in distance), the effect persists from one season to the next, and subjective counts (hits, giveaways, takeaways, blocked shots) differ by building beyond what the clubs involved explain. Expected size in the 2021+ feed: a standard deviation across arenas near 1 ft, far below the 3-5 ft effects published for the 2007-2013 feed. | meanConsecutiveR_2021on 0.519 · meanConsecutiveR_allSeasons 0.591 · consecutiveR 2018->2019: 0.803, 2019->2020: 0.763, 2020->2021: 0.494, 2021->2022: 0.57, 2022->2023: 0.561, 2023->2024: 0.64, 2024->2025: 0.306 · shareKsPBelow0.01_2025 0.25 · placeboShareKsPBelow0.01 0.035 · sdOfArenaMeanDiff_ft 2018: 1.698, 2019: 1.67, 2020: 1.814, 2021: 1.713, 2022: 1.174, 2023: 0.801, 2024: 0.832, 2025: 0.996 | |
| score-dist-challengers-v1 | 2026-09-29 | rejected | Given the same walk-forward Kalman goal rates, a low-score dependence (Dixon-Coles), a shared component (bivariate Poisson) or per-side over-dispersion (negative binomial) improves the regulation-score log score over independent Poisson on 2025-26; NHL regulation scores are close to independent Poisson so any gain is expected to be small (under 0.01 nats per game) and the negative binomial may fit worse than Poisson. | fits dixonColes: (rho: -0.116), bivariatePoisson: (lambda3: 0), negativeBinomial: (r: 522.36, impliedVarianceInflationAtLambda3: 1.006) · test2025 poissonLeagueAverage: (scoreLogScore: 3.851, crps: 1.307, sixPlus: (predicted: 0.546, actual: 0.573), vsPoissonRange95: -0.009, 0.022), poissonIndependent: (scoreLogScore: 3.844, crps: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), vsPoissonRange95: —), dixonColes: (scoreLogScore: 3.844, crps: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), vsPoissonRange95: -0.003, 0.002), bivariatePoisson: (scoreLogScore: 3.844, crps: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), vsPoissonRange95: 0, 0), negativeBinomial: (scoreLogScore: 3.845, crps: 1.312, sixPlus: (predicted: 0.54, actual: 0.573), vsPoissonRange95: 0, 0) · decision keep independent Poisson (no family beats it with the range excluding zero and no worse 6+ share) · evaluatedAt 2026-09-29T03:13:03Z | |
| momentum-hawkes-v1 | 2026-09-29 | run | Within NHL regulation play, goals cluster in time beyond what a Poisson process with a rate varying by game minute and by score state predicts (a "reply goal" / momentum effect); if it exists, the Hawkes branching ratio fitted on data exceeds the ratio the same fit produces on simulated null games, and short inter-goal intervals (under 60 s) are more common than in the null. | alphaData 0 · alphaBootstrap95 0, 0 · alphaNull95 0.001, 0.014 · alphaMcP 1 · under60Data 0.085 · under60Null95 0.113, 0.119 | |
| state-space-kalman-v2 | 2026-09-29 | run | A Kalman random-walk attack/defence model on regulation goals (linear-Gaussian observation, daily process noise, season-start inflation, home advantage and scoring level as slowly drifting states) predicts home wins walk-forward about as well as the champion; it is expected NOT to beat the champion on log loss because it sees only goals, but it may add information in a blend and gives a transparent rating. | tuned q: 0, sigma2: 3.5, seasonInflation: 0.02, logLoss: 0.657, games: 3491 · tunedOnGridEdge q: no, sigma2: yes, seasonInflation: no · proof 2023: (n: 1230, logLossStateSpace: 0.659, logLossChampion: 0.654, blend5050LogLoss: 0.654, diffRange95: -0.003, 0.013), 2024: (n: 1235, logLossStateSpace: 0.664, logLossChampion: 0.653, blend5050LogLoss: 0.656, diffRange95: 0.002, 0.019), 2025: (n: 1311, logLossStateSpace: 0.695, logLossChampion: 0.681, blend5050LogLoss: 0.685, diffRange95: 0.006, 0.022), pooled: (n: 3776, logLossStateSpace: 0.673, logLossChampion: 0.663, blend5050LogLoss: 0.665, diffRange95: 0.005, 0.015) · decision display rating only; remains a challenger (pooled proof does not beat the champion with the range excluding zero) · evaluatedAt 2026-09-29T03:06:56Z | |
| state-space-kalman-v1 | 2026-09-29 | run | A Kalman random-walk attack/defence model on regulation goals (linear-Gaussian observation, daily process noise, season-start inflation, home advantage and scoring level as slowly drifting states) predicts home wins walk-forward about as well as the champion; it is expected NOT to beat the champion on log loss because it sees only goals, but it may add information in a blend and gives a transparent rating. | tuned q: 0.001, sigma2: 4.5, seasonInflation: 0.02, logLoss: 0.657, games: 3491 · tunedOnGridEdge q: yes, sigma2: yes, seasonInflation: yes · proof 2023: (n: 1230, logLossStateSpace: 0.66, logLossChampion: 0.654, blend5050LogLoss: 0.655, diffRange95: -0.002, 0.014), 2024: (n: 1235, logLossStateSpace: 0.664, logLossChampion: 0.653, blend5050LogLoss: 0.656, diffRange95: 0.003, 0.019), 2025: (n: 1311, logLossStateSpace: 0.695, logLossChampion: 0.681, blend5050LogLoss: 0.685, diffRange95: 0.006, 0.022), pooled: (n: 3776, logLossStateSpace: 0.673, logLossChampion: 0.663, blend5050LogLoss: 0.666, diffRange95: 0.006, 0.015) · decision display rating only; remains a challenger (pooled proof does not beat the champion with the range excluding zero) · evaluatedAt 2026-09-29T03:06:09Z | |
| isotonic-calibrator-v1 | 2026-09-29 | rejected | The champion home-win probabilities carry a miscalibration that a monotone (isotonic) map fitted on earlier seasons corrects out of sample; because the walk-forward reliability is already close to the diagonal the expected gain is small (under 0.003 log loss) and may be zero or negative. | fit2020to2024 logLossBefore: 0.681, logLossAfter: 0.683, diffRange95: 0, 0.005 · fit2023to2024 logLossBefore: 0.681, logLossAfter: 0.683, diffRange95: -0.001, 0.006 · decision rejected · evaluatedAt 2026-09-29T02:49:22Z | |
| hist-2026-09-26-game-model-v2 | 2026-09-26 | adopted | Even-strength xG Elo + goals Elo + PP/PK EWMA + rest and back-to-back lowers held-out log loss versus ratings only. | ratingsOnly 0.675 · plusRest 0.672 · plusSpecialTeams no gain · note 60-70% bucket ~4 points over-confident | |
| hist-2026-09-26-starting-goalie-feature | 2026-09-26 | rejected | Adding the starting goalie's GSAx as a pre-game feature lowers log loss. | before 0.678 · after 0.679 · note worse; dropped | |
| hist-2026-09-26-tournament | 2026-09-26 | adopted | Boosted trees or added feature families beat a ridge logistic on ratings + rest. | winner logistic, everything, ridge 50 · avgLogLoss 0.651 · accuracy 0.614 · ratingsPlusRest 0.654 · trees worse on every fold · note numbers before shootout results were added; later comparable figures ~0.656-0.667 | |
| hist-2026-09-26-search-3000 | 2026-09-26 | rejected | One of 3,000 feature-set versions beats the champion on a proof season it was not selected on. | top10SelectionLL 0.658 · proofLL 0.682-0.687 vs champion 0.6806 · luckGap 0.027 · note every finalist lost on proof; champion holds | |
| hist-2026-09-26-block-search-49k | 2026-09-26 | rejected | Some combination of 14 input families (special teams, finishing+goalie, lineup, rivalry, travel, goalie rotation, schedule spot, GM context, discipline+referees, puck play, grudge scrums, league scoring) beats the champion. | versions 49152 · championProof 0.664 · bestFinalist 0.664 · range -0.001, 0 · note not better; model saturated on public box data. Top-200 family shares: special teams 100%, lineup 84%, finishing+goalie 62% | |
| hist-2026-09-26-travel | 2026-09-26 | rejected | Travel distance, time zones crossed, three-in-four and road-trip length carry information beyond rest and back-to-backs. | all 0.656 · minusTravel 0.656 · note no gain | |
| hist-2026-09-26-goalie-rotation-trap | 2026-09-26 | rejected | Goalie rotation (backup/tired) and schedule-spot (trap game) features lower log loss. | all 0.657 · minusTrap 0.657 · minusRotation 0.657 · note neither adopted | |
| hist-2026-09-26-score-engines | 2026-09-26 | rejected | Monte Carlo minute-level or Markov score engines beat the closed-form Poisson on score and finish log loss. | win indistinguishable · score indistinguishable (3.648 vs 3.650) · finish closed form clearly best; MC and Markov under-predict ties · note late-tied caution measured: last 5 minutes tied x0.77 | |
| hist-2026-09-26-lineups | 2026-09-26 | adopted | Walk-forward skater relative-xG ratings plus starter GSAx and missing-vs-usual lineup lower log loss. | before 0.678 · after 0.677 · note small gain; confident (>=70%) games ~78% right on ~70 of 1,301 | |
| hist-2026-09-26-props | 2026-09-26 | adopted | A learned logistic layer over hand features (shot rate, finishing, opponent shots allowed, venue, ice time, individual xG rate, opposing starter form) beats the base rate for anytime-goal and anytime-point. | goal logistic: 0.381, trees: 0.381, baseline: 0.386 · point blend: 0.59, baselinePtsPerGame: 0.597 · note stars (>=40%) ~3.5 points over-confident; training on 2019-24 was WORSE for points than 2022-24 (scoring drift) | |
| hist-2026-09-26-confidence | 2026-09-26 | rejected | Bootstrap spread across 100 refits marks the games the model gets right more often. | medianSpread 0.024 · note loose estimates were right MORE often; shrink-by-spread gave no gain. Confidence = lean tier, calibrated within ~2 points | |
| hist-2026-09-22-xg-v1 | 2026-09-22 | adopted | A logistic expected-goals model on distance, angle, shot type, strength, rebound, rush and empty-net flags beats the league-average goal rate on an unseen season. | logLoss 0.228 · baseline 0.259 · auc 0.757 · note top decile over-predicted by ~15%; bottom decile under-predicted |
Every job has a written contract
What each job reads and writes, when it runs, what makes it green, what it may spend, how it is stopped, and whether it can be run dry or replayed. Green means it produced its artefact on time. The digests are taken from this morning's content-addressed snapshot, so what a job read and wrote on a given day can be cited exactly.
| Job | State | Kind | Schedule | Green when | Budget | Details |
|---|---|---|---|---|---|---|
| League record pull and nightly model state ·critical | green | Pulls | daily 08:00 Eastern (task: CapOrCup ledger morning) | the model state was rebuilt in the last 36 hours | network: api-web.nhle.com; calls per day not fixed; no spend. one request at a time; never beside another league pull | |
| Licensed feeds pull ·critical | green | Pulls | daily 08:15 Eastern (task: CapOrCup sportradar trial) | roster transactions reach yesterday's league day | network: api.sportradar.com; 900 calls in all; calls per day not fixed; no spend. 900 calls for the whole trial, counted in the client's own ledger; the client refuses the call that would pass the budget | |
| Base-salary pull: Sportradar team profiles (trial) | manual | Pulls | by hand (no scheduled task yet) | its salary rows were written | network: api.sportradar.com; 32 calls a day at most; no spend. 32 calls a pull, counted against the trial's 900 in the shared client's ledger | |
| Contract fields check: league public endpoints | manual | Pulls | by hand (no scheduled task yet) | it re-checked that the league's endpoints still carry no contract fields | network: api-web.nhle.com, api.nhle.com; 2000 calls in all; 100 calls a day at most; no spend. one landing per open target, 1 call/s | |
| Contract reference pull: ESPN core API (reference only, never shown) | manual | Pulls | by hand (no scheduled task yet) | its rows file was written | network: site.api.espn.com, sports.core.api.espn.com; 3000 calls in all; 200 calls a day at most; no spend. rosters for the target clubs plus one contracts call per target, 1 call/s | |
| Contract reference pull: TheSportsDB (reference only until a commercial key) | manual | Pulls | by hand (no scheduled task yet) | its rows file was written | network: www.thesportsdb.com; 3000 calls in all; 200 calls a day at most; no spend. two calls per target at 2.2 s (free tier: 30 a minute) | |
| Licensed contract API pull: cap vendor A (off until keyed) | manual | Pulls | by hand (no scheduled task yet) | off until the owner adds the vendor's key and confirms the licence permits display | no network; 3000 calls a day at most; $0.92 a day at most. vendor limit 4 requests/s and one full import a day; client paces 350 ms and stops at 3,000 calls a UTC day; $25 a month | |
| Licensed contract API pull: cap vendor B (off until keyed and contracted) | manual | Pulls | by hand (no scheduled task yet) | off until a commercial licence, its API base URL and the key exist | no network; 2000 calls a day at most; spend not fixed. client paces 1 call/s and stops at 2,000 a UTC day; price on request | |
| Lineup watch for the ledger | green | Pulls | every 2 hours (task: CapOrCup ledger lineups) | the run log was written in the last 3 hours | network: api-web.nhle.com; calls per day not fixed; no spend. a handful of requests per run, only for games due today | |
| Starting goalie probes | green | Pulls | started by the 08:15 job; probes at 95, 35 and 5 minutes before puck drop and 3 hours after | the watch log was written in the last 36 hours | network: api.sportradar.com, api-web.nhle.com; 8 calls a day at most; no spend. at most 8 licensed calls a day, for at most 2 games | |
| Data contracts (shape, ranges, text encoding, ledger consistency) ·critical | green | Checks | morning assay | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| Game passports ·critical | green | Checks | morning assay | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| Career arcs: aging curves, draft projections, arc labels | green | Checks | morning assay (optional step); rebuilds only when an input changed, so weekly or when a season lands | its artefact carries a stamp inside the cadence | no network; no spend | |
| Content-addressed snapshot | green | Checks | morning assay | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| Contract agreement matrix | green | Checks | morning assay | its artefact carries a stamp inside the cadence | no network; no spend | |
| Control charts and changepoints | green | Checks | morning assay | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| External timestamp of the ledger | green | Checks | morning assay | a proof was written in the last 36 hours | network: freetsa.org (RFC 3161); 1 calls a day at most; no spend | |
| Game model bias audit | green | Checks | 08:00 ledger job and the morning assay | its artefact carries a stamp inside the cadence | no network; no spend | |
| Golden replay | green | Checks | morning assay (skipped with --fast) | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| Identity spine | green | Checks | morning assay | its artefact carries a stamp inside the cadence | no network; no spend | |
| Lineage and stale builds | green | Checks | morning assay | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| Model drift | green | Checks | morning assay | its artefact carries a stamp inside the cadence | no network; no spend | |
| Mutation tests of the checks | green | Checks | morning assay (skipped with --fast) | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| Per-game anomaly scores and duplicate rows | green | Checks | nightly | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| QA crawler: every number on a page against the data | green | Checks | nightly | it wrote its heartbeat with an artefact, on time | network: caporcup.com; 450 calls a day at most; no spend. one page at a time, at least 1.5 seconds apart | |
| Red team: published values re-derived by an independent path | green | Checks | nightly | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| Source verifier | green | Checks | morning assay | its artefact carries a stamp inside the cadence | network: the cited public pages, allowed hosts only; 40 calls a day at most; no spend. 40 cited pages a morning, oldest check first | |
| Contracts, rights and budgets ·critical | green | Governs | morning assay, before the gate | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| Morning assay and the gate ·critical | green | Governs | daily 08:15 Eastern, inside the feeds job, before the data deploy | it wrote its heartbeat with an artefact, on time | no network; no spend | |
| Cost governor (model account) | green | Governs | morning assay | it wrote its heartbeat with an artefact, on time | network: openrouter.ai; 1 calls a day at most; no spend. reads the balance; spends nothing | |
| Hub controls (pull and apply) | green | Governs | first step of the morning assay | it wrote its heartbeat with an artefact, on time | network: hub database; 4 calls a day at most; no spend | |
| Decision cards for the owner | green | Notifies | morning assay | it wrote its heartbeat with an artefact, on time | network: hub database; 6 calls a day at most; no spend. one card per decision, keyed by content, never repeated | |
| Morning note to the hub inbox | green | Notifies | morning assay, only when something needs a person | it wrote its heartbeat with an artefact, on time | network: the estate's mail relay; 1 calls a day at most; no spend | |
| Daily data deploy ·critical | green | Deploys | after the morning assay | a build was promoted in the last 36 hours | network: vercel.com; calls per day not fixed; spend not fixed. one staged build a day; promoted only after its pages answer | |
| Guarded code deploy | manual | Deploys | by hand, when a change is ready | run by hand; its record is the deploy state | network: vercel.com; calls per day not fixed; spend not fixed | |
| Build abroad.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build advanced.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build champion-state.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build champion.json | unverified | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build coaching.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build contracts | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build deployment.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build draft-class.json | unverified | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build durability.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build game-model.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build goal-archive.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build impact.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build injuries.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build integrity.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build involvement.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build league-ladder.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build ledger-public.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build matchups.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build measures-2025.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build model-cards.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build officials.json | unverified | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build player-props.json | unverified | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build player-ratings.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build projected-lines.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build prospect-rankings.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build schedule-strength.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build shotmaps.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build stat-cards.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build transactions.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build video-reads.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build white-noise.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend | |
| Build xg-2025-26.json | green | Builds | when its inputs change | its output is not older than any of its inputs and holds every game its inputs hold | no network; no spend |
What may be shown, and until when
The licence is stored with the data, not remembered. Every source says what it supplies, what the site shows from it and what it never shows. A licence with an end date is counted down here; when it ends, its files are emptied by the build and the gate closes if they are not.
| Source | State | Tier | Licence | Supplies | What the site shows | Never shown | Details |
|---|---|---|---|---|---|---|---|
| The league's public statistical record | ok | measured | no end date | play-by-play, shift charts, schedules, rosters, game logs, season summaries, draft records | facts and statistics computed from them, with the league named as the origin | the league's video, audio, photographs, logos and written text | |
| The league's goal clips, condensed games and stories | ok | measured | no end date | links to the league's own clips and stories; counts of stories | links that play on the league's own player, and counts | rehosted video or audio; frames sampled for a model read are deleted after the read; story text | |
| Sportradar NHL feed (trial) | ok | licensed | ends 2026-10-28 (28 days left) | roster transactions, injury listings, depth charts, starting goalies; base salary per player (team profile), kept as a salary reference, never a cap hit | the feed's facts, named as a licensed feed, while the trial runs | anything from the feed after the trial ends without a contract | |
| Public pages read by the contract desk | ok | researched | no end date | contract terms, cap hits and clauses stated on the page | the fact, its tier and a link to the page it was seen on | the article's text beyond the short quote that carries the fact | |
| Other leagues' public statistics | ok | measured | no end date | season totals by player in the leagues prospects and free agents play in | statistics computed from them, with the league named | media and text | |
| Attention counts (Wikimedia page views, public social posts, news volume) | ok | measured | no end date | how many views, posts and stories, by day | counts and indices only | the posts and articles themselves | |
| Broadcast audio (the ear) | not-connected | judgement | not connected | counts of mentions in broadcasts | nothing until written permission exists; then counts only, never words | transcripts, quotes, audio | |
| Betting market prices | not-connected | licensed | not connected | closing prices, as a benchmark for the game model | nothing until a licence exists | prices without a licence that allows display | |
| ESPN core API (contracts collection) | ok | licensed | no end date | per-athlete contracts collection (populated for the NBA; empty for the NHL when checked 2026-09-29) | nothing: undocumented API, terms do not licence commercial display; used only to raise disagreements | every value | |
| TheSportsDB (contracts lookup) | ok | licensed | no end date | crowd-sourced player contracts (lookupcontracts) | nothing on the free test key (non-commercial); reference only | every value until a commercial key and a check of its accuracy |
Sourcing rule: 5 commercial cap sites are never scraped or cited, and an article that credits one of them for a fact is not a source for that fact. Their data may appear only through a paid licensed API whose licence allows display, named as the provider where the licence requires it. Every published file is scanned each morning and again by each deploy; any other mention closes the gate. Fine for finding leads, never cited for a fact: aggregators, fan blogs, cap tables.
Budgets against their ceilings
A job that would pass its ceiling is stopped and the stop is logged. Model work by any desk refuses to start while the account reads exhausted or the hub does not allow spend.
| Budget | State | Used | Ceiling | Note |
|---|---|---|---|---|
| Licensed feed calls (whole trial) | ok | 82 | 900 | 35 today |
| Model account credit | exhausted | $55.12 | $55 | hub ceiling $1 a day, $10 a month; spend not allowed by the hub |
Managed from the BaldwIndustries hub
This business's Assay tab in the hub is the master switch and spend ceiling; its decision queue holds every choice that needs a person; the morning result arrives as an in-app notification and, when it needs action, as a note in the hub inbox.
Open the Assay tab in the hub. Standing owner approvals (the daily data deploy, the licensed feed pulls) are not governed by this switch; it governs the new automated actions listed below.
- Master switch (autonomy_enabled). Whether the assay may take any automated action beyond measuring and publishing status: model spend, challenger adoption, verifier volume. Off = measure and report only.
- Pause. Stops every automated action immediately, with the reason shown on this page and in the morning note.
- Authority. suggest-only: every adoption is a card for you; low-risk: calibrators may be adopted on their pre-registered rule; all-except-email: challengers may be adopted on their rule.
- Daily and monthly $ ceilings. The cost governor's limits for model calls (video reads, extraction, the assistant). Spend to date is shown beside them.
- Decision queue cards. Register entries with proof results, correction tickets, identity batches and budget requests. Approve or reject in the hub; the morning job applies it and records who decided.
Waiting for a decision:
- assay_identity_review 6 player-identity pairs need a person (identity:2026-09-29, queued 2026-09-29)
- assay_source_review 13 cap-hit facts not shown on their cited page (source-review:2026-09-29, queued 2026-09-29)
- assay_budget Model account exhausted: top up or leave dark? (budget:openrouter-2026-09, queued 2026-09-29)
- assay_ledger_defects 29 contract rows contradict themselves (ledger-defects:0320e75814, queued 2026-09-29)
- assay_xg_clock Expected goals: repair the event clock the model reads? (xg-clock:7ef5afd2f8, queued 2026-09-29)
- assay_adopt_prospect Prospects: formula blended with draft position? (prospect:prospect-pick-blend-v1, queued 2026-09-29)
- assay_sourcing_rule Sourcing rule applied: 18 cap hits and 123 clause facts withdrawn (sourcing-rule:916acd966f, queued 2026-09-29)
Applied decisions:
- prospect:prospect-pick-blend-v1 approved: register prospect-pick-blend-v1 -> adopted
- sourcing-rule:916acd966f approved: sourcing-rule 916acd966f -> approved (recorded for the desk in desk-decisions.json)
- xg-clock:7ef5afd2f8 approved: xg-clock 7ef5afd2f8 -> approved (recorded for the desk in desk-decisions.json)
Model spend against the ceiling
The account is out of credit: the assistant answers extractively, video reads and extraction are paused. Topping up is the owner's step; the ceiling stays in the hub panel.
What the assay found in the record
Real defects and gaps in the data on disk, as measured this morning, with what each does to the measures.
- shift-sums (2020-21): 35 games whose shift charts fail the sums check (doubled or half-missing charts); their minutes are still counted where the shooter check passes
- shift-sums (2021-22): 36 games whose shift charts fail the sums check (doubled or half-missing charts); their minutes are still counted where the shooter check passes
- shift-sums (2023-24): 92 games whose shift charts fail the sums check (doubled or half-missing charts); their minutes are still counted where the shooter check passes
- exposure (2024-25): shift charts cover 95.7% of expected minutes
- shift-sums (2025-26): 36 games whose shift charts fail the sums check (doubled or half-missing charts); their minutes are still counted where the shooter check passes
- events-gap (2025-26): 80 games without an events file: no zone starts, sequences or puck battles for them
Video read coverage
Video-derived numbers rank no one across clubs until coverage is balanced; human verdicts are the ground truth for read precision. Coverage by club: CGY 3, PIT 3, BUF 3, NYR 3, PHI 3, OTT 3, SEA 2, CHI 2, MIN 2, NJD 2, DAL 2, STL 2, EDM 2, CBJ 2, DET 2, CAR 2, FLA 2, BOS 2, NSH 2, TBL 2, LAK 2, UTA 2, VAN 2, SJS 2, VGK 2, ANA 2, TOR 2, WSH 1, MTL 1, WPG 1, NYI 1.
Reported numbers
Anyone can report a number from its provenance popover or the corrections page; the ticket goes to the hub inbox and appears here with its resolution.
No reports on file. Report a number.
The assay line
Every morning: pull the hub policy and apply decisions → game passports → data contracts → control charts → golden replay → mutation tests → identity, agreement matrix, verifier, timestamp → bias audit and drift → the gate → stat cards, model cards, this page → hub cards and notification → the data deploy, which reads the gate. Seasons on disk: 2018-19, 2019-20, 2020-21, 2021-22, 2022-23, 2023-24, 2024-25, 2025-26.