HoopQ.
Open NBA analytics, built possession-by-possession.
Independent NBA analytics. Six ways to measure player impact, per-game career trajectories, and how well any lineup actually fits together. 25 seasons of data, free to use.
Six ways to measure impact
All in points per 100 possessions. When they agree, we're confident. When they disagree, that's usually the interesting part.Explore
Featured Players
Top 15 by 2025-26 Q-Rating Total Impact · click any card for the full profile.Leaderboards
Top 3 by each metric · click any name for the profile, "See all →" for the full ranking.The classical impact metric of NBA analytics. For every player, RAPM (Regularized Adjusted Plus-Minus) asks: how many more or fewer points does the team score per 100 possessions when this player is on the floor, once you have adjusted for the quality of everyone else who was on the floor at the same time? Positive means the player helps their team's scoring margin; negative means they hurt it.
How to read it. Values are pooled across 26 seasons (2000-01 through 2025-26, 5.49M possessions), so this is a career-scale number, not a single-season one. Click any column header to sort. Click any row to open the full player profile.
How it's built. A ridge regression where every possession is one data point, every player gets one offensive and one defensive coefficient, and the model finds the coefficients that best explain observed point margins. A Bayesian prior derived from box-score stats stabilizes the estimate for players with fewer possessions, so a rookie with 200 minutes doesn't swing to +30 by accident. See About for the full methodology.
Clutch = Q4 or overtime, clock ≤5 minutes, |margin| ≤5. Two independent measurements combined into one leaderboard:
Action Q (pts above league expectation per 100 clutch actor events): for each clutch event, we compare the observed reward to the per-season bucket mean. Playmakers get credit on assisted makes so passing value doesn't disappear into the shooter alone. EB-shrunk κ=500 multi / κ=100 single season.
Clutch ORAPM / DRAPM: Ridge regression on ~230k clutch possessions (44k WNBA) with α=25,000. Captures lineup-level offensive and defensive presence in the clutch subset (rim deterrence, gravity, off-ball value). Sign conventions match regular RAPM: ORAPM positive = good, DRAPM negative = good.
Total = Action + ORAPM − DRAPM. Positive = good overall clutch player.
Each projection combines a recency-weighted average of the player's most recent three seasons (weights 0.2 / 0.3 / 0.5) with the age-indexed league aging-curve delta from their current age to the target season's age. Applied independently to RAPM, Q-Rating, offensive presence, and defensive presence.
Aging curves are possession-weighted means by actual age, from birthdates when available and career year otherwise. The 95% credible interval reflects year-over-year drift plus the variance of the player's recent seasons.
League aging curves
Per-action aging
Coach effect isolated via a two-step fit. Step 1 fixes each player's expected contribution using a player-quality prior. Step 2 regresses the per-possession residual (actual minus player-implied points) on the offensive and defensive head coach one-hots plus season fixed effects. The coach coefficient captures whatever pts/100 the players + league scoring era couldn't explain, read as coach + team-baseline effect for the coach's tenure.
Two priors, two views: Coach Q-Rating (default) uses the full deep-RL Q-Rating, action head plus presence, as the player prior. Coach RAPM swaps in box-score-prior player RAPM instead, with the defensive term as a 50/50 blend of def_rapm and Q-Rating def presence so rim-deterrence gets credit. The two agree overall at Pearson r=0.76 on the ≥250-game qualified pool; their defense signals correlate at 0.97 but offense signals diverge (r=0.49). When both metrics agree the coach's effect is robust; when they disagree the ranking is method-sensitive on that coach (typically an offensive-scheme case Q-Rating captures better than box-score RAPM).
Beyond the ratings. Click any coach row for their profile: Kalman-smoothed season trajectory (raw shrunk + smoothed line), aggregate shot chart of teams under this coach (frequency + efficiency), and a "Coach style" widget showing which action buckets they over- or under-index on relative to league mean.
Sign convention: off positive = raises team pts/100; def negative = lowers opponent pts/100. Total = off − def (positive is good, mirroring player RAPM). Career shrunk with κ=40000 possessions; per-season with κ=12500 plus a 1-D Kalman smoother across seasons for trajectory. Click a coach row for their season-by-season chart.
Win Probability Added credits each event by its contribution to the offensive team's win probability. A clutch shot that moved WP by 5 percentage points contributes +0.05 WPA. Per-season models are trained across 2000-01 through 2025-26; the selector defaults to the most recent season.
Total WPA is cumulative and rewards volume as well as efficiency; WPA per event isolates efficiency. Defensive WPA only credits explicit defensive events (steals, blocks, defensive rebounds, fouls). Rim deterrence and off-ball spacing effects are not captured, so elite rim protectors read systematically low here.
Playstyle vectors are the per-player distribution over action classes for the 2024-25 season. Each column shows the fraction of that player's actions that fall in the given category.
Click a row to open the player's signature bar chart and their ten most-similar players by cosine similarity.
Per-player play-type distribution from a hierarchical RL options framework. Each offensive action is labeled as one of twelve play types: pick-and-roll handler, pick-and-roll finisher, isolation, post-up, cut, putback, spot-up, off-screen, transition, free throw, rebound, or other.
Columns show the fraction of a player's offensive actions in each option. Defensive events are tracked separately as per-100-poss steals, blocks, and fouls. Click a player row to see their per-option success breakdown.
For each award, we assemble a candidate pool per season, compute features from our existing metrics, and fit a conditional-logit (softmax within season) on historical voting shares scraped from Basketball-Reference 2000-2025. Predictions cover all seasons including the current one, so you can compare historical model picks against actual winners.
Features: MVP uses Q-Rating total impact, PPG, AST/G, team wins. DPOY uses def_presence (flipped), blocks/g, steals/g, team wins. ROY same as MVP filtered to first-season players. COY uses coach Q-Rating + wins-over-projected (from championship model) + team wins. All-NBA / All-Defensive / All-Rookie use the same features as their singleton counterparts but train on top-15 / top-10 / top-10 team selections.
Retrospective accuracy (top-1 hit rate): ROY 77%, All-Rookie 92%, All-NBA 69%, All-Defensive 65%, MVP 46%, DPOY 35%, COY 23%. COY is voter-narrative-driven so hardest to predict from stats.
End of regular season is the default view. Team strength = per-player Q-Rating presence weighted by possession share + career pair-chemistry contributions weighted by shared possessions. Raw strength is calibrated to observed net rating, converted to projected wins, then Top-8 seeds by projected wins run a fixed bracket (1v8, 4v5, 2v7, 3v6 → semis → conf finals → Finals). Series outcomes are sampled from a logistic model fit on ~390 historical NBA playoff series (156 WNBA). 30,000 trials per season. Retrospective accuracy (2005-06 → 2024-25): champion in top-1 25%, top-3 60%, top-5 95%, top-8 100%.
Preseason view estimates player values ONLY from data through the prior season. Chain: Kalman-preferred Q-Rating presence + V-Rating ensemble (0.7/0.3) + Kalman age drift + draft-slot rookie prior + injury-availability discount + newcomer skill / role compression + share-weighted pair chemistry. Raw strength is centered within each season (era-normalized) and calibrated to observed net rating, with a continuity + prior-year-residual adjustment for the portion of last year's result the roster-sum misses. Same downstream 30k-trial bracket sim. For past seasons it uses each team's opening-night roster (reconstructed from their first eight games), answering "how would we have projected the season before it started?" For the upcoming, not-yet-played 2026-27 season there are no games yet, so it reads current rosters, assigns minutes with a position-aware depth-chart model, and applies known long-term injuries (scraped from ESPN). Retrospective accuracy (2005-06 → 2024-25): champion in top-1 45%, top-3 60%, top-5 75%, top-8 95% (avg rank 3.3). Point projections are noisier than end-of-season (~7 wins MAE vs ~4.5). Misses concentrate on mid-summer trades and steep breakout years the model can't extrapolate.
See the Championship odds glossary entry for Brier score, log-loss, per-conference Spearman rank correlation, reliability diagram, and per-round hit rates.
Unlike a seed-your-own playoff sim, March Madness has a given 64-team bracket. We take the real bracket (teams, seeds, matchups from ESPN) and simulate it forward 20,000 times. Team strength = per-player Q-Rating presence weighted by possession share + career pair-chemistry contributions, calibrated to net rating. Single games (not series) are sampled from a logistic fit on ~1,080 historical tournament games. First Four play-in games are simulated too.
Retrospective accuracy (2008-2026, 18 tournaments): champion in top-1 39%, top-3 72%, top-5 89%, top-10 100%; average champion rank 2.6 of ~64; Final Four recall 69%. The model's strongest team by our rating reached the Final Four every year it was clear chalk (UConn 2009-2016). Misses are the Cinderellas (2023 LSU, a 3-seed, ranked 8th).
Odds shown are pre-tournament (computed from the season's roster + our team strength), not updated as games are played. The bracket view shows the actual matchups with our advancement probabilities; actual results are marked. Coverage 2008-2026 (2020 cancelled). The Odds toggle switches between In-season (rated on this season's play) and Preseason (forward-looking, rated only on players' prior-season value, the way a real bracket forecast works); the Preseason vs actual view lines the two up so over- and under-performers stand out. See the March Madness glossary entry for the win model, bracket reconstruction, and full retrospective accuracy.
Each team's strength is our March Madness chain (per-player Q-Rating presence weighted by possession share + pair chemistry, calibrated to net rating), then adjusted for strength of schedule: a team that earned its rating against strong opponents gets a boost, against weak opponents a penalty. Without this, mid-majors that dominate a weak schedule look elite. Schedule comes from the full regular-season game log; the adjustment iterates each team's rating against its opponents' adjusted ratings.
This is a model rating and can disagree with a team's record or seed (that is the point). The W-L column is shown so you can see where the model likes a team more (or less) than its results. Covers all 360+ Division I teams per season, 2006-2026. Powered by the same team strength as the March Madness odds.
Almost no public site can build this: it needs both college and pro players scored with the same impact metrics, which we have. We train three models on college inputs (Q-Rating, V-Rating, iRAPM):
Reach = probability a college player reaches the WNBA with real minutes. Projected impact = expected WNBA per-100 impact if she reaches (the target is observed pro impact, not our shrunk career rating, which compresses the range and kills the signal). Projected outcome = the five-part bar (Never sticks, then Fringe/Depth, Rotation, Starter, All-WNBA), a calibrated probability distribution that sums to 100%; hover a segment for the exact %. Comp = the WNBA player whose playstyle vector is closest (cosine similarity on the shared action buckets). Everything is validated out-of-sample with grouped-by-player cross-validation so a player's own seasons never leak into her prediction, and the probabilities are calibrated (expected calibration error 0.00 reach, 0.02 tiers).
Honest limits. The impact signal is real but modest (college dominance only partly translates, and the model is deliberately conservative, so it will not crown the next superstar outright). Coverage is US-college only: international players never played NCAA ball, so they are unmatchable and absent. Reach is the strongest piece; impact and longevity are directional. See the WNBA Draft Projection glossary entry for the full method and validation.
1. Per-player contribution =
off_presence_shrunk + (−def_presence_shrunk). These come from the Q model's leave-one-out lineup-swap procedure, and they capture rim deterrence, spacing, gravity, and off-ball value (the lineup-level effects that pure action-bucket ratings miss). Both sides are shown "good-direction" (positive = good).
2. Pair chemistry = Σ across all C(5,2)=10 pairs. We use the observed coefficient from the multi-year ridge fit (residual above each player's individual contribution). Pairs that haven't played together show 0.00. A Q-model-based predictor for hypothetical pairs is on the roadmap but not yet wired into this page.
3. Net pts/100 = Σ per-player contributions + Σ pair chemistry.
Caveats. Triple-and-higher chemistry isn't modeled. Multi-year ratings are the most stable input; single-season ratings are noisier and overweight 2025-26 small samples. Presence-based defense does capture rim deterrence and off-ball value (that's the whole point of the leave-one-out lineup swap), but the magnitude depends on how much between-lineup variance the Q model saw during training. Players with very repetitive teammates may have their deterrence partly absorbed into "average team defense."
off_presence_shrunk + (−def_presence_shrunk) from the Q model's leave-one-out lineup swap. Shared floor time is approximated as (mini × minj) / 48 assuming independent rotations. Wins = 41 + net_rating × 2.7 (Basketball-Reference rule of thumb, ~2.7 wins per +1 net rating).
Confidence band. The ± range reflects uncertainty from small-sample players and unmeasured factors (coach, health, schedule strength). Wider band = more of your roster is small-sample or hypothetical.
Caveats. No schedule strength. No injury / DNP modeling. Assumes coach distributes minutes as configured. Rotation-independence assumption for pair chemistry ignores the fact that starters typically overlap much more than starter-bench pairs, so pair contributions of your top starters are slightly under-counted.
For each side, we recompute the calibrated net rating using the same chain as Team Builder: player contribution = Σ (mini/48) × (off_presence − def_presence), pair chemistry = Σ shared-minute-weighted (chem_off − chem_def), raw = player + chem, calibrated =
−6.15 + 0.80 × raw (NBA) / −2.83 + 1.27 × raw (WNBA). Projected wins = 41 + 2.7 × calibrated_net (NBA) or 22 + 1.45 × calibrated_net (WNBA). Championship odds run a joint 5,000-trial Monte Carlo bracket sim where BOTH swapped teams update simultaneously; the remaining teams are held at their current-season baseline.
Caveats. Uneven trades (N-for-M) will leave rosters with total minutes ≠ team target; the projection is still computed but the roster note will flag it. No salary cap or contract logic. No coach fit, no positional balance check.
Every pair of players who have shared the floor gets a number here. Positive means the pair lifts each other's game (elite chemistry, like Curry and Draymond's screen-and-roll partnership). Negative means they undermine each other. The value is in points per 100 possessions on the same scale as RAPM, on top of what each player already contributes individually.
Season filter. "Any season" gives a stable career-long average, one number per pair. Pick a specific year and you get a rolling 3-year window that captures how a duo has been playing lately. Curry and Draymond dip through their injury years; Gobert and Naz Reid emerge as an elite defensive pair in 2024-25.
Context filter. Splits the data into three game situations: clutch (last 5 minutes of Q4 or OT, within 5 points), regular (within 15 points, not clutch), and blowout (winning or losing by more than 15). Curry and Draymond drop from +4.1 in regular time to +1.4 in clutch; defenses take away their favorite hand-off late. Dirk Nowitzki and Jason Terry surface as one of the best clutch pairs of their era. Fewer pairs qualify in the clutch subset (about 2,500 vs 34,000 in regular time) because each pair needs enough clutch minutes to register. The context split is only available on the career-long view, so combining it with a specific season falls back to the career pool.
How it's built. A ridge regression across 26 seasons of NBA play-by-play (5.49M possessions). The model has one coefficient for each player and one for each same-team pair, fit jointly, so the pair coefficient is the extra effect on top of what the two players do individually. Held-out R² improves over a player-only baseline, so the pair signal validates out of sample (not just in-sample overfitting).
Presence Impact: Methodology
V₀-based on/off with cluster-robust CIs, multi-team stint filtering, and cross-model comparison. Ships in Rankings and on every player profile.
About HoopQ
HoopQ is an independent NBA analytics site. Six ways to measure player impact: Multi-season RAPM (the classical baseline), Q-Rating (grades every on-court decision), V-Rating (grades whether the decision actually helped), Interpretable RAPM and Interpretable WPA (both break impact down into categories like scoring, playmaking, and defense), and Presence Impact (how much better your team plays when you're on the floor). Plus per-game career trajectories and lineup fit. Every metric is built possession-by-possession from raw NBA play-by-play (every shot attempt, assist, rebound, foul, turnover, block, and steal), covering 2000-01 through 2025-26: 13.0M events, 5.49M possessions, 30,815 games.
The Glossary explains every metric on the site, what it measures, how it's computed, how to read it, and where it falls short. The Evaluation page reports how well each metric behaves under stress: year-over-year stability, sample-size sensitivity, agreement across metrics.
Not affiliated with the NBA. Independent hobby project. Data is derived from public play-by-play feeds. Built and maintained by Ben Jenkins. Contact: hoopqhq@gmail.com. Corrections, methodology questions, and collaboration ideas welcome.
The models extend research that won Best Overall Paper at the 2026 MIT Sloan Sports Analytics Conference Research Paper Competition: Deep Reinforcement Learning for NBA Player Valuation: A Temporal Difference Approach with Shapley Attribution.
Multi-season RAPM (2000-2026)
What it is. Regularized Adjusted Plus-Minus, the standard "all-in-one" basketball impact metric. Each player gets two coefficients: an offensive rating (points contributed per 100 possessions when they're on offense) and a defensive rating (points conceded per 100 possessions when they're on defense, where more negative is better). Total RAPM = ORAPM − DRAPM.
How it's computed. For every NBA possession from 2000-01 through 2025-26 (5.49M possessions across 30,815 games), we record the 5-on-5 lineup and the points scored. A sparse linear regression then solves for each player's marginal contribution while controlling for who they shared the floor with. We use a Bayesian prior built from each player's box-score profile (per-100 stats: points, assists, turnovers, rebounds, steals, blocks, etc.) so that low-minute players shrink toward what their box stats predict rather than toward zero. Ridge regularization strength (λ) was chosen by 3-fold cross-validation on held-out possessions.
Strengths. Captures impact you can't see in the box score: screens, spacing, defensive positioning. Multi-season pooling reduces single-season noise by ~10× (the GOAT-era list is dominated by Jokić, LeBron, Embiid, Chris Paul, Curry, Giannis, Kawhi, Garnett, Luka, Nash, consensus-correct).
Limitations. No causal claim, RAPM is correlational. Assumes linear and additive player effects (no interactions or context terms, so pair chemistry and lineup fit fall outside the model). Small-sample players (Wembanyama, Chet Holmgren: 2 seasons each) still appear in the top 25 with wide uncertainty. Doesn't separate role from talent (a great defender on a great defense gets partial credit for teammates' help).
Q-Rating (RL Q-Model)
What it is. A deep-RL player rating with two complementary flavors per side, both in points per 100 possessions (RAPM units):
- Off / Def (action), sum of per-action-bucket contributions. Each event is credited to its actor; tells you which buckets (3pt pullups, rim finishes, blocks, rebounds, etc.) drove the value.
- Off / Def (presence), neural ORAPM / DRAPM via a leave-one- out lineup swap on the trained Q model. Captures value not tied to any single event: spacing, screen-setting, rim deterrence, screen navigation. Closer analog to RAPM.
How it's computed. Each NBA possession is an RL episode; each
event (shot, rebound, turnover, foul) is a step. A neural network learns
Q(state, action, player), the expected remaining points in the
possession given the game context, action, and actor. Player identity is a learned
embedding; lineups are pooled by a 4-head self-attention layer.
Action ratings = average advantage (Q with actual embedding −
Q with league-average embedding) × usage per 100 possessions, summed per side.
Presence ratings = average Q change at the start of each possession
when the player is replaced with a neutral baseline (the mean of all real player
embeddings), averaged across every possession they were in the off / def lineup, × 100.
Empirical-Bayes shrinkage. Per-100-poss rates have high variance
for low-sample players, so every rating is shrunk by n / (n + κ) with
κ = 8000 on-floor events (~1 season for a half-time player). Veterans keep ~90%
of their raw rating; small-sample noise players collapse toward 0.
How to read it. Presence is the headline number; action shows where the value came from. For elite defenders the two diverge, Def (action) is mostly blocks/steals/rebounds, Def (presence) includes the value of just being out there. Player embeddings also cluster meaningfully (Curry near Lillard near Trae; Jokić near Sabonis), so the model isn't just memorizing role.
Presence Impact (V₀ on-off)
What it is. A per-100-possession on-off rating built from V₀, the possession-start value scored by the V-Rating causal transformer. Not the same as Q-Rating LOO presence: that estimator swaps one player out of a trained neural lineup; Presence Impact is a pure observational split of V₀ by whether the player is on the floor.
- Δ Off = mean(V₀ | on court, their team on offense) − mean(V₀ | off court, their team on offense). Positive = team's expected value is higher when they're on the floor.
- Δ Def = mean(V₀ | off court, their team on defense) − mean(V₀ | on court, their team on defense). Sign is flipped so positive = good defender (they reduce opponent value).
- Δ Total = Δ Off + Δ Def.
Cluster-robust 95% CI. Possessions within the same game are correlated (opponent, tempo, night's shooting variance). Treating them as independent understates variance by ~4×. We group possessions by game and use the sandwich-formula SE. Empirical post-fix coverage: 68% (NBA) / 84% (WNBA) against a 3-year rolling anchor. On the UI, green CI brackets mean the interval excludes zero (statistically distinguishable from average); gray means the sample can't tell them apart from average.
Multi-team stint filter. A mid-season trade doesn't count team A's post-trade games as "off floor" opportunities. Each (player, team) pair's game range is derived from actual in-lineup appearances; without-samples outside that range are dropped. Before the fix, 12% of NBA rows and 70% of WNBA rows had inflated without-pools. Post-fix example: Nikola Vučević 2024-25 went from Δ Total −3.9 (worst in the top 15) to +0.3 (neutral).
Teammate Lift. The chem-adjusted variant subtracts the expected pair-chemistry bump from raw Δ, isolating individual value from teammate lift. On player profiles you see chem_bump_off, chem_bump_def, and the adjusted total. Jamal Murray 2024-25 raw Δ Total +6.4 collapses to +0.35 after subtracting the Jokić-driven pair bump; the exposure itself is a story no other public site publishes. Available for both NBA (2000-2026) and WNBA (2002-2026) with the caveat that pre-2004 pair-chem values are damped by degenerate rolling-window sample sizes.
On-off blind spot. Presence under-credits stars on deep teams because their bench faces the opponent's bench and looks fine without them. Poster child: A'ja Wilson ranks #669 on the career pool despite being a 3× WNBA MVP. Same pattern in NBA: SGA (deep OKC), Tatum (Celtics), Bam Adebayo (Heat). This is a fundamental limit of any on-off estimand, not a bug. Cross-verify against RAPM, Q-Rating, V-Rating, and iRAPM on the Impact Metric Comparison strip on every player profile; when Presence disagrees with the rest, the disagreement is usually the deep-team blind spot in action.
How to read the leaderboard. Two views. "All-time single-season peaks" pools every (player, season) row so LeBron 2008 and Jokić 2022 co-appear as historical peak seasons. "Single year" filters to one season, one row per player. WNBA defaults to a 3-year rolling pool because single-season sample sizes are too small for tight CIs. Details on the Presence Methodology page.
V-Rating
What it is. The state-value sibling to Q-Rating. A causal transformer estimates the state-value function V(s) at every prefix within a possession, where V(s) is the model's expected total possession reward given the events observed so far. For each event, ΔV = V(s_t) − V(s_{t−1}) is the value change attributable to that event and is credited to the actor. Aggregated across a season, this is a player's contribution to their team's expected offensive value (EPV, à la Cervone/D'Amour/Bornn 2016).
How it's computed. A lineup-context token seeds V₀; positions 1..T are event tokens (event type, action bucket, actor, is_assisted, is_made, shot location, clock, period, margin). Trained on 2000-01 → 2025-26 pooled possessions (~5.5M possessions). Attribution: shot events split into "choice" (bucket-conditional expected reward) plus "luck" (residual vs bucket mean, EB-shrunk); assists take 60% of shooter's choice; rebounds get 30% credit (contested); turnovers get 40% actor share; on-floor lineup shares the V₀ gap so passers/creators get credit for setup value without needing explicit pass events, and elite defensive lineups get credit for rim deterrence.
How to read it. Same points-per-100 scale as Q-Rating and RAPM, MVPs score +6 to +10 per 100 actions. Q-Rating is Q(state, action, player), action-conditional. V-Rating is V(state) with per-event attribution, state-based, so it captures value that isn't tied to a specific action-value bucket (setup, spacing, defensive presence). Cross-verifying: when Q-Rating and V-Rating agree the player's impact is well-priced; when they disagree, the gap is informative about what each metric is measuring.
Limitations. No tracking data, so off-ball motion (screens, cuts, spacing) is captured only via the on-floor lineup share (Fix 3), not as explicit events. Shot luck is heavily shrunk to avoid single-season shot-making noise.
Clutch Q-Rating
What it is. Q-Rating restricted to clutch moments: Q4 or overtime, clock ≤5 minutes, |score margin| ≤5. Combines two independent measurements into one leaderboard.
- Action Q: per-100 shot quality above league expectation on clutch actor events. Playmakers get credit on assisted makes so passing value doesn't disappear into the shooter alone. Empirical-Bayes shrunk with κ=500 (multi) / κ=100 (single season).
- Clutch ORAPM / DRAPM: Ridge regression on ~230,000 clutch possessions (~45,000 for WNBA) with α=25,000 aggressive regularization. Captures lineup-level offensive and defensive presence, rim deterrence, off-ball value, and creation gravity that per-action doesn't see.
How to read it. Sign conventions match regular RAPM: Action Q and Clutch ORAPM positive = good; Clutch DRAPM negative = good (allowed fewer pts / 100, DRAPM convention). Total = Action + ORAPM − DRAPM. Positive total = good overall clutch player.
Why this fixes a real gap. Action-only clutch rewards shooters (Curry, Nash, Redick, Peja, Reggie Miller, Ray Allen top out) but underrates high-usage playmakers whose clutch impact is via others (LeBron, CP3, Draymond), rim protectors whose value is deterrence (Wemby, Gobert, Fowles), and off-ball gravity (Klay). Adding Clutch ORAPM/DRAPM via lineup ridge captures those effects and produces a balanced two-way clutch leaderboard: LeBron went from action-rank #296 to total-rank #28 (action −0.46, ORAPM +1.39, DRAPM −0.83). Dirk Nowitzki is #1 all-time by total.
Cross-verify against iWPA. Interpretable WPA measures leverage-weighted win probability change per clutch event. Clutch Q-Rating measures pts per 100 clutch events + lineup ridge. Different units, similar concept. Correlation ≈ 0.31 on shared players (positive but distinct), because iWPA weights by state leverage while Clutch Q-Rating weights by event count.
Limitations. Only 4.7% of events are clutch, so per-season RAPM at ~2-10k possessions is too noisy and the CSV emits action-only for single-year. Multi-season pool (229k NBA / 45k WNBA possessions) is where the ORAPM/DRAPM signal lives. Very-low-clutch-min players (rookies, deep bench) are dropped by the min-events filter.
Interpretable RAPM
What it is. A ridge regression variant of RAPM that decomposes each player's impact into 8 interpretable roles: Scoring, Playmaking, Off Reb, Def Reb, Def Actions, Def Presence, Off Presence, Turnovers. Unlike standard RAPM which gives one opaque coefficient per player, Interpretable RAPM tells you where that impact comes from.
How it's computed. One design matrix row per PBP event, target y = points scored on that event (0, 2, or 3 typically). Each row's columns are: (a) offensive lineup on floor (5 indicators), (b) defensive lineup on floor (5 indicators), (c) actor's specific (player, bucket) column with sign matching offense/defense, (d) assister's column on assisted shots, (e) blocker's column on blocked shots, (f) stealer's column on stolen turnovers. Sparse ridge (α = 25,000) with per-event actor-per-bucket coefficients rolled up post-fit into the 8 role categories.
How to read it. Multi-season 2000-2025 pooled, so a Total of ~+40 for MVPs equals ~+8 per season on the classical RAPM scale. Scoring, Playmaking, Off Reb, Def Reb, Off Presence, Turnovers: positive = good. Def Actions and Def Presence: NEGATIVE = good (matches DRAPM tradition, the player reduces opponent expected value). UI applies visual color inversion on the def columns so green always renders when the player is above-average defensively.
What each role captures. Scoring: actor value on shot events (rim, midrange, 3PT, FT buckets). Playmaking: assister-specific coefficient on assisted shots, cleanly isolates pass value from finish value. Off Reb: β_actor[oreb] with y augmented by same-possession points scored after the rebound — elite offensive rebounders whose boards lead to points get positive credit. Def Reb: β_actor[def_reb] sign-flipped, with y augmented by points scored by the rebounding team on their NEXT possession — rewards elite rim protection + outlet passing that fuels transition. Def Actions: blocker + stealer + defensive fouls. Def Presence: on-floor defensive lineup coefficient (rim deterrence, help defense). Off Presence: on-floor offensive lineup coefficient (spacing, gravity). Turnovers: actor value on TO buckets + offensive fouls.
Limitations. One target per event; non-scoring events (rebounds, fouls, blocks) have y=0 so those coefficients are estimated via the fit's global adjustment rather than direct point-scoring signal. Multi-season pooling means changing players (rookie growth, aging veterans) get averaged into one number. Sample-per-parameter ratio (~40:1) is fine for ridge but individual per-bucket cells with <500 events are noisy, the role rollup helps but raw per-bucket values should be read with the sample-size context.
Interpretable WPA
What it is. Same ridge design as Interpretable RAPM but the target is 100 × Δ win probability rather than points scored. Same 8-role decomposition (Scoring, Playmaking, Off Reb, Def Reb, Def Actions, Def Presence, Off Presence, Turnovers). Reads as "wins added per 100 events" broken down by role.
How it's computed. One design matrix row per PBP event; y = 100 × (WP_after, WP_before). Identical column layout to iRAPM (offensive lineup + defensive lineup + actor-per-bucket + assister + blocker + stealer). Sparse ridge (α = 25,000). No explicit clutch weighting is applied, but because WP shifts are naturally larger in close, late-game moments (a tied-4Q shot moves WP by ~0.05 vs ~0.001 in a blowout), leverage is baked into the target directly.
How to read it. Sign conventions match iRAPM: Scoring, Playmaking, Off Reb, Def Reb, Off Presence, Turnovers are positive = good; Def Actions and Def Presence are NEGATIVE = good (matches DRAPM tradition). Multi-season 2000-2025 pool, top-15 spans Wade, Kobe, LeBron, Paul, Ginobili, Lillard, Nash, Westbrook, the "clutch pantheon" is exactly who you'd expect.
Limitations. WP shifts have larger per-event variance than points (0-3), so on-floor lineup coefficients are noisier. At half-season sample sizes the ridge cannot cleanly separate player from lineup and the leaderboard becomes dominated by presence terms rather than actor skill. For that reason we only ship the multi-season pool; single-season iWPA is not user-facing. Correlation with 25-year RAPM is moderate (ρ ≈ 0.5), because iWPA rewards a narrower "big-moment" signal.
QR-DQN Skill Advantage
What it is. A distributional-RL companion to the Q-Rating model. For every
offensive event, we predict the per-event reward distribution above a no-actor baseline
Q(s, a, teammates, defenders, context). The per-player skill advantage
is the mean of that residual, points per event the actor adds beyond what a league-average
player would have produced in the same situation. The 95% CI shown in the Player Profile is
a proper SE-of-the-mean confidence interval, not the spread of individual outcomes.
How it's computed. Two-stage: first a baseline mean-Q model is trained with the actor masked out of the offensive lineup; then a Quantile-Regression DQN with 51 quantile heads is fit (pinball Huber loss) on advantage = actual return − baseline prediction. Per-player aggregation averages the distributional mean across each player's events.
Strengths. Interpretable units (pts/event) and clean residualization that removes the action-outcome variance dominating a raw RTG estimate. Top players by mean advantage are consistently the league's high-usage per-event creators.
Limitations. Per-event efficiency systematically undersells cumulative creators (Jokić, Giannis show small positives) because action-by-action credit can't capture possession-spanning value the way RAPM does. Best read alongside RAPM, not as a replacement.
Kalman Career Trajectory (game-by-game)
What it is. A DARKO-style Kalman filter over per-game Q-Rating, producing a smoothed skill trajectory across every game a player has played. The observation is per-game Q-Rating in points per 100 team possessions, not a Game Score composite: it's the neural action-bucket decomposition redistributed per event, so the trajectory tracks actual skill rather than box-score production.
How it's computed. Each season's per-player action-bucket coefficients are decomposed into per-event credits, aggregated per (player, game) as a per-100 rate, and fed one at a time into a Kalman filter. Aging drift comes from a smoothed age curve fit on all observations pooled by age. The rookie prior is seeded from a draft-slot prior (median career-year-1 Q-Rating by pick). Career-average presence is folded in as a season-constant so elite defenders and off-ball creators surface at their real magnitudes.
How to read it. The bold line is the posterior mean; the shaded band is a ±1.96σ 95% confidence interval; small dots are per-game observations. For players drafted since 2000-01 the x-axis is career game number so trajectories are apples-to-apples across eras; for earlier careers the axis falls back to date since our data window starts mid-career. The Compare-With input overlays a second player.
Limitations. Career-average presence gives the trajectory its right magnitude but is constant within a career, so within-season presence changes (injury, role shift) don't move the line beyond what action-bucket noise contributes. Rookies have wide bands until they've played enough games.
Player Projections
What it is. 1-, 2-, and 3-year-ahead forecasts of each per-100 metric (points, assists, rebounds, etc.) using DARKO-style aging curves and a Kalman-flavored update that blends a player's own trajectory with a prior built from comparable players at the same age. Surfaced on the Projections page and in the Player Profile career trajectory chart.
Playstyle vectors
What it is. Each player's action-frequency vector projected into a low-dim space, then used to find their nearest neighbors in style (not in impact). Useful for "who plays like X?" comparisons. Distinct from the Play Types view: that one shows the raw distribution, this one shows similarity.
Play Types (Hierarchical Options)
What it is. Per-player distribution over 12 offensive "options" inferred heuristically from the action stream: pnr_handler, pnr_finisher, isolation, post_up, cut, putback, spot_up, off_screen, transition, free_throw, rebound, other. Tells you what a player does, not how well they do it. Click a player for offensive distribution + per-option success (pts/action) + defensive rates per 100 def-poss (steals, blocks, fouls).
How it's computed. Each event is labeled by a rule combining action
bucket, is_assisted, possession step, and seconds-into-possession. Isolation vs
PnR handler split uses the assisted flag (assisted pullup/drive → PnR; unassisted →
isolation). Defensive credits are computed separately: stealer/blocker via player3_id
from steal/block events, normalized by the player's defensive possessions on the floor.
Limitations. Heuristic, not learned. Without tracking data, "iso vs PnR"
is an imperfect proxy; missed shots default to isolation since assist credits only attach to
makes. Transition uses a clock-only proxy (no start_type available in the
possessions parquet). Use it for shape/role inference, not for skill grading.
Revealed Preferences (Inverse RL)
What it is. For each (player, action, game-context) triple, the log-ratio of the player's frequency against the league baseline, what actions a player chooses reveals their inferred reward weights. Reveals e.g. DeMar DeRozan's strong midrange preference, Curry's elevated clutch-3 preference. State contexts: clutch, blowout, early-clock, late-clock, transition, normal.
Behavior Policy (π_b)
What it is. A neural classifier that learns π_b(a | state, player): which action does this specific player tend to choose in this specific state? Used as an action-choice prior across several other models on the site.
How it's computed. Same lineup-attention encoder as the Q model, but the output head is a softmax over the action vocabulary instead of a scalar Q. Trained on cross-entropy against the observed action; achieves ~38% top-1 and ~68% top-3 accuracy on held-out events.
Shot Quality (xFG)
What it is. A LightGBM classifier that estimates expected FG% for each shot given the shot location, action type, and shot-clock state. Per-player it aggregates to two axes: Shot Making (actual points minus expected points, how much the player beats their location-baseline) and Shot Selection (mean xFG per attempt, how high-quality the shots they take are on average).
How it's computed. Gradient-boosted trees trained on every field-goal attempt with the shot's location coordinates, primary action bucket, and clock features. Per-shot xFG is used to compute per-player pts and selection aggregates, then compared against actual pts scored on those shots.
How to read it. Shot Making of +150 pts means the player scored 150 more points on their attempts than xFG expected. Shot Selection of 0.55 means the average shot they attempted had a 55% baseline FG probability. Two axes let a low-volume rim finisher (high selection, modest making) and a high-volume creator taking hard shots (lower selection, positive making) end up on different quadrants.
Limitations. The model doesn't see defender proximity or contested-vs-open state (no tracking data). That context lives in V-Rating instead.
Shot Charts
What it is. Side-by-side frequency + efficiency SVG heatmaps over a half-court projection, available on every player profile and every coach profile. Each bin is 1 foot × 1 foot; the shooter's location is bucketed into a 50 × 47 grid over the standard NBA half-court coordinate space (LOC_X ∈ [−250, 250], LOC_Y ∈ [0, 470]).
Frequency panel. Color intensity encodes how often the subject shoots from that bin. Log-normalized per-subject so star spots don't drown out the rest of the chart.
Efficiency panel. Color encodes FG% at that bin minus league baseline FG% at the same bin. Green above (better than league), red below (worse), clamped at ±20pp so individual hot/cold cells don't dominate.
Coach charts aggregate every shot taken by teams under that coach across their full career or a selected season. Kerr's coach shot chart lights up above the break; D'Antoni's spreads across the arc and rim; Popovich's late-Spurs chart shows the corner-three shift.
Player charts support side-by-side compare-with (type a second player's name to overlay) and shot-type toggle (All / 2PT / 3PT). Hover a bin for shot count, FG%, and delta vs league.
Pair Chemistry (Pair-Augmented RAPM)
What it is. Extra points per 100 possessions a specific pair generates (offense) or prevents (defense) when they're on the floor together, beyond what their individual RAPMs predict. A measure of two-player chemistry / fit / synergy.
How it's computed. Same ridge regression as RAPM but with an extra binary feature for every player pair that played enough possessions together. The pair coefficient is what's left after the individual player effects are accounted for: the "team-up bonus." Positive on offense = synergy. Negative on defense = synergy (less points allowed).
Context filter. The same ridge is also fit on three possession subsets, clutch (Q4, last 5 min, |margin| ≤ 5), regular (|margin| ≤ 15 and not clutch), and blowout (|margin| > 15), surfacing pair chemistry that varies by game state. Curry+Draymond generate +4.13 pts/100 in regular offense but only +1.35 in clutch (defenses take away the DHO late); Nowitzki+Terry surface as a top-3 clutch offensive pair, matching their real legacy. The Context dropdown on the Chemistry page switches between the four views.
Limitations. Identifiable pairs need a meaningful number of possessions together; one-game cameos shrink to zero. Pair coefficients are still correlational, not causal, and can't separate "lineup synergy" from "shared strategy" or "shared coach." Clutch context has ~2,500 qualified pairs vs ~34,000 in regular because each pair needs enough clutch minutes to fit.
Lineup GNN + Matchup Influence
What it is. A heterogeneous graph neural network over each 5-vs-5 possession, treated as a 10-node graph with three edge types: within-offense (chemistry among the offensive 5), within-defense (chemistry among the defensive 5), and cross (offense-vs-defense matchup edges). Two hops of message passing let signal like "my teammate has a favorable matchup, which frees me up" propagate through the graph.
How it's computed. Node features are learned player embeddings plus a role token (offense or defense). Each layer aggregates three separate messages (within side, within side, cross) into an update. Two layers are applied, node representations are pooled by side, and a standard Q head predicts expected reward on the possession. Downstream artifacts: per-player GNN presence via leave-one-out lineup swaps, a per-player synthetic matchup sweep (X vs candidate defender in an isolated 5-vs-5 context), and team-vs-team predictions aggregated over each team's most-used 5-man lineups.
Where it shows up. The Matchup Influence table on every player profile (defenders whose presence lowers or raises the player's expected Q per possession), the vs-team preview widget on team profiles, and the GNN presence stat card on the Lineup Sim page.
Limitations. Play-by-play has no per-shot defender attribution, so the matchup signal is on-court concurrence not literal 1-on-1 assignment. Same-era filtering restricts candidate defenders to seasons a player actually played in so anachronistic pairings don't populate the table.
Lineup Optimization
What it is. For each team, exhaustively scores every 5-man lineup over its rotation players and returns the best by predicted net rating.
How it's computed. Lineup score = sum of the 5 players' Q-Rating presence (Off − Def, neg-good convention) + offensive pair chemistry summed over all 10 pairs − defensive pair chemistry summed over all 10 pairs. Q-Rating presence is the leave-one-out swap on the trained Q model with attention pooling over lineups (closer analog to RAPM than the action decomposition).
Team RAPM (schedule-adjusted)
What it is. A schedule-adjusted per-team-season net rating in pts/100 possessions. Same possession-level ridge regression skeleton as player RAPM, but the features are team-season one-hots (30 offense + 30 defense per season) rather than player indicators. The coefficient answers: "How many pts/100 does this team score above league average, holding opponent quality constant?"
How it differs from Net Rating. Net Rating = ORtg − DRtg observed on the floor, no adjustment for opponent quality. Team RAPM controls for schedule strength via the def-team-season indicators on every possession. Typical correlation between the two: ~0.95. Where they diverge (5-10% of team-seasons), the divergence tells you whose schedule was abnormal that year.
How to read it. Both Team RAPM and Net Rating are shown on the team profile page alongside Team Q-Rating (the roster-aggregate estimate). Three parallel team-strength estimates lets you cross-check: schedule-adjusted (Team RAPM), observed (Net Rating), and roster-quality (Team Q-Rating).
Career floor tests. Top-10 all-time (2000-01 through 2025-26): 2015-16 Warriors +4.38, 2014-15 Warriors +4.02, 2023-24 Thunder +3.99, 2014-15 Spurs +3.81. All defensible.
Team Q-Rating (detrended)
What it is. A team-season summary of the roster's Q-Rating quality, centered so the sign is meaningful. Raw team Q-Rating is a possession-weighted sum of every roster player's Q-Rating (action + presence). Detrended = raw minus the season league mean. Positive means the roster grades above the league that year, negative means below.
How it's computed. For each (team, season): weighted mean over the roster of each player's total-impact (action + presence − opponent-side presence, good-direction on both sides), weighted by that player's possessions with that team that year. The result is multiplied by 5 to put it on per-100-poss scale (five players share each possession). Detrending subtracts the season-wide league mean so ~50% of teams end up on each side of zero.
How to read it. Semantically parallels Team RAPM (already zero-centered). +2 means "roster grades ~2 pts/100 poss above the league that year," −3 means "3 below." The trajectory chart on the team page shows total plus off/def breakdown, with Team RAPM overlaid as a reference line.
Limitations. Aggregates the whole roster equally by possessions played, so it doesn't distinguish starter-heavy vs bench-heavy usage. Presence uses career-average magnitudes broadcast across seasons, so team-season snapshots reflect the long-run quality of the players on the roster rather than that specific year's form.
Team Action Profile
What it is. Per (team, season, action bucket) rate of how much value each team generated from each of the ~30 offensive and defensive action buckets, centered against the season league mean per bucket. Surfaces two views on the team page: "Season strengths" ranks the current-season buckets from most above to most below league, and "By season" shows a heatmap of every season the team has on record.
How it's computed. For each (team, season, bucket): sum every roster player's per-bucket Q-Rating coefficient weighted by their possessions with that team that year, divided by total team possessions. Detrend by subtracting the season league mean per bucket. Green cells indicate the team was above league on that bucket that year; red indicates below.
How to read it. Bucket rows in the heatmap are sorted by how identity-defining the bucket is over the team's history, so signature strengths and weaknesses float to the top. Low-volume specialty buckets (3pt corner pullup, midrange post hook, foul-take, etc.) swing more than high-volume ones (rim finishes, midrange catches).
Limitations. Only sees action-level credit, not presence. Two teams with identical action profiles but different rim protection would look identical here even though their defense is not.
Team Builder (projected W-L)
What it is. Interactive roster projector: pick 8-12 players and set per-player minutes, get a projected 82-game W-L with a 90% confidence interval bell curve. Useful for trade / free-agency what-ifs and for stress-testing whether the presence + chemistry stack agrees with real team outcomes.
How it's computed. Player contribution =
Σ (mini/48) × (off_presence_shrunk + (−def_presence_shrunk)), the same
Q-Rating presence used by the Lineup Simulator, weighted by each player's share of a
48-minute team floor time. Pair chemistry contribution =
Σ shared-minute-weighted (chem_off − chem_def) from the multi-year ridge, where
shared minutes ≈ mini × minj / 48 (independence approximation).
Raw net rating (player + chemistry) is then calibrated to actual team net ratings via a
linear rescale, calibrated = −6.15 + 0.80 × raw, fit on
773 historical team-seasons (2000-2025, R² = 0.74). The calibration
corrects the systematic overstate that comes from presence values already partially
capturing pair-chemistry effects, so a naïve sum double-counts. Wins =
41 + calibrated_net × 2.7 (Basketball-Reference rule of thumb: ~2.7 wins per +1 net rating,
41-41 baseline).
How to read it. The Projected W-L (μ) stat card shows the mean of the wins distribution. The bell curve below shades the 90% CI band (μ ± 1.645σ). Total σ combines three variance sources: 82-game binomial noise (~4-5 wins for a typical team), model residual after calibration (σ ≈ 6.7 wins), and a small-sample penalty of ~1.5 wins per rotation player with < 1000 possessions of presence data. Total σ is typically 7-10 wins for a healthy rotation; wider bands mean more of your roster is small-sample.
Where it falls short. No schedule strength, no coach effects, no injury / DNP modeling. Independence assumption for shared minutes overstates bench-pair overlap and understates starter-pair overlap. Small-sample players widen the confidence band but don't fix mean bias for teams with unusual chemistry patterns (e.g., 2024-25 OKC significantly overperformed roster-projected wins). Sanity checks: 2024-25 Boston → 60-22 (actual 61-21) ✓, 2016-17 Warriors → 72-10 (actual 67-15) ✓.
Championship odds (end-of-regular-season + preseason)
What it is. Two Monte Carlo bracket simulators sharing the same series-win logistic. End-of-regular-season uses final rosters + observed per-season Q-Rating presence. Preseason uses prior-season player value: Kalman-preferred prior-3-year presence with V-Rating ensemble, age drift, draft-slot rookie prior, availability discount, newcomer skill discount, role compression, and share-weighted pair chemistry. Raw strength is era-centered (per-season) and calibrated to net rating with a continuity + prior-year-residual term. For past seasons the roster is each team's opening-night lineup (reconstructed from first 8 games); for the upcoming 2026-27 season (no games yet) it is the current roster with a position-aware depth-chart minutes model and scraped long-term injuries. Both feed the same raw → net rating → Pythagorean W-L → seed → 30k-trial bracket (1v8 / 4v5 / 2v7 / 3v6 → semis → conf finals → Finals). Retrospective evaluation on 2005-06 through 2024-25 (20 NBA seasons, 600 team-seasons).
Top-K champion accuracy. How often the actual champion is in the model's top K by predicted p_champion:
| Preseason | End of reg season | |
|---|---|---|
| Top-1 | 45% | 25% |
| Top-3 | 60% | 60% |
| Top-5 | 75% | 95% |
| Top-8 | 95% | 100% |
Probabilistic calibration. Brier score on the (team, season) → won-championship binary: preseason 0.0290, end-of-season 0.0287 (uniform baseline 0.0322, lower is better). Both are essentially as well-calibrated as probabilistic forecasters. Log-loss: preseason 0.139, end-of-season 0.103, because end-of-season concentrates probability more tightly on real contenders, which log-loss rewards.
Ranking accuracy. Per-conference Spearman rank correlation between predicted calibrated_net and observed net rating: preseason 0.74, end-of-season 0.90. This directly drives seed accuracy: the higher end-of-season correlation is why its top-5 champion hit rate is 95% vs preseason's 75%.
Playoff-round hit rates (predicted top-K by p_champion vs actual teams reaching that round):
| Round | Preseason | End of reg season |
|---|---|---|
| Playoff pool (top-16) | 79% | 72% |
| Conf semis (top-8) | 65% | 54% |
| Conf finals (top-4) | 46% | 45% |
| Finals (top-2) | 43% | 38% |
Preseason wins the playoff-pool hit rate because it doesn't overreact to mid-season injuries + trades that reshape the end-of-season model's top-16 away from the actual playoff seeds.
Wins projection MAE (against 41 + 2.7 × observed_net): preseason 6.67 wins, end-of-season 4.49 wins. This is the most user-facing accuracy number; preseason projections are ~50% noisier than end-of-season.
Reliability of preseason p_champion (predicted decile vs observed champion rate): predicted 5-10% → observed 5.8% ✓, predicted 10-20% → observed 13.3% ✓, predicted 20-35% → observed 50% (n=8, small-sample; model is slightly underconfident at the very top). Well-calibrated in the middle bands.
Where it falls short. Preseason can't extrapolate mid-summer trades (Kawhi to TOR 2018-19, AD to LAL 2019-20) or steep breakout years (Jokić 2022-23, SGA 2024-25). End-of-season misses concentrate on Cinderella playoff runs (2011 Mavs, 2019 Raptors) where regular-season net rating didn't predict the playoff run.
March Madness odds
What it is. A bracket-seeded Monte Carlo sim over the ACTUAL NCAA tournament bracket. Unlike a seed-your-own playoff model, March Madness has a fixed 64-team bracket, so we take the real field (teams, seeds, matchups scraped from ESPN) and simulate it forward 20,000 times to get each team's odds of reaching every round.
Team strength. Same chain as the pro Championship model: per-player Q-Rating presence (off minus def) weighted by possession share, plus career pair-chemistry contributions weighted by shared possessions, calibrated to net rating. This orders the field well even though raw college net rating is schedule-inflated: the Spearman correlation between our team strength and the committee's seeding is -0.76 (better strength maps to a lower seed number), and the eventual champion is our top-8 team in 100% of tournaments.
Win model. Single games (not best-of-series) sampled from a logistic fit on ~1,080 historical tournament games on neutral courts: P(A beats B) = sigmoid(0.13 x (strength_A minus strength_B)), so a +10 strength edge is about a 79% favorite and +5 is about 66%. First Four play-in games are simulated too.
Bracket reconstruction. ESPN's region labels are inconsistent (a region's early rounds are sometimes tagged by host-city pods), so the bracket tree is rebuilt from results: a round's two participants each won a prior-round game, which become its child nodes. This sidesteps region parsing and the Final Four pairing entirely.
Retrospective accuracy (2008 through 2026, 18 tournaments; the champion's rank is by our predicted p_champion within the field of ~64):
| Champion in top-1 | 39% |
| Champion in top-3 | 72% |
| Champion in top-5 | 89% |
| Champion in top-10 | 100% |
| Avg champion rank | 2.6 of ~64 |
| Final Four recall (top-4) | 69% |
How to read it. Odds are pre-tournament (from the season's roster + team strength), not updated as games are played. The Odds table gives each team's chance to reach each round; the Bracket view shows the actual matchups with per-game win probabilities, marks the real winner, and flags the champion. Coverage is 2008 onward (2020 cancelled for COVID; 2006-2007 predate ESPN's bracket tagging).
Where it falls short. Misses are the Cinderellas that a strength model can't foresee: 2023 LSU (a 3-seed) ranked 8th by our model before winning it all. Chalk years are nailed (UConn ranks 1st across most of its 2009-2016 run).
Preseason (forward-looking) mode. The Odds toggle switches team strength to a purely retrospective input: each player is rated only by a recency-weighted blend of their prior three seasons (no current-season play), newcomers get a neutral prior, and the win logistic is refit on those preseason strengths. This is what a real bracket forecast can use before a season starts. It is naturally less accurate than the in-season model: over the same 18 tournaments the champion lands in our preseason top-1 33% / top-3 72% / top-5 83% / top-10 89% (avg rank 3.9), Final Four recall 60%. Its blind spot is deliberate and important: with no recruiting signal it cannot see a freshman or transfer breakout, so it under-rates newcomer-driven teams, most glaringly 2023 champion LSU (Angel Reese plus transfers), which it had at 0.1%. The Preseason vs actual view lines the preseason and in-season numbers up with the real result so those over- and under-performers stand out.
WNBA Draft Projection
What it is. A translation model that projects current college players to the WNBA. It is trained on the players who appear in both our college and pro datasets (493 of them), each scored with the same impact metrics on both sides. That shared-metric bridge is what makes the projection possible, and almost no public site has it: college and pro numbers usually come from different systems that cannot be compared directly.
Matching the two leagues. Each player's full college career is joined to their full WNBA career via shared ESPN player ids and name matching (with maiden-name and nickname recovery). Internationals who never played NCAA ball are unmatchable, so coverage is US college only (2006-2026).
Three models, all on college inputs (Q-Rating presence, V-Rating, iRAPM, plus per-season volume and peak-season RAPM):
- Reach (logistic) = probability the player reaches the WNBA with real minutes. Out-of-sample AUC 0.90.
- Projected impact (ridge) = expected WNBA per-100 impact if she reaches. The target is observed pro impact, not our shrunk career rating: shrinkage compresses the range and kills the signal. Out-of-sample correlation +0.45.
- Projected longevity (ridge) = expected career length, out-of-sample correlation +0.32.
Reach features are deliberately volume-balanced. An earlier version leaned on raw cumulative college possessions, which flattered four-year role players and under-rated dominant underclassmen (JuJu Watkins, Sarah Strong). The current model uses per-season volume plus unshrunk peak-season RAPM, so young stars keep a sane reach probability without hurting AUC.
Projected outcome distribution. Each prospect gets a five-part probability bar, summing to 100%: Never sticks (= 1 minus reach) then, conditional on reaching, Fringe/Depth, Rotation, Starter, and All-WNBA. The tier split comes from a Normal centered on projected impact with the model's out-of-sample residual spread, so a wider bar honestly reflects a less certain projection. Hover any segment for its exact percentage.
Comp. The current WNBA player whose playstyle vector is closest to the prospect's (cosine similarity on the shared action buckets), a similarity read, not a ceiling.
How it is validated. Everything is out-of-sample with grouped-by-player cross-validation, so a player's own seasons never leak into her own prediction. The probabilities are calibrated: expected calibration error is 0.00 for reach and 0.02 for the outcome tiers (0 is perfect), meaning a stated 60% happens about 60% of the time. On a held-out retrospective of recent debutants, the projected order matches actual WNBA impact at rank correlation +0.43.
How to read the page. Draft prospects ranks current college players by draft score (reach times projected impact) and shows the outcome bar, reach %, projected impact with its percentile, and the closest comp. Retrospective shows recent debutants with both projected and actual tier, and a marker for whether they beat, hit, or fell short of the projection.
Where it falls short. The impact signal is real but modest: college dominance only partly translates and the model is deliberately conservative, so it will not crown the next superstar outright. Reach is the strongest piece; impact and longevity are directional. It also cannot see things outside box-adjacent impact (injuries, motor, off-court development), and the training set is capped by the size of the two-league population, not by how much college data we have.
Trade Analyzer
What it is. Two-team roster simulator. Pick two teams, select players on each side, hit Swap: both rosters get recomputed side-by-side with new calibrated net ratings, projected wins, and championship odds for both teams.
How it's computed. Same chain as the Team Builder, applied to both
sides in parallel. Each team's calibrated net rating =
a + b × (Σ (mini/48) × per-player impact + Σ shared-min pair chemistry)
with league-specific (a, b), fit on 773 historical NBA team-seasons (R² = 0.74) and 318 WNBA
team-seasons (R² = 0.60). Projected wins follow the Basketball-Ref Pythagorean rule.
Championship odds run a joint 5,000-trial Monte Carlo bracket sim where BOTH swapped teams
update simultaneously (so a trade that moves talent from a title contender to a lottery team
shifts both trajectories at once); every other team is held at its current-season baseline.
How to read it. Green Δ = the team improved; red Δ = they got worse. The most meaningful comparisons are 1-for-1 star swaps (each side stays at the correct team-minutes total). Uneven trades leave rosters mis-summed, projections are still shown but the minutes note will flag it. Sanity checks: SGA (OKC) ↔ LeBron (LAL) hits OKC for ~12 wins and lifts LAL by ~6, Jokić (DEN) ↔ Tatum (BOS) drops DEN by ~10 wins, both matching the intuition that MVP-tier stars carry ~8-12 wins of value.
Where it falls short. No salary or contract logic, so all trades are cap-legal here even when they wouldn't be. No positional balance check. No coach fit or chemistry-with-teammates interaction (though pair-chemistry deltas do fold in). Minutes transfer with each player; if you send a starter, the receiving team's rotation absorbs their minutes even if it pushes past 240.
Coach Ratings
What it is. Coach impact isolated as the possession-level residual after controlling for player quality. Two parallel metrics are shown side-by-side per coach: Coach RAPM and Coach Q-Rating. Both are in points per 100 possessions, same sign convention as player RAPM (off positive = good, def negative = good, total = off − def). Covers every head coach who worked a regular-season game from 2000-01 through 2025-26.
Coach-by-game attribution. Basketball-Reference publishes W-L splits per coach per team-season (e.g., Bulls 2003-04: Cartwright 4-10, Myers 0-2, Skiles 19-47). We scrape those splits, then walk each team's game log chronologically and allocate games in order, first N to coach A, next M to coach B, and so on. Franchise transitions handled explicitly (SEA → OKC, NJN → BKN, CHH → NOH → NOP, VAN → MEM). Output: 57,402 (game, team, coach) rows across 776 team-seasons.
Two-step residual. For every qualifying possession
(~5.2M after garbage-time filter), we first compute the expected pts/100 using a
player prior held constant, either box-score-prior player RAPM (with defensive
term as a 50/50 blend of def_rapm and Q-Rating defensive presence, so rim deterrence and
off-ball help are credited) or the full Q-Rating (action + presence). We then subtract that
expectation from actually observed points to get a per-possession residual, and finally
ridge-regress the residual on [off_coach, def_coach, season] one-hots. The
coach coefficient is what pts/100 the players plus league era couldn't explain, the
coach's persistent contribution above player baseline.
Shrinkage and Kalman smoothing. Career values are empirical-Bayes shrunk with κ = 40,000 possessions (~500 games per side). Per-season values use κ = 12,500 plus a 1-D Kalman smoother that treats each season as a noisy observation of a slowly drifting latent coach quality, so trajectories don't jump on noise.
Cross-metric agreement. The two priors produce two independent estimates. On the ≥250-games qualified pool: Pearson r = 0.76 total, r = 0.97 defense (near perfect agreement), r = 0.49 offense (real divergence). When both metrics agree the coach's effect is robust; when they disagree the ranking is method-sensitive on that coach (typically an offensive-scheme case Q-Rating captures better than box-score RAPM).
What the coach profile also shows. Kalman-smoothed season trajectory with off/def breakdown; aggregate shot chart of every shot taken by teams under that coach (frequency and efficiency versus league baseline, per season and career); Coach style widget showing signature action-bucket tendencies weighted by their share of team-season games, reveals things like Kerr over-indexed on above-the-break pull-up threes (+13.9% vs league), D'Antoni on catch-and-shoot threes (+7.3%), Popovich on corner threes.
Limitations. Lifers whose star players carry an outsized share of team identity, Popovich with Duncan/Ginobili/Parker/Kawhi, Kerr with Curry/Klay/Draymond, have a partial-credit ceiling. If a player's box-score-prior RAPM already accounts for their team's success, there's less residual variance left for the coach coefficient to absorb. This is a fundamental identifiability limit of the two-step framing, not a bug.
How to read the numbers
RAPM-style numbers are in points per 100 possessions. A +5 RAPM player is roughly +5 net points per 100 possessions compared to a league-average player on a neutral team. League scoring runs ~115 per 100, so +5 is genuinely elite (top-10ish in any given season). For defense, more negative is better: a −3 DRAPM player concedes 3 fewer points per 100 than average. QR-DQN skill advantage is in different units (points per event, not per 100 possessions). Play Types percentages are shares of a player's offensive actions; defensive rates are per 100 defensive possessions they were on the floor for.
Data & pipeline
NBA play-by-play covering the 2000-01 through 2025-26 regular seasons, 5.49M possessions across 30,815 games, parsed into per-event tuples with the 5-on-5 lineup, action, actor, and reward. Models implemented in scikit-learn (ridge with box-score prior, K-fold CV), PyTorch (Q-network with lineup attention, distributional QR-DQN, lineup GNN, causal transformer for V-Rating, Kalman filter for career trajectories), LightGBM (win probability, expected FG%), and a closed-form Gaussian-Gaussian Bayesian solve for the RAPM prior step. Ratings regenerate end-to-end from raw play-by-play, no hand curation.
Evaluation
Empirical checks on the four headline player-impact ratings: how much do the numbers change from one season to the next, and how does that stability scale with sample size? A metric that flips wildly year over year isn't measuring the player, it's measuring noise. This page reports Pearson correlation between each player's value in season N and season N+1, broken down by how much on-court sample the player had.