HoopQ

Open NBA Analytics
Since the 2000-2001 season · 5.5M possessions

HoopQ.
Open NBA analytics, built possession-by-possession.

Independent NBA analytics. Six ways to measure player impact, per-game career trajectories, and how well any lineup actually fits together. 25 seasons of data, free to use.

Explore

Clutch
Action + presence RAPM on clutch possessions (Q4/OT ≤5min, |margin| ≤5). Who elevates in the moment?
Lineup Optim
Best 5-man lineups per team via Q-Rating presence + pair chemistry.
Lineup Sim
Swap any player into a 5-man lineup and see the projected net rating, chemistry included.
Projections
DARKO-style 1/2/3-year forecasts with aging-curve adjustments.
Coaches
Two-step residual RAPM + Q-Rating: coach effect above player baseline. Kalman trajectory, shot chart, style profile.
Playstyle
Action-frequency vectors and nearest stylistic neighbors.
Play Types
Twelve offensive option buckets via hierarchical RL.
Revealed Prefs
Inverse-RL action × context heatmap of decision preferences.

Featured Players

Top 15 by 2025-26 Q-Rating Total Impact · click any card for the full profile.

Leaderboards

Top 3 by each metric · click any name for the profile, "See all →" for the full ranking.
When a player is on the court, does their team outscore opponents? RAPM works that out across 26 seasons, cancelling out the effect of teammates and opponents. In points per 100 possessions.

The classical impact metric of NBA analytics. For every player, RAPM (Regularized Adjusted Plus-Minus) asks: how many more or fewer points does the team score per 100 possessions when this player is on the floor, once you have adjusted for the quality of everyone else who was on the floor at the same time? Positive means the player helps their team's scoring margin; negative means they hurt it.

How to read it. Values are pooled across 26 seasons (2000-01 through 2025-26, 5.49M possessions), so this is a career-scale number, not a single-season one. Click any column header to sort. Click any row to open the full player profile.

How it's built. A ridge regression where every possession is one data point, every player gets one offensive and one defensive coefficient, and the model finds the coefficients that best explain observed point margins. A Bayesian prior derived from box-score stats stabilizes the estimate for players with fewer possessions, so a rookie with 200 minutes doesn't swing to +30 by accident. See About for the full methodology.

Single-season RAPM leaderboard. Scrub the season dropdown to see the top of each era.
Best RAPM players in a single season. Each row is one (player, season) pair, computed from a prior-informed Ridge fit on just that season's possessions. Use the season dropdown to scrub through the era.
Clutch Q-Rating: shot-quality above expectation on clutch actor events, plus Clutch RAPM presence.

Clutch = Q4 or overtime, clock ≤5 minutes, |margin| ≤5. Two independent measurements combined into one leaderboard:

Action Q (pts above league expectation per 100 clutch actor events): for each clutch event, we compare the observed reward to the per-season bucket mean. Playmakers get credit on assisted makes so passing value doesn't disappear into the shooter alone. EB-shrunk κ=500 multi / κ=100 single season.

Clutch ORAPM / DRAPM: Ridge regression on ~230k clutch possessions (44k WNBA) with α=25,000. Captures lineup-level offensive and defensive presence in the clutch subset (rim deterrence, gravity, off-ball value). Sign conventions match regular RAPM: ORAPM positive = good, DRAPM negative = good.

Total = Action + ORAPM − DRAPM. Positive = good overall clutch player.

Projected 2026-27, 2027-28, and 2028-29 player metrics. Recency-weighted recent seasons plus league aging curve. 95% CI shown.

Each projection combines a recency-weighted average of the player's most recent three seasons (weights 0.2 / 0.3 / 0.5) with the age-indexed league aging-curve delta from their current age to the target season's age. Applied independently to RAPM, Q-Rating, offensive presence, and defensive presence.

Aging curves are possession-weighted means by actual age, from birthdates when available and career year otherwise. The 95% credible interval reflects year-over-year drift plus the variance of the player's recent seasons.

Click any player row above to see their per-metric projection curves.
League aging curves
Possession-weighted mean per age (player-seasons ≥500 poss, age 19-42). Toggle the metric to see how it ages across the league. Lower chart breaks RAPM down by Q-Rating action bucket. Pick a skill to see when in a career it peaks.

Per-action aging

Team Q-Rating = poss-weighted sum of player Q-Rating (action + presence). Split into offense + defense. Sign convention (DRAPM-style): Off positive = good, Def NEGATIVE = good (fewer points allowed). Team RAPM shown alongside for reference.
Per-season team strength from two independent estimators. Team Q-Rating (total) is the possession-weighted sum of each player's Q-Rating (action + presence combined), capturing both event-attributed contribution and lineup-swap presence (rim deterrence, spacing, gravity). Split into offense (off action + off presence) and defense (def action − def presence, both flipped to good-direction). Team RAPM = net points per 100 possessions (offense minus defense), the independent ridge-regression estimator. When Q-Rating and RAPM disagree, the delta is informative. Q-Rating covers 2005-26; pre-2005 team-seasons show blank (retrain in progress will extend this to 2000-26). Click a team row for its season-by-season trajectory.
Coach effect isolated as the possession-level residual after controlling for player quality. Two metrics side by side, Coach Q-Rating (default) and Coach RAPM (cross-check), plus shot chart and action-style profile on the coach profile page.

Coach effect isolated via a two-step fit. Step 1 fixes each player's expected contribution using a player-quality prior. Step 2 regresses the per-possession residual (actual minus player-implied points) on the offensive and defensive head coach one-hots plus season fixed effects. The coach coefficient captures whatever pts/100 the players + league scoring era couldn't explain, read as coach + team-baseline effect for the coach's tenure.

Two priors, two views: Coach Q-Rating (default) uses the full deep-RL Q-Rating, action head plus presence, as the player prior. Coach RAPM swaps in box-score-prior player RAPM instead, with the defensive term as a 50/50 blend of def_rapm and Q-Rating def presence so rim-deterrence gets credit. The two agree overall at Pearson r=0.76 on the ≥250-game qualified pool; their defense signals correlate at 0.97 but offense signals diverge (r=0.49). When both metrics agree the coach's effect is robust; when they disagree the ranking is method-sensitive on that coach (typically an offensive-scheme case Q-Rating captures better than box-score RAPM).

Beyond the ratings. Click any coach row for their profile: Kalman-smoothed season trajectory (raw shrunk + smoothed line), aggregate shot chart of teams under this coach (frequency + efficiency), and a "Coach style" widget showing which action buckets they over- or under-index on relative to league mean.

Sign convention: off positive = raises team pts/100; def negative = lowers opponent pts/100. Total = off − def (positive is good, mirroring player RAPM). Career shrunk with κ=40000 possessions; per-season with κ=12500 plus a 1-D Kalman smoother across seasons for trajectory. Click a coach row for their season-by-season chart.

Win Probability Added per event. Clutch = 4th quarter or later, under 5 minutes, score within 5.

Win Probability Added credits each event by its contribution to the offensive team's win probability. A clutch shot that moved WP by 5 percentage points contributes +0.05 WPA. Per-season models are trained across 2000-01 through 2025-26; the selector defaults to the most recent season.

Total WPA is cumulative and rewards volume as well as efficiency; WPA per event isolates efficiency. Defensive WPA only credits explicit defensive events (steals, blocks, defensive rebounds, fouls). Rim deterrence and off-ball spacing effects are not captured, so elite rim protectors read systematically low here.

Action-frequency distribution per player. Click a row for signature bars and stylistic nearest neighbors.

Playstyle vectors are the per-player distribution over action classes for the 2024-25 season. Each column shows the fraction of that player's actions that fall in the given category.

Click a row to open the player's signature bar chart and their ten most-similar players by cosine similarity.

Twelve offensive play types via hierarchical RL options. Click a player for per-option success rates.

Per-player play-type distribution from a hierarchical RL options framework. Each offensive action is labeled as one of twelve play types: pick-and-roll handler, pick-and-roll finisher, isolation, post-up, cut, putback, spot-up, off-screen, transition, free throw, rebound, or other.

Columns show the fraction of a player's offensive actions in each option. Defensive events are tracked separately as per-100-poss steals, blocks, and fouls. Click a player row to see their per-option success breakdown.

Inverse RL: action × context heatmap. Positive = player picks that action more than league baseline.
Inverse RL, Revealed Preferences. Treats each player's action choices as a softmax over a personal reward function and recovers the per-player reward weights via log-ratio against the league baseline, conditioned on game state (clutch / blowout / early or late shot clock / transition / normal). Positive cell = player chooses that action far more than expected; negative = chooses it less. Pick a player to see their action × context heatmap.
MVP / DPOY / ROY / COY + All-NBA / All-Defensive / All-Rookie probabilities. Conditional-logit trained on 26 seasons of actual voting shares.

For each award, we assemble a candidate pool per season, compute features from our existing metrics, and fit a conditional-logit (softmax within season) on historical voting shares scraped from Basketball-Reference 2000-2025. Predictions cover all seasons including the current one, so you can compare historical model picks against actual winners.

Features: MVP uses Q-Rating total impact, PPG, AST/G, team wins. DPOY uses def_presence (flipped), blocks/g, steals/g, team wins. ROY same as MVP filtered to first-season players. COY uses coach Q-Rating + wins-over-projected (from championship model) + team wins. All-NBA / All-Defensive / All-Rookie use the same features as their singleton counterparts but train on top-15 / top-10 / top-10 team selections.

Retrospective accuracy (top-1 hit rate): ROY 77%, All-Rookie 92%, All-NBA 69%, All-Defensive 65%, MVP 46%, DPOY 35%, COY 23%. COY is voter-narrative-driven so hardest to predict from stats.

Monte Carlo playoff bracket sim. Two views: End of regular season (final rosters + observed presence) or Preseason (forward-looking, prior-season player value). The upcoming 2026-27 season is projected from current rosters + a depth-chart minutes model + injuries.

End of regular season is the default view. Team strength = per-player Q-Rating presence weighted by possession share + career pair-chemistry contributions weighted by shared possessions. Raw strength is calibrated to observed net rating, converted to projected wins, then Top-8 seeds by projected wins run a fixed bracket (1v8, 4v5, 2v7, 3v6 → semis → conf finals → Finals). Series outcomes are sampled from a logistic model fit on ~390 historical NBA playoff series (156 WNBA). 30,000 trials per season. Retrospective accuracy (2005-06 → 2024-25): champion in top-1 25%, top-3 60%, top-5 95%, top-8 100%.

Preseason view estimates player values ONLY from data through the prior season. Chain: Kalman-preferred Q-Rating presence + V-Rating ensemble (0.7/0.3) + Kalman age drift + draft-slot rookie prior + injury-availability discount + newcomer skill / role compression + share-weighted pair chemistry. Raw strength is centered within each season (era-normalized) and calibrated to observed net rating, with a continuity + prior-year-residual adjustment for the portion of last year's result the roster-sum misses. Same downstream 30k-trial bracket sim. For past seasons it uses each team's opening-night roster (reconstructed from their first eight games), answering "how would we have projected the season before it started?" For the upcoming, not-yet-played 2026-27 season there are no games yet, so it reads current rosters, assigns minutes with a position-aware depth-chart model, and applies known long-term injuries (scraped from ESPN). Retrospective accuracy (2005-06 → 2024-25): champion in top-1 45%, top-3 60%, top-5 75%, top-8 95% (avg rank 3.3). Point projections are noisier than end-of-season (~7 wins MAE vs ~4.5). Misses concentrate on mid-summer trades and steep breakout years the model can't extrapolate.

See the Championship odds glossary entry for Brier score, log-loss, per-conference Spearman rank correlation, reliability diagram, and per-round hit rates.

Bracket-seeded Monte Carlo sim over the ACTUAL NCAA tournament bracket. Each team's odds to reach every round, driven by Q-Rating + chemistry team strength.

Unlike a seed-your-own playoff sim, March Madness has a given 64-team bracket. We take the real bracket (teams, seeds, matchups from ESPN) and simulate it forward 20,000 times. Team strength = per-player Q-Rating presence weighted by possession share + career pair-chemistry contributions, calibrated to net rating. Single games (not series) are sampled from a logistic fit on ~1,080 historical tournament games. First Four play-in games are simulated too.

Retrospective accuracy (2008-2026, 18 tournaments): champion in top-1 39%, top-3 72%, top-5 89%, top-10 100%; average champion rank 2.6 of ~64; Final Four recall 69%. The model's strongest team by our rating reached the Final Four every year it was clear chalk (UConn 2009-2016). Misses are the Cinderellas (2023 LSU, a 3-seed, ranked 8th).

Odds shown are pre-tournament (computed from the season's roster + our team strength), not updated as games are played. The bracket view shows the actual matchups with our advancement probabilities; actual results are marked. Coverage 2008-2026 (2020 cancelled). The Odds toggle switches between In-season (rated on this season's play) and Preseason (forward-looking, rated only on players' prior-season value, the way a real bracket forecast works); the Preseason vs actual view lines the two up so over- and under-performers stand out. See the March Madness glossary entry for the win model, bracket reconstruction, and full retrospective accuracy.

Schedule-adjusted team power ranking. Q-Rating presence + chemistry team strength, adjusted for strength of schedule.

Each team's strength is our March Madness chain (per-player Q-Rating presence weighted by possession share + pair chemistry, calibrated to net rating), then adjusted for strength of schedule: a team that earned its rating against strong opponents gets a boost, against weak opponents a penalty. Without this, mid-majors that dominate a weak schedule look elite. Schedule comes from the full regular-season game log; the adjustment iterates each team's rating against its opponents' adjusted ratings.

This is a model rating and can disagree with a team's record or seed (that is the point). The W-L column is shown so you can see where the model likes a team more (or less) than its results. Covers all 360+ Division I teams per season, 2006-2026. Powered by the same team strength as the March Madness odds.

Projecting college players to the WNBA, trained on the ~500 players who appear in both leagues with our own consistent metrics.

Almost no public site can build this: it needs both college and pro players scored with the same impact metrics, which we have. We train three models on college inputs (Q-Rating, V-Rating, iRAPM):

Reach = probability a college player reaches the WNBA with real minutes. Projected impact = expected WNBA per-100 impact if she reaches (the target is observed pro impact, not our shrunk career rating, which compresses the range and kills the signal). Projected outcome = the five-part bar (Never sticks, then Fringe/Depth, Rotation, Starter, All-WNBA), a calibrated probability distribution that sums to 100%; hover a segment for the exact %. Comp = the WNBA player whose playstyle vector is closest (cosine similarity on the shared action buckets). Everything is validated out-of-sample with grouped-by-player cross-validation so a player's own seasons never leak into her prediction, and the probabilities are calibrated (expected calibration error 0.00 reach, 0.02 tiers).

Honest limits. The impact signal is real but modest (college dominance only partly translates, and the model is deliberately conservative, so it will not crown the next superstar outright). Coverage is US-college only: international players never played NCAA ball, so they are unmatchable and absent. Reach is the strongest piece; impact and longevity are directional. See the WNBA Draft Projection glossary entry for the full method and validation.

Best 5-man lineups per team. Score = Q-Rating presence sum + pair chemistry.
Best 5-man lineups per team (default 2025-26; pick a season via the dropdown), found by exhaustive search over rotation players. Score = sum of player Q-Rating presence (Off − Def, neg-good convention) + offensive pair chemistry − defensive pair chemistry. The player term uses neural-lineup-aware Q-Rating presence (from a leave-one-out swap on the trained Q model with attention pooling over lineups) instead of career RAPM, which puts scores in a more realistic range (~+15 to +25 per 100 poss instead of the previous +30 to +40).
Lineup Simulator: swap any player into a 5-man lineup and see the projected net pts/100, combining per-player Q-Rating presence + observed pair chemistry.
How the score is built.
1. Per-player contribution = off_presence_shrunk + (−def_presence_shrunk). These come from the Q model's leave-one-out lineup-swap procedure, and they capture rim deterrence, spacing, gravity, and off-ball value (the lineup-level effects that pure action-bucket ratings miss). Both sides are shown "good-direction" (positive = good).
2. Pair chemistry = Σ across all C(5,2)=10 pairs. We use the observed coefficient from the multi-year ridge fit (residual above each player's individual contribution). Pairs that haven't played together show 0.00. A Q-model-based predictor for hypothetical pairs is on the roadmap but not yet wired into this page.
3. Net pts/100 = Σ per-player contributions + Σ pair chemistry.

Caveats. Triple-and-higher chemistry isn't modeled. Multi-year ratings are the most stable input; single-season ratings are noisier and overweight 2025-26 small samples. Presence-based defense does capture rim deterrence and off-ball value (that's the whole point of the leave-one-out lineup swap), but the magnitude depends on how much between-lineup variance the Q model saw during training. Players with very repetitive teammates may have their deterrence partly absorbed into "average team defense."
Team Builder: pick 8-12 players + their minutes per game, get a projected W-L.
How it works. Team net rating = Σ (minutesi/48) × per-player impact + Σ pair chemistry weighted by shared floor time. Per-player impact = off_presence_shrunk + (−def_presence_shrunk) from the Q model's leave-one-out lineup swap. Shared floor time is approximated as (mini × minj) / 48 assuming independent rotations. Wins = 41 + net_rating × 2.7 (Basketball-Reference rule of thumb, ~2.7 wins per +1 net rating).

Confidence band. The ± range reflects uncertainty from small-sample players and unmeasured factors (coach, health, schedule strength). Wider band = more of your roster is small-sample or hypothetical.

Caveats. No schedule strength. No injury / DNP modeling. Assumes coach distributes minutes as configured. Rotation-independence assumption for pair chemistry ignores the fact that starters typically overlap much more than starter-bench pairs, so pair contributions of your top starters are slightly under-counted.
Trade Analyzer: swap players between two teams. Recomputes calibrated net rating, projected wins, and championship odds for both sides.
How it works. Pick two teams and a season. Each side loads its default 8-player rotation with possession-weighted minutes (the same roster the Team Builder starts from). Click any player to select them for the trade; then hit Swap Selected to move them to the other team, carrying their minutes along.

For each side, we recompute the calibrated net rating using the same chain as Team Builder: player contribution = Σ (mini/48) × (off_presence − def_presence), pair chemistry = Σ shared-minute-weighted (chem_off − chem_def), raw = player + chem, calibrated = −6.15 + 0.80 × raw (NBA) / −2.83 + 1.27 × raw (WNBA). Projected wins = 41 + 2.7 × calibrated_net (NBA) or 22 + 1.45 × calibrated_net (WNBA). Championship odds run a joint 5,000-trial Monte Carlo bracket sim where BOTH swapped teams update simultaneously; the remaining teams are held at their current-season baseline.

Caveats. Uneven trades (N-for-M) will leave rosters with total minutes ≠ team target; the projection is still computed but the roster note will flag it. No salary cap or contract logic. No coach fit, no positional balance check.
How much better (or worse) two players do together than the sum of what they'd do apart. In points per 100 possessions.

Every pair of players who have shared the floor gets a number here. Positive means the pair lifts each other's game (elite chemistry, like Curry and Draymond's screen-and-roll partnership). Negative means they undermine each other. The value is in points per 100 possessions on the same scale as RAPM, on top of what each player already contributes individually.

Season filter. "Any season" gives a stable career-long average, one number per pair. Pick a specific year and you get a rolling 3-year window that captures how a duo has been playing lately. Curry and Draymond dip through their injury years; Gobert and Naz Reid emerge as an elite defensive pair in 2024-25.

Context filter. Splits the data into three game situations: clutch (last 5 minutes of Q4 or OT, within 5 points), regular (within 15 points, not clutch), and blowout (winning or losing by more than 15). Curry and Draymond drop from +4.1 in regular time to +1.4 in clutch; defenses take away their favorite hand-off late. Dirk Nowitzki and Jason Terry surface as one of the best clutch pairs of their era. Fewer pairs qualify in the clutch subset (about 2,500 vs 34,000 in regular time) because each pair needs enough clutch minutes to register. The context split is only available on the career-long view, so combining it with a specific season falls back to the career pool.

How it's built. A ridge regression across 26 seasons of NBA play-by-play (5.49M possessions). The model has one coefficient for each player and one for each same-team pair, fit jointly, so the pair coefficient is the extra effect on top of what the two players do individually. Held-out R² improves over a player-only baseline, so the pair signal validates out of sample (not just in-sample overfitting).

V-Rating: per-event ΔV attribution from a multi-season neural state-value function. Points per 100 actions, split into offensive and defensive components with per-bucket decomposition. Sibling to Q-Rating (state-based rather than action-conditional).
A causal transformer scores every prefix within a possession. For each event, ΔV = V(post) − V(pre) is credited to the actor. Shots split into "choice" (bucket-conditional expected value) and "luck" (residual vs bucket mean); assists take 60% of choice; rebounds count at 30% (contested); on-floor lineup shares the V₀ gap so passers and elite defenders get proper credit for setup / rim deterrence. All values are in points-per-100 actions, comparable to Q-Rating. Season totals in raw points are shown alongside for volume context.

Presence Impact: Methodology

V₀-based on/off with cluster-robust CIs, multi-team stint filtering, and cross-model comparison. Ships in Rankings and on every player profile.

What Presence Impact measures

For every possession in the data (5.49M NBA / 951K WNBA), the V-Rating causal state-value transformer scores V₀: the expected point value of that possession conditioned on state, lineup, and pre-tip context. Presence Impact answers a simple question: how does V₀ change when this player is on the floor vs off?

Per 100 possessions:

  • Δ Off = mean(V₀ | player on court, their team on offense) − mean(V₀ | off court, their team on offense).
  • Δ Def = mean(V₀ | off court, their team on defense) − mean(V₀ | on court, their team on defense). Sign is flipped so positive = good defender (reduces opponent value).
  • Δ Total = Δ Off + Δ Def.

Cluster-robust 95% confidence interval (the key methodological win)

The naive standard error of a difference-in-means (Welch's) treats every possession as independent. Basketball possessions within a single game are anything but — they share the opponent, the tempo, the shooting-variance of the night, and often the same lineup stint. We measured the underestimate empirically on the NBA pool: the IID SE is ≈ 2× too narrow. Reported 95% CIs from IID SE actually cover ~32% of the anchor value, not 95%.

We fix this with the cluster-robust ("sandwich") estimator. Group possessions by game_id. For each player, the influence-function contribution of game g to Δ is

s_g = (Σᵢ∈W∩g (yᵢ − ȳ_w)) / n_w  −  (Σᵢ∈O∩g (yᵢ − ȳ_o)) / n_o

Cluster-robust variance is Σ_g s_g² × (G/(G−1)). We use this SE for the CIs shown in the app. Post-fix coverage rises to 68% (NBA) / 84% (WNBA). The residual gap is genuine year-over-year drift in the 3-year anchor, not sampling error.

What the CI colors mean: [+2.1, +6.4] green = interval excludes zero → statistically distinguishable from average. [−1.3, +1.9] gray = interval straddles zero → we can't tell them apart from average with this sample.

Multi-team stint filtering

A player traded from team A to team B mid-season can't be off-floor for team A after they leave, because they're on team B. Yet naive on-off using an end-of-season roster would happily count team A's post-trade games as "off floor" opportunities, inflating the without-sample.

Fix: for each (player, team) pair, derive the game_id range where the player actually appeared in a lineup for that team from the possession data itself. Any "without" contribution outside that range is dropped.

Before/after: NBA 2024-25 Nikola Vučević (traded CHI → BOS). Pre-fix Δ Total = −3.9 (worst in top-15). Post-fix Δ Total = +0.3 (neutral). The −3.9 was entirely an artifact of comparing his on-court possessions against Chicago's late-season games he had no relationship to. In WNBA where roster churn is higher, 70% of player-seasons were affected pre-fix.

The on-off blind spot (A'ja Wilson case study)

A'ja Wilson is a 3× WNBA MVP. On our career-pooled Presence leaderboard she ranks #669. This is not a bug.

Reason: Aces' starters (Wilson, Chelsea Gray, Kelsey Plum, Jackie Young) play against opposing starters. Their bench, which is deep, plays against the opposing bench. Bench units are weaker on both sides, so V₀ with A'ja on court is naturally higher than V₀ with the bench in (because the opponent's V₀ generator is also stronger). Presence has no way to know this. On-off metrics fundamentally can't recover deep-team defensive anchors.

Same failure mode in NBA: SGA (2024-25 OKC), Jayson Tatum (Celtics), Bam Adebayo (Heat). All are top-10 by RAPM but bottom-30 by Presence.

Cross-verification, not replacement. Presence is one estimator of impact; RAPM (via box priors), Q-Rating (via LOO on a neural model), V-Rating (via state-value transformer), and iRAPM (via 8-role decomposition) are others. The comparison strip on the player profile shows all five side-by-side with an agreement flag. When they agree, confidence is high; when Presence disagrees with the rest, it's usually the deep-team blind spot in action.

Teammate Lift (chem-adjusted Δ)

How much of a player's on-court signal is actually their teammates? We subtract the expected pair-chemistry contribution from the raw Δ:

Δ_adjusted = Δ_raw − Σ_partners chem(X, Y) × poss_together(X, Y) / poss_X_on

Example (NBA 2024-25): Jamal Murray's raw Δ Total is +6.4 — top-10 in the league. His chem_bump_off is +3.6 and chem_bump_def is +2.4 — Jokić contributes most of that. Chem-adjusted Δ Total drops to +0.35. He's correctly identified as "elite on-court signal comes almost entirely from playing with Jokić."

The chem_bump exposure itself is the interesting stat. Nobody else publishes this: how much of your reputation is your best teammate?

WNBA — why the default is 3-year rolling

WNBA regular seasons are 44 games (NBA: 82). Per-player possession samples are ~5× smaller. Single-season CIs are correspondingly wide, so most rankings are indistinguishable from noise. To keep the WNBA leaderboard useful, the default WNBA source is the 3-year rolling pool: for each label year Y, presence is fit on possessions from Y−2 through Y. Single-season data is still available under the hood but not the default leaderboard view.

Data pipeline

  1. Load V-Rating possessions parquet (V₀ + lineups) and merge season/period/margin from possessions parquet.
  2. Build per-(player, team) game_id ranges from actual in-lineup appearances (stint filter).
  3. For each possession, accumulate per-player with-sums / without-sums (V₀, V₀²), per-(player, game) cluster contributions, and situational flags (clutch: period≥4 & |margin|≤5; blowout: |margin|≥20).
  4. Compute Δ Off, Δ Def, Δ Total and their cluster-robust SE via the sandwich formula.
  5. Emit per-season, 3-year rolling (WNBA), and career-pooled CSVs.

The NBA and WNBA pipelines share the same core estimator; the WNBA build differs only in defaulting to a 3-year rolling pool for the leaderboard.

About HoopQ

HoopQ is an independent NBA analytics site. Six ways to measure player impact: Multi-season RAPM (the classical baseline), Q-Rating (grades every on-court decision), V-Rating (grades whether the decision actually helped), Interpretable RAPM and Interpretable WPA (both break impact down into categories like scoring, playmaking, and defense), and Presence Impact (how much better your team plays when you're on the floor). Plus per-game career trajectories and lineup fit. Every metric is built possession-by-possession from raw NBA play-by-play (every shot attempt, assist, rebound, foul, turnover, block, and steal), covering 2000-01 through 2025-26: 13.0M events, 5.49M possessions, 30,815 games.

The Glossary explains every metric on the site, what it measures, how it's computed, how to read it, and where it falls short. The Evaluation page reports how well each metric behaves under stress: year-over-year stability, sample-size sensitivity, agreement across metrics.

Not affiliated with the NBA. Independent hobby project. Data is derived from public play-by-play feeds. Built and maintained by Ben Jenkins. Contact: hoopqhq@gmail.com. Corrections, methodology questions, and collaboration ideas welcome.

The models extend research that won Best Overall Paper at the 2026 MIT Sloan Sports Analytics Conference Research Paper Competition: Deep Reinforcement Learning for NBA Player Valuation: A Temporal Difference Approach with Shapley Attribution.

How to read this glossary

Each entry below explains one metric on the site: what it is, how it's computed, how to read it, and where it falls short. Use the search box to filter, or the Expand all / Collapse all buttons. For a shorter intro to the site, see About. For empirical validation of the metrics, see Evaluation.

Multi-season RAPM (2000-2026)

What it is. Regularized Adjusted Plus-Minus, the standard "all-in-one" basketball impact metric. Each player gets two coefficients: an offensive rating (points contributed per 100 possessions when they're on offense) and a defensive rating (points conceded per 100 possessions when they're on defense, where more negative is better). Total RAPM = ORAPM − DRAPM.

How it's computed. For every NBA possession from 2000-01 through 2025-26 (5.49M possessions across 30,815 games), we record the 5-on-5 lineup and the points scored. A sparse linear regression then solves for each player's marginal contribution while controlling for who they shared the floor with. We use a Bayesian prior built from each player's box-score profile (per-100 stats: points, assists, turnovers, rebounds, steals, blocks, etc.) so that low-minute players shrink toward what their box stats predict rather than toward zero. Ridge regularization strength (λ) was chosen by 3-fold cross-validation on held-out possessions.

Strengths. Captures impact you can't see in the box score: screens, spacing, defensive positioning. Multi-season pooling reduces single-season noise by ~10× (the GOAT-era list is dominated by Jokić, LeBron, Embiid, Chris Paul, Curry, Giannis, Kawhi, Garnett, Luka, Nash, consensus-correct).

Limitations. No causal claim, RAPM is correlational. Assumes linear and additive player effects (no interactions or context terms, so pair chemistry and lineup fit fall outside the model). Small-sample players (Wembanyama, Chet Holmgren: 2 seasons each) still appear in the top 25 with wide uncertainty. Doesn't separate role from talent (a great defender on a great defense gets partial credit for teammates' help).

Q-Rating (RL Q-Model)

What it is. A deep-RL player rating with two complementary flavors per side, both in points per 100 possessions (RAPM units):

  • Off / Def (action), sum of per-action-bucket contributions. Each event is credited to its actor; tells you which buckets (3pt pullups, rim finishes, blocks, rebounds, etc.) drove the value.
  • Off / Def (presence), neural ORAPM / DRAPM via a leave-one- out lineup swap on the trained Q model. Captures value not tied to any single event: spacing, screen-setting, rim deterrence, screen navigation. Closer analog to RAPM.

How it's computed. Each NBA possession is an RL episode; each event (shot, rebound, turnover, foul) is a step. A neural network learns Q(state, action, player), the expected remaining points in the possession given the game context, action, and actor. Player identity is a learned embedding; lineups are pooled by a 4-head self-attention layer. Action ratings = average advantage (Q with actual embedding − Q with league-average embedding) × usage per 100 possessions, summed per side. Presence ratings = average Q change at the start of each possession when the player is replaced with a neutral baseline (the mean of all real player embeddings), averaged across every possession they were in the off / def lineup, × 100.

Empirical-Bayes shrinkage. Per-100-poss rates have high variance for low-sample players, so every rating is shrunk by n / (n + κ) with κ = 8000 on-floor events (~1 season for a half-time player). Veterans keep ~90% of their raw rating; small-sample noise players collapse toward 0.

How to read it. Presence is the headline number; action shows where the value came from. For elite defenders the two diverge, Def (action) is mostly blocks/steals/rebounds, Def (presence) includes the value of just being out there. Player embeddings also cluster meaningfully (Curry near Lillard near Trae; Jokić near Sabonis), so the model isn't just memorizing role.

Presence Impact (V₀ on-off)

What it is. A per-100-possession on-off rating built from V₀, the possession-start value scored by the V-Rating causal transformer. Not the same as Q-Rating LOO presence: that estimator swaps one player out of a trained neural lineup; Presence Impact is a pure observational split of V₀ by whether the player is on the floor.

  • Δ Off = mean(V₀ | on court, their team on offense) − mean(V₀ | off court, their team on offense). Positive = team's expected value is higher when they're on the floor.
  • Δ Def = mean(V₀ | off court, their team on defense) − mean(V₀ | on court, their team on defense). Sign is flipped so positive = good defender (they reduce opponent value).
  • Δ Total = Δ Off + Δ Def.

Cluster-robust 95% CI. Possessions within the same game are correlated (opponent, tempo, night's shooting variance). Treating them as independent understates variance by ~4×. We group possessions by game and use the sandwich-formula SE. Empirical post-fix coverage: 68% (NBA) / 84% (WNBA) against a 3-year rolling anchor. On the UI, green CI brackets mean the interval excludes zero (statistically distinguishable from average); gray means the sample can't tell them apart from average.

Multi-team stint filter. A mid-season trade doesn't count team A's post-trade games as "off floor" opportunities. Each (player, team) pair's game range is derived from actual in-lineup appearances; without-samples outside that range are dropped. Before the fix, 12% of NBA rows and 70% of WNBA rows had inflated without-pools. Post-fix example: Nikola Vučević 2024-25 went from Δ Total −3.9 (worst in the top 15) to +0.3 (neutral).

Teammate Lift. The chem-adjusted variant subtracts the expected pair-chemistry bump from raw Δ, isolating individual value from teammate lift. On player profiles you see chem_bump_off, chem_bump_def, and the adjusted total. Jamal Murray 2024-25 raw Δ Total +6.4 collapses to +0.35 after subtracting the Jokić-driven pair bump; the exposure itself is a story no other public site publishes. Available for both NBA (2000-2026) and WNBA (2002-2026) with the caveat that pre-2004 pair-chem values are damped by degenerate rolling-window sample sizes.

On-off blind spot. Presence under-credits stars on deep teams because their bench faces the opponent's bench and looks fine without them. Poster child: A'ja Wilson ranks #669 on the career pool despite being a 3× WNBA MVP. Same pattern in NBA: SGA (deep OKC), Tatum (Celtics), Bam Adebayo (Heat). This is a fundamental limit of any on-off estimand, not a bug. Cross-verify against RAPM, Q-Rating, V-Rating, and iRAPM on the Impact Metric Comparison strip on every player profile; when Presence disagrees with the rest, the disagreement is usually the deep-team blind spot in action.

How to read the leaderboard. Two views. "All-time single-season peaks" pools every (player, season) row so LeBron 2008 and Jokić 2022 co-appear as historical peak seasons. "Single year" filters to one season, one row per player. WNBA defaults to a 3-year rolling pool because single-season sample sizes are too small for tight CIs. Details on the Presence Methodology page.

V-Rating

What it is. The state-value sibling to Q-Rating. A causal transformer estimates the state-value function V(s) at every prefix within a possession, where V(s) is the model's expected total possession reward given the events observed so far. For each event, ΔV = V(s_t) − V(s_{t−1}) is the value change attributable to that event and is credited to the actor. Aggregated across a season, this is a player's contribution to their team's expected offensive value (EPV, à la Cervone/D'Amour/Bornn 2016).

How it's computed. A lineup-context token seeds V₀; positions 1..T are event tokens (event type, action bucket, actor, is_assisted, is_made, shot location, clock, period, margin). Trained on 2000-01 → 2025-26 pooled possessions (~5.5M possessions). Attribution: shot events split into "choice" (bucket-conditional expected reward) plus "luck" (residual vs bucket mean, EB-shrunk); assists take 60% of shooter's choice; rebounds get 30% credit (contested); turnovers get 40% actor share; on-floor lineup shares the V₀ gap so passers/creators get credit for setup value without needing explicit pass events, and elite defensive lineups get credit for rim deterrence.

How to read it. Same points-per-100 scale as Q-Rating and RAPM, MVPs score +6 to +10 per 100 actions. Q-Rating is Q(state, action, player), action-conditional. V-Rating is V(state) with per-event attribution, state-based, so it captures value that isn't tied to a specific action-value bucket (setup, spacing, defensive presence). Cross-verifying: when Q-Rating and V-Rating agree the player's impact is well-priced; when they disagree, the gap is informative about what each metric is measuring.

Limitations. No tracking data, so off-ball motion (screens, cuts, spacing) is captured only via the on-floor lineup share (Fix 3), not as explicit events. Shot luck is heavily shrunk to avoid single-season shot-making noise.

Clutch Q-Rating

What it is. Q-Rating restricted to clutch moments: Q4 or overtime, clock ≤5 minutes, |score margin| ≤5. Combines two independent measurements into one leaderboard.

  • Action Q: per-100 shot quality above league expectation on clutch actor events. Playmakers get credit on assisted makes so passing value doesn't disappear into the shooter alone. Empirical-Bayes shrunk with κ=500 (multi) / κ=100 (single season).
  • Clutch ORAPM / DRAPM: Ridge regression on ~230,000 clutch possessions (~45,000 for WNBA) with α=25,000 aggressive regularization. Captures lineup-level offensive and defensive presence, rim deterrence, off-ball value, and creation gravity that per-action doesn't see.

How to read it. Sign conventions match regular RAPM: Action Q and Clutch ORAPM positive = good; Clutch DRAPM negative = good (allowed fewer pts / 100, DRAPM convention). Total = Action + ORAPM − DRAPM. Positive total = good overall clutch player.

Why this fixes a real gap. Action-only clutch rewards shooters (Curry, Nash, Redick, Peja, Reggie Miller, Ray Allen top out) but underrates high-usage playmakers whose clutch impact is via others (LeBron, CP3, Draymond), rim protectors whose value is deterrence (Wemby, Gobert, Fowles), and off-ball gravity (Klay). Adding Clutch ORAPM/DRAPM via lineup ridge captures those effects and produces a balanced two-way clutch leaderboard: LeBron went from action-rank #296 to total-rank #28 (action −0.46, ORAPM +1.39, DRAPM −0.83). Dirk Nowitzki is #1 all-time by total.

Cross-verify against iWPA. Interpretable WPA measures leverage-weighted win probability change per clutch event. Clutch Q-Rating measures pts per 100 clutch events + lineup ridge. Different units, similar concept. Correlation ≈ 0.31 on shared players (positive but distinct), because iWPA weights by state leverage while Clutch Q-Rating weights by event count.

Limitations. Only 4.7% of events are clutch, so per-season RAPM at ~2-10k possessions is too noisy and the CSV emits action-only for single-year. Multi-season pool (229k NBA / 45k WNBA possessions) is where the ORAPM/DRAPM signal lives. Very-low-clutch-min players (rookies, deep bench) are dropped by the min-events filter.

Interpretable RAPM

What it is. A ridge regression variant of RAPM that decomposes each player's impact into 8 interpretable roles: Scoring, Playmaking, Off Reb, Def Reb, Def Actions, Def Presence, Off Presence, Turnovers. Unlike standard RAPM which gives one opaque coefficient per player, Interpretable RAPM tells you where that impact comes from.

How it's computed. One design matrix row per PBP event, target y = points scored on that event (0, 2, or 3 typically). Each row's columns are: (a) offensive lineup on floor (5 indicators), (b) defensive lineup on floor (5 indicators), (c) actor's specific (player, bucket) column with sign matching offense/defense, (d) assister's column on assisted shots, (e) blocker's column on blocked shots, (f) stealer's column on stolen turnovers. Sparse ridge (α = 25,000) with per-event actor-per-bucket coefficients rolled up post-fit into the 8 role categories.

How to read it. Multi-season 2000-2025 pooled, so a Total of ~+40 for MVPs equals ~+8 per season on the classical RAPM scale. Scoring, Playmaking, Off Reb, Def Reb, Off Presence, Turnovers: positive = good. Def Actions and Def Presence: NEGATIVE = good (matches DRAPM tradition, the player reduces opponent expected value). UI applies visual color inversion on the def columns so green always renders when the player is above-average defensively.

What each role captures. Scoring: actor value on shot events (rim, midrange, 3PT, FT buckets). Playmaking: assister-specific coefficient on assisted shots, cleanly isolates pass value from finish value. Off Reb: β_actor[oreb] with y augmented by same-possession points scored after the rebound — elite offensive rebounders whose boards lead to points get positive credit. Def Reb: β_actor[def_reb] sign-flipped, with y augmented by points scored by the rebounding team on their NEXT possession — rewards elite rim protection + outlet passing that fuels transition. Def Actions: blocker + stealer + defensive fouls. Def Presence: on-floor defensive lineup coefficient (rim deterrence, help defense). Off Presence: on-floor offensive lineup coefficient (spacing, gravity). Turnovers: actor value on TO buckets + offensive fouls.

Limitations. One target per event; non-scoring events (rebounds, fouls, blocks) have y=0 so those coefficients are estimated via the fit's global adjustment rather than direct point-scoring signal. Multi-season pooling means changing players (rookie growth, aging veterans) get averaged into one number. Sample-per-parameter ratio (~40:1) is fine for ridge but individual per-bucket cells with <500 events are noisy, the role rollup helps but raw per-bucket values should be read with the sample-size context.

Interpretable WPA

What it is. Same ridge design as Interpretable RAPM but the target is 100 × Δ win probability rather than points scored. Same 8-role decomposition (Scoring, Playmaking, Off Reb, Def Reb, Def Actions, Def Presence, Off Presence, Turnovers). Reads as "wins added per 100 events" broken down by role.

How it's computed. One design matrix row per PBP event; y = 100 × (WP_after, WP_before). Identical column layout to iRAPM (offensive lineup + defensive lineup + actor-per-bucket + assister + blocker + stealer). Sparse ridge (α = 25,000). No explicit clutch weighting is applied, but because WP shifts are naturally larger in close, late-game moments (a tied-4Q shot moves WP by ~0.05 vs ~0.001 in a blowout), leverage is baked into the target directly.

How to read it. Sign conventions match iRAPM: Scoring, Playmaking, Off Reb, Def Reb, Off Presence, Turnovers are positive = good; Def Actions and Def Presence are NEGATIVE = good (matches DRAPM tradition). Multi-season 2000-2025 pool, top-15 spans Wade, Kobe, LeBron, Paul, Ginobili, Lillard, Nash, Westbrook, the "clutch pantheon" is exactly who you'd expect.

Limitations. WP shifts have larger per-event variance than points (0-3), so on-floor lineup coefficients are noisier. At half-season sample sizes the ridge cannot cleanly separate player from lineup and the leaderboard becomes dominated by presence terms rather than actor skill. For that reason we only ship the multi-season pool; single-season iWPA is not user-facing. Correlation with 25-year RAPM is moderate (ρ ≈ 0.5), because iWPA rewards a narrower "big-moment" signal.

QR-DQN Skill Advantage

What it is. A distributional-RL companion to the Q-Rating model. For every offensive event, we predict the per-event reward distribution above a no-actor baseline Q(s, a, teammates, defenders, context). The per-player skill advantage is the mean of that residual, points per event the actor adds beyond what a league-average player would have produced in the same situation. The 95% CI shown in the Player Profile is a proper SE-of-the-mean confidence interval, not the spread of individual outcomes.

How it's computed. Two-stage: first a baseline mean-Q model is trained with the actor masked out of the offensive lineup; then a Quantile-Regression DQN with 51 quantile heads is fit (pinball Huber loss) on advantage = actual return − baseline prediction. Per-player aggregation averages the distributional mean across each player's events.

Strengths. Interpretable units (pts/event) and clean residualization that removes the action-outcome variance dominating a raw RTG estimate. Top players by mean advantage are consistently the league's high-usage per-event creators.

Limitations. Per-event efficiency systematically undersells cumulative creators (Jokić, Giannis show small positives) because action-by-action credit can't capture possession-spanning value the way RAPM does. Best read alongside RAPM, not as a replacement.

Kalman Career Trajectory (game-by-game)

What it is. A DARKO-style Kalman filter over per-game Q-Rating, producing a smoothed skill trajectory across every game a player has played. The observation is per-game Q-Rating in points per 100 team possessions, not a Game Score composite: it's the neural action-bucket decomposition redistributed per event, so the trajectory tracks actual skill rather than box-score production.

How it's computed. Each season's per-player action-bucket coefficients are decomposed into per-event credits, aggregated per (player, game) as a per-100 rate, and fed one at a time into a Kalman filter. Aging drift comes from a smoothed age curve fit on all observations pooled by age. The rookie prior is seeded from a draft-slot prior (median career-year-1 Q-Rating by pick). Career-average presence is folded in as a season-constant so elite defenders and off-ball creators surface at their real magnitudes.

How to read it. The bold line is the posterior mean; the shaded band is a ±1.96σ 95% confidence interval; small dots are per-game observations. For players drafted since 2000-01 the x-axis is career game number so trajectories are apples-to-apples across eras; for earlier careers the axis falls back to date since our data window starts mid-career. The Compare-With input overlays a second player.

Limitations. Career-average presence gives the trajectory its right magnitude but is constant within a career, so within-season presence changes (injury, role shift) don't move the line beyond what action-bucket noise contributes. Rookies have wide bands until they've played enough games.

Player Projections

What it is. 1-, 2-, and 3-year-ahead forecasts of each per-100 metric (points, assists, rebounds, etc.) using DARKO-style aging curves and a Kalman-flavored update that blends a player's own trajectory with a prior built from comparable players at the same age. Surfaced on the Projections page and in the Player Profile career trajectory chart.

Playstyle vectors

What it is. Each player's action-frequency vector projected into a low-dim space, then used to find their nearest neighbors in style (not in impact). Useful for "who plays like X?" comparisons. Distinct from the Play Types view: that one shows the raw distribution, this one shows similarity.

Play Types (Hierarchical Options)

What it is. Per-player distribution over 12 offensive "options" inferred heuristically from the action stream: pnr_handler, pnr_finisher, isolation, post_up, cut, putback, spot_up, off_screen, transition, free_throw, rebound, other. Tells you what a player does, not how well they do it. Click a player for offensive distribution + per-option success (pts/action) + defensive rates per 100 def-poss (steals, blocks, fouls).

How it's computed. Each event is labeled by a rule combining action bucket, is_assisted, possession step, and seconds-into-possession. Isolation vs PnR handler split uses the assisted flag (assisted pullup/drive → PnR; unassisted → isolation). Defensive credits are computed separately: stealer/blocker via player3_id from steal/block events, normalized by the player's defensive possessions on the floor.

Limitations. Heuristic, not learned. Without tracking data, "iso vs PnR" is an imperfect proxy; missed shots default to isolation since assist credits only attach to makes. Transition uses a clock-only proxy (no start_type available in the possessions parquet). Use it for shape/role inference, not for skill grading.

Revealed Preferences (Inverse RL)

What it is. For each (player, action, game-context) triple, the log-ratio of the player's frequency against the league baseline, what actions a player chooses reveals their inferred reward weights. Reveals e.g. DeMar DeRozan's strong midrange preference, Curry's elevated clutch-3 preference. State contexts: clutch, blowout, early-clock, late-clock, transition, normal.

Behavior Policy (π_b)

What it is. A neural classifier that learns π_b(a | state, player): which action does this specific player tend to choose in this specific state? Used as an action-choice prior across several other models on the site.

How it's computed. Same lineup-attention encoder as the Q model, but the output head is a softmax over the action vocabulary instead of a scalar Q. Trained on cross-entropy against the observed action; achieves ~38% top-1 and ~68% top-3 accuracy on held-out events.

Shot Quality (xFG)

What it is. A LightGBM classifier that estimates expected FG% for each shot given the shot location, action type, and shot-clock state. Per-player it aggregates to two axes: Shot Making (actual points minus expected points, how much the player beats their location-baseline) and Shot Selection (mean xFG per attempt, how high-quality the shots they take are on average).

How it's computed. Gradient-boosted trees trained on every field-goal attempt with the shot's location coordinates, primary action bucket, and clock features. Per-shot xFG is used to compute per-player pts and selection aggregates, then compared against actual pts scored on those shots.

How to read it. Shot Making of +150 pts means the player scored 150 more points on their attempts than xFG expected. Shot Selection of 0.55 means the average shot they attempted had a 55% baseline FG probability. Two axes let a low-volume rim finisher (high selection, modest making) and a high-volume creator taking hard shots (lower selection, positive making) end up on different quadrants.

Limitations. The model doesn't see defender proximity or contested-vs-open state (no tracking data). That context lives in V-Rating instead.

Shot Charts

What it is. Side-by-side frequency + efficiency SVG heatmaps over a half-court projection, available on every player profile and every coach profile. Each bin is 1 foot × 1 foot; the shooter's location is bucketed into a 50 × 47 grid over the standard NBA half-court coordinate space (LOC_X ∈ [−250, 250], LOC_Y ∈ [0, 470]).

Frequency panel. Color intensity encodes how often the subject shoots from that bin. Log-normalized per-subject so star spots don't drown out the rest of the chart.

Efficiency panel. Color encodes FG% at that bin minus league baseline FG% at the same bin. Green above (better than league), red below (worse), clamped at ±20pp so individual hot/cold cells don't dominate.

Coach charts aggregate every shot taken by teams under that coach across their full career or a selected season. Kerr's coach shot chart lights up above the break; D'Antoni's spreads across the arc and rim; Popovich's late-Spurs chart shows the corner-three shift.

Player charts support side-by-side compare-with (type a second player's name to overlay) and shot-type toggle (All / 2PT / 3PT). Hover a bin for shot count, FG%, and delta vs league.

Pair Chemistry (Pair-Augmented RAPM)

What it is. Extra points per 100 possessions a specific pair generates (offense) or prevents (defense) when they're on the floor together, beyond what their individual RAPMs predict. A measure of two-player chemistry / fit / synergy.

How it's computed. Same ridge regression as RAPM but with an extra binary feature for every player pair that played enough possessions together. The pair coefficient is what's left after the individual player effects are accounted for: the "team-up bonus." Positive on offense = synergy. Negative on defense = synergy (less points allowed).

Context filter. The same ridge is also fit on three possession subsets, clutch (Q4, last 5 min, |margin| ≤ 5), regular (|margin| ≤ 15 and not clutch), and blowout (|margin| > 15), surfacing pair chemistry that varies by game state. Curry+Draymond generate +4.13 pts/100 in regular offense but only +1.35 in clutch (defenses take away the DHO late); Nowitzki+Terry surface as a top-3 clutch offensive pair, matching their real legacy. The Context dropdown on the Chemistry page switches between the four views.

Limitations. Identifiable pairs need a meaningful number of possessions together; one-game cameos shrink to zero. Pair coefficients are still correlational, not causal, and can't separate "lineup synergy" from "shared strategy" or "shared coach." Clutch context has ~2,500 qualified pairs vs ~34,000 in regular because each pair needs enough clutch minutes to fit.

Lineup GNN + Matchup Influence

What it is. A heterogeneous graph neural network over each 5-vs-5 possession, treated as a 10-node graph with three edge types: within-offense (chemistry among the offensive 5), within-defense (chemistry among the defensive 5), and cross (offense-vs-defense matchup edges). Two hops of message passing let signal like "my teammate has a favorable matchup, which frees me up" propagate through the graph.

How it's computed. Node features are learned player embeddings plus a role token (offense or defense). Each layer aggregates three separate messages (within side, within side, cross) into an update. Two layers are applied, node representations are pooled by side, and a standard Q head predicts expected reward on the possession. Downstream artifacts: per-player GNN presence via leave-one-out lineup swaps, a per-player synthetic matchup sweep (X vs candidate defender in an isolated 5-vs-5 context), and team-vs-team predictions aggregated over each team's most-used 5-man lineups.

Where it shows up. The Matchup Influence table on every player profile (defenders whose presence lowers or raises the player's expected Q per possession), the vs-team preview widget on team profiles, and the GNN presence stat card on the Lineup Sim page.

Limitations. Play-by-play has no per-shot defender attribution, so the matchup signal is on-court concurrence not literal 1-on-1 assignment. Same-era filtering restricts candidate defenders to seasons a player actually played in so anachronistic pairings don't populate the table.

Lineup Optimization

What it is. For each team, exhaustively scores every 5-man lineup over its rotation players and returns the best by predicted net rating.

How it's computed. Lineup score = sum of the 5 players' Q-Rating presence (Off − Def, neg-good convention) + offensive pair chemistry summed over all 10 pairs − defensive pair chemistry summed over all 10 pairs. Q-Rating presence is the leave-one-out swap on the trained Q model with attention pooling over lineups (closer analog to RAPM than the action decomposition).

Team RAPM (schedule-adjusted)

What it is. A schedule-adjusted per-team-season net rating in pts/100 possessions. Same possession-level ridge regression skeleton as player RAPM, but the features are team-season one-hots (30 offense + 30 defense per season) rather than player indicators. The coefficient answers: "How many pts/100 does this team score above league average, holding opponent quality constant?"

How it differs from Net Rating. Net Rating = ORtg − DRtg observed on the floor, no adjustment for opponent quality. Team RAPM controls for schedule strength via the def-team-season indicators on every possession. Typical correlation between the two: ~0.95. Where they diverge (5-10% of team-seasons), the divergence tells you whose schedule was abnormal that year.

How to read it. Both Team RAPM and Net Rating are shown on the team profile page alongside Team Q-Rating (the roster-aggregate estimate). Three parallel team-strength estimates lets you cross-check: schedule-adjusted (Team RAPM), observed (Net Rating), and roster-quality (Team Q-Rating).

Career floor tests. Top-10 all-time (2000-01 through 2025-26): 2015-16 Warriors +4.38, 2014-15 Warriors +4.02, 2023-24 Thunder +3.99, 2014-15 Spurs +3.81. All defensible.

Team Q-Rating (detrended)

What it is. A team-season summary of the roster's Q-Rating quality, centered so the sign is meaningful. Raw team Q-Rating is a possession-weighted sum of every roster player's Q-Rating (action + presence). Detrended = raw minus the season league mean. Positive means the roster grades above the league that year, negative means below.

How it's computed. For each (team, season): weighted mean over the roster of each player's total-impact (action + presence − opponent-side presence, good-direction on both sides), weighted by that player's possessions with that team that year. The result is multiplied by 5 to put it on per-100-poss scale (five players share each possession). Detrending subtracts the season-wide league mean so ~50% of teams end up on each side of zero.

How to read it. Semantically parallels Team RAPM (already zero-centered). +2 means "roster grades ~2 pts/100 poss above the league that year," −3 means "3 below." The trajectory chart on the team page shows total plus off/def breakdown, with Team RAPM overlaid as a reference line.

Limitations. Aggregates the whole roster equally by possessions played, so it doesn't distinguish starter-heavy vs bench-heavy usage. Presence uses career-average magnitudes broadcast across seasons, so team-season snapshots reflect the long-run quality of the players on the roster rather than that specific year's form.

Team Action Profile

What it is. Per (team, season, action bucket) rate of how much value each team generated from each of the ~30 offensive and defensive action buckets, centered against the season league mean per bucket. Surfaces two views on the team page: "Season strengths" ranks the current-season buckets from most above to most below league, and "By season" shows a heatmap of every season the team has on record.

How it's computed. For each (team, season, bucket): sum every roster player's per-bucket Q-Rating coefficient weighted by their possessions with that team that year, divided by total team possessions. Detrend by subtracting the season league mean per bucket. Green cells indicate the team was above league on that bucket that year; red indicates below.

How to read it. Bucket rows in the heatmap are sorted by how identity-defining the bucket is over the team's history, so signature strengths and weaknesses float to the top. Low-volume specialty buckets (3pt corner pullup, midrange post hook, foul-take, etc.) swing more than high-volume ones (rim finishes, midrange catches).

Limitations. Only sees action-level credit, not presence. Two teams with identical action profiles but different rim protection would look identical here even though their defense is not.

Team Builder (projected W-L)

What it is. Interactive roster projector: pick 8-12 players and set per-player minutes, get a projected 82-game W-L with a 90% confidence interval bell curve. Useful for trade / free-agency what-ifs and for stress-testing whether the presence + chemistry stack agrees with real team outcomes.

How it's computed. Player contribution = Σ (mini/48) × (off_presence_shrunk + (−def_presence_shrunk)), the same Q-Rating presence used by the Lineup Simulator, weighted by each player's share of a 48-minute team floor time. Pair chemistry contribution = Σ shared-minute-weighted (chem_off − chem_def) from the multi-year ridge, where shared minutes ≈ mini × minj / 48 (independence approximation). Raw net rating (player + chemistry) is then calibrated to actual team net ratings via a linear rescale, calibrated = −6.15 + 0.80 × raw, fit on 773 historical team-seasons (2000-2025, R² = 0.74). The calibration corrects the systematic overstate that comes from presence values already partially capturing pair-chemistry effects, so a naïve sum double-counts. Wins = 41 + calibrated_net × 2.7 (Basketball-Reference rule of thumb: ~2.7 wins per +1 net rating, 41-41 baseline).

How to read it. The Projected W-L (μ) stat card shows the mean of the wins distribution. The bell curve below shades the 90% CI band (μ ± 1.645σ). Total σ combines three variance sources: 82-game binomial noise (~4-5 wins for a typical team), model residual after calibration (σ ≈ 6.7 wins), and a small-sample penalty of ~1.5 wins per rotation player with < 1000 possessions of presence data. Total σ is typically 7-10 wins for a healthy rotation; wider bands mean more of your roster is small-sample.

Where it falls short. No schedule strength, no coach effects, no injury / DNP modeling. Independence assumption for shared minutes overstates bench-pair overlap and understates starter-pair overlap. Small-sample players widen the confidence band but don't fix mean bias for teams with unusual chemistry patterns (e.g., 2024-25 OKC significantly overperformed roster-projected wins). Sanity checks: 2024-25 Boston → 60-22 (actual 61-21) ✓, 2016-17 Warriors → 72-10 (actual 67-15) ✓.

Championship odds (end-of-regular-season + preseason)

What it is. Two Monte Carlo bracket simulators sharing the same series-win logistic. End-of-regular-season uses final rosters + observed per-season Q-Rating presence. Preseason uses prior-season player value: Kalman-preferred prior-3-year presence with V-Rating ensemble, age drift, draft-slot rookie prior, availability discount, newcomer skill discount, role compression, and share-weighted pair chemistry. Raw strength is era-centered (per-season) and calibrated to net rating with a continuity + prior-year-residual term. For past seasons the roster is each team's opening-night lineup (reconstructed from first 8 games); for the upcoming 2026-27 season (no games yet) it is the current roster with a position-aware depth-chart minutes model and scraped long-term injuries. Both feed the same raw → net rating → Pythagorean W-L → seed → 30k-trial bracket (1v8 / 4v5 / 2v7 / 3v6 → semis → conf finals → Finals). Retrospective evaluation on 2005-06 through 2024-25 (20 NBA seasons, 600 team-seasons).

Top-K champion accuracy. How often the actual champion is in the model's top K by predicted p_champion:

Preseason End of reg season
Top-145%25%
Top-360%60%
Top-575%95%
Top-895%100%

Probabilistic calibration. Brier score on the (team, season) → won-championship binary: preseason 0.0290, end-of-season 0.0287 (uniform baseline 0.0322, lower is better). Both are essentially as well-calibrated as probabilistic forecasters. Log-loss: preseason 0.139, end-of-season 0.103, because end-of-season concentrates probability more tightly on real contenders, which log-loss rewards.

Ranking accuracy. Per-conference Spearman rank correlation between predicted calibrated_net and observed net rating: preseason 0.74, end-of-season 0.90. This directly drives seed accuracy: the higher end-of-season correlation is why its top-5 champion hit rate is 95% vs preseason's 75%.

Playoff-round hit rates (predicted top-K by p_champion vs actual teams reaching that round):

Round Preseason End of reg season
Playoff pool (top-16)79%72%
Conf semis (top-8)65%54%
Conf finals (top-4)46%45%
Finals (top-2)43%38%

Preseason wins the playoff-pool hit rate because it doesn't overreact to mid-season injuries + trades that reshape the end-of-season model's top-16 away from the actual playoff seeds.

Wins projection MAE (against 41 + 2.7 × observed_net): preseason 6.67 wins, end-of-season 4.49 wins. This is the most user-facing accuracy number; preseason projections are ~50% noisier than end-of-season.

Reliability of preseason p_champion (predicted decile vs observed champion rate): predicted 5-10% → observed 5.8% ✓, predicted 10-20% → observed 13.3% ✓, predicted 20-35% → observed 50% (n=8, small-sample; model is slightly underconfident at the very top). Well-calibrated in the middle bands.

Where it falls short. Preseason can't extrapolate mid-summer trades (Kawhi to TOR 2018-19, AD to LAL 2019-20) or steep breakout years (Jokić 2022-23, SGA 2024-25). End-of-season misses concentrate on Cinderella playoff runs (2011 Mavs, 2019 Raptors) where regular-season net rating didn't predict the playoff run.

March Madness odds

What it is. A bracket-seeded Monte Carlo sim over the ACTUAL NCAA tournament bracket. Unlike a seed-your-own playoff model, March Madness has a fixed 64-team bracket, so we take the real field (teams, seeds, matchups scraped from ESPN) and simulate it forward 20,000 times to get each team's odds of reaching every round.

Team strength. Same chain as the pro Championship model: per-player Q-Rating presence (off minus def) weighted by possession share, plus career pair-chemistry contributions weighted by shared possessions, calibrated to net rating. This orders the field well even though raw college net rating is schedule-inflated: the Spearman correlation between our team strength and the committee's seeding is -0.76 (better strength maps to a lower seed number), and the eventual champion is our top-8 team in 100% of tournaments.

Win model. Single games (not best-of-series) sampled from a logistic fit on ~1,080 historical tournament games on neutral courts: P(A beats B) = sigmoid(0.13 x (strength_A minus strength_B)), so a +10 strength edge is about a 79% favorite and +5 is about 66%. First Four play-in games are simulated too.

Bracket reconstruction. ESPN's region labels are inconsistent (a region's early rounds are sometimes tagged by host-city pods), so the bracket tree is rebuilt from results: a round's two participants each won a prior-round game, which become its child nodes. This sidesteps region parsing and the Final Four pairing entirely.

Retrospective accuracy (2008 through 2026, 18 tournaments; the champion's rank is by our predicted p_champion within the field of ~64):

Champion in top-139%
Champion in top-372%
Champion in top-589%
Champion in top-10100%
Avg champion rank2.6 of ~64
Final Four recall (top-4)69%

How to read it. Odds are pre-tournament (from the season's roster + team strength), not updated as games are played. The Odds table gives each team's chance to reach each round; the Bracket view shows the actual matchups with per-game win probabilities, marks the real winner, and flags the champion. Coverage is 2008 onward (2020 cancelled for COVID; 2006-2007 predate ESPN's bracket tagging).

Where it falls short. Misses are the Cinderellas that a strength model can't foresee: 2023 LSU (a 3-seed) ranked 8th by our model before winning it all. Chalk years are nailed (UConn ranks 1st across most of its 2009-2016 run).

Preseason (forward-looking) mode. The Odds toggle switches team strength to a purely retrospective input: each player is rated only by a recency-weighted blend of their prior three seasons (no current-season play), newcomers get a neutral prior, and the win logistic is refit on those preseason strengths. This is what a real bracket forecast can use before a season starts. It is naturally less accurate than the in-season model: over the same 18 tournaments the champion lands in our preseason top-1 33% / top-3 72% / top-5 83% / top-10 89% (avg rank 3.9), Final Four recall 60%. Its blind spot is deliberate and important: with no recruiting signal it cannot see a freshman or transfer breakout, so it under-rates newcomer-driven teams, most glaringly 2023 champion LSU (Angel Reese plus transfers), which it had at 0.1%. The Preseason vs actual view lines the preseason and in-season numbers up with the real result so those over- and under-performers stand out.

WNBA Draft Projection

What it is. A translation model that projects current college players to the WNBA. It is trained on the players who appear in both our college and pro datasets (493 of them), each scored with the same impact metrics on both sides. That shared-metric bridge is what makes the projection possible, and almost no public site has it: college and pro numbers usually come from different systems that cannot be compared directly.

Matching the two leagues. Each player's full college career is joined to their full WNBA career via shared ESPN player ids and name matching (with maiden-name and nickname recovery). Internationals who never played NCAA ball are unmatchable, so coverage is US college only (2006-2026).

Three models, all on college inputs (Q-Rating presence, V-Rating, iRAPM, plus per-season volume and peak-season RAPM):

  • Reach (logistic) = probability the player reaches the WNBA with real minutes. Out-of-sample AUC 0.90.
  • Projected impact (ridge) = expected WNBA per-100 impact if she reaches. The target is observed pro impact, not our shrunk career rating: shrinkage compresses the range and kills the signal. Out-of-sample correlation +0.45.
  • Projected longevity (ridge) = expected career length, out-of-sample correlation +0.32.

Reach features are deliberately volume-balanced. An earlier version leaned on raw cumulative college possessions, which flattered four-year role players and under-rated dominant underclassmen (JuJu Watkins, Sarah Strong). The current model uses per-season volume plus unshrunk peak-season RAPM, so young stars keep a sane reach probability without hurting AUC.

Projected outcome distribution. Each prospect gets a five-part probability bar, summing to 100%: Never sticks (= 1 minus reach) then, conditional on reaching, Fringe/Depth, Rotation, Starter, and All-WNBA. The tier split comes from a Normal centered on projected impact with the model's out-of-sample residual spread, so a wider bar honestly reflects a less certain projection. Hover any segment for its exact percentage.

Comp. The current WNBA player whose playstyle vector is closest to the prospect's (cosine similarity on the shared action buckets), a similarity read, not a ceiling.

How it is validated. Everything is out-of-sample with grouped-by-player cross-validation, so a player's own seasons never leak into her own prediction. The probabilities are calibrated: expected calibration error is 0.00 for reach and 0.02 for the outcome tiers (0 is perfect), meaning a stated 60% happens about 60% of the time. On a held-out retrospective of recent debutants, the projected order matches actual WNBA impact at rank correlation +0.43.

How to read the page. Draft prospects ranks current college players by draft score (reach times projected impact) and shows the outcome bar, reach %, projected impact with its percentile, and the closest comp. Retrospective shows recent debutants with both projected and actual tier, and a marker for whether they beat, hit, or fell short of the projection.

Where it falls short. The impact signal is real but modest: college dominance only partly translates and the model is deliberately conservative, so it will not crown the next superstar outright. Reach is the strongest piece; impact and longevity are directional. It also cannot see things outside box-adjacent impact (injuries, motor, off-court development), and the training set is capped by the size of the two-league population, not by how much college data we have.

Trade Analyzer

What it is. Two-team roster simulator. Pick two teams, select players on each side, hit Swap: both rosters get recomputed side-by-side with new calibrated net ratings, projected wins, and championship odds for both teams.

How it's computed. Same chain as the Team Builder, applied to both sides in parallel. Each team's calibrated net rating = a + b × (Σ (mini/48) × per-player impact + Σ shared-min pair chemistry) with league-specific (a, b), fit on 773 historical NBA team-seasons (R² = 0.74) and 318 WNBA team-seasons (R² = 0.60). Projected wins follow the Basketball-Ref Pythagorean rule. Championship odds run a joint 5,000-trial Monte Carlo bracket sim where BOTH swapped teams update simultaneously (so a trade that moves talent from a title contender to a lottery team shifts both trajectories at once); every other team is held at its current-season baseline.

How to read it. Green Δ = the team improved; red Δ = they got worse. The most meaningful comparisons are 1-for-1 star swaps (each side stays at the correct team-minutes total). Uneven trades leave rosters mis-summed, projections are still shown but the minutes note will flag it. Sanity checks: SGA (OKC) ↔ LeBron (LAL) hits OKC for ~12 wins and lifts LAL by ~6, Jokić (DEN) ↔ Tatum (BOS) drops DEN by ~10 wins, both matching the intuition that MVP-tier stars carry ~8-12 wins of value.

Where it falls short. No salary or contract logic, so all trades are cap-legal here even when they wouldn't be. No positional balance check. No coach fit or chemistry-with-teammates interaction (though pair-chemistry deltas do fold in). Minutes transfer with each player; if you send a starter, the receiving team's rotation absorbs their minutes even if it pushes past 240.

Coach Ratings

What it is. Coach impact isolated as the possession-level residual after controlling for player quality. Two parallel metrics are shown side-by-side per coach: Coach RAPM and Coach Q-Rating. Both are in points per 100 possessions, same sign convention as player RAPM (off positive = good, def negative = good, total = off − def). Covers every head coach who worked a regular-season game from 2000-01 through 2025-26.

Coach-by-game attribution. Basketball-Reference publishes W-L splits per coach per team-season (e.g., Bulls 2003-04: Cartwright 4-10, Myers 0-2, Skiles 19-47). We scrape those splits, then walk each team's game log chronologically and allocate games in order, first N to coach A, next M to coach B, and so on. Franchise transitions handled explicitly (SEA → OKC, NJN → BKN, CHH → NOH → NOP, VAN → MEM). Output: 57,402 (game, team, coach) rows across 776 team-seasons.

Two-step residual. For every qualifying possession (~5.2M after garbage-time filter), we first compute the expected pts/100 using a player prior held constant, either box-score-prior player RAPM (with defensive term as a 50/50 blend of def_rapm and Q-Rating defensive presence, so rim deterrence and off-ball help are credited) or the full Q-Rating (action + presence). We then subtract that expectation from actually observed points to get a per-possession residual, and finally ridge-regress the residual on [off_coach, def_coach, season] one-hots. The coach coefficient is what pts/100 the players plus league era couldn't explain, the coach's persistent contribution above player baseline.

Shrinkage and Kalman smoothing. Career values are empirical-Bayes shrunk with κ = 40,000 possessions (~500 games per side). Per-season values use κ = 12,500 plus a 1-D Kalman smoother that treats each season as a noisy observation of a slowly drifting latent coach quality, so trajectories don't jump on noise.

Cross-metric agreement. The two priors produce two independent estimates. On the ≥250-games qualified pool: Pearson r = 0.76 total, r = 0.97 defense (near perfect agreement), r = 0.49 offense (real divergence). When both metrics agree the coach's effect is robust; when they disagree the ranking is method-sensitive on that coach (typically an offensive-scheme case Q-Rating captures better than box-score RAPM).

What the coach profile also shows. Kalman-smoothed season trajectory with off/def breakdown; aggregate shot chart of every shot taken by teams under that coach (frequency and efficiency versus league baseline, per season and career); Coach style widget showing signature action-bucket tendencies weighted by their share of team-season games, reveals things like Kerr over-indexed on above-the-break pull-up threes (+13.9% vs league), D'Antoni on catch-and-shoot threes (+7.3%), Popovich on corner threes.

Limitations. Lifers whose star players carry an outsized share of team identity, Popovich with Duncan/Ginobili/Parker/Kawhi, Kerr with Curry/Klay/Draymond, have a partial-credit ceiling. If a player's box-score-prior RAPM already accounts for their team's success, there's less residual variance left for the coach coefficient to absorb. This is a fundamental identifiability limit of the two-step framing, not a bug.

How to read the numbers

RAPM-style numbers are in points per 100 possessions. A +5 RAPM player is roughly +5 net points per 100 possessions compared to a league-average player on a neutral team. League scoring runs ~115 per 100, so +5 is genuinely elite (top-10ish in any given season). For defense, more negative is better: a −3 DRAPM player concedes 3 fewer points per 100 than average. QR-DQN skill advantage is in different units (points per event, not per 100 possessions). Play Types percentages are shares of a player's offensive actions; defensive rates are per 100 defensive possessions they were on the floor for.

Data & pipeline

NBA play-by-play covering the 2000-01 through 2025-26 regular seasons, 5.49M possessions across 30,815 games, parsed into per-event tuples with the 5-on-5 lineup, action, actor, and reward. Models implemented in scikit-learn (ridge with box-score prior, K-fold CV), PyTorch (Q-network with lineup attention, distributional QR-DQN, lineup GNN, causal transformer for V-Rating, Kalman filter for career trajectories), LightGBM (win probability, expected FG%), and a closed-form Gaussian-Gaussian Bayesian solve for the RAPM prior step. Ratings regenerate end-to-end from raw play-by-play, no hand curation.

Evaluation

Empirical checks on the four headline player-impact ratings: how much do the numbers change from one season to the next, and how does that stability scale with sample size? A metric that flips wildly year over year isn't measuring the player, it's measuring noise. This page reports Pearson correlation between each player's value in season N and season N+1, broken down by how much on-court sample the player had.

Year-over-year stability

Pearson r between (player, season N) and (player, season N+1) values. NBA pairs span 2000-01 → 2001-02 through 2024-25 → 2025-26 (25 pair-years). Higher = more stable. Per-event ratings (Q-Rating, V-Rating) use event count as the sample-size proxy; per-possession ratings (RAPM, iRAPM) use possession count.

Scatter: season N vs season N+1

Each point is one player. X-axis: their metric value in season N. Y-axis: the same metric in season N+1. Perfect stability = points on the y=x line. Wider scatter = more year-to-year variance. Filter by sample bucket to see how stability scales.

Metric agreement matrix

Pairwise Pearson r between each metric's per-(player, season) values, computed on the intersection of players where both metrics have a value. Higher = the two metrics rank players similarly. Low correlations between metrics mean they're measuring genuinely different things.

Multi-season pooling reliability

Year-over-year Pearson r at three pooling depths: single-season, 2-year rolling mean, 3-year rolling mean. Pooling averages out season-level noise. The dramatic jump in the RAPM row (0.26 → 0.83 for NBA at 3-year) is why career-shrunk values on player profiles are more trustworthy than single-season ratings.

How long each metric takes to stabilize

Reliability curve for each metric: rolling window sorted by sample size, EMA-smoothed. Y-axis is Pearson r between (season N) and (season N+1) values within that window. Higher = more stable. The table below shows the sample size needed to hit r = 0.7 and r = 0.8. Note: iRAPM has a data floor of ≈500 events per player-season (below that the metric isn't computed), so its curve legitimately starts at N ≈ 500 while the other three extend down closer to zero.

Biggest year-over-year moves

Individual year-over-year pairs across all four metrics, sorted by biggest absolute change. Useful for seeing which players' ratings move the most and why (rookies, injury seasons, mid-season trades).
No game selected. Pick one from the strip above.