NBAValuation

Model Lab

Gradient boosting (selected on validation) · statistically tied with random forest

How well the salary model holds up on seasons it never saw. Its estimates miss actual pay by $3.82M on average, and 80% of salaries land inside the 80% range it states. Everything below is evidence for or against those two numbers.

How to read the Lab
MAE
Mean absolute error: the average dollar gap between the model's estimate and the actual salary. Lower is better.
Validation seasons
2017-18 → 2020-21, each predicted by a model trained only on earlier seasons. Used to tune the models and choose one.
Held-out seasons
2021-22: kept aside and scored once, after the model was chosen. The fairest test of accuracy.
OLS
Ordinary linear regression, the baseline the tree models are measured against. It gets the same features they do.
Ablation
Rerunning the whole evaluation with one group of features removed, to see how much of the model's accuracy depended on it.
R²
The share of the variation in salaries the model accounts for (1 would be perfect).
80% range · coverage
Every estimate carries a range built to contain 80% of actual salaries. Coverage is how often it really did, on seasons that played no part in setting the width.
SHAP
Splits one estimate into how much each input pushed it up or down, in dollars.
Held-out MAE
$3.82M 95% CI $3.37M – $4.28M
Lower error than OLS
17.3% 12.6% – 21.9%
Held-out R²
0.628
Coverage of the 80% range
80.4% held out

Actual vs predicted

The first thing to check is whether the model is right on average at every level of pay, not just overall. Each held-out player-season sits against the diagonal of perfect prediction, with every tenth of the prediction range summarized. The summaries bend away from the diagonal at both ends, which is the classic sign of a model hedging: it pulls cheap players up and expensive players down, toward the middle of the pay scale.

$1M$1M$5M$5M$10M$10M$20M$20M$40M$40M$60M$60MPredicted →↑ ActualDecile 1 (n = 39): mean prediction $1.78M, mean actual $1.95M; middle half of actual $1.4M – $2.2MDecile 2 (n = 38): mean prediction $2.27M, mean actual $2.76M; middle half of actual $1.6M – $3.8MDecile 3 (n = 38): mean prediction $2.86M, mean actual $3.12M; middle half of actual $1.8M – $4.2MDecile 4 (n = 38): mean prediction $3.64M, mean actual $4.60M; middle half of actual $2.0M – $5.7MDecile 5 (n = 39): mean prediction $4.41M, mean actual $4.07M; middle half of actual $1.7M – $4.4MDecile 6 (n = 38): mean prediction $5.55M, mean actual $6.04M; middle half of actual $1.7M – $8.7MDecile 7 (n = 38): mean prediction $7.55M, mean actual $10.33M; middle half of actual $4.9M – $13.6MDecile 8 (n = 38): mean prediction $10.45M, mean actual $12.26M; middle half of actual $7.1M – $16.0MDecile 9 (n = 38): mean prediction $14.67M, mean actual $16.89M; middle half of actual $9.8M – $22.6MDecile 10 (n = 39): mean prediction $21.90M, mean actual $28.23M; middle half of actual $20.0M – $35.4M
  • 0 held-out player-seasons (Gradient boosting)
  • Mean actual per decile of Gradient boosting's prediction, bar = middle half

Residual distribution

How the misses are spread: actual minus predicted on the held-out seasons, on one shared axis for every model. Gradient boosting lands almost on top of the typical player (−$8K median error) while running +$1.52M low on average. Both are true at once because NBA pay is heavily right-skewed and the model predicts in log space, then exponentiates — which lands near the median rather than the mean (a retransformation effect; a smearing correction would shift it).

0%10%20%30%40%−$30M−$20M−$10M0+$10M+$20M+$30M+$40M−$30.0M to −$27.5M: 0 player-seasons (0.0%)−$27.5M to −$25.0M: 0 player-seasons (0.0%)−$25.0M to −$22.5M: 0 player-seasons (0.0%)−$22.5M to −$20.0M: 0 player-seasons (0.0%)−$20.0M to −$17.5M: 0 player-seasons (0.0%)−$17.5M to −$15.0M: 0 player-seasons (0.0%)−$15.0M to −$12.5M: 0 player-seasons (0.0%)−$12.5M to −$10.0M: 2 player-seasons (0.5%)−$10.0M to −$7.5M: 5 player-seasons (1.3%)−$7.5M to −$5.0M: 14 player-seasons (3.7%)−$5.0M to −$2.5M: 45 player-seasons (11.7%)−$2.5M to $0: 126 player-seasons (32.9%)$0 to +$2.5M: 75 player-seasons (19.6%)+$2.5M to +$5.0M: 44 player-seasons (11.5%)+$5.0M to +$7.5M: 28 player-seasons (7.3%)+$7.5M to +$10.0M: 13 player-seasons (3.4%)+$10.0M to +$12.5M: 9 player-seasons (2.3%)+$12.5M to +$15.0M: 5 player-seasons (1.3%)+$15.0M to +$17.5M: 7 player-seasons (1.8%)+$17.5M to +$20.0M: 3 player-seasons (0.8%)+$20.0M to +$22.5M: 4 player-seasons (1.0%)+$22.5M to +$25.0M: 2 player-seasons (0.5%)+$25.0M to +$27.5M: 0 player-seasons (0.0%)+$27.5M to +$30.0M: 0 player-seasons (0.0%)+$30.0M to +$32.5M: 1 player-seasons (0.3%)+$32.5M to +$35.0M: 0 player-seasons (0.0%)+$35.0M to +$37.5M: 0 player-seasons (0.0%)+$37.5M to +$40.0M: 0 player-seasons (0.0%)← paid less than predictedpaid more than predicted →

Gradient boosting: median residual −$8K · 52% within ±$2.5M · 76% within ±$5M · middle 80% from −$3.7M to +$9.0M

Where the model errs

Errors are not spread evenly across the pay scale, and where they cluster says something about the league. Over all out-of-sample seasons, minimum-level contracts are paid $1.76M less than gradient boosting predicts and max-level contracts $14.78M more. These groups use salary thresholds, not verified contract types. The differences describe residual patterns; they do not establish that contract rules caused the errors.

← paid less than predictedpaid more →n−$20M−$10M0+$10M+$20M<2% (min-level)<2% (min-level): paid $1.76M less than predicted on average (median −$1.13M, MAE $1.79M, n = 555)−$1.8M5552–5%2–5%: paid $704K less than predicted on average (median −$140K, MAE $1.73M, n = 474)−$704K4745–10%5–10%: paid $905K more than predicted on average (median +$1.70M, MAE $3.52M, n = 338)+$905K33810–20%10–20%: paid $4.36M more than predicted on average (median +$4.73M, MAE $5.58M, n = 325)+$4.4M32520–30%20–30%: paid $9.89M more than predicted on average (median +$8.69M, MAE $9.95M, n = 157)+$9.9M15730%+ (max-level)30%+ (max-level): paid $14.78M more than predicted on average (median +$13.65M, MAE $14.78M, n = 54)+$14.8M54

Error by subgroup, every model

A single headline number can hide a model that is much better in some places and worse in others, so the same comparison is repeated inside each group. The gain is not uniform. On the held-out seasons gradient boosting is worse than OLS in the 5–10% tier (−11.3%, n = 67). These subgroup results qualify the overall average.

MAEvs OLS$0$5M$10M$15M$20M$25M<2% (min-level)<2% (min-level) · Lasso: MAE $2.46M (n = 114)<2% (min-level) · OLS: MAE $2.46M (n = 114)<2% (min-level) · Random forest: MAE $2.03M (n = 114)<2% (min-level) · Gradient boosting: MAE $1.87M (n = 114)+24.0%<2% (min-level): n = 1142–5%2–5% · Lasso: MAE $2.16M (n = 98)2–5% · OLS: MAE $2.16M (n = 98)2–5% · Random forest: MAE $1.57M (n = 98)2–5% · Gradient boosting: MAE $1.63M (n = 98)+24.7%2–5%: n = 985–10%5–10% · Lasso: MAE $3.16M (n = 67)5–10% · OLS: MAE $3.16M (n = 67)5–10% · Random forest: MAE $3.62M (n = 67)5–10% · Gradient boosting: MAE $3.52M (n = 67)−11.3%5–10%: n = 6710–20%10–20% · Lasso: MAE $6.39M (n = 63)10–20% · OLS: MAE $6.41M (n = 63)10–20% · Random forest: MAE $5.44M (n = 63)10–20% · Gradient boosting: MAE $5.37M (n = 63)+16.1%10–20%: n = 6320–30%*20–30% · Lasso: MAE $13.71M (n = 26)20–30% · OLS: MAE $13.70M (n = 26)20–30% · Random forest: MAE $10.85M (n = 26)20–30% · Gradient boosting: MAE $10.29M (n = 26)+24.8%20–30%: n = 2630%+ (max-level)*30%+ (max-level) · Lasso: MAE $20.44M (n = 15)30%+ (max-level) · OLS: MAE $20.41M (n = 15)30%+ (max-level) · Random forest: MAE $16.71M (n = 15)30%+ (max-level) · Gradient boosting: MAE $16.65M (n = 15)+18.4%30%+ (max-level): n = 15
  • Lasso
  • OLS
  • Random forest
  • Gradient boosting
  • * fewer than 30 player-seasons

Prediction intervals

A single number is a false promise, so every valuation ships with a range. The range is meant to contain the real salary 80% of the time, and on the held-out seasons it contained it 80.4% of the time: close to the promise, and slightly on the safe side. The width comes from split-conformal prediction, where the size of the model's past misses on the 1,520 walk-forward validation predictions sets one width in log salary-share space, which widens into bigger dollar ranges for expensive players. The held-out seasons played no part in setting it. The typical range is $8.92M wide, which is wide enough to be a real statement of doubt rather than decoration.

50%60%70%80%90%100%80% target<2% (min-level)<2% (min-level): 83.3% of 114 held-out actual salaries inside their range83% · 1142–5%2–5%: 94.9% of 98 held-out actual salaries inside their range95% · 985–10%5–10%: 85.1% of 67 held-out actual salaries inside their range85% · 6710–20%10–20%: 73.0% of 63 held-out actual salaries inside their range73% · 6320–30%*20–30%: 53.8% of 26 held-out actual salaries inside their range54% · 2630%+ (max-level)*30%+ (max-level): 20.0% of 15 held-out actual salaries inside their range20% · 15

Right column: coverage · held-out player-seasons. * fewer than 30 player-seasons. Orange: the range is too narrow for the group; teal: wider than needed.

Coverage by season

Held outCalibration (in-sample)60%70%80%90%100%’18’19’20’21’222017-18: 81.5% of 372 inside the range (calibration season, in-sample); median width $9.1M2018-19: 77.2% of 381 inside the range (calibration season, in-sample); median width $9.5M2019-20: 81.8% of 384 inside the range (calibration season, in-sample); median width $8.9M2020-21: 80.2% of 383 inside the range (calibration season, in-sample); median width $8.9M2021-22: 80.4% of 383 inside the range (held out); median width $8.9M

Nominal vs held-out coverage

50%50%60%60%70%70%80%80%90%90%100%100%Nominal →↑ Held-out50% nominal → 50.9% held-out coverage; median width $4.4M60% nominal → 59.5% held-out coverage; median width $5.6M70% nominal → 70.8% held-out coverage; median width $7.1M80% nominal → 80.4% held-out coverage; median width $8.9M90% nominal → 89.3% held-out coverage; median width $11.8M95% nominal → 94.5% held-out coverage; median width $15.0M

Stress tests

Four attempts to break the evidence rather than confirm it: whether the range would have held if it had been run live season by season, whether one width per position group would work better, whether the held-out seasons are too different from the training years to trust, and whether the model could be learning salary from somewhere other than its features.

Would the range have held in real time?

The saved ranges were calibrated once, at the end. Rebuilding them season by season, each using only the seasons before it, the 80% range would have covered 76%–83% of salaries as it went. No season fell far below target, so the promise would have held in real time too.

70%80%90%’19’20’21’2218-19: 76.1% of 381 (calibrated on 372 earlier rows)19-20: 83.1% of 384 (calibrated on 753 earlier rows)20-21: 80.2% of 383 (calibrated on 1137 earlier rows)21-22: 80.4% of 383 (held out, saved interval)

validation season, calibrated on earlier seasons held out

Would one width per position group be better?

Giving each position group its own width (group-conditional, or Mondrian, conformal) and scoring it on the held-out seasons: the worst group misses 80% by 6 points instead of 4, and C ranges shrink from $10.35M to $10.89M. Overall coverage is 81%. The group-specific variant was evaluated but does not replace the saved ranges.

PositionSavedPer groupWidth, saved → per group
C81%86%$10.3M → $10.9M
PF80%80%$9.4M → $8.9M
PG76%79%$9.3M → $9.8M
SF84%84%$8.4M → $8.7M
SG81%78%$8.5M → $7.8M

The held-out seasons look different

The held-out season are not a random sample of the data; they are the most recent ones, in a league that keeps changing. They are in fact easy to tell apart from the training years: a classifier trained to spot them succeeds with AUC 0.58. The shifts below show which inputs changed. Mean error changes by +$230K per season (SE $122K); neither this trend nor the classifier establishes future performance.

  • Prior-season minutes−0.17 SD
  • Prior-season WAR−0.08 SD
  • PREDATOR defense+0.06 SD
  • Pace impact+0.06 SD
  • RAPTOR defense+0.06 SD
  • PREDATOR offense+0.05 SD

Could it be learning salary some other way?

If the pipeline leaked pay through some back door, a model trained on scrambled salaries would still score well. Refit on development seasons with salaries shuffled within each season, its held-out error collapses to roughly what guessing one number for everybody achieves. This is consistent with a loss of predictive signal, though it cannot prove the absence of every form of leakage.

  • Real salaries$3.82M
  • Shuffled within season (3×)$6.93M
  • Constant prediction$6.64M

Held-out MAE

Leakage audit

The most common way a model like this gets flattering results is by seeing something it should not. These checks run in code (src/models/audit.py) rather than being asserted in prose: no feature uses salary, prior-season features really are season t − 1, the evaluation model never touched the held-out seasons, and refitting reproduces the saved predictions exactly.

12 passed — show all 12 checks

Leakage

PASSno target or salary field appears in the feature registry
offenders: none
PASSno feature is a near-copy of the target (|rho| < 0.95)
strongest: prev_mp rho=0.631

identity

PASSone row per player-season
duplicates: 0

target

PASSsalary target is positive and finite
range: $268,959-$45,780,966
PASSno suspect release salary remains in the modeled sample
suspect rows modeled: 0

temporal

PASSprior-season features equal the raw source's season t-1
rows checked: 3055; rows with a prior season: 2644
PASSevery validation fold trains only on earlier seasons
<=2017 -> 2018; <=2018 -> 2019; <=2019 -> 2020; <=2020 -> 2021
PASSheld-out season is absent from validation predictions
test seasons: [2022]

features

PASSexperience equals years since debut in historical RAPTOR
range: 0-21
PASSno retained pair has |r| > 0.99 (development seasons)
max |r| 0.972 (predator_offense / raptor_offense)
PASSno retained feature has R2 >= 0.999 on the others (development seasons)
max R2 0.9540 (predator_offense); rank 13/13

predictions

PASSall model families score the same out-of-sample rows
{'test': {'gradient_boosting': 383, 'lasso': 383, 'ols': 383, 'random_forest': 383}, 'validation': {'gradient_boosting': 1520, 'lasso': 1520, 'ols': 1520, 'random_forest': 1520}}