Model Lab
Gradient boosting (selected on validation) · statistically tied with random forest
How well the salary model holds up on seasons it never saw. Its estimates miss actual pay by $3.82M on average, and 80% of salaries land inside the 80% range it states. Everything below is evidence for or against those two numbers.
How to read the Lab
- MAE
- Mean absolute error: the average dollar gap between the model's estimate and the actual salary. Lower is better.
- Validation seasons
- 2017-18 → 2020-21, each predicted by a model trained only on earlier seasons. Used to tune the models and choose one.
- Held-out seasons
- 2021-22: kept aside and scored once, after the model was chosen. The fairest test of accuracy.
- OLS
- Ordinary linear regression, the baseline the tree models are measured against. It gets the same features they do.
- Ablation
- Rerunning the whole evaluation with one group of features removed, to see how much of the model's accuracy depended on it.
- R²
- The share of the variation in salaries the model accounts for (1 would be perfect).
- 80% range · coverage
- Every estimate carries a range built to contain 80% of actual salaries. Coverage is how often it really did, on seasons that played no part in setting the width.
- SHAP
- Splits one estimate into how much each input pushed it up or down, in dollars.
- Held-out MAE
- $3.82M 95% CI $3.37M – $4.28M
- Lower error than OLS
- 17.3% 12.6% – 21.9%
- Held-out R²
- 0.628
- Coverage of the 80% range
- 80.4% held out
Actual vs predicted
The first thing to check is whether the model is right on average at every level of pay, not just overall. Each held-out player-season sits against the diagonal of perfect prediction, with every tenth of the prediction range summarized. The summaries bend away from the diagonal at both ends, which is the classic sign of a model hedging: it pulls cheap players up and expensive players down, toward the middle of the pay scale.
- 0 held-out player-seasons (Gradient boosting)
- Mean actual per decile of Gradient boosting's prediction, bar = middle half
Residual distribution
How the misses are spread: actual minus predicted on the held-out seasons, on one shared axis for every model. Gradient boosting lands almost on top of the typical player (−$8K median error) while running +$1.52M low on average. Both are true at once because NBA pay is heavily right-skewed and the model predicts in log space, then exponentiates — which lands near the median rather than the mean (a retransformation effect; a smearing correction would shift it).
Gradient boosting: median residual −$8K · 52% within ±$2.5M · 76% within ±$5M · middle 80% from −$3.7M to +$9.0M
Where the model errs
Errors are not spread evenly across the pay scale, and where they cluster says something about the league. Over all out-of-sample seasons, minimum-level contracts are paid $1.76M less than gradient boosting predicts and max-level contracts $14.78M more. These groups use salary thresholds, not verified contract types. The differences describe residual patterns; they do not establish that contract rules caused the errors.
Error by subgroup, every model
A single headline number can hide a model that is much better in some places and worse in others, so the same comparison is repeated inside each group. The gain is not uniform. On the held-out seasons gradient boosting is worse than OLS in the 5–10% tier (−11.3%, n = 67). These subgroup results qualify the overall average.
- Lasso
- OLS
- Random forest
- Gradient boosting
- * fewer than 30 player-seasons
Prediction intervals
A single number is a false promise, so every valuation ships with a range. The range is meant to contain the real salary 80% of the time, and on the held-out seasons it contained it 80.4% of the time: close to the promise, and slightly on the safe side. The width comes from split-conformal prediction, where the size of the model's past misses on the 1,520 walk-forward validation predictions sets one width in log salary-share space, which widens into bigger dollar ranges for expensive players. The held-out seasons played no part in setting it. The typical range is $8.92M wide, which is wide enough to be a real statement of doubt rather than decoration.
Right column: coverage · held-out player-seasons. * fewer than 30 player-seasons. Orange: the range is too narrow for the group; teal: wider than needed.
Coverage by season
Nominal vs held-out coverage
Stress tests
Four attempts to break the evidence rather than confirm it: whether the range would have held if it had been run live season by season, whether one width per position group would work better, whether the held-out seasons are too different from the training years to trust, and whether the model could be learning salary from somewhere other than its features.
Would the range have held in real time?
The saved ranges were calibrated once, at the end. Rebuilding them season by season, each using only the seasons before it, the 80% range would have covered 76%–83% of salaries as it went. No season fell far below target, so the promise would have held in real time too.
validation season, calibrated on earlier seasons held out
Would one width per position group be better?
Giving each position group its own width (group-conditional, or Mondrian, conformal) and scoring it on the held-out seasons: the worst group misses 80% by 6 points instead of 4, and C ranges shrink from $10.35M to $10.89M. Overall coverage is 81%. The group-specific variant was evaluated but does not replace the saved ranges.
| Position | Saved | Per group | Width, saved → per group |
|---|---|---|---|
| C | 81% | 86% | $10.3M → $10.9M |
| PF | 80% | 80% | $9.4M → $8.9M |
| PG | 76% | 79% | $9.3M → $9.8M |
| SF | 84% | 84% | $8.4M → $8.7M |
| SG | 81% | 78% | $8.5M → $7.8M |
The held-out seasons look different
The held-out season are not a random sample of the data; they are the most recent ones, in a league that keeps changing. They are in fact easy to tell apart from the training years: a classifier trained to spot them succeeds with AUC 0.58. The shifts below show which inputs changed. Mean error changes by +$230K per season (SE $122K); neither this trend nor the classifier establishes future performance.
- Prior-season minutes−0.17 SD
- Prior-season WAR−0.08 SD
- PREDATOR defense+0.06 SD
- Pace impact+0.06 SD
- RAPTOR defense+0.06 SD
- PREDATOR offense+0.05 SD
Could it be learning salary some other way?
If the pipeline leaked pay through some back door, a model trained on scrambled salaries would still score well. Refit on development seasons with salaries shuffled within each season, its held-out error collapses to roughly what guessing one number for everybody achieves. This is consistent with a loss of predictive signal, though it cannot prove the absence of every form of leakage.
Held-out MAE
Leakage audit
The most common way a model like this gets flattering results is by seeing something it should not. These checks run in code (src/models/audit.py) rather than being asserted in prose: no feature uses salary, prior-season features really are season t − 1, the evaluation model never touched the held-out seasons, and refitting reproduces the saved predictions exactly.
12 passed — show all 12 checks
Leakage
offenders: none
strongest: prev_mp rho=0.631
identity
duplicates: 0
target
range: $268,959-$45,780,966
suspect rows modeled: 0
temporal
rows checked: 3055; rows with a prior season: 2644
<=2017 -> 2018; <=2018 -> 2019; <=2019 -> 2020; <=2020 -> 2021
test seasons: [2022]
features
range: 0-21
max |r| 0.972 (predator_offense / raptor_offense)
max R2 0.9540 (predator_offense); rank 13/13
predictions
{'test': {'gradient_boosting': 383, 'lasso': 383, 'ols': 383, 'random_forest': 383}, 'validation': {'gradient_boosting': 1520, 'lasso': 1520, 'ols': 1520, 'random_forest': 1520}}