NBAValuation

Model Lab

Gradient boosting (selected on validation) · statistically tied with random forest

How well the salary model holds up on seasons it never saw. Its estimates miss actual pay by $3.82M on average, and 80% of salaries land inside the 80% range it states. Everything below is evidence for or against those two numbers.

How to read the Lab
MAE
Mean absolute error: the average dollar gap between the model's estimate and the actual salary. Lower is better.
Validation seasons
2017-18 → 2020-21, each predicted by a model trained only on earlier seasons. Used to tune the models and choose one.
Held-out seasons
2021-22: kept aside and scored once, after the model was chosen. The fairest test of accuracy.
OLS
Ordinary linear regression, the baseline the tree models are measured against. It gets the same features they do.
Ablation
Rerunning the whole evaluation with one group of features removed, to see how much of the model's accuracy depended on it.
R²
The share of the variation in salaries the model accounts for (1 would be perfect).
80% range · coverage
Every estimate carries a range built to contain 80% of actual salaries. Coverage is how often it really did, on seasons that played no part in setting the width.
SHAP
Splits one estimate into how much each input pushed it up or down, in dollars.
Held-out MAE
$3.82M 95% CI $3.37M – $4.28M
Lower error than OLS
17.3% 12.6% – 21.9%
Held-out R²
0.628
Coverage of the 80% range
80.4% held out

What moves the estimate

The selected gradient boosting model is explained with SHAP values, exact in its square root space and allocated back to dollars. Each contribution shows what an input added to or subtracted from the average player's valuation, and the pieces add back to the estimate exactly (computed from the same fitted model that produced the valuation). Averaging the size of those pieces over the held-out seasons (2021-22), Years since NBA debut alone accounts for 33.7% of all the movement. Read these as patterns in how the league has paid people, not as causes. This build shows aggregate feature importance; individual feature-value plots are available in the full local build.

Where the model's attention goes

Context and prior season account for 70% of the total movement, in these explanations. Minutes, WAR and RAPTOR are correlated, so attribution can shift between them; these shares are not independent measures of basketball importance, and none is a wage premium. No verified age, draft or rookie-contract feature is available in this dataset.

  • Context3 features
  • Prior season4 features
  • PREDATOR2 features
  • Role1 feature
  • Wins above replacement2 features
  • Overall RAPTOR2 features