NBAValuation

Model Lab

Gradient boosting (selected on validation) · statistically tied with random forest

How well the salary model holds up on seasons it never saw. Its estimates miss actual pay by $3.82M on average, and 80% of salaries land inside the 80% range it states. Everything below is evidence for or against those two numbers.

How to read the Lab
MAE
Mean absolute error: the average dollar gap between the model's estimate and the actual salary. Lower is better.
Validation seasons
2017-18 → 2020-21, each predicted by a model trained only on earlier seasons. Used to tune the models and choose one.
Held-out seasons
2021-22: kept aside and scored once, after the model was chosen. The fairest test of accuracy.
OLS
Ordinary linear regression, the baseline the tree models are measured against. It gets the same features they do.
Ablation
Rerunning the whole evaluation with one group of features removed, to see how much of the model's accuracy depended on it.
R²
The share of the variation in salaries the model accounts for (1 would be perfect).
80% range · coverage
Every estimate carries a range built to contain 80% of actual salaries. Coverage is how often it really did, on seasons that played no part in setting the width.
SHAP
Splits one estimate into how much each input pushed it up or down, in dollars.
Held-out MAE
$3.82M 95% CI $3.37M – $4.28M
Lower error than OLS
17.3% 12.6% – 21.9%
Held-out R²
0.628
Coverage of the 80% range
80.4% held out

How the models were evaluated

This is a pricing model, not a forecast. It asks what the market paid for a season of production once that season was over, so a player's own stats are fair inputs. Every model gets the same 14 features, the same folds and its own tuned target transform, and every season is predicted by a model trained only on earlier seasons.

Hover a season to see what its model was trained on.

Target
Salary as a share of that season's cap (square root transform chosen by validation for the selected model). Errors are reported in 2021-22 cap dollars (share × $112.41M).
Eligible sample
Player-seasons with a matched annual salary above the partial-payment threshold, 200+ minutes: 3,055 modeled of 4,203 integrated.
Validation
Expanding-window walk-forward: 4 validation seasons (2017-18 → 2020-21), each predicted by a model trained on every earlier season back to 2014-15. Used for tuning and model selection.
Held-out test
2021-22: 383 player-seasons, scored once by a model trained on 2,672 development rows.
Uncertainty
2,000-draw bootstrap that resamples players (383 in the held-out seasons), preserving player-level dependence.

Four models, two stages

Validation selected gradient boosting before the final season was scored; it is statistically tied with random forest. It also has the lowest held-out MAE, which confirms the choice but played no part in it. All four families receive the same 14 inputs and temporal splits.

Validation (selection)mean of 4 seasons · ±1 SD$3M$3.5M$4M$4.5M$5M$5.5MOLS$4.42MOLS: validation MAE $4.42M (±1 SD across seasons $4.13M – $4.71M)Lasso$4.41MLasso: validation MAE $4.41M (±1 SD across seasons $4.13M – $4.70M)Random forest$3.76MRandom forest: validation MAE $3.76M (±1 SD across seasons $3.45M – $4.07M)Gradient boosting$3.76MGradient boosting: validation MAE $3.76M (±1 SD across seasons $3.44M – $4.07M)Held-out (evaluation)1 season · 95% bootstrap CI$3M$3.5M$4M$4.5M$5M$5.5M$4.62MOLS: held-out MAE $4.62M (95% CI $4.08M – $5.14M)$4.62MLasso: held-out MAE $4.62M (95% CI $4.08M – $5.14M)$3.92MRandom forest: held-out MAE $3.92M (95% CI $3.45M – $4.38M)$3.82MGradient boosting: held-out MAE $3.82M (95% CI $3.37M – $4.28M)
Model comparison: validation and held-out error
ModelTargetConfigs tunedValidation MAEHeld-out MAE (95% CI)Held-out RMSEHeld-out R²vs OLS
OLS (linear regression)√ cap share1$4.42M$4.62M $4.08M – $5.14M$7.00M0.486—
Lasso regression√ cap share5$4.41M$4.62M $4.08M – $5.14M$7.01M0.4860.0%
Random foresttied on validation√ cap share8$3.76M$3.92M $3.45M – $4.38M$6.08M0.61315.1%
Gradient boostingreference√ cap share8$3.76M$3.82M $3.37M – $4.28M$5.96M0.62817.3%

Selection history, in order

The order of events matters, so it is shown as it happened.

  1. Development

    Selected on validation

    Tuning and model choice used the walk-forward validation seasons (2017-18 → 2020-21).

    Gradient boosting had the lowest validation MAE ($3.76M vs random forest $3.76M) and was selected before the held-out seasons were scored.

  2. Held-out evaluation

    Scored once

    2021-22, never used for tuning or selection, were then scored with the chosen specification.

    Gradient boosting $3.82M, random forest $3.92M. These results were then inspected.

  3. Later audit
    Audit rule, added in the final review

    Statistically tied

    A player-level bootstrap on the validation seasons only found the model families statistically tied: the validation difference spans −$64K – $76K.

    The held-out difference (−$23K – $213K) is within noise too.

  4. Reference model

    Why gradient boosting stays

    Switching models because of held-out MAE would choose the model from the evaluation results.

    All four model families are reported.

The tie · Random forest MAE minus gradient boosting MAE, 95% player-bootstrap intervals

← random forest lower errorgradient boosting lower error →−$400K−$200K0+$200K+$400KValidationValidation: +$4K (95% CI −$64K to $76K)Held outHeld out: +$101K (95% CI −$23K to $213K)

Both intervals include zero: the two ensembles are tied.

Gradient boosting's validation score is the best of 8 tuned configurations against 8 for random forest, so the comparison includes tuning uncertainty; the small validation margin should not be read as a real difference.

Error season by season

Each season is scored using only earlier training seasons. Gradient boosting beats OLS in 5 of 5 out-of-sample seasons, with median gain 14.8%. The smallest gain is 13.6% in 2020-21.

Held out$3M$3.5M$4M$4.5M$5M17-1818-1919-2020-2121-22OLSLassoRandom forestGradient boosting

Not memorization

Players absent from training provide a separate check on generalization. The comparison below reports their errors alongside returning players.

Never seen in training

18.4%

lower MAE than OLS · 62 player-seasons · Gradient boosting $1.28M vs OLS $1.57M

Seen in training

17.2%

lower MAE than OLS · 321 player-seasons · Gradient boosting $4.31M vs OLS $5.21M

Robustness and what the model is really using

Each variant changes one feature group or sample rule, keeps the tuned settings fixed, and is evaluated on the final season. Removing prior-season performance or context raises error. Sample variants show how partial payments and playing-time filters affect the result.

MAE vs main specificationGain over OLS−10%−5%0+5%+10%+15%+20%+25%+30%Main specification (all features)—17.3%Main specification (all features): MAE $3.82M (OLS $4.62M), 2,672 training rowsWithout prior season (4 features)+5.1%17.0%Without prior season (4 features): MAE $4.02M (OLS $4.84M), 2,672 training rowsWithout context (3 features)+24.2%3.1%Without context (3 features): MAE $4.75M (OLS $4.90M), 2,672 training rowsNo playing-time filter (full-season salaries)−3.0%17.4%No playing-time filter (full-season salaries): MAE $3.71M (OLS $4.49M), 2,870 training rowsPartial-season payments kept−0.7%17.1%Partial-season payments kept: MAE $3.80M (OLS $4.58M), 2,733 training rowsSuspect release salaries kept−0.6%17.5%Suspect release salaries kept: MAE $3.80M (OLS $4.60M), 2,710 training rows

Sample variants change the held-out population as well as the training rows, so their MAE is not directly comparable with the main specification; the gain over OLS on the same rows is the like-for-like comparison. OLS on the main specification: $4.62M. Held-out coverage of the 80% range is on the Diagnostics tab.