Model Lab
Gradient boosting (selected on validation) · statistically tied with random forest
How well the salary model holds up on seasons it never saw. Its estimates miss actual pay by $3.82M on average, and 80% of salaries land inside the 80% range it states. Everything below is evidence for or against those two numbers.
How to read the Lab
- MAE
- Mean absolute error: the average dollar gap between the model's estimate and the actual salary. Lower is better.
- Validation seasons
- 2017-18 → 2020-21, each predicted by a model trained only on earlier seasons. Used to tune the models and choose one.
- Held-out seasons
- 2021-22: kept aside and scored once, after the model was chosen. The fairest test of accuracy.
- OLS
- Ordinary linear regression, the baseline the tree models are measured against. It gets the same features they do.
- Ablation
- Rerunning the whole evaluation with one group of features removed, to see how much of the model's accuracy depended on it.
- R²
- The share of the variation in salaries the model accounts for (1 would be perfect).
- 80% range · coverage
- Every estimate carries a range built to contain 80% of actual salaries. Coverage is how often it really did, on seasons that played no part in setting the width.
- SHAP
- Splits one estimate into how much each input pushed it up or down, in dollars.
- Held-out MAE
- $3.82M 95% CI $3.37M – $4.28M
- Lower error than OLS
- 17.3% 12.6% – 21.9%
- Held-out R²
- 0.628
- Coverage of the 80% range
- 80.4% held out
How the models were evaluated
This is a pricing model, not a forecast. It asks what the market paid for a season of production once that season was over, so a player's own stats are fair inputs. Every model gets the same 14 features, the same folds and its own tuned target transform, and every season is predicted by a model trained only on earlier seasons.
Hover a season to see what its model was trained on.
- Target
- Salary as a share of that season's cap (square root transform chosen by validation for the selected model). Errors are reported in 2021-22 cap dollars (share × $112.41M).
- Eligible sample
- Player-seasons with a matched annual salary above the partial-payment threshold, 200+ minutes: 3,055 modeled of 4,203 integrated.
- Validation
- Expanding-window walk-forward: 4 validation seasons (2017-18 → 2020-21), each predicted by a model trained on every earlier season back to 2014-15. Used for tuning and model selection.
- Held-out test
- 2021-22: 383 player-seasons, scored once by a model trained on 2,672 development rows.
- Uncertainty
- 2,000-draw bootstrap that resamples players (383 in the held-out seasons), preserving player-level dependence.
Four models, two stages
Validation selected gradient boosting before the final season was scored; it is statistically tied with random forest. It also has the lowest held-out MAE, which confirms the choice but played no part in it. All four families receive the same 14 inputs and temporal splits.
| Model | Target | Configs tuned | Validation MAE | Held-out MAE (95% CI) | Held-out RMSE | Held-out R² | vs OLS |
|---|---|---|---|---|---|---|---|
| OLS (linear regression) | √ cap share | 1 | $4.42M | $4.62M $4.08M – $5.14M | $7.00M | 0.486 | — |
| Lasso regression | √ cap share | 5 | $4.41M | $4.62M $4.08M – $5.14M | $7.01M | 0.486 | 0.0% |
| Random foresttied on validation | √ cap share | 8 | $3.76M | $3.92M $3.45M – $4.38M | $6.08M | 0.613 | 15.1% |
| Gradient boostingreference | √ cap share | 8 | $3.76M | $3.82M $3.37M – $4.28M | $5.96M | 0.628 | 17.3% |
Selection history, in order
The order of events matters, so it is shown as it happened.
- Development
Selected on validation
Tuning and model choice used the walk-forward validation seasons (2017-18 → 2020-21).
Gradient boosting had the lowest validation MAE ($3.76M vs random forest $3.76M) and was selected before the held-out seasons were scored.
- Held-out evaluation
Scored once
2021-22, never used for tuning or selection, were then scored with the chosen specification.
Gradient boosting $3.82M, random forest $3.92M. These results were then inspected.
- Later auditAudit rule, added in the final review
Statistically tied
A player-level bootstrap on the validation seasons only found the model families statistically tied: the validation difference spans −$64K – $76K.
The held-out difference (−$23K – $213K) is within noise too.
- Reference model
Why gradient boosting stays
Switching models because of held-out MAE would choose the model from the evaluation results.
All four model families are reported.
The tie · Random forest MAE minus gradient boosting MAE, 95% player-bootstrap intervals
Both intervals include zero: the two ensembles are tied.
Gradient boosting's validation score is the best of 8 tuned configurations against 8 for random forest, so the comparison includes tuning uncertainty; the small validation margin should not be read as a real difference.
Error season by season
Each season is scored using only earlier training seasons. Gradient boosting beats OLS in 5 of 5 out-of-sample seasons, with median gain 14.8%. The smallest gain is 13.6% in 2020-21.
Not memorization
Players absent from training provide a separate check on generalization. The comparison below reports their errors alongside returning players.
Never seen in training
18.4%lower MAE than OLS · 62 player-seasons · Gradient boosting $1.28M vs OLS $1.57M
Seen in training
17.2%lower MAE than OLS · 321 player-seasons · Gradient boosting $4.31M vs OLS $5.21M
Robustness and what the model is really using
Each variant changes one feature group or sample rule, keeps the tuned settings fixed, and is evaluated on the final season. Removing prior-season performance or context raises error. Sample variants show how partial payments and playing-time filters affect the result.
Sample variants change the held-out population as well as the training rows, so their MAE is not directly comparable with the main specification; the gain over OLS on the same rows is the like-for-like comparison. OLS on the main specification: $4.62M. Held-out coverage of the 80% range is on the Diagnostics tab.