Model Lab
Gradient boosting (selected on validation) · statistically tied with random forest
How well the salary model holds up on seasons it never saw. Its estimates miss actual pay by $3.82M on average, and 80% of salaries land inside the 80% range it states. Everything below is evidence for or against those two numbers.
How to read the Lab
- MAE
- Mean absolute error: the average dollar gap between the model's estimate and the actual salary. Lower is better.
- Validation seasons
- 2017-18 → 2020-21, each predicted by a model trained only on earlier seasons. Used to tune the models and choose one.
- Held-out seasons
- 2021-22: kept aside and scored once, after the model was chosen. The fairest test of accuracy.
- OLS
- Ordinary linear regression, the baseline the tree models are measured against. It gets the same features they do.
- Ablation
- Rerunning the whole evaluation with one group of features removed, to see how much of the model's accuracy depended on it.
- R²
- The share of the variation in salaries the model accounts for (1 would be perfect).
- 80% range · coverage
- Every estimate carries a range built to contain 80% of actual salaries. Coverage is how often it really did, on seasons that played no part in setting the width.
- SHAP
- Splits one estimate into how much each input pushed it up or down, in dollars.
- Held-out MAE
- $3.82M 95% CI $3.37M – $4.28M
- Lower error than OLS
- 17.3% 12.6% – 21.9%
- Held-out R²
- 0.628
- Coverage of the 80% range
- 80.4% held out
Does the salary benchmark say anything about later pay?
Reading under rules fixed before the estimates: The benchmark is associated with later pay at the same current salary, with little out-of-sample forecast gain.
For every modeled player-season from 2017-18 through 2021-22, the salary one and two seasons later is looked up by exact season in the salary records, including players who no longer meet the playing-time rule. The benchmark is the model's estimate for that season from a model whose fitted parameters never saw that season's salaries. Its settings were chosen on validation seasons that include them; re-selecting the settings with earlier seasons only gives a coefficient of 0.47.
A later salary was found for 1,651 of 1,903 origins one season ahead and 1,443 two seasons ahead. The rest are unknown, not zero: the salary workbook lists about 450 players a season, fewer than are paid. Every estimate below is conditional on a later salary being observed.
At the same current salary, a higher benchmark goes with higher later pay
One season ahead, each extra point of cap in the benchmark goes with 0.45 more points of cap the next season on average (95% CI 0.37 to 0.55; n = 1,651), holding current cap share, career stage and season fixed.
Per point of benchmark the coefficient is 1.78 for players 0–3 years from their debut against 0.40 and 0.34 for 4–7 and 8+ years. Early-career benchmarks vary much less (standard deviation 1.2 points against 3.8 and 4.3), so per standard deviation the contrast is smaller: 2.15 against 1.52 and 1.44 points of later cap share (post hoc). On both scales the association is largest early in careers. Two seasons ahead the overall coefficient is 0.86 (95% CI 0.72 to 1.01). Years since debut are not contract status: rookie-scale years, options and extensions are not in the public data.
Why current salary must be held fixed: growth and the gap (benchmark minus salary) both contain current salary. Regressing growth on the gap alone gives 0.27, which also carries current pay's own relationship with growth. With current salary controlled, the gap coefficient equals the one above. It measures extra information in the benchmark, not a share of any “mispricing” that later closes.
How sturdy the one-season coefficient is
The sign and an interval excluding zero survive every pre-specified check listed below. It is an average that leans on large pay changes. Dropping the 111 most influential rows moves it to 0.28, and the median response (a post hoc median regression) is 0.09 (95% CI 0.06 to 0.13). For the median player the association is much weaker, though still above zero; most of the average comes from the minority whose pay changes a lot. Refitting the benchmark inside the bootstrap gives 0.34 to 0.55.
It adds little to a one-season forecast
Adding the benchmark to a regression on current salary and career stage moves mean absolute error from 2.66 to 2.65 points of cap (difference −0.003, 95% CI −0.065 to 0.061). Repeating this season's cap share scores 2.33 on the same metric.
Each forecast is made at the end of the origin season and trained only on pairs whose later salary was already known then. All four are scored on the same 1,325 player-seasons (2018-19 to 2021-22). The regressions are fitted by least squares, which targets the mean, while mean absolute error rewards the median. Refitted by median regression (post hoc) they score 2.35 (salary and stage), 2.31 (with the benchmark), 2.30 (with minutes and WAR), at or below repetition, so repetition's edge largely reflects the loss function. In that refit the benchmark lowers mean absolute error by 0.035 points (95% CI 0.022 to 0.049), a gain that is consistent but small. On root mean squared error, which weights large misses, the benchmark forecast scores 4.43 against 4.68 without it (a supplementary comparison added after the first run; 95% CI of the difference −0.33 to −0.15).
One season ahead
Two seasons ahead
Two seasons ahead, adding the benchmark to the salary-and-stage regression changes mean absolute error by −0.31 points (95% CI −0.51 to −0.09); against simply repeating current pay the difference is −0.15 (95% CI −0.42 to 0.13). Adding minutes and WAR instead of the benchmark changes it by −0.51 (95% CI −0.68 to −0.34).
Pay and production by years since debut
Among the most productive third of players each season, those 0–3 years from their debut averaged 4.5% of the cap; those 8+ years in averaged 17.9%.
- Bottom third of WAR (within season)
- Middle third of WAR (within season)
- Top third of WAR (within season)
Within a season, each win above replacement goes with 0.37 points of cap for players 0–3 years in (95% CI 0.28 to 0.47) and 1.52 for 8+ years (95% CI 1.36 to 1.67). This is descriptive: consistent with rookie-scale and restricted free agency rules, not proof of them. Because the salary model uses years since debut, it prices early-career players at early-career rates.
Where the 80% range fails: expensive players
On the held-out season the range held 14 of 26 salaries at 20–30% of the cap and 3 of 15 at 30% or more. In both tiers every miss was a salary above the range.
| Actual salary tier (% of cap) | Inside the saved range | Above the range | Per-quartile variant |
|---|---|---|---|
| <2% (min-level) | 95 of 114 | 0 | 91 of 114 |
| 2-5% | 93 of 98 | 0 | 82 of 98 |
| 5-10% | 57 of 67 | 8 | 50 of 67 |
| 10-20% | 46 of 63 | 17 | 48 of 63 |
| 20-30% | 14 of 26 | 12 | 17 of 26 |
| 30%+ (max-level) | 3 of 15 | 12 | 5 of 15 |
A variant with one width per quartile of the predicted salary (known at prediction time, calibrated on the validation seasons only) gives 17 and 5 in those tiers, with overall coverage 76.5%. The saved ranges are unchanged. That maximum contracts explain the misses is a hypothesis, not a fix. Actual-salary tiers are a diagnostic, never a calibration group.
What this does not show
- No causal effect: nothing here identifies how pay responds to production, or whether any player was mispriced.
- Selection: later salaries are observed for 87% of origins one season ahead; players who leave the data are missing, not zero.
- The prospective forecast uses a benchmark re-selected each year with earlier seasons only; the association estimates use the site's benchmark, whose settings were tuned on seasons that overlap later outcomes.
- Contract status, options, waivers and dead money are not observed. See Methodology & Data.