Methodology & Data
How the benchmark and the later-pay analysis are built, what the data cover, and where both break down.
How to read a valuation
A valuation has three parts: the salary the player actually earned, the model's estimate of what NBA teams have paid for production like his, and an 80% range around that estimate showing how sure the model is. The label compares the salary with the range. It is a statement about the market's pricing, not about how good the player is or what a team should have offered him.
- Model-estimated underpayment
- Actual salary below the 80% range: the production is consistent with higher pay than was received.
- In line with model
- Actual salary inside the range. Most player-seasons land here (1,526 of 1,903).
- Model-estimated overpayment
- Actual salary above the range: pay exceeds what the model associates with the production.
A label is a prompt to investigate — contract timing, injury, role, the pay scale's floor and ceiling — not a finding about the player.
The question
NBA pay is not a free market. A salary cap, a rookie scale, minimum deals and maximum deals bound what anyone can be paid, and a player's price often has more to do with where the player sits in those rules than with last season's box score. The project asks how much of a salary measurable on-court production can explain, and whose pay sits furthest from it. Contracts are signed before the seasons they pay for, so this describes how production and pay line up after the fact; it does not predict what anyone will sign for next.
Data and eligibility
Mendeley annual athlete salaries are joined by normalized player name and season to FiveThirtyEight RAPTOR. A small reviewed alias list resolves common spelling variants. 98.5% of salary rows match.
3,055 player-seasons cover 2014-15 through 2021-22. The sample requires a matching positive salary, at least 200 minutes, and pay of at least half the season's rookie minimum. A frozen rule also drops multi-team seasons paid under half the prior season's salary, likely a released player's final-team pay. Neither rule establishes contract type.
Traded players have one total-season performance row and one published annual salary. 1,903 out-of-sample valuations cover 645 players from 2017-18 onward. Earlier seasons train the first model.
Later salaries
For the later-pay analysis every modeled player-season from 2017-18 on is linked to the same player's salary one, two and three seasons later, through 2023-24. The link joins on the exact target season in SQL. A “next record” is not used, because a player who skips a season would be matched to the wrong year. Outcomes come from all salary records, so a player need not meet the playing-time rule later.
A later salary is observed for 1,651 of 1,903 player-seasons one season ahead and 1,443 two seasons ahead. The workbook lists about 450 players a season, fewer than are paid, and even omits some stars in some seasons. A missing salary is recorded as unknown, with the reason where the data allow (played that season, not in RAPTOR, after RAPTOR ends, after the last salary season), and never as zero. No player is marked as having no NBA earnings: no permitted source that could verify it was found. Estimates use complete cases and are conditional on an observed later salary.
Salary-cap normalization
The cap rose from $63.1M to $112.4M, so nominal salaries aren't comparable across seasons. The target is salary as a share of that season's cap — the unit in which the CBA defines maximum contracts. Errors are reported in 2021-22 cap dollars (share × $112.4M). Each model family chose between the raw share, its square root and its log by cross-validation; the selected gradient boosting uses the square root.
Features
14 inputs describe minutes, RAPTOR offense and defense, regular-season and playoff wins above replacement, PREDATOR, pace impact, position, years since NBA debut, and prior-season minutes, RAPTOR and WAR. No salary enters the valuation features. Totals and near-duplicates (possessions, minutes share, component ratings) are dropped by a rule fixed in advance, so each concept enters once.
RAPTOR totals include the regular season and playoffs. Years since debut come from the full RAPTOR history back to 1977, and prior-season values come from the full source file, including seasons excluded from the modeled sample. Years since debut are not credited service or contract status. Age, draft history and traditional box-score fields are not in the product model's sources and are omitted.
Pipeline and architecture
Acquire
- Mendeley salaries through 2023-24
- FiveThirtyEight RAPTOR through 2021-22
src/data/public_data.pyNormalize
- Reviewed names and season joins
- Player-season and salary-record tables in SQLite
src/data/Outcomes
- Exact t + h salary joins in SQL
- Unknown kept apart from zero; SQL checks
src/data/salary_outcomes.pyModel
- 14 features, four model families
- Walk-forward validation, 2021-22 test
src/models/Analyze
- Benchmark vs later pay, chronological backtest
- Cluster bootstrap, pre-specified checks
src/analysis/pay_adjustment.pyExport & web
- Typed JSON, schema and checksum checks
- Next.js static site; aggregate-only option
src/web/ · web/
Python loads the public files into SQLite. Tables hold player-seasons, every salary record, caps and minimums, a name crosswalk, the later-salary outcomes and the benchmark predictions. Relational checks cover duplicate keys, orphan joins, exact horizons and unknown-versus-zero salaries. Feature engineering, chronological validation, model comparison, interval calibration and the later-pay analysis produce saved artifacts. The exporter checks them against a typed JSON contract before the Next.js static build.
Models and validation
OLS, Lasso, random forest and gradient boosting get the same features and folds. Walk-forward validation: each of the 4 validation seasons (2017-18 → 2020-21) is predicted by a model trained only on earlier seasons, which exposes drift such as the 2016-17 cap spike. The held-out season (2021-22; 383 player-seasons) were scored once. Uncertainty comes from a bootstrap that resamples players, to retain within-player dependence in repeated observations.
Held-out MAE: gradient boosting $3.82M against $4.62M for OLS, 17.3% lower (95% CI 12.6%–21.9%). Evidence, season by season and by subgroup, is in the Model Lab.
Removing pace impact, position and years since debut increases held-out MAE by 24.2% to $4.75M. Removing prior-season performance also raises error. The ablation measures dependence on these inputs without retuning.
Model selection history
- Development. Tuning and model choice used the walk-forward validation seasons (2017-18 → 2020-21). Gradient boosting had the lowest validation MAE ($3.76M vs random forest $3.76M) and was selected before the held-out seasons were scored.
- Held-out evaluation. 2021-22, never used for tuning or selection, were then scored with the chosen specification. Gradient boosting $3.82M, random forest $3.92M. These results were then inspected.
- Later audit. (Audit rule, added in the final review.) A player-level bootstrap on the validation seasons only found the model families statistically tied: the validation difference spans −$64K – $76K. The held-out difference (−$23K – $213K) is within noise too.
- Reference model. Switching models because of held-out MAE would choose the model from the evaluation results. All four model families are reported.
Ranges and SHAP
The 80% ranges use absolute residuals in the selected model's square root space from 1,520 walk-forward validation predictions. Held-out coverage is 80.4% overall and 3 of 15 for salaries at 30% of the cap or more, where every miss is a salary above the range. The ranges are marginal, not conditional on salary level, and temporal drift means this is empirical calibration, not a guarantee.
For the selected gradient boosting, SHAP values are exact in the square root space and are converted to dollars with an exact additive allocation. They explain the fitted model. They are not causal effects, wage premiums or a decomposition of pay.
Comparables
Statistical comparables use 10 production and role statistics in 6 groups, standardized within each season, with small-sample rates shrunk toward average (300-minute prior). No salary or model output enters the profile; pay is compared only afterwards, in cap share. As a check, a player's previous-season profile is his single closest match 4% of the time, against 3% for a minutes-and-RAPTOR baseline. “Similar profiles paid less” lists statistical matches, not players a team could have signed instead.
Later pay analysis
The question: at the same current salary, does a higher benchmark go with higher pay later? The specification, units, samples and reading rules were written down and committed before any estimate (docs/PAY_ADJUSTMENT_DESIGN.md). Later cap share is regressed on current cap share, the benchmark, career stage and season effects. Because growth and the gap both contain current salary, the benchmark coefficient measures extra information, not a causal correction.
One season ahead the coefficient is 0.45 (95% CI 0.37 to 0.55), and 1.78 per point for players 0–3 years from their debut; it is an average that leans on large pay changes (the median response, post hoc, is 0.09). As a forecast, adding the benchmark to a regression on current salary and career stage moves mean absolute error from 2.66 to 2.65 points of cap (−0.003); repeating current cap share scores 2.33 on that metric, an edge that largely disappears when the regressions are fitted by median regression. Intervals come from a player-cluster bootstrap that keeps repeated draws, and a second bootstrap refits the benchmark itself. Full results are in Model Lab · Later pay.
Testing
- Python (pytest): cleaning and name rules, player-season identity checks, feature timing, validation folds, evaluation views, comparables and every web-export invariant.
- Later salaries: a hand-built database checks skipped seasons, outcomes without future eligibility, partial pay, invalid records, shared names and censoring. The bootstrap is tested for repeated draws, and the gap coefficient for its identity with the benchmark coefficient.
- An independent script re-derives every headline number from the saved predictions; refitting reproduces the saved predictions exactly.
- Web (Vitest): formatting and search parity with Python via generated fixtures, comparables ranking parity, scales, rankings, Compare state and the product's copy rules.
- End-to-end smoke tests (Playwright) cover the main journeys in a production build.
Limitations
- These are retrospective salary estimates. Contracts generally precede the performance they pay for.
- Published salary accounting is imperfect; the partial-payment filter is a practical approximation.
- The model is unconstrained: estimates and intervals can exceed NBA contract limits. Large estimates should be read as extrapolation.
- Offense, defense, WAR and minutes are correlated, so individual SHAP contribution sizes are unstable even when their sum is not.
- The latest test is one season. The selected gradient boosting is statistically tied with random forest on validation; the choice follows the fixed rule, not the test.
- Later salaries are missing for some players; later-pay estimates are conditional on observing them, and years since debut stand in for contract status.
- Missing salary matches and the playing-time threshold limit who appears.
Sources and coverage
Performance ends in 2021-22: FiveThirtyEight RAPTOR stops there. Salaries run through 2023-24, the last sheet of the Mendeley workbook. Two source reviews under a $0 budget looked for newer data across research repositories, open data packages, free APIs, official league and union publications and public salary sites. They found a performance source through 2025-26 whose terms allow private, non-commercial use (SportsDataverse box scores from NBA Stats), so a box-score comparison runs locally and its results are not published here. They found no 2024-25 or 2025-26 salary source this project may collect: the sites that publish salaries either bar machine-learning use (Basketball-Reference), block AI agents (ESPN, HoopsHype, RealGM), or refuse automated access (Spotrac), and the free APIs checked offer no season salary history at $0. Modernization through 2025-26 is therefore deferred. The review is in docs/SOURCE_FEASIBILITY.md.
Data provenance and terms
Salary: Annual salaries of athletes from 10 professional sports (2014–2023), Thitithep Sitthiyot, Mendeley Data v1, CC BY 4.0, sheets NBA14–NBA23 (2014-15 to 2023-24). The author identifies Spotrac as upstream provenance. This project uses the published workbook and does not scrape Spotrac. Upstream permissions are not independently established. Salary caps come from NBA Communications announcements.
Performance: FiveThirtyEight NBA RAPTOR: modern and historical player files and the team-stint file, CC BY 4.0. Data are normalized, joined, filtered and used to train new models. Accessed September 20, 2026.
The source licenses and provenance are documented under the project's portfolio-use standard; this is not a claim of complete legal clearance.
Media credits
Player photos come only from Wikimedia Commons. Each is matched to the player through Wikidata's Basketball-Reference id (never by name alone), and kept only if the file's own license is CC0, public domain, CC BY or CC BY-SA and it names an author. Photos are resized and cropped for display. Everyone else keeps the monogram. Team logos are not used: their reuse terms aren't clear for this project, so teams stay abbreviations.
Player photos are omitted from this build.