Eli Baltz · Economics & Data Science project
NBA Salary Benchmark and Later Pay Analysis
An end-to-end analysis of how NBA production relates to pay, and whether a salary benchmark adds information about later earnings.
What I built
A reproducible Python and SQLite pipeline integrates 4,203 player-seasons into a 3,055-row modeling dataset with 14 features. Four regression methods are compared using walk-forward validation. A second analysis links the resulting benchmark to salaries one, two and three seasons later.
The public Next.js and D3 dashboard presents aggregate findings, SHAP explanations, diagnostics and uncertainty. The full repository stays private; the selected implementation excerpts below show two decisions that protect the analysis from subtle errors.
- Held-out MAE
- $3.82M
- Reduction versus OLS
- 17.3%
- Coverage of 80% range
- 80.4%
Data and evaluation
FiveThirtyEight RAPTOR performance records and Mendeley annual salary records are joined by reviewed name mappings and season. Salary is normalized to each season’s cap. Historical performance coverage ends in 2021–22; later salary outcomes extend through 2023–24. Missing salary records remain unknown rather than becoming zero earnings.
Development uses earlier seasons, with 4 chronological validation seasons and 1 untouched test season. Gradient boosting was selected by validation MAE before evaluating the held-out season. SHAP describes the fitted benchmark; conformal ranges show its uncertainty. The model describes market pay associated with observed production, not intrinsic player value or a contract recommendation.
SQL: join the exact future season
A “next observed salary” can belong to the wrong year when a record is missing. The outcome panel computes target_season = origin_season + horizon, then joins by that exact season. Fourteen SQL quality checks cover uniqueness, timing, identity and missingness.
LEFT JOIN salary_record r
ON r.season = l.target_season AND r.name_key = l.name_key AND r.valid_amount = 1Excerpt from src/data/salary_outcomes.py. The LEFT JOIN preserves origins without a valid future record; later status logic records why the outcome is unknown.
Python: keep repeated bootstrap draws
Several observations belong to the same player. Resampling entire player clusters respects that dependence. A player drawn twice must contribute two copies of every row; a membership filter would silently discard that multiplicity.
def cluster_draws(clusters: np.ndarray, n_boot: int, seed: int) -> list[np.ndarray]:
"""Row indices for each bootstrap draw. Clusters are drawn with replacement and a cluster drawn k times
contributes all of its rows k times (the multiplicity that ``isin(sampled_ids)`` would silently drop)."""
codes, uniques = pd.factorize(pd.Series(clusters), sort=True)
order = np.argsort(codes, kind="stable")
bounds = np.searchsorted(codes[order], np.arange(len(uniques) + 1))
members = [order[bounds[k]:bounds[k + 1]] for k in range(len(uniques))]
rng = np.random.default_rng(seed)
draws = []
for _ in range(n_boot):
pick = rng.integers(0, len(uniques), len(uniques))
draws.append(np.concatenate([members[k] for k in pick]))
return drawsFunction from src/analysis/pay_adjustment.py. Seeded draws support reproducibility. A separate two-stage bootstrap refits the salary benchmark to assess fitting uncertainty.
What the later-pay analysis found
The benchmark is associated with later salary after adjusting for current pay, career stage and season. That association is conditional on having an observed later salary and is not causal evidence that teams correct mispricing.
For the pre-specified one-year forecast comparison, adding the benchmark changes MAE by −0.003 salary-cap percentage points (95% interval −0.065 to 0.061). The interval includes zero. A descriptive association therefore does not establish a useful forecasting improvement.
Limits and next work
The salary source omits some paid players, complete-case outcomes can be selective, and rookie and maximum contract rules constrain market pay. Overall interval coverage hides weak coverage at the top of the salary scale. Historical data support a transparent benchmark, while a more current release requires a suitable free source with clear reuse terms.
This public build contains aggregate artifacts and selected source excerpts. Player-level records and private pilot results are excluded. The methodological documentation distinguishes pre-specified analysis from post hoc checks.