WBB — additional Python functions — Models and calculators: build_wbb–win_prob
build_wbb_season_wp
build_wbb_season_wp(season: 'int', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"
A WBB season's play-by-play with win-probability columns joined in.
Delegates to sportsdataverse.mbb.mbb_win_prob.build_mbb_season_wp
with league="womens" (WBB loaders + women's constants).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
season | int | Season year (e.g. 2024); bounded by load_wbb_pbp release availability. | |
return_as_pandas | bool | False | Return a pandas DataFrame instead of polars. |
Returns
The season's load_wbb_pbp frame with pregame_home_prob + home_win_prob appended -- see the mbb core for the contract.
| col_name | type | description |
|---|---|---|
game_play_number | integer | Game play number |
id | integer | Unique play identification number |
sequence_number | integer | Sequence number representing a shot-possession (V3 PBP). |
type_id | integer | Type identifier (numeric). |
type_text | character | Display text for the type field. |
text | character | Text description of the play / record. |
away_score | integer | Away team score at the time of the play. |
home_score | integer | Home team score at the time of the play. |
period_number | integer | Numeric period (1-4 for quarters; 5+ for OT). |
period_display_value | character | Period display label (e.g. '1st Quarter', 'OT'). |
clock_display_value | character | Game clock display string (e.g. '8:32'). |
scoring_play | logical | TRUE if the play resulted in points scored. |
score_value | integer | Point value of the play (2 / 3 / 1). |
team_id | integer | Unique team identifier. |
athlete_id_1 | integer | Primary athlete identifier (e.g. shooter). |
wallclock | character | Wallclock. |
shooting_play | logical | TRUE if the play was a shooting attempt. |
athlete_id_2 | integer | Secondary athlete identifier (e.g. assister / fouler). |
game_id | integer | Unique game identifier. |
season | integer | Season identifier (4-digit year or 'YYYY-YY' string). |
season_type | integer | Season type (1=pre-season, 2=regular season, 3=postseason, 4=off-season for ESPN; or string label for WNBA Stats). |
home_team_id | integer | Unique identifier for the home team. |
home_team_name | character | Home team name. |
home_team_mascot | character | Home team mascot. |
home_team_abbrev | character | Home team three-letter abbreviation. |
home_team_name_alt | character | Alternate versions of the home team abbreviation |
away_team_id | integer | Unique identifier for the away team. |
away_team_name | character | Away team name. |
away_team_mascot | character | Away team mascot. |
away_team_abbrev | character | Away team three-letter abbreviation. |
away_team_name_alt | character | Alternate versions of the away team abbreviation |
game_spread | double | Game spread in (-X Team) format. There are almost none, I would recommend not trusting any of these three columns |
home_favorite | logical | Logical (TRUE/FALSE) indicating whether the home team is favored |
game_spread_available | logical | Logical (TRUE/FALSE) indicating whether the spread was available from ESPN. Basically, I would just not recommend using any of the spread information, I think I defaulted a lot of them to -2.5 for the home team. Most games probably do not have spread information. This column should really be listed first |
home_team_spread | double | The game spread with respect to the home team |
qtr | integer | Quarter of the game |
time | character | Time left within the period |
clock_minutes | integer | Clock minutes split from seconds for developer convenience |
clock_seconds | integer | Clock seconds split from minutes for developer convenience |
home_timeout_called | logical | |
away_timeout_called | logical | |
half | integer | Half of the game |
game_half | integer | Half of the game |
lag_qtr | integer | A lag column on the quarter |
lead_qtr | integer | A lead column on the quarter |
lag_half | integer | A lag column on the half |
lead_half | integer | A lead column on the half |
start_quarter_seconds_remaining | integer | Quarter seconds remaining at the start of the play (these are more or less code artifacts from other sports, but may eventually be used more seriously) |
start_half_seconds_remaining | integer | Game half seconds remaining at the start of the play (these are more or less code artifacts from other sports, but may eventually be used more seriously) |
start_game_seconds_remaining | integer | Game seconds remaining at the start of the play (''') |
end_quarter_seconds_remaining | integer | Quarter seconds remaining at the end of the play (''') |
end_half_seconds_remaining | integer | Game half seconds remaining at the end of the play (''') |
end_game_seconds_remaining | integer | Game seconds remaining at the end of the play (''') |
period | integer | Period of the game (1-4 quarters; 5+ for OT). |
game_date | character | Game date (YYYY-MM-DD). |
game_date_time | character | Game start date/time (ISO 8601). |
coordinate_x | double | X coordinate on the court (half-court layout). |
coordinate_y | double | Y coordinate on the court (half-court layout). |
coordinate_x_raw | double | X coordinate as returned by the API before any adjustment. |
coordinate_y_raw | double | Y coordinate as returned by the API before any adjustment. |
athlete_name_1 | character | |
athlete_name_2 | character | |
athlete_name_3 | character | |
media_id | integer | Where did you come from |
pregame_home_prob | double | |
home_win_prob | double |
Example
from sportsdataverse.wbb import build_wbb_season_wp
wp = build_wbb_season_wp(2024)
build_weak_prior_from_rapm
build_weak_prior_from_rapm(rapm_results: 'list[float]', off_or_def: 'str') -> 'list[dict[str, float]]'
Wrap a flat RAPM-estimate vector into playersWeak-shaped dicts.
Faithful port of RapmUtils.buildWeakPriorFromRapm (RapmUtils.ts:410-419),
used only by pick_ridge_regression's use_recursive_weak_prior
branch to substitute the just-computed (pre-strong-prior) RAPM values as
the weak prior for a follow-up apply_weak_priors call -- "the
recursive prior" per the upstream /** For "recursive" prior */ comment.
Uncovered by the oracle -- semiRealRapmResults.testContext.priorInfo .useRecursiveWeakPrior is false, so RapmUtils.test.ts's
"pickRidgeRegression" test never calls this function. Ported
faithfully from TS regardless (per "TS governs"); flagged as a documented
gap rather than backed by a synthetic test, matching this module's
existing convention for other upstream-untested branches (e.g. the
"Task 3.3 coverage gap" note above).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
rapm_results | list[float] | A flat per-player RAPM estimate vector, e.g. pick_ridge_regression's own results_pre_prior. | |
off_or_def | str | "off" or "def" -- selects the output key, f"{off_or_def}_adj_ppp". |
Returns
One {f"{off_or_def}_adj_ppp": rapm} dict per input element, index-aligned with rapm_results.
Example
from sportsdataverse.mbb.mbb_rapm import build_weak_prior_from_rapm
weak_prior = build_weak_prior_from_rapm([5.0, 4.5], "off")
print(weak_prior[0]) # {"off_adj_ppp": 5.0}
calc_collinearity_diag
calc_collinearity_diag(weight_matrix: 'NDArray[np.float64]', ctx: 'RapmPlayerContext') -> 'RapmPreProcDiagnostics'
Multi-collinearity diagnostic between the players in an off/def design matrix.
Faithful port of RapmUtils.calcCollinearityDiag (RapmUtils.ts:1629-1760).
Runs an SVD of weight_matrix, builds condition indices ("lineup
combos") from the ratio of the largest to each singular value, and a
variance-decomposition-proportions ("VDP") matrix identifying which
players load onto which collinear combo -- the classic Belsley-Kuh-Welsch
collinearity-diagnostics recipe (see the upstream comment's
colldiag.m
citation). Also builds a plain Pearson player/player correlation matrix
(calc_player_correlations) and folds it into a possession -weighted adaptive_correl_weights` summary per player.
numpy.linalg.svd(weight_matrix, full_matrices=False) replaces
svd-js's SVD(weightMatrix, false). Both are the standard
Golub-Kahan-Reinsch decomposition (A = U @ diag(S) @ Vᵀ); numpy's
Vh return value already is Vᵀ (what the TS code separately
computes via transpose(matrix(v))), so this port skips that
transpose. The TS code (and this port) never reads u/the first SVD
return -- only q/S (singular values) and v/Vᵀ. Singular
-vector sign is immaterial here: every place V is used
(phiMatrix/phi_matrix) squares each entry (val * val), and a
per-singular-value sign flip on U/V together is a valid SVD
regardless -- so any U/V sign convention difference between
svd-js and LAPACK (numpy's backend) cannot change this function's
output. Singular-value ordering is likewise immaterial: both this
port and the TS source explicitly re-sort q (ascending, carrying the
original index along) before using it, so whichever order either SVD
implementation returns values in, the final result only depends on the
values themselves (up to the explicit resort), not on numpy's native
descending convention vs whatever order svd-js happens to return.
correl_matrix/poss_correl_matrix stay numpy.ndarray (see
the module docstring's "Task 3.6 notes" for why this doesn't hit the
Task 3.5 "ndarray breaks deep ==" concern).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
weight_matrix | NDArray[float64] | An off/def design matrix, shape (num_lineups, ctx["num_players"]) (e.g. calc_player_weights's first return value, or a hand-built matrix for isolated testing). | |
ctx | RapmPlayerContext | A RapmPlayerContext. ctx["num_players"] sizes every per-player structure; ctx["col_to_player"] keys player_combos. |
Returns
A RapmPreProcDiagnostics.
Example
from sportsdataverse.mbb.mbb_rapm import calc_collinearity_diag, calc_player_weights
off_weights, _ = calc_player_weights(ctx)
diag = calc_collinearity_diag(off_weights, ctx)
print(diag["lineup_combos"][0]) # the worst-conditioned combo
calc_lineup_outputs
calc_lineup_outputs(field: 'str', off_offset: 'float', def_offset: 'float', ctx: 'RapmPlayerContext', adaptive_correl_weights: 'list[float] | None' = None, use_old_val_if_possible: 'tuple[bool, bool]' = (False, False)) -> 'list[NDArray[np.float64]]'
Build the off/def target vectors the RAPM design matrices are fit against.
Faithful port of RapmUtils.calcLineupOutputs (RapmUtils.ts:598-751).
For each filtered lineup, computes a possession-weighted residual: the
lineup's own stat value, plus any global luck adjustment, minus the
accumulated "prior offset" contributed by every player on the lineup
(a strong-prior blend for kept players -- see get_strong_weight`
-- or a fixed baseline contribution for removed players).
Upstream keeps this as a plain Array<Array<number>> (not a mathjs
Matrix, unlike calc_player_weights's offWeights/
defWeights -- RapmUtils.test.ts's own tidyResults helper for
this function has a visibly different shape, see the classification map
in tests/fixtures/hoop_explorer/README.md). This port still
materializes both output vectors as numpy.ndarray for consistency
with calc_player_weights at the same dict -> array boundary --
Task 3.4's ridge-regression solve consumes both as arrays regardless of
the upstream distinction.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
field | str | The stat suffix to read off each lineup, e.g. "adj_ppp" (read as {prefix}_{field}, e.g. "off_adj_ppp"). | |
off_offset | float | The D1-average offensive value for field (the regression's starting/baseline value on the RHS). | |
def_offset | float | The D1-average defensive value for field. | |
ctx | RapmPlayerContext | A RapmPlayerContext, e.g. from build_player_context. | |
adaptive_correl_weights | list[float] | None | None | Optional per-player adaptive-correlation weights (index-aligned with ctx["col_to_player"]), used as the strong-prior blend fallback when ctx["prior_info"] ["strong_weight"] < 0 -- see get_strong_weight`. |
use_old_val_if_possible | tuple[bool, bool] | (False, False) | (use_old_val_for_off, use_old_val_for_def) -- whether to prefer each lineup/team stat's luck-adjusted old_value over its raw value when present. This is the luck-adjustment hook Task 3.1's classification map flags as an inherited coverage gap: the vendored oracle fixture has old_value == value on every field (via insertOldValues), so neither jest nor this port's replay test ever observes this flag change the resulting numbers -- only that passing it doesn't crash. See the module docstring's "Task 3.3 coverage gap" note. |
Returns
[off_outputs, def_outputs] -- two 1-D numpy.ndarray target vectors, index-aligned with ctx["filtered_lineups"]("off"/"def") (plus one extra element each when ctx["unbias_weight"] > 0, an "unbiasing observation" target -- always unreached in production, same as calc_player_weights's extra row).
Example
from sportsdataverse.mbb.mbb_rapm import calc_lineup_outputs
off_outputs, def_outputs = calc_lineup_outputs(
"adj_ppp", 100.0, 100.0, ctx
)
print(off_outputs.shape) # (num_off_lineups,)
# Luck-adjusted variant (reads ``old_value`` where present)
off_luck, def_luck = calc_lineup_outputs(
"adj_ppp", 100.0, 100.0, ctx, use_old_val_if_possible=(True, True)
)
calc_player_weights
calc_player_weights(ctx: 'RapmPlayerContext') -> 'list[NDArray[np.float64]]'
Build the off/def player-weight (design) matrices for the RAPM solve.
Faithful port of RapmUtils.calcPlayerWeights (RapmUtils.ts:544-595).
One row per (filtered) lineup, one column per remaining player; each
filled cell is sqrt(lineup_possessions / total_side_possessions) --
the possession-weighted design-matrix entry the ridge regression (Task
3.4) solves against. This is the first function in the module where a
dict-shaped RapmPlayerContext gets materialized into a
numpy.ndarray -- see the module docstring's "dict -> numpy.ndarray
boundary" note.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
ctx | RapmPlayerContext | A RapmPlayerContext, e.g. from build_player_context. |
Returns
[off_weights, def_weights] -- two numpy.ndarray matrices of shape (num_{off,def}_lineups [+1 if ctx["unbias_weight"] > 0], ctx["num_players"]). The optional extra row (only emitted when ctx["unbias_weight"] > 0 -- always 0.0 in production per build_player_context's hardcoded local, but settable directly on the returned context dict, as the oracle test does) holds each column's unbias_weight-scaled sum-of-squares, an "unbiasing observation" row (RapmUtils.ts:578-593).
Example
from sportsdataverse.mbb.mbb_rapm import calc_player_weights
off_weights, def_weights = calc_player_weights(ctx)
print(off_weights.shape) # (num_off_lineups, num_players)
calc_slow_pseudo_inverse
calc_slow_pseudo_inverse(player_weight_matrix: 'NDArray[np.float64]', ridge_lambda: 'float', ctx: 'RapmPlayerContext') -> 'NDArray[np.float64]'
Per-parameter variance terms for the ridge-regression standard errors.
Faithful port of the private RapmUtils.calcSlowPseudoInverse
(RapmUtils.ts:1544-1557): the same (XᵀX + ridge_lambda·I)⁻¹ as
slow_regression's bottomInv, but this function returns the
square root of its diagonal instead of the full solver matrix -- the
paramErrs term consumed by the standard-error formula (see
calculate_sd_rapm).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
player_weight_matrix | NDArray[float64] | The off/def design matrix, same shape as slow_regression's. | |
ridge_lambda | float | The Tikhonov regularization strength (must match the ridge_lambda used to build the corresponding slow_regression solver, for the SEs to be meaningful). | |
ctx | RapmPlayerContext | A RapmPlayerContext -- only ctx["num_players"] is read. |
Returns
A length-num_players array, sqrt(diag((XᵀX + λI)⁻¹)).
Example
from sportsdataverse.mbb.mbb_rapm import calc_slow_pseudo_inverse
param_errs = calc_slow_pseudo_inverse(x, 1.0, ctx)
calculate_predicted_out
calculate_predicted_out(player_weight_matrix: 'NDArray[np.float64]', regressed_players: 'list[float]', ctx: 'RapmPlayerContext') -> 'NDArray[np.float64]'
Predict per-lineup outputs from fitted per-player RAPM values.
Faithful port of RapmUtils.calculatePredictedOut (RapmUtils.ts:1559-1567).
ctx is accepted for signature parity with the TS source but unused in
the body (ported verbatim -- upstream's own ctx param is likewise
dead in this function).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
player_weight_matrix | NDArray[float64] | The off/def design matrix, shape (num_lineups, num_players). | |
regressed_players | list[float] | The fitted per-player values (e.g. the final, strong-prior-blended RAPM from Task 3.5's pickRidgeRegression, or a raw calculate_rapm output), length num_players. | |
ctx | RapmPlayerContext | A RapmPlayerContext (unused). |
Returns
The predicted per-lineup value, length num_lineups -- feed into calculate_residual_error alongside the actual lineup outputs.
Example
from sportsdataverse.mbb.mbb_rapm import calculate_predicted_out
predicted = calculate_predicted_out(x, [0.875, 1.375], ctx)
calculate_rapm
calculate_rapm(regression_matrix: 'NDArray[np.float64]', player_outputs: 'list[float]') -> 'NDArray[np.float64]'
Apply a regression solver matrix to a target-outputs vector.
Faithful port of RapmUtils.calculateRapm (RapmUtils.ts:772-775).
Note the TS signature carries no ctx parameter (unlike its solve-layer
siblings) -- ported verbatim, param-for-param.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
regression_matrix | NDArray[float64] | The (num_players, num_lineups) solver from slow_regression. | |
player_outputs | list[float] | The per-lineup target vector, length num_lineups (e.g. calc_lineup_outputs's off_outputs/def_outputs). |
Returns
The per-player RAPM estimate, length num_players.
Example
from sportsdataverse.mbb.mbb_rapm import calculate_rapm
rapm = calculate_rapm(solver, [1.0, 2.0, 3.0])
print(rapm.shape) # (num_players,)
calculate_residual_error
calculate_residual_error(player_outs: 'list[float]', regressed_outs: 'list[float]', ctx: 'RapmPlayerContext') -> 'float'
Sum of squared residuals between actual and predicted lineup outputs.
Faithful port of RapmUtils.calculateResidualError (RapmUtils.ts:1569-1579).
ctx is accepted for signature parity but unused in the body (dead
upstream too).
NaN/shape regime (landmine 7): TS zips the two arrays via lodash
.zip(pads the shorter side withundefined, so a length mismatch silently contributes NaNto the running sum viaundefined - number) then reduces with plain +. This port instead subtracts the two as numpyarrays: a length mismatch **raises**ValueError(numpy broadcast rules), rather than the TS silent-NaN behavior -- not reachable via either language's own call sites (both arguments are always index-aligned to the same lineup count in production), so this is a divergence in dead territory, not a fixed bug. ANaN*value already present* inside either input (as opposed to a length mismatch) propagates through thenumpy` subtraction/sum
exactly as it would through the JS arithmetic (both regimes:
numpy-propagate).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
player_outs | list[float] | The actual per-lineup target values (e.g. calc_lineup_outputs's output). | |
regressed_outs | list[float] | The predicted per-lineup values (e.g. calculate_predicted_out's output). | |
ctx | RapmPlayerContext | A RapmPlayerContext (unused). |
Returns
sum((player_outs[i] - regressed_outs[i]) ** 2) -- the errSq term consumed by calculate_sd_rapm.
Example
from sportsdataverse.mbb.mbb_rapm import calculate_residual_error
err_sq = calculate_residual_error([1.0, 2.0, 3.0], [0.875, 1.375, 2.25], ctx)
calculate_sd_rapm
calculate_sd_rapm(param_errs: 'NDArray[np.float64]', err_sq: 'float', num_lineups: 'int', num_players: 'int') -> 'NDArray[np.float64]'
Per-player RAPM standard errors.
Faithful port of the inline sdRapm computation in
RapmUtils.pickRidgeRegression (RapmUtils.ts:1373-1390, not itself
a named TS function -- promoted to a standalone, independently testable
helper here since Task 3.4's brief calls out the formula explicitly).
Cites arXiv:1509.09169.
Two NaN/error regimes (landmines 8-9):
dof_inv = 1.0 / (num_lineups - num_players)-- ifnum_lineups == num_playersexactly, JS silently producesInfinity(float division by zero); this port instead raisesZeroDivisionError(Python float division by zero), matching this module's already-established landmine-2 convention (unguarded division, Python-raises vs JS-Infinity/NaN). Not reachable via the oracle fixtures (num_off_lineups/num_def_lineupsalways comfortably exceednum_playersthere).sqrt(sqrt(param_errs) * err_sq * dof_inv)-- a negativeparam_errsentry (only possible ifXᵀX + λIisn't actually positive-definite, e.g.ridge_lambda < 0) silently numpy-propagates toNaN(matching JSMath.sqrt(negative) -> NaN, with aRuntimeWarningrather than a raise) -- both language regimes agree here, unlike landmine 8.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
param_errs | NDArray[float64] | Per-player variance terms from calc_slow_pseudo_inverse, length num_players. | |
err_sq | float | The residual sum of squares from calculate_residual_error. | |
num_lineups | int | ctx["num_off_lineups"] or ctx["num_def_lineups"] (whichever side param_errs/err_sq were computed for). | |
num_players | int | ctx["num_players"]. |
Returns
A length-num_players array of per-player RAPM standard errors.
Example
from sportsdataverse.mbb.mbb_rapm import calculate_sd_rapm
sd_rapm = calculate_sd_rapm(param_errs, err_sq, num_lineups=3, num_players=2)
calibration_table
calibration_table(y_true: 'np.ndarray', p_pred: 'np.ndarray', n_bins: 'int' = 10) -> 'pl.DataFrame'
Bucket predicted probabilities into bins and compare to actual outcome rates.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
y_true | ndarray | Array of binary outcomes (0/1). | |
p_pred | ndarray | Array of predicted probabilities in [0, 1]. | |
n_bins | int | 10 | Number of equal-width probability bins. |
Returns
A polars.DataFrame with columns bin_mid, mean_pred, mean_actual, n (one row per non-empty bin).
| col_name | type | description |
|---|---|---|
bin_mid | double | |
mean_pred | double | |
mean_actual | double | |
n | integer |
Example
import numpy as np
from sportsdataverse._common.metrics import calibration_table
calibration_table(np.array([1, 0, 1, 0]), np.array([0.9, 0.1, 0.8, 0.2]))
fit_shrinkage_k
fit_shrinkage_k(scored: 'pl.DataFrame', *, seed: 'int' = 0) -> 'float'
Fit the talent shrinkage k split-half (see module docstring).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
scored | DataFrame | mbb_shot_quality output. | |
seed | int | 0 | Split seed (deterministic fit). |
Returns
The k in [1, 5000] minimizing talent_split_mse.
Example
from sportsdataverse.mbb.mbb_shooter_talent import fit_shrinkage_k
k = fit_shrinkage_k(scored)
get_constants
get_constants(league: 'str') -> 'ShotQualityConstants'
League constants bundle for the shot-quality spine.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
league | str | "mens" or "womens". |
Returns
The frozen ShotQualityConstants for that league.
Example
from sportsdataverse.mbb.mbb_shot_quality_constants import get_constants
get_constants("mens").rim_radius_ft
get_player_value_constants
get_player_value_constants(league: 'str') -> 'PlayerValueConstants'
Return the PlayerValueConstants for a league.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
league | str | "mens" or "womens". |
Returns
The league's PlayerValueConstants.
Example
from sportsdataverse.mbb.mbb_player_value_constants import get_player_value_constants
get_player_value_constants("mens").bundle_prefix
in_game_features
in_game_features(pbp: 'pl.DataFrame', pregame_home_prob: 'float') -> 'pl.DataFrame'
Per-play in-game win-probability features from a load_mbb_pbp frame.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
pbp | DataFrame | Play-by-play frame with start_game_seconds_remaining, home_score, away_score, team_id (event team) and home_team_id (the load_mbb_pbp schema). | |
pregame_home_prob | float | The pregame home win probability (e.g. from win_prob_from_margin), encoded as a constant logit column. Clipped to [1e-6, 1 - 1e-6] so a saturated CDF (exact 0/1) cannot crash the logit. |
Returns
One row per input play: score_diff (home - away), sec_left (clipped at 0 -- overtime plays count as 0 seconds left), sqrt_sec_left, pregame_logit, home_has_ball (Int8; dead-ball / unknown-team plays are 0).
| col_name | type | description |
|---|---|---|
score_diff | double | |
sec_left | double | |
sqrt_sec_left | double | |
pregame_logit | double | |
home_has_ball | integer |
Example
from sportsdataverse.mbb.mbb_game_predict import in_game_features
from sportsdataverse.mbb.mbb_loaders import load_mbb_pbp
pbp = load_mbb_pbp([2024]).filter(pl.col("game_id") == 401638643)
feats = in_game_features(pbp, 0.62)
inject_rapm_into_players
inject_rapm_into_players(players: 'list[PlayerOnOffStats]', off_rapm_input: 'RapmProcessingInputs', def_rapm_input: 'RapmProcessingInputs', stats_averages: 'PureStatSet', ctx: 'RapmPlayerContext', adaptive_correl_weights: 'list[float] | None', read_value_keys: 'tuple[ValueKey, ValueKey]' = ('value', 'value'), write_value_key: 'ValueKey' = 'value') -> 'None'
Write pick_ridge_regression's RAPM predictions back onto each player.
Faithful port of RapmUtils.injectRapmIntoPlayers (RapmUtils.ts:781-916).
For every onOffReportReplacement field (minus the possession/title/
separator/adj_opp housekeeping keys -- see landmine 11 for the exact,
faithfully-ported omit-key quirk), re-derives that field's off/def target
vectors via calc_lineup_outputs, applies each side's
calculate_rapm solver, blends in the strong prior (mirroring
pick_ridge_regression's own blend, except for adj_ppp which
reuses off_rapm_input["rapm_adj_ppp"]/def_rapm_input["rapm_adj_ppp"]
directly rather than recomputing), then writes {playerId}.rapm[field] = {write_value_key: result, "override": ...} onto every player not in
ctx["removed_players"].
NOTE (upstream comment, verbatim): when write_value_key == "old_value", this must be called after an initial write_value_key == "value" call on the same players list -- the old_value
pass .merges (lodash_merge) its results into each player's
existing rapm dict rather than replacing it, so a player's
rapm["field"] ends up carrying both a value (from the first
call) and an old_value (from the second) side by side.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
players | list[PlayerOnOffStats] | The players to write RAPM results onto (mutated in place -- each qualifying player gets a "rapm" key set/merged). | |
off_rapm_input | RapmProcessingInputs | pick_ridge_regression's offensive output. | |
def_rapm_input | RapmProcessingInputs | pick_ridge_regression's defensive output. | |
stats_averages | PureStatSet | League/context average stat set -- consulted for each field's off/def offset before ctx["team_info"]. | |
ctx | RapmPlayerContext | A RapmPlayerContext (the same one pick_ridge_regression was called with). | |
adaptive_correl_weights | list[float] | None | Optional per-player adaptive-correlation weights, forwarded to calc_lineup_outputs / get_strong_weightexactly aspick_ridge_regression` does. | |
read_value_keys | tuple[ValueKey, ValueKey] | ('value', 'value') | (off_key, def_key) -- which key ("value"/"old_value") to prefer when reading stats_averages/ctx["team_info"] offsets and when calling calc_lineup_outputs (forwarded as its use_old_val_if_possible flag). |
write_value_key | ValueKey | 'value' | "value" or "old_value" -- which key each written field carries its result under. |
Example
from sportsdataverse.mbb.mbb_rapm import inject_rapm_into_players
inject_rapm_into_players(players, off_results, def_results, {}, ctx, None)
print(players[0]["rapm"]["off_adj_ppp"]) # {"value": ..., "override": None}
# Luck-adjusted two-call sequence (``"value"`` first, THEN ``"old_value"``)
inject_rapm_into_players(
players, off_results, def_results, {}, ctx, None, ("value", "old_value"), "value"
)
inject_rapm_into_players(
players, off_results, def_results, {}, ctx, None, ("old_value", "old_value"), "old_value"
)
kmeans_fit
kmeans_fit(X: 'np.ndarray', k: 'int', seed: 'int', n_init: 'int' = 10, max_iter: 'int' = 100) -> "'tuple[np.ndarray, np.ndarray]'"
Seeded Lloyd's KMeans, best-of-n_init by inertia.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
X | ndarray | Feature matrix (n, d) (standardize first). | |
k | int | Number of clusters. | |
seed | int | RNG seed (deterministic output). | |
n_init | int | 10 | Independent restarts. |
max_iter | int | 100 | Lloyd iterations per restart. |
Returns
(centers[k, d], labels[n]).
Example
centers, labels = kmeans_fit(Z, k=8, seed=0)
load_artifact
load_artifact(name: 'str') -> 'dict'
Read a bundled player-value artifact (mbb/models/<name>.json).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
name | str | Artifact stem, e.g. "mbb_box_bpm". |
Returns
The parsed JSON dict.
Example
from sportsdataverse.mbb.mbb_player_value_constants import load_artifact
art = load_artifact("mbb_box_bpm")
log_loss_score
log_loss_score(y_true: 'np.ndarray', p_pred: 'np.ndarray', eps: 'float' = 1e-15) -> 'float'
Binary cross-entropy loss between predicted probabilities and outcomes.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
y_true | ndarray | Array of binary outcomes (0/1). | |
p_pred | ndarray | Array of predicted probabilities in [0, 1]. | |
eps | float | 1e-15 | Clipping bound to avoid log(0). |
Returns
The mean log loss.
Example
import numpy as np
from sportsdataverse._common.metrics import log_loss_score
log_loss_score(np.array([1, 0]), np.array([0.9, 0.1]))
logistic_fit
logistic_fit(X: 'np.ndarray', y: 'np.ndarray', lam: 'float' = 1.0) -> 'np.ndarray'
L2-penalized logistic regression via L-BFGS (intercept unpenalized).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
X | ndarray | Feature matrix (n, d). | |
y | ndarray | Binary outcomes (0/1). | |
lam | float | 1.0 | L2 penalty on the non-intercept coefficients. |
Returns
Coefficient vector of length d + 1 (intercept first).
Example
coef = logistic_fit(X, drafted, lam=1.0)
mae
mae(a: 'np.ndarray', b: 'np.ndarray') -> 'float'
Mean absolute error between two arrays.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
a | ndarray | First array of values. | |
b | ndarray | Second array of values (same length as a). |
Returns
The mean absolute error.
Example
import numpy as np
from sportsdataverse._common.metrics import mae
mae(np.array([1.0, 2.0]), np.array([1.5, 2.5]))
pick_ridge_regression
pick_ridge_regression(off_weights: 'NDArray[np.float64]', def_weights: 'NDArray[np.float64]', ctx: 'RapmPlayerContext', adaptive_correl_weights: 'list[float] | None', diag_mode: 'bool', agg_value_key: 'ValueKey' = 'value', lineup_value_keys: 'tuple[ValueKey, ValueKey]' = ('value', 'value')) -> 'tuple[RapmProcessingInputs, RapmProcessingInputs]'
Adaptively pick a ridge-regression lambda and blend in the RAPM priors.
Faithful port of RapmUtils.pickRidgeRegression (RapmUtils.ts:1001-1540)
-- the top-level driver that, per off/def side: scales a dimensionless
lambda_range by the design matrix's mean singular value
(avg_eigen_val) into an actual ridge strength, solves via
slow_regression/calculate_rapm, blends in each player's
strong prior (get_strong_weight), reconciles the possession -weighted team total against the actual team efficiency ([IMPORTANT-EQUATION-01], see below), nudges the result back towards the weak priors on any remaining error (apply_weak_priors), and decides whether to keep sweeping lambda` upward, roll back to the
previous step, or stop.
[IMPORTANT-EQUATION-01] (RapmUtils.ts:1306-1314/:1325-1333):
combined_adj_eff = sum(pct_by_player[i] * rapm[i] for i) + add_low_volume_adj_rtg, compared against actual_eff[off_or_def]
(the team's actual, prior-basis-adjusted efficiency, including
bench/removed-player possessions) to derive adj_eff_err -- the error
signal both the weak-prior nudge and the stopping rule react to.
Stopping rule (checked once per lambda step, in order): (1) once a
second step has run (not_first_step) and, unless in diag_mode,
the current step is past lambda_range_to_use[3], roll back to the
previous step's soln_matrix/ridge_lambda (but not
rapm_adj_ppp/rapm_raw_adj_ppp/sd_rapm, which stay at the
current, over-threshold step's values -- a faithful, non-obvious TS
asymmetry, RapmUtils.ts:1443-1448 vs :1483-1484) when
adj_eff_err >= error_exit_thresh (1.35 for the low-possession
-count offense special case, else 1.05) and the error is still
increasing (>= last_error); else (2) stop in place once
mean_diff (the mean per-player RAPM change since the previous step)
drops below pick_ridge_thresh (0.061 off / 0.091 def --
"more confident in offensive priors"); else (3) keep sweeping.
Adaptive-weight / prior asymmetry (the deep-equality oracle's load
-bearing behavior): the per-player strong-prior blend
(get_strong_weight(ctx["prior_info"], adaptive_correl_weights[i])) only consults adaptive_correl_weightswhenctx["prior_info"]["strong_weight"] < 0(adaptive mode) -- a fixed, non-negativestrong_weightalways wins. A fixture whoseplayers_strongentries carry nodef_adj_pppkey makes the blend'sstat.get(f"{off_or_def}_adj_ppp") or 0.0term (and, transitively,calc_lineup_outputs's own strong_valterm) contribute exactly0on the def side regardless ofstrong_weightoradaptive_correl_weights-- see the oracle test'sdef_results1/ def_results2` invariance assertions.
svd is numpy.linalg.svd(..., compute_uv=False), singular values
only. Upstream's SVD(weights[side].valueOf()) (svd-js) also
computes u/v, but only svd.q (the singular values, via
mean(svd.off.q)/mean(svd.def.q) at avg_eigen_val,
RapmUtils.ts:1077) is ever read -- u/v are dead. Skipping them
is an efficiency-only deviation with an identical result (singular
values are unique to a matrix regardless of the underlying SVD
implementation).
Dead-debug computation promoted to a real output (Python-side
addition, not upstream's own shape): upstream also computes
residuals/errSq/paramErrs/sdRapm at this point
(RapmUtils.ts:1363-1394) purely to feed a console.log gated
behind the same hardcoded-False debugMode as
apply_weak_priors -- none of the four is ever stored on
acc.output upstream (RapmProcessingInputs has no sdRapm
field there either). Since Task 3.4 built
calculate_predicted_out/calculate_residual_error/
calc_slow_pseudo_inverse/calculate_sd_rapm specifically
so this task could surface real standard errors, this port keeps
calling all four (matching TS's actual computation, which reuses the
exact same XᵀX + ridge_lambda·I inverse slow_regression
already computed -- so no new failure mode is introduced by keeping
this) and additionally stores the result on sd_rapm -- a superset
of, not a divergence from, the upstream return shape.
soln_matrix/sd_rapm are nested Python lists, not
NDArrays. Every field on the returned RapmProcessingInputs
is a plain (possibly nested) Python list/float specifically so
the whole dict stays comparable via plain == -- the oracle's deep
-equality assertions (e.g. off_results1 == off_results) would
otherwise raise ValueError: truth value of an array with more than one element is ambiguous the moment Python's dict/list equality machinery
tried to bool() a multi-element ndarray comparison.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
off_weights | NDArray[float64] | The offensive design matrix (e.g. calc_player_weights's first return value). | |
def_weights | NDArray[float64] | The defensive design matrix. | |
ctx | RapmPlayerContext | A RapmPlayerContext. | |
adaptive_correl_weights | list[float] | None | Optional per-player adaptive-correlation weights (index-aligned with ctx["col_to_player"]) -- see the "adaptive-weight / prior asymmetry" note above. | |
diag_mode | bool | If True, keeps sweeping every remaining lambda step (collecting prev_attempts diagnostics for all of them) even after a stopping condition has already fired, and relaxes the rollback/pick eligibility guards for the first few (< lambda_range_to_use[3]) diagnostic-only steps. Not exercised by this task's oracle (always called with False) -- ported faithfully from TS, uncovered by test. | |
agg_value_key | ValueKey | 'value' | "value" or "old_value" -- which key team/aggregate-level reads (actual_eff, the low-volume player adjustment) prefer when present. |
lineup_value_keys | tuple[ValueKey, ValueKey] | ('value', 'value') | (off_key, def_key) -- forwarded to calc_lineup_outputs as its use_old_val_if_possible flag (translated: key == "old_value"). |
Returns
(off_results, def_results) -- two RapmProcessingInputs.
Example
from sportsdataverse.mbb.mbb_rapm import pick_ridge_regression
off_results, def_results = pick_ridge_regression(
off_weights, def_weights, ctx, None, False
)
print(off_results["ridge_lambda"], off_results["rapm_adj_ppp"][:3])
player_per100_features
player_per100_features(season_stats: 'pl.DataFrame') -> 'pl.DataFrame'
Per-100 / rate features for every (player_id, season).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
season_stats | DataFrame | One row per player-season with the canonical counting columns (minutes, field_goals_made, field_goals_attempted, three_point_field_goals_made, free_throws_attempted, turnovers, points, fga_rim, fga_mid, fga_three, offensive_rebounds, defensive_rebounds, assists, blocks, steals) -- built from the player-boxscore aggregation (see the Phase-0 fitters). |
Returns
One row per (player_id, season): ids as Utf8 plus the 17 rate / per-100 features. Empty input returns the schema with zero rows.
| col_name | type | description |
|---|---|---|
player_id | character | Unique player identifier. |
season | integer | Season identifier (4-digit year or 'YYYY-YY' string). |
team_id | character | Unique team identifier. |
min | double | Minutes played. |
usage | double | Usage rate. |
ts_pct | double | True shooting percentage (0-1). |
efg_pct | double | |
ast_pct | double | Assist percentage. |
tov_pct | double | |
oreb_pct | double | |
dreb_pct | double | |
blk_pct | double | Blocks percentage (0-1 decimal). |
stl_pct | double | Steals percentage (0-1 decimal). |
ftr | double | Free-throw rate (offense). |
rim_share | double | |
mid_share | double | |
three_share | double | |
pts_per100 | double | |
reb_per100 | double | |
ast_per100 | double |
Example
from sportsdataverse.mbb.mbb_player_value_constants import player_per100_features
feats = player_per100_features(season_stats)
poss_calc_fragment_sum
poss_calc_fragment_sum(a: 'PossCalcFragment', b: 'PossCalcFragment') -> 'PossCalcFragment'
Field-wise add two PossCalcFragment\ s
(PossCalcFragment.sum, PossessionUtils.scala:146-153).
The Scala original uses shapeless.Generic to zip the two case
classes' fields and sum pairwise; since every field is a plain Int,
a plain zip over dataclasses.astuple reproduces the same
behavior without the generic-programming machinery.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
a | PossCalcFragment | The left-hand fragment. | |
b | PossCalcFragment | The right-hand fragment. |
Returns
A new PossCalcFragment with each field summed.
Example
from sportsdataverse.mbb.mbb_ncaa_models import (
PossCalcFragment,
poss_calc_fragment_sum,
)
frag1 = PossCalcFragment(1, 2, 3, 4, 5, 6, 7, 8)
frag2 = PossCalcFragment(1, 3, 5, 7, 9, 11, 13, 15)
poss_calc_fragment_sum(frag1, frag2)
# PossCalcFragment(2, 5, 8, 11, 14, 17, 20, 23)
predict_margin
predict_margin(home_adj_em: 'float', away_adj_em: 'float', neutral: 'bool' = False) -> 'float'
Women's expected margin.
Delegates to sportsdataverse.mbb.mbb_game_predict.predict_margin with league="womens".
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
home_adj_em | float | Home team's adjusted efficiency margin. | |
away_adj_em | float | Away team's adjusted efficiency margin. | |
neutral | bool | False | True for a neutral-site game. |
Returns
Expected margin in points (positive favors the home team).
Example
from sportsdataverse.wbb.wbb_game_predict import predict_margin
predict_margin(20.0, 10.0)
predict_total
predict_total(home_adj_o: 'float', home_adj_d: 'float', away_adj_o: 'float', away_adj_d: 'float', home_tempo: 'float', away_tempo: 'float') -> 'float'
Women's expected total points.
Delegates to sportsdataverse.mbb.mbb_game_predict.predict_total with league="womens".
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
home_adj_o | float | Home adjusted offensive efficiency (points / 100 poss). | |
home_adj_d | float | Home adjusted defensive efficiency. | |
away_adj_o | float | Away adjusted offensive efficiency. | |
away_adj_d | float | Away adjusted defensive efficiency. | |
home_tempo | float | Home adjusted tempo (possessions / game). | |
away_tempo | float | Away adjusted tempo. |
Returns
Expected combined points scored by both teams.
Example
from sportsdataverse.wbb.wbb_game_predict import predict_total
predict_total(100.0, 88.0, 96.0, 92.0, 72.0, 70.0)
raw_game_efficiency
raw_game_efficiency(schedule: 'pl.DataFrame', team_box: 'pl.DataFrame') -> 'pl.DataFrame'
Per-team, per-game possessions + raw offensive/defensive efficiency.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
schedule | DataFrame | Frame with game_id, season, date, home_team_id, away_team_id, neutral_site (ids as strings or ints; cast to Utf8 here). | |
team_box | DataFrame | Per-team boxscore with game_id, team_id, field_goals_attempted, offensive_rebounds, turnovers, free_throws_attempted, team_score. When it also carries total_turnovers / team_turnovers (the loaders do), a row whose turnovers and total_turnovers are both 0 takes its count from team_turnovers -- where ESPN's 2009-2012 women's box files it. |
Returns
One row per (game_id, team_id): game_id, season, date, team_id, opp_team_id, is_home, neutral_site, poss, off_eff, def_eff. Empty input returns that schema with zero rows. Team-game rows whose possession estimate is non-positive (an all-zero ESPN boxscore shell) are dropped with a UserWarning -- their efficiency is undefined, and one of them poisons the whole season's fixed point. Games in which either team still has 0 turnovers are dropped the same way: the possession estimate would miss its turnover term.
| col_name | type | description |
|---|---|---|
game_id | character | Unique game identifier. |
season | integer | Season identifier (4-digit year or 'YYYY-YY' string). |
date | character | Date in YYYY-MM-DD format. |
team_id | character | Unique team identifier. |
opp_team_id | character | |
is_home | logical | Whether the team was home. |
neutral_site | logical | Neutral site. |
poss | double | Poss. |
off_eff | double | Off eff. |
def_eff | double | Def eff. |
Example
from sportsdataverse.mbb.mbb_loaders import load_mbb_schedule, load_mbb_team_boxscore
from sportsdataverse.mbb.mbb_team_ratings import raw_game_efficiency
eff = raw_game_efficiency(load_mbb_schedule([2024]), load_mbb_team_boxscore([2024]))
ridge_cv_lambda
ridge_cv_lambda(X: 'np.ndarray', y: 'np.ndarray', groups: 'np.ndarray', lams: "'list[float]'") -> 'float'
Pick lambda by leave-one-group-out CV (groups = seasons/classes).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
X | ndarray | Feature matrix. | |
y | ndarray | Targets. | |
groups | ndarray | Group label per row (e.g. season); each held out once. | |
lams | list[float] | Candidate penalties. |
Returns
The candidate with the lowest mean held-out MSE.
Example
lam = ridge_cv_lambda(X, y, seasons, [0.1, 1, 10, 100])
ridge_fit
ridge_fit(X: 'np.ndarray', y: 'np.ndarray', lam: 'float') -> 'np.ndarray'
Closed-form ridge with an unpenalized intercept (coefficient 0).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
X | ndarray | Feature matrix (n, d). | |
y | ndarray | Targets (n,). | |
lam | float | L2 penalty on the non-intercept coefficients. |
Returns
Coefficient vector of length d + 1 (intercept first).
Example
import numpy as np
from sportsdataverse.mbb.mbb_player_value_constants import ridge_fit
beta = ridge_fit(np.random.rand(50, 3), np.random.rand(50), lam=1.0)
roc_auc
roc_auc(y_true: 'np.ndarray', score: 'np.ndarray') -> 'float'
Area under the ROC curve via the rank-sum (Mann-Whitney) identity.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
y_true | ndarray | Binary outcomes (0/1). | |
score | ndarray | Predicted scores (any monotone scale). |
Returns
AUC in [0, 1]; nan when only one class is present.
Example
import numpy as np
from sportsdataverse.mbb.mbb_player_value_constants import roc_auc
roc_auc(np.array([0, 1]), np.array([0.2, 0.9]))
save_artifact
save_artifact(name: 'str', obj: 'dict') -> 'None'
Write a bundled artifact (dev/fitter use -- writes into the source tree).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
name | str | Artifact stem, e.g. "mbb_box_bpm". | |
obj | dict | JSON-serializable artifact payload. |
Example
save_artifact("mbb_box_bpm", {"league": "mens", "coef": [0.1]})
score_to_tuple
score_to_tuple(s: 'str') -> 'tuple[int, int]'
Parse a "scored-allowed" score string (ExtractorUtils.score_to_tuple,
ExtractorUtils.scala:107-113).
Scala's str match { case regex(s1, s2) => ... } on a compiled
Regex requires the ENTIRE string to match (Regex.unapplySeq calls
Matcher.matches(), not find()) -- ported here as
re.fullmatch, not re.match/re.search.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
s | str | The raw score string, e.g. "55-68". |
Returns
(scored, allowed) as a tuple of ints, or (0, 0) if s doesn't fully match ([0-9]+)-([0-9]+).
Example
from sportsdataverse.mbb.mbb_ncaa_models import score_to_tuple
score_to_tuple("55-68") # (55, 68)
score_to_tuple("garbage") # (0, 0)
simulate_game
simulate_game(home_em: 'float', away_em: 'float', neutral: 'bool', rng: 'np.random.Generator') -> 'bool'
Sample one women's game outcome (women's sigma/HFA/em_scale).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
home_em | float | The home team's efficiency margin. | |
away_em | float | The away team's efficiency margin. | |
neutral | bool | True for a neutral-site game. | |
rng | Generator | The random generator to draw from. |
Returns
True if the home team wins the sampled game.
Example
import numpy as np
from sportsdataverse.wbb.wbb_season_sim import simulate_game
simulate_game(20.0, 5.0, False, np.random.default_rng(0))
slow_regression
slow_regression(player_weight_matrix: 'NDArray[np.float64]', ridge_lambda: 'float', ctx: 'RapmPlayerContext') -> 'NDArray[np.float64]'
Build the Tikhonov (ridge) regression solver matrix.
Faithful port of the private RapmUtils.slowRegression
(RapmUtils.ts:756-769): (XᵀX + ridge_lambda·I)⁻¹Xᵀ, where X
is player_weight_matrix (one row per lineup, one column per player --
see calc_player_weights). See the section banner above for why
this is a plain matrix inverse (numpy.linalg.inv), not an SVD.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
player_weight_matrix | NDArray[float64] | The off/def design matrix, shape (num_lineups, ctx["num_players"]). | |
ridge_lambda | float | The Tikhonov regularization strength. | |
ctx | RapmPlayerContext | A RapmPlayerContext -- only ctx["num_players"] is read (sizes the identity matrix). |
Returns
The (num_players, num_lineups) solver matrix; apply it to a target vector via calculate_rapm.
Example
import numpy as np
from sportsdataverse.mbb.mbb_rapm import slow_regression, calculate_rapm
x = np.array([[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]])
solver = slow_regression(x, 1.0, ctx) # ctx["num_players"] == 2
rapm = calculate_rapm(solver, [1.0, 2.0, 3.0])
spearman_corr
spearman_corr(a: 'np.ndarray', b: 'np.ndarray') -> 'float'
Spearman rank correlation between two arrays.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
a | ndarray | First array of values. | |
b | ndarray | Second array of values (same length as a). |
Returns
The Spearman rank correlation coefficient.
Example
import numpy as np
from sportsdataverse._common.metrics import spearman_corr
spearman_corr(np.array([1, 2, 3]), np.array([3, 1, 2]))
talent_split_mse
talent_split_mse(scored: 'pl.DataFrame', *, k: 'float', seed: 'int' = 0) -> 'float'
Weighted MSE of the k-regressed first half predicting the raw second half.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
scored | DataFrame | mbb_shot_quality output. | |
k | float | Shrinkage pseudo-shots to evaluate. | |
seed | int | 0 | Split seed. |
Returns
sum(n_h2 * (oe_h1 * n_h1/(n_h1+k) - oe_h2)^2) / sum(n_h2).
Example
from sportsdataverse.mbb.mbb_shooter_talent import talent_split_mse
talent_split_mse(scored, k=200.0)
three_point_radius
three_point_radius(league: 'str', season: 'int') -> 'float'
Three-point arc radius (feet) for a league x season.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
league | str | "mens" or "womens". | |
season | int | Season-ending year (e.g. 2020 = 2019-20). |
Returns
Arc radius in feet.
Example
from sportsdataverse.mbb.mbb_shot_quality_constants import three_point_radius
three_point_radius("mens", 2020)
transfer_cohort
transfer_cohort(rosters: 'pl.DataFrame') -> 'pl.DataFrame'
One row per transfer: same player_id, different team_id in
consecutive seasons.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
rosters | DataFrame | Frame with player_id, team_id, season (extra columns ignored; one row per player-season-team). |
Returns
player_id: Utf8, from_team_id:Utf8, to_team_id:Utf8, from_season:Int64, to_season:Int64 -- a player transferring twice appears twice.
| col_name | type | description |
|---|---|---|
player_id | character | Unique player identifier. |
from_team_id | character | Unique identifier for from team. |
to_team_id | character | Unique identifier for to team. |
from_season | integer | |
to_season | integer |
Example
from sportsdataverse.mbb import mbb_box_bpm, transfer_cohort
bpm = mbb_box_bpm([2025, 2026]).filter(pl.col("min") >= 150)
moves = transfer_cohort(bpm.select("player_id", "team_id", "season"))
wbb_bracket_sim
wbb_bracket_sim(seeded_field: 'pl.DataFrame', ratings: 'pl.DataFrame', *, n_sims: 'int' = 10000, seed: 'int' = 0, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'
Women's single-elimination bracket Monte Carlo.
Delegates to sportsdataverse.mbb.mbb_season_sim.mbb_bracket_sim with league="womens".
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
seeded_field | DataFrame | Bracket-ordered rows with team_id (adjacent rows meet in round 1). | |
ratings | DataFrame | One row per team (team_id, adj_em). | |
n_sims | int | 10000 | Number of simulated brackets. |
seed | int | 0 | RNG seed (deterministic output). |
return_as_pandas | bool | False | Return a pandas DataFrame instead of polars. |
Returns
One row per field team: team_id, seed?, reach_r32 .. champion probabilities -- see the mbb core for the full contract.
No returns table is published for this function: no capture: it needs a bracket-ordered field whose size is a power of two, and no package function returns one.
Example
from sportsdataverse.wbb import wbb_bracket_sim
odds = wbb_bracket_sim(field_64, ratings, n_sims=20000, seed=42)
wbb_in_game_win_prob
wbb_in_game_win_prob(pbp: 'pl.DataFrame', pregame_home_prob: 'float', *, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'
Women's per-play in-game win probability (bundled wbb_in_game_wp.ubj).
Delegates to sportsdataverse.mbb.mbb_game_predict.mbb_in_game_win_prob with league="womens".
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
pbp | DataFrame | One game's plays in the load_wbb_pbp schema. | |
pregame_home_prob | float | Pregame home win probability. | |
return_as_pandas | bool | False | Return a pandas DataFrame instead of polars. |
Returns
One row per play: the five feature columns plus home_win_prob -- see the mbb core for the full contract.
| col_name | type | description |
|---|---|---|
score_diff | double | |
sec_left | double | |
sqrt_sec_left | double | |
pregame_logit | double | |
home_has_ball | integer | |
home_win_prob | double |
Example
from sportsdataverse.wbb import wbb_in_game_win_prob
wp = wbb_in_game_win_prob(pbp, 0.62)
wbb_predict_games
wbb_predict_games(games: 'pl.DataFrame', ratings: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'
Women's vectorized pregame predictions over a schedule.
Delegates to sportsdataverse.mbb.mbb_game_predict.mbb_predict_games with league="womens".
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
games | DataFrame | One row per game (game_id, home_team_id, away_team_id and optionally neutral_site). | |
ratings | DataFrame | One row per team (the wbb_team_ratings output). | |
return_as_pandas | bool | False | Return a pandas DataFrame instead of polars. |
Returns
One row per input game: game_id, home_team_id, away_team_id, exp_margin, home_win_prob, exp_total -- see the mbb core for the full contract.
No returns table is published for this function: no capture: it raises a join-key dtype mismatch on every real season (the schedule's home_team_id is Int32, the ratings' team_id is String).
Example
from sportsdataverse.wbb import wbb_predict_games, wbb_team_ratings
preds = wbb_predict_games(games, wbb_team_ratings(2024))
wbb_season_sim
wbb_season_sim(ratings: 'pl.DataFrame', remaining_schedule: 'pl.DataFrame', *, n_sims: 'int' = 10000, seed: 'int' = 0, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'
Women's remaining-schedule Monte Carlo.
Delegates to sportsdataverse.mbb.mbb_season_sim.mbb_season_sim with league="womens".
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
ratings | DataFrame | One row per team (season, team_id, adj_em + optional conference / current_wins). | |
remaining_schedule | DataFrame | Games to simulate. | |
n_sims | int | 10000 | Number of simulated seasons. |
seed | int | 0 | RNG seed (deterministic output). |
return_as_pandas | bool | False | Return a pandas DataFrame instead of polars. |
Returns
One row per team: season, team_id, exp_wins, playoff_prob, conf_title_prob -- see the mbb core for the full contract.
No returns table is published for this function: no capture: it raises KeyError on a real schedule (the schedule's team ids are integers, the ratings' are strings).
Example
from sportsdataverse.wbb import wbb_season_sim
odds = wbb_season_sim(ratings, remaining, n_sims=5000, seed=42)
wbb_team_ratings
wbb_team_ratings(seasons: 'Union[int, list[int]]', *, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'
Women's opponent-adjusted team ratings (AdjO/AdjD/AdjEM/AdjTempo).
Delegates to sportsdataverse.mbb.mbb_team_ratings.mbb_team_ratings
with league="womens" (WBB loaders + women's fitted constants).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
seasons | Union[int, list[int]] | A season (e.g. 2024) or list of seasons. | |
return_as_pandas | bool | False | Return a pandas frame instead of polars. |
Returns
One row per (season, team_id) -- see the mbb core for the schema.
| col_name | type | description |
|---|---|---|
season | integer | Season identifier (4-digit year or 'YYYY-YY' string). |
team_id | character | Unique team identifier. |
adj_o | double | Adj o. |
adj_d | double | Adj d. |
adj_em | double | Adj em. |
adj_tempo | double | |
raw_o | double | Raw o. |
raw_d | double | Raw d. |
games | integer | Games played. |
rank | integer | Rank. |
adj_em_z | double |
Example
from sportsdataverse.wbb import wbb_team_ratings
wbb_team_ratings(2024).sort("rank").head()
win_prob_from_margin
win_prob_from_margin(exp_margin: 'float') -> 'float'
Women's home win probability from an expected margin.
Delegates to sportsdataverse.mbb.mbb_game_predict.win_prob_from_margin with league="womens".
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
exp_margin | float | Expected home-minus-away margin in points. |
Returns
Probability the home team wins, in (0, 1).
Example
from sportsdataverse.wbb.wbb_game_predict import win_prob_from_margin
win_prob_from_margin(5.0)