Skip to main content

WBB — additional Python functions — Other: right_kind–win_prob

right_kind_of_shot​

right_kind_of_shot(shot: 'ShotEvent', pbp_event: 'MiscGameEvent', strict: 'bool') -> 'bool'

Whether pbp_event's shot type is compatible with shot's

distance and make/miss (ShotEnrichmentUtils.right_kind_of_shot, PlayByPlayUtils.scala:659-679).

The distance-in-the-data is approximate, so exact 2-vs-3 discrimination is impossible; this only rules out the obvious mismatches (a clearly-short shot matched to a 3, or vice versa) and always requires make/miss agreement.

Parameters

ParameterTypeDefaultDescription
shotShotEventThe shot being enriched (pts/dist read).
pbp_eventMiscGameEventThe candidate play-by-play event.
strictboolIf True, also apply the distance gate; if False, only the make/miss agreement is required.

Returns

True if the event could plausibly be this shot.

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import right_kind_of_shot
right_kind_of_shot(shot, pbp_event, strict=True)

roc_auc​

roc_auc(y_true: 'np.ndarray', score: 'np.ndarray') -> 'float'

Area under the ROC curve via the rank-sum (Mann-Whitney) identity.

Parameters

ParameterTypeDefaultDescription
y_truendarrayBinary outcomes (0/1).
scorendarrayPredicted scores (any monotone scale).

Returns

AUC in [0, 1]; nan when only one class is present.

Example

import numpy as np
from sportsdataverse.mbb.mbb_player_value_constants import roc_auc
roc_auc(np.array([0, 1]), np.array([0.2, 0.9]))

run_iterative_adjustment_with_hca​

run_iterative_adjustment_with_hca(teams: 'Sequence[TeamDetail]', team_by_name: 'dict[str, TeamDetail]', fields: 'Sequence[str]', league_averages: 'LeagueAverages', poss_splits: 'dict[str, PossessionSplits]', *, max_iterations: 'int' = 100, tolerance: 'float' = 1e-06) -> 'IterationResult'

KenPom-style SoS + HCA fixed-point solver (runIterativeAdjustmentWithHCA, ts:306-527).

Each iteration (Jacobi -- all teams read the previous iteration's adjustments, then commit together):

  1. Per team/field, adjust every game adj_game = raw_game * (league / (opp_adj +/- hca)) and take the weighted mean; a field with no valid games keeps its current value.
  2. Re-estimate per-field HCA from home/away possession-imbalance residuals hca = sum((raw - pred) * |imbalance|) / sum(|imbalance|) over teams with |imbalance| >= IMBALANCE_MIN.

Stops when the max per-team/field change drops below tolerance or after max_iterations sweeps (the HCA re-estimate still runs on the final sweep). The cross-guard on the per-game branch, the asymmetric residual prediction, and the cross-named opponent strengths are all preserved -- see the module docstring's landmine list.

Parameters

ParameterTypeDefaultDescription
teamsSequence[TeamDetail]The teams to solve over.
team_by_namedict[str, TeamDetail]{team_name: team_detail} for opponent lookup.
fieldsSequence[str]The stat fields to solve.
league_averagesLeagueAveragesOutput of compute_league_averages_from_per_game.
poss_splitsdict[str, PossessionSplits]{team_name: PossessionSplits }.
max_iterationsint100Iteration cap (default MAX_ITERATIONS; pin to 1 to inspect a single sweep).
tolerancefloat1e-06Convergence tolerance (default TOLERANCE).

Returns

An IterationResult (adj_values, hca_per_field).

Example

from sportsdataverse.mbb.mbb_ncaa_strength import (
STRENGTH_ADJUSTED_FIELDS,
compute_league_averages_from_per_game,
compute_possession_splits,
run_iterative_adjustment_with_hca,
)

by_name = {t["team_name"]: t for t in teams}
league = compute_league_averages_from_per_game(teams)
splits = {t["team_name"]: compute_possession_splits(t) for t in teams}
result = run_iterative_adjustment_with_hca(
teams, by_name, STRENGTH_ADJUSTED_FIELDS, league, splits,
)
print(result.hca_per_field["3p"]["hca_off"])

save_artifact​

save_artifact(name: 'str', obj: 'dict') -> 'None'

Write a bundled artifact (dev/fitter use -- writes into the source tree).

Parameters

ParameterTypeDefaultDescription
namestrArtifact stem, e.g. "mbb_box_bpm".
objdictJSON-serializable artifact payload.

Example

save_artifact("mbb_box_bpm", {"league": "mens", "coef": [0.1]})

score_to_tuple​

score_to_tuple(s: 'str') -> 'tuple[int, int]'

Parse a "scored-allowed" score string (ExtractorUtils.score_to_tuple,

ExtractorUtils.scala:107-113).

Scala's str match { case regex(s1, s2) => ... } on a compiled Regex requires the ENTIRE string to match (Regex.unapplySeq calls Matcher.matches(), not find()) -- ported here as re.fullmatch, not re.match/re.search.

Parameters

ParameterTypeDefaultDescription
sstrThe raw score string, e.g. "55-68".

Returns

(scored, allowed) as a tuple of ints, or (0, 0) if s doesn't fully match ([0-9]+)-([0-9]+).

Example

from sportsdataverse.mbb.mbb_ncaa_models import score_to_tuple

score_to_tuple("55-68") # (55, 68)
score_to_tuple("garbage") # (0, 0)

scoreboard_event_parsing​

scoreboard_event_parsing(event)

No description available.

Parameters

ParameterTypeDefaultDescription
event

select_contains​

select_contains(root: 'Tag', selector: 'str', text: 'str') -> 'list[Tag]'

JSoup root.select(sel + ":contains(text)"): candidates whose full

text (own + every descendant's) case-insensitively CONTAINS text as a plain substring -- not a regex (Task 5e.2 addition; see the module docstring's "Critical divergence" note).

JSoup's :contains() is documented case-insensitive substring containment; soupsieve's :-soup-contains() (the non-deprecated spelling of its :contains()) is case-SENSITIVE, with no case-insensitive variant of its own. Reproducing JSoup's actual semantics therefore needs this helper rather than :-soup-contains().

Parameters

ParameterTypeDefaultDescription
rootTagThe element to search within.
selectorstrA plain (soupsieve-legal) CSS selector for the structural part of the match (everything before :contains).
textstrThe plain substring each candidate's collapsed text must case-insensitively contain.

Returns

Every selector match whose jsoup_text case-insensitively contains text, in document order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_contains
soup = parse_html("<td>game date:</td><td>Location:</td>")
select_contains(soup, "td", "Game Date:") # [<td>game date:</td>]

select_matching​

select_matching(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'

JSoup root.select(sel + ":matches(regex)"): candidates whose full

text (own + every descendant's) matches regex.

Soupsieve has no :matches() pseudo-class equivalent, so this runs the plain structural selector first, then filters by re.search over each candidate's jsoup_text (own text plus descendants', matching JSoup's :matches() semantics -- as opposed to select_matching_own's own-text-only :matchesOwn()).

Parameters

ParameterTypeDefaultDescription
rootTagThe element to search within.
selectorstrA plain (soupsieve-legal) CSS selector.
regexstrThe pattern each candidate's collapsed text must re.search-match.

Returns

Every selector match whose jsoup_text contains a regex match, in document order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching
soup = parse_html("<div><p>Home Team</p><p>Away Team</p></div>")
select_matching(soup, "p", r"^Home") # [<p>Home Team</p>]

select_matching_own​

select_matching_own(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'

JSoup root.select(sel + ":matchesOwn(regex)"): candidates whose

OWN text only (excluding descendant elements' text) matches regex.

JSoup's Element.ownText() walks only the element's direct TextNode children, not text nested inside child elements -- the same distinction bs4 draws between a tag's direct bs4.NavigableString children and its full .get_text().

Parameters

ParameterTypeDefaultDescription
rootTagThe element to search within.
selectorstrA plain (soupsieve-legal) CSS selector.
regexstrThe pattern each candidate's own (whitespace-collapsed) text must re.search-match.

Returns

Every selector match whose own text contains a regex match, in document order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching_own
soup = parse_html('<div class="card-header">Coach <b>Info</b></div>')
select_matching_own(soup, "div.card-header", r"^Coach")
# [<div class="card-header">Coach <b>Info</b></div>]

shot_events_to_frame​

shot_events_to_frame(events: 'list[ShotEvent]', *, season: 'int', league: 'str' = 'mens') -> 'pl.DataFrame'

Flatten NCAA HTML ShotEvent objects to the canonical frame.

The NCAA SVG shot maps carry location + made/miss but no shot-type label (shot_type = "unknown"); point_value/shot_zone come from the geometry classifiers. The parser-phase pts field is the MADE flag (1/0), not the point value.

Parameters

ParameterTypeDefaultDescription
eventslist[ShotEvent]Parsed shot events (create_shot_event_data output).
seasonintSeason-ending year the events belong to.
leaguestr'mens'"mens" or "womens".

Returns

The canonical shot frame (CANONICAL_SHOT_SCHEMA); empty input returns the zero-row schema.

Example

from sportsdataverse.mbb.mbb_shots_adapter import shot_events_to_frame
df = shot_events_to_frame(events, season=2025)

shot_js_to_html​

shot_js_to_html(js: 'str') -> 'list[Tag]'

Converts client-side addShot(...) JS calls into parseable

circle.shot HTML, for pages where the shot map is built on the fly rather than baked into the initial HTML (ShotEventParser .shot_js_to_html, :266-283). See the module docstring's "Scala idiom decision" note -- the Scala's builders/browser parameters are dropped here since the Scala body never actually uses them.

Parameters

ParameterTypeDefaultDescription
jsstrThe concatenated <script> text containing one or more addShot(x, y, ..., 'title', ...) calls, one per line.

Returns

The circle.shot elements reconstructed from every matching line (non-matching lines, e.g. the addShot function definition line itself, are silently skipped).

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import shot_js_to_html
js = "addShot(27.0, 77.0, 392, false, 1, 'title text', 'class', false);"
circles = shot_js_to_html(js)

simulate_game​

simulate_game(home_em: 'float', away_em: 'float', neutral: 'bool', rng: 'np.random.Generator') -> 'bool'

Sample one women's game outcome (women's sigma/HFA/em_scale).

Parameters

ParameterTypeDefaultDescription
home_emfloat
away_emfloat
neutralbool
rngGenerator

Example

import numpy as np
from sportsdataverse.wbb.wbb_season_sim import simulate_game
simulate_game(20.0, 5.0, False, np.random.default_rng(0))

slow_regression​

slow_regression(player_weight_matrix: 'NDArray[np.float64]', ridge_lambda: 'float', ctx: 'RapmPlayerContext') -> 'NDArray[np.float64]'

Build the Tikhonov (ridge) regression solver matrix.

Faithful port of the private RapmUtils.slowRegression (RapmUtils.ts:756-769): (XᵀX + ridge_lambda·I)⁻¹Xᵀ, where X is player_weight_matrix (one row per lineup, one column per player -- see calc_player_weights). See the section banner above for why this is a plain matrix inverse (numpy.linalg.inv), not an SVD.

Parameters

ParameterTypeDefaultDescription
player_weight_matrixNDArray[float64]The off/def design matrix, shape (num_lineups, ctx["num_players"]).
ridge_lambdafloatThe Tikhonov regularization strength.
ctxRapmPlayerContextA RapmPlayerContext -- only ctx["num_players"] is read (sizes the identity matrix).

Returns

The (num_players, num_lineups) solver matrix; apply it to a target vector via calculate_rapm.

Example

import numpy as np
from sportsdataverse.mbb.mbb_rapm import slow_regression, calculate_rapm

x = np.array([[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]])
solver = slow_regression(x, 1.0, ctx) # ctx["num_players"] == 2
rapm = calculate_rapm(solver, [1.0, 2.0, 3.0])

spearman_corr​

spearman_corr(a: 'np.ndarray', b: 'np.ndarray') -> 'float'

Spearman rank correlation between two arrays.

Parameters

ParameterTypeDefaultDescription
andarrayFirst array of values.
bndarraySecond array of values (same length as a).

Returns

The Spearman rank correlation coefficient.

Example

import numpy as np
from sportsdataverse._common.metrics import spearman_corr
spearman_corr(np.array([1, 2, 3]), np.array([3, 1, 2]))

start_time_from_period​

start_time_from_period(period: 'int', is_women_game: 'bool') -> 'float'

The game-clock time (minutes elapsed) a period starts at

(ExtractorUtils.scala:272-281).

Women's games play four 10-minute quarters then 5-minute overtimes; men's games play two 20-minute halves then 5-minute overtimes.

Parameters

ParameterTypeDefaultDescription
periodintThe 1-indexed period number (1/2 = halves for men, 1-4 = quarters for women, 5+ = overtimes for both).
is_women_gameboolWhether to use the women's (quarters) or men's (halves) period schedule.

Returns

The game-clock minute the period begins at.

Example

from sportsdataverse.mbb.mbb_ncaa_stints import start_time_from_period
start_time_from_period(2, is_women_game=False) # 20.0 (men's 2nd half)
start_time_from_period(1, is_women_game=True) # 0.0 (women's 1st quarter)
start_time_from_period(6, is_women_game=False) # 45.0 (men's 2nd OT)

strength_of_schedule​

strength_of_schedule(results: 'pl.DataFrame', ratings: 'pl.DataFrame', *, league: 'str' = 'mens') -> 'pl.DataFrame'

Per-team SoS + Quad 1-4 record + WAB from completed games and ratings.

Parameters

ParameterTypeDefaultDescription
resultsDataFrameCompleted games with game_id, season, home_team_id, away_team_id, home_score, away_score, neutral_site.
ratingsDataFrameOne row per team with season, team_id, adj_em, rank (the mbb_team_ratings output). Team-id dtype must match results.
leaguestr'mens'"mens" or "womens" (quad thresholds, HFA, bubble EM).

Returns

One row per (season, team_id): season, team_id, sos, sos_rank, wab, quad1_w .. quad4_l, quality_wins. sos is the mean opponent adj_em (rank 1 = hardest schedule); quads follow the NET venue-adjusted opponent-rank thresholds; quality_wins is Quad-1 + Quad-2 wins; wab is actual wins minus a bubble-quality team's expected wins against the same schedule. Empty input returns the schema with zero rows.

Example

from sportsdataverse.mbb.mbb_strength_of_schedule import strength_of_schedule
resume = strength_of_schedule(results, ratings)

sum_event_stats​

sum_event_stats(lhs: 'LineupEventStats', rhs: 'LineupEventStats') -> 'LineupEventStats'

Field-wise add two :class:`~sportsdataverse.mbb.mbb_ncaa_models

.LineupEventStats (protected def sum_event_stats, LineupUtils.scala :1534-1622, debug-only -- the Scala's own docstring says "just used for debug"). The Scala builds this via shapeless.Generic` field-zipping; this port is an explicit field-by-field call since Python has no equivalent generic-programming machinery.

Parameters

ParameterTypeDefaultDescription
lhsLineupEventStatsThe left-hand stat tree.
rhsLineupEventStatsThe right-hand stat tree.

Returns

A new ~sportsdataverse.mbb.mbb_ncaa_models.LineupEventStats with every field summed (see the module's private sum_*helpers for theOptional`/nested-field summing rules).

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_event_stats
from sportsdataverse.mbb.mbb_ncaa_models import LineupEventStats

sum_event_stats(LineupEventStats.empty(), LineupEventStats.empty()).num_events

sum_shot_infos​

sum_shot_infos(shot_infos: 'list[PlayerShotInfo]') -> 'Optional[PlayerShotInfo]'

Field-wise sum a list of :class:`~sportsdataverse.mbb.mbb_ncaa_models

.PlayerShotInfo\ s (sum_shot_infos, LineupUtils.scala:1625-1655`, debug-only).

Parameters

ParameterTypeDefaultDescription
shot_infoslist[PlayerShotInfo]The list to combine, in order.

Returns

None if shot_infos is empty; the single element if there's exactly one; otherwise a left-fold of pairwise field-wise sums (reduceOption).

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_shot_infos
from sportsdataverse.mbb.mbb_ncaa_models import PlayerShotInfo

sum_shot_infos([PlayerShotInfo(ast_3pm=(1, 0, 0, 0, 0)), PlayerShotInfo(ast_3pm=(0, 1, 0, 0, 0))])

talent_split_mse​

talent_split_mse(scored: 'pl.DataFrame', *, k: 'float', seed: 'int' = 0) -> 'float'

Weighted MSE of the k-regressed first half predicting the raw second half.

Parameters

ParameterTypeDefaultDescription
scoredDataFramembb_shot_quality output.
kfloatShrinkage pseudo-shots to evaluate.
seedint0Split seed.

Returns

sum(n_h2 * (oe_h1 * n_h1/(n_h1+k) - oe_h2)^2) / sum(n_h2).

Example

from sportsdataverse.mbb.mbb_shooter_talent import talent_split_mse
talent_split_mse(scored, k=200.0)

td_at​

td_at(row: 'Tag', n: 'int') -> 'Optional[Tag]'

JSoup row >?> element("td:eq(n)"): the n-th <td> child.

Soupsieve has no :eq() positional pseudo-class (unlike JSoup), so this is a plain 0-indexed lookup into row.find_all("td"), guarded against an out-of-range index (JSoup's >?> returns None rather than raising when the selector matches nothing).

Parameters

ParameterTypeDefaultDescription
rowTagThe row (or other container) element to search.
nintThe 0-indexed <td> position.

Returns

The n-th <td> descendant, or None if row has fewer than n + 1 of them.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, td_at
soup = parse_html("<tr><td>A</td><td>B</td></tr>")
row = soup.find("tr")
td_at(row, 1).get_text() # "B"
td_at(row, 5) # None

test_positional_aware_filter​

test_positional_aware_filter(sorted_to_test: 'list[dict[str, str]]', pve_frags: 'list[dict[str, Any]]', nve_frags: 'list[dict[str, Any]]') -> 'bool'

Check a positional-aware filter (from build_positional_aware_filter)

against a sorted (order_lineup-ordered) lineup array.

Faithful port of PositionUtils.testPositionalAwareFilter (PositionUtils.ts:831-858). A fragment matches if any of its position-restricted slots (or, when pos is empty, any slot at all) has a code/id containing the fragment's filter text (case-insensitive substring match). Every positive fragment must match (vacuously true if there are none); no negative fragment may match (vacuously true if there are none).

Parameters

ParameterTypeDefaultDescription
sorted_to_testlist[dict[str, str]]The ordered lineup, each a {"id": ..., "code": ...} dict (as returned by order_lineup).
pve_fragslist[dict[str, Any]]Positive-filter fragments (must ALL match).
nve_fragslist[dict[str, Any]]Negative-filter fragments (NONE may match).

Returns

Whether the lineup satisfies both the positive and negative filters.

Example

::

from sportsdataverse.mbb.mbb_positions import test_positional_aware_filter
lineup = [{"code": "AnCowan", "id": "Cowan, Anthony"}]
test_positional_aware_filter(lineup, [{"filter": "cowan", "pos": []}], [])

three_point_radius​

three_point_radius(league: 'str', season: 'int') -> 'float'

Three-point arc radius (feet) for a league x season.

Parameters

ParameterTypeDefaultDescription
leaguestr"mens" or "womens".
seasonintSeason-ending year (e.g. 2020 = 2019-20).

Returns

Arc radius in feet.

Example

from sportsdataverse.mbb.mbb_shot_quality_constants import three_point_radius
three_point_radius("mens", 2020)

tidy_player​

tidy_player(p_in: 'str', ctx: 'TidyPlayerContext') -> 'tuple[str, TidyPlayerContext]'

Resolve a raw play-by-play name to its box-score full name, via an

ordered fallback chain (LineupErrorAnalysisUtils.tidy_player, :76-144). Order is semantic -- ported as ordered first-non-None:

  1. Cache hit (see the module docstring's asymmetric-cache note).
  2. Exact box-score code match.
  3. Unique truncated-code match (TidyPlayerContext.alt_all_players_map).
  4. Double-barrel-surname strip retry ("Smith-Jones" -> "Jones"), recursing into this same function.
  5. Initials (convert_from_initials).
  6. Jersey number (convert_from_digits, against ctx.box_lineup.players_out).
  7. Truncated-code + inserted-"j"-for-"junior" retry.
  8. Fuzzy match (fuzzy_box_match) -- skipped (and normalized to "Team") for the four team-stat-row sentinels.
  9. Identity fallthrough (the input is returned unresolved; a later, out-of-scope validation pass is expected to reject it).

Parameters

ParameterTypeDefaultDescription
p_instrThe raw play-by-play name.
ctxTidyPlayerContextThe lookup context (see build_tidy_player_context).

Returns

(resolved_name, updated_ctx) -- updated_ctx carries the new cache entry.

Example

from sportsdataverse.mbb.mbb_ncaa_names import build_tidy_player_context, tidy_player
ctx = build_tidy_player_context(box_lineup)
resolved_name, ctx = tidy_player("MITCHELL,M", ctx)

transfer_cohort​

transfer_cohort(rosters: 'pl.DataFrame') -> 'pl.DataFrame'

One row per transfer: same player_id, different team_id in

consecutive seasons.

Parameters

ParameterTypeDefaultDescription
rostersDataFrameFrame with player_id, team_id, season (extra columns ignored; one row per player-season-team).

Returns

player_id: Utf8, from_team_id:Utf8, to_team_id:Utf8, from_season:Int64, to_season:Int64 -- a player transferring twice appears twice.

Example

from sportsdataverse.mbb import mbb_box_bpm, transfer_cohort
bpm = mbb_box_bpm([2025, 2026]).filter(pl.col("min") >= 150)
moves = transfer_cohort(bpm.select("player_id", "team_id", "season"))

transform_shot_location​

transform_shot_location(x: 'float', y: 'float', second_half_switch: 'bool', team_shooting_left_in_first_period: 'bool', is_offensive: 'bool') -> 'tuple[float, float, float, float]'

Transforms a raw SVG pixel location into feet from the basket, always

oriented as if shooting towards the left goal (ShotEventParser .transform_shot_location, :588-620).

Parameters

ParameterTypeDefaultDescription
xfloatRaw SVG cx pixel coordinate.
yfloatRaw SVG cy pixel coordinate.
second_half_switchboolWhether this shot is in the "other" half of the game from team_shooting_left_in_first_period (each False factor below flips which side is treated as "left").
team_shooting_left_in_first_periodboolWhether the team under analysis shot towards the left goal in the first period (see is_team_shooting_left_to_start).
is_offensiveboolWhether the team under analysis is shooting (an opponent shot flips the expected side again).

Returns

(x, y, alt_x, alt_y) in feet -- the believed-correct location, then the alternative (mirror-image) location, both relative to the goal the shot is (believed to be) attacking.

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import transform_shot_location
transform_shot_location(310.2, 235, False, False, True)

update_config​

update_config(**kwargs: 'object') -> 'NcaaFetchConfig'

Update the active config in place.

Returns

The (mutated) global config object.

Example

from sportsdataverse.mbb.mbb_ncaa_fetch import update_config
update_config(proxy_url="http://user:pass@1.2.3.4:8080")

using_roster_pos​

using_roster_pos(pos_class: 'str', roster_pos: 'str | None') -> 'tuple[str, str | None]'

Reconcile a stats-derived position class against roster metadata.

Faithful port of PositionUtils.usingRosterPos (PositionUtils.ts:583-626). When the classifier landed on an "unsure" bucket ("G?"/"F/C?"), roster info narrows it (a roster "C" always wins outright); otherwise an obviously-wrong stats classification is compromised toward the roster-implied side, gated by pos_class_to_score thresholds.

Parameters

ParameterTypeDefaultDescription
pos_classstrThe stats-derived position class.
roster_posstr | NoneThe roster-reported position ("G"/"F"/"C"), or None/"" when unknown. if (rosterPos) (ts:587) is a plain JS truthiness check on a string -- "" and None behave identically (both mean "no correction"), so if not roster_pos is the faithful Python mirror, not an is None landmine.

Returns

A (position, info) tuple. info is None when no correction/explanation applies (matches the TS undefined), else a human-readable note on why the position was adjusted.

Example

from sportsdataverse.mbb.mbb_positions import using_roster_pos
using_roster_pos("G?", "C")

validate_box_score​

validate_box_score(team: 'TeamId', lineup: 'list[str]') -> 'Union[list[PlayerCodeId], ParseError]'

Checks there are no duplicates in the lineup (``BoxscoreParser

.validate_box_score, :388-404``).

Parameters

ParameterTypeDefaultDescription
teamTeamIdThe team the lineup belongs to (feeds ~sportsdataverse.mbb.mbb_ncaa_stints.build_player_code's team-scoped misspelling corrections).
lineuplist[str]The raw player-name strings, in whatever order they were assembled by inject_validated_players.

Returns

lineup, mapped to ~sportsdataverse.mbb.mbb_ncaa_models.PlayerCodeId (same order, no sort -- see the module docstring's "not sorted" note). When two teammates collide on the {first-two-letters}{Surname} scheme -- siblings, in practice -- only the colliding players are re-coded to {First}{Last} by disambiguate_sibling_codes; every other player keeps the Scala-faithful code. This is a DELIBERATE divergence from ExtractorUtils.scala, which rejects the game: since a team's roster is the same all season, one sibling pair cost the team its ENTIRE season of lineups. A ~sportsdataverse.mbb.mbb_ncaa_data_quality.ParseErroris returned only when widening cannot separate them, i.e. two players with the SAME full name -- genuinely ambiguous, so still an error. Callers must not re-derive a code from a name after this point:build_player_codewould undo the widening and silently drop one twin. Use~sportsdataverse.mbb.mbb_ncaa_names.code_from_box`, which resolves against this roster.

Example

from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import validate_box_score
from sportsdataverse.mbb.mbb_ncaa_models import TeamId
validate_box_score(TeamId("Team"), ["Player One", "Player Two"])

validate_lineup​

validate_lineup(lineup_event: 'LineupEvent', box_lineup: 'LineupEvent', valid_player_codes: 'set[str]') -> 'list[ValidationError]'

Flags a lineup stint as internally inconsistent, via 3 independent

checks (LineupErrorAnalysisUtils.validate_lineup, :181-218).

Parameters

ParameterTypeDefaultDescription
lineup_eventLineupEventThe lineup stint to validate.
box_lineupLineupEventThe team's box-score lineup event (players is the full roster) -- used both to build the name-resolution context (see ~sportsdataverse.mbb.mbb_ncaa_names.build_tidy_player_context) and, indirectly, as the source of players_out for jersey-number resolution inside ~sportsdataverse.mbb .mbb_ncaa_names.tidy_player.
valid_player_codesset[str]Every player code that's actually on the box score / roster for this team-season.

Returns

The failing ValidationError\ s, in declaration order (see the module docstring's "Return shape" note) -- empty if lineup_event is clean. * ValidationError.WRONG_NUMBER_OF_PLAYERS -- lineup_event doesn't have exactly 5 players on the floor. * ValidationError.UNKNOWN_PLAYERS -- some player on the floor isn't in valid_player_codes. * ValidationError.INACTIVE_PLAYERS -- some player mentioned in lineup_event's own (team-side) raw game events resolves to a code not in valid_player_codes (i.e. isn't on the floor, per the lineup being validated).

Example

from sportsdataverse.mbb.mbb_ncaa_stint_validation import validate_lineup
errors = validate_lineup(lineup_event, box_lineup, {"MiMitchell", "BbBob"})
assert not errors # a clean lineup returns []

weighted_avg​

weighted_avg(mutable_acc: 'LineupStatSet', obj: 'LineupStatSet') -> 'None'

Merge obj into mutable_acc with possession weighting.

Faithful port of LineupUtils.weightedAvg (LineupUtils.ts:645). Mutates mutable_acc in place (matching the upstream mutable-state contract) and returns None. Each call accumulates a weighted sum, not a weighted average -- the companion completeWeightedAvg (upstream LineupUtils.ts:752, not yet ported) divides by the accumulated weight totals to finish the average. The per-field weight used at each merge step is derived from obj's own totals (e.g. that single lineup's total_off_fga), not from any running total on mutable_acc -- callers accumulating many lineups must call weighted_avg once per lineup so every lineup contributes its own weight.

Parameters

ParameterTypeDefaultDescription
mutable_accLineupStatSetThe running accumulator (LineupStatSet). Mutated in place; fields absent from the accumulator are initialized to {"value": 0.0} (plus old_value / override when obj's field carries a luck-adjustment override marker) before obj's contribution is added.
objLineupStatSetThe per-lineup LineupStatSet document to merge in.

Returns

None. mutable_acc is mutated in place.

Example

from sportsdataverse.mbb.mbb_lineup_stats import weighted_avg

acc: dict = {}
weighted_avg(acc, lineup_a)
weighted_avg(acc, lineup_b)
print(acc["off_poss"]["value"]) # plain sum (SUM_FIELDS)

# Two-lineup possession-weighted merge

acc = {}
for lineup in three_lineups:
weighted_avg(acc, lineup)
# acc now holds weighted SUMS; complete_weighted_avg (not yet
# ported) is required to turn these into rate-stat averages.

win_prob_from_margin​

win_prob_from_margin(exp_margin: 'float') -> 'float'

Women's home win probability from an expected margin.

Delegates to sportsdataverse.mbb.mbb_game_predict.win_prob_from_margin with league="womens".

Parameters

ParameterTypeDefaultDescription
exp_marginfloatExpected home-minus-away margin in points.

Returns

Probability the home team wins, in (0, 1).

Example

from sportsdataverse.wbb.wbb_game_predict import win_prob_from_margin
win_prob_from_margin(5.0)