MBB — additional Python functions — Other: reorder_and–win_prob
reorder_and_reverse
reorder_and_reverse(reversed_partial_events: 'Iterable[PlayByPlayEvent]') -> 'list[PlayByPlayEvent]'
Orders same-minute play-by-play events so subs never enclose the plays
they logically precede/follow (ExtractorUtils.scala:435-599).
Groups consecutive events sharing the same min into a block (the
input arrives in descending/reverse-chronological order, so blocks are
discovered and internally accumulated in reverse too), then -- for any
block containing a sub -- reorders it via inner_sort: events
referencing a subbed-OUT player (or scoring no higher than the sub) land
in a pre-sub group, the subs themselves come next (in ascending-score
order), and events referencing a subbed-IN player (or scoring higher
than the sub) land in a trailing post-sub group. Free-throw attempts
sharing the sub's inferred "direction" (team vs. opponent, inferred from
the nearest preceding shot/FT/foul) are pulled into the pre-sub group
unless the shooter is one of the players being subbed in. Blocks with no
sub are returned unchanged apart from the initial score-based sort.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
reversed_partial_events | Iterable[PlayByPlayEvent] | Events for one lineup event, in reverse-chronological (descending-time) order -- the natural order encountered walking play-by-play text bottom-up. |
Returns
The same events, forward-chronological (ascending time), with each same-minute block internally reordered so no sub encloses a play it logically shouldn't.
Example
from sportsdataverse.mbb.mbb_ncaa_models import Score
from sportsdataverse.mbb.mbb_ncaa_stints import (
OtherTeamEvent,
SubInEvent,
reorder_and_reverse,
)
events = [
SubInEvent(0.4, Score(0, 0), "player1"),
OtherTeamEvent(0.4, Score(0, 0), "rebound"),
]
reorder_and_reverse(events)
# [OtherTeamEvent(...), SubInEvent(...)]
reset_config
reset_config() -> 'NcaaFetchConfig'
Reset the active config to its env-var-derived defaults.
Example
from sportsdataverse.mbb.mbb_ncaa_fetch import update_config, reset_config
update_config(timeout=5)
reset_config()
resolve_ncaa_team_id
resolve_ncaa_team_id(team: 'str', season: 'str', league: 'str' = 'mbb') -> 'Optional[int]'
Resolve a school name + season to its stats.ncaa.org team id.
Exact (team, season) match first (bigballR semantics), then a
case-insensitive fallback.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
team | str | School name as it appears in the crosswalk ("Illinois", not "Illinois Fighting Illini"). | |
season | str | Season string, e.g. "2025-26". | |
league | str | 'mbb' | "mbb" or "wbb" -- which league's crosswalk to search. (Deliberate fix of wbigballR, which always searched the men's table.) |
Returns
The team id as int, or None when no row matches.
Example
from sportsdataverse.mbb import resolve_ncaa_team_id
resolve_ncaa_team_id("Illinois", "2025-26")
# Women's league lookup
resolve_ncaa_team_id("South Carolina", "2024-25", league="wbb")
ridge_cv_lambda
ridge_cv_lambda(X: 'np.ndarray', y: 'np.ndarray', groups: 'np.ndarray', lams: "'list[float]'") -> 'float'
Pick lambda by leave-one-group-out CV (groups = seasons/classes).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
X | ndarray | Feature matrix. | |
y | ndarray | Targets. | |
groups | ndarray | Group label per row (e.g. season); each held out once. | |
lams | list[float] | Candidate penalties. |
Returns
The candidate with the lowest mean held-out MSE.
Example
lam = ridge_cv_lambda(X, y, seasons, [0.1, 1, 10, 100])
ridge_fit
ridge_fit(X: 'np.ndarray', y: 'np.ndarray', lam: 'float') -> 'np.ndarray'
Closed-form ridge with an unpenalized intercept (coefficient 0).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
X | ndarray | Feature matrix (n, d). | |
y | ndarray | Targets (n,). | |
lam | float | L2 penalty on the non-intercept coefficients. |
Returns
Coefficient vector of length d + 1 (intercept first).
Example
import numpy as np
from sportsdataverse.mbb.mbb_player_value_constants import ridge_fit
beta = ridge_fit(np.random.rand(50, 3), np.random.rand(50), lam=1.0)
right_kind_of_shot
right_kind_of_shot(shot: 'ShotEvent', pbp_event: 'MiscGameEvent', strict: 'bool') -> 'bool'
Whether pbp_event's shot type is compatible with shot's
distance and make/miss (ShotEnrichmentUtils.right_kind_of_shot,
PlayByPlayUtils.scala:659-679).
The distance-in-the-data is approximate, so exact 2-vs-3 discrimination is impossible; this only rules out the obvious mismatches (a clearly-short shot matched to a 3, or vice versa) and always requires make/miss agreement.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
shot | ShotEvent | The shot being enriched (pts/dist read). | |
pbp_event | MiscGameEvent | The candidate play-by-play event. | |
strict | bool | If True, also apply the distance gate; if False, only the make/miss agreement is required. |
Returns
True if the event could plausibly be this shot.
Example
from sportsdataverse.mbb.mbb_ncaa_pbp_glue import right_kind_of_shot
right_kind_of_shot(shot, pbp_event, strict=True)
roc_auc
roc_auc(y_true: 'np.ndarray', score: 'np.ndarray') -> 'float'
Area under the ROC curve via the rank-sum (Mann-Whitney) identity.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
y_true | ndarray | Binary outcomes (0/1). | |
score | ndarray | Predicted scores (any monotone scale). |
Returns
AUC in [0, 1]; nan when only one class is present.
Example
import numpy as np
from sportsdataverse.mbb.mbb_player_value_constants import roc_auc
roc_auc(np.array([0, 1]), np.array([0.2, 0.9]))
run_iterative_adjustment_with_hca
run_iterative_adjustment_with_hca(teams: 'Sequence[TeamDetail]', team_by_name: 'dict[str, TeamDetail]', fields: 'Sequence[str]', league_averages: 'LeagueAverages', poss_splits: 'dict[str, PossessionSplits]', *, max_iterations: 'int' = 100, tolerance: 'float' = 1e-06) -> 'IterationResult'
KenPom-style SoS + HCA fixed-point solver (runIterativeAdjustmentWithHCA, ts:306-527).
Each iteration (Jacobi -- all teams read the previous iteration's adjustments, then commit together):
- Per team/field, adjust every game
adj_game = raw_game * (league / (opp_adj +/- hca))and take the weighted mean; a field with no valid games keeps its current value. - Re-estimate per-field HCA from home/away possession-imbalance residuals
hca = sum((raw - pred) * |imbalance|) / sum(|imbalance|)over teams with|imbalance| >= IMBALANCE_MIN.
Stops when the max per-team/field change drops below tolerance or after
max_iterations sweeps (the HCA re-estimate still runs on the final
sweep). The cross-guard on the per-game branch, the asymmetric residual
prediction, and the cross-named opponent strengths are all preserved -- see
the module docstring's landmine list.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
teams | Sequence[TeamDetail] | The teams to solve over. | |
team_by_name | dict[str, TeamDetail] | {team_name: team_detail} for opponent lookup. | |
fields | Sequence[str] | The stat fields to solve. | |
league_averages | LeagueAverages | Output of compute_league_averages_from_per_game. | |
poss_splits | dict[str, PossessionSplits] | {team_name: PossessionSplits }. | |
max_iterations | int | 100 | Iteration cap (default MAX_ITERATIONS; pin to 1 to inspect a single sweep). |
tolerance | float | 1e-06 | Convergence tolerance (default TOLERANCE). |
Returns
An IterationResult (adj_values, hca_per_field).
Example
from sportsdataverse.mbb.mbb_ncaa_strength import (
STRENGTH_ADJUSTED_FIELDS,
compute_league_averages_from_per_game,
compute_possession_splits,
run_iterative_adjustment_with_hca,
)
by_name = {t["team_name"]: t for t in teams}
league = compute_league_averages_from_per_game(teams)
splits = {t["team_name"]: compute_possession_splits(t) for t in teams}
result = run_iterative_adjustment_with_hca(
teams, by_name, STRENGTH_ADJUSTED_FIELDS, league, splits,
)
print(result.hca_per_field["3p"]["hca_off"])
same_school
same_school(a: 'str', b: 'str') -> 'bool'
Whether two team-name spellings denote the same school.
Exact match, or both names inside one team_name_equivalents class.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
a | str | One team-name spelling. | |
b | str | The other team-name spelling. |
Returns
True if the two names refer to the same school.
Example
from sportsdataverse.mbb.mbb_ncaa_data_quality import same_school
same_school("New Orleans", "LSU New Orleans") # True
same_school("Miami (FL)", "Miami (OH)") # False
save_artifact
save_artifact(name: 'str', obj: 'dict') -> 'None'
Write a bundled artifact (dev/fitter use -- writes into the source tree).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
name | str | Artifact stem, e.g. "mbb_box_bpm". | |
obj | dict | JSON-serializable artifact payload. |
Example
save_artifact("mbb_box_bpm", {"league": "mens", "coef": [0.1]})
score_to_tuple
score_to_tuple(s: 'str') -> 'tuple[int, int]'
Parse a "scored-allowed" score string (ExtractorUtils.score_to_tuple,
ExtractorUtils.scala:107-113).
Scala's str match { case regex(s1, s2) => ... } on a compiled
Regex requires the ENTIRE string to match (Regex.unapplySeq calls
Matcher.matches(), not find()) -- ported here as
re.fullmatch, not re.match/re.search.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
s | str | The raw score string, e.g. "55-68". |
Returns
(scored, allowed) as a tuple of ints, or (0, 0) if s doesn't fully match ([0-9]+)-([0-9]+).
Example
from sportsdataverse.mbb.mbb_ncaa_models import score_to_tuple
score_to_tuple("55-68") # (55, 68)
score_to_tuple("garbage") # (0, 0)
scoreboard_event_parsing
scoreboard_event_parsing(event)
No description available.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
event |
select_contains
select_contains(root: 'Tag', selector: 'str', text: 'str') -> 'list[Tag]'
JSoup root.select(sel + ":contains(text)"): candidates whose full
text (own + every descendant's) case-insensitively CONTAINS text as
a plain substring -- not a regex (Task 5e.2 addition; see the module
docstring's "Critical divergence" note).
JSoup's :contains() is documented case-insensitive substring
containment; soupsieve's :-soup-contains() (the non-deprecated
spelling of its :contains()) is case-SENSITIVE, with no
case-insensitive variant of its own. Reproducing JSoup's actual
semantics therefore needs this helper rather than :-soup-contains().
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
root | Tag | The element to search within. | |
selector | str | A plain (soupsieve-legal) CSS selector for the structural part of the match (everything before :contains). | |
text | str | The plain substring each candidate's collapsed text must case-insensitively contain. |
Returns
Every selector match whose jsoup_text case-insensitively contains text, in document order.
Example
from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_contains
soup = parse_html("<td>game date:</td><td>Location:</td>")
select_contains(soup, "td", "Game Date:") # [<td>game date:</td>]
select_matching
select_matching(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'
JSoup root.select(sel + ":matches(regex)"): candidates whose full
text (own + every descendant's) matches regex.
Soupsieve has no :matches() pseudo-class equivalent, so this runs the
plain structural selector first, then filters by re.search
over each candidate's jsoup_text (own text plus descendants',
matching JSoup's :matches() semantics -- as opposed to
select_matching_own's own-text-only :matchesOwn()).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
root | Tag | The element to search within. | |
selector | str | A plain (soupsieve-legal) CSS selector. | |
regex | str | The pattern each candidate's collapsed text must re.search-match. |
Returns
Every selector match whose jsoup_text contains a regex match, in document order.
Example
from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching
soup = parse_html("<div><p>Home Team</p><p>Away Team</p></div>")
select_matching(soup, "p", r"^Home") # [<p>Home Team</p>]
select_matching_own
select_matching_own(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'
JSoup root.select(sel + ":matchesOwn(regex)"): candidates whose
OWN text only (excluding descendant elements' text) matches regex.
JSoup's Element.ownText() walks only the element's direct
TextNode children, not text nested inside child elements -- the
same distinction bs4 draws between a tag's direct
bs4.NavigableString children and its full .get_text().
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
root | Tag | The element to search within. | |
selector | str | A plain (soupsieve-legal) CSS selector. | |
regex | str | The pattern each candidate's own (whitespace-collapsed) text must re.search-match. |
Returns
Every selector match whose own text contains a regex match, in document order.
Example
from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching_own
soup = parse_html('<div class="card-header">Coach <b>Info</b></div>')
select_matching_own(soup, "div.card-header", r"^Coach")
# [<div class="card-header">Coach <b>Info</b></div>]
shot_events_to_frame
shot_events_to_frame(events: 'list[ShotEvent]', *, season: 'int', league: 'str' = 'mens') -> 'pl.DataFrame'
Flatten NCAA HTML ShotEvent objects to the canonical frame.
The NCAA SVG shot maps carry location + made/miss but no shot-type label
(shot_type = "unknown"); point_value/shot_zone come from the
geometry classifiers. The parser-phase pts field is the MADE flag
(1/0), not the point value.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
events | list[ShotEvent] | Parsed shot events (create_shot_event_data output). | |
season | int | Season-ending year the events belong to. | |
league | str | 'mens' | "mens" or "womens". |
Returns
The canonical shot frame (CANONICAL_SHOT_SCHEMA); empty input returns the zero-row schema.
Example
from sportsdataverse.mbb.mbb_shots_adapter import shot_events_to_frame
df = shot_events_to_frame(events, season=2025)
shot_js_to_html
shot_js_to_html(js: 'str') -> 'list[Tag]'
Converts client-side addShot(...) JS calls into parseable
circle.shot HTML, for pages where the shot map is built on the fly
rather than baked into the initial HTML (ShotEventParser .shot_js_to_html, :266-283). See the module docstring's "Scala
idiom decision" note -- the Scala's builders/browser parameters
are dropped here since the Scala body never actually uses them.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
js | str | The concatenated <script> text containing one or more addShot(x, y, ..., 'title', ...) calls, one per line. |
Returns
The circle.shot elements reconstructed from every matching line (non-matching lines, e.g. the addShot function definition line itself, are silently skipped).
Example
from sportsdataverse.mbb.mbb_ncaa_shot_parser import shot_js_to_html
js = "addShot(27.0, 77.0, 392, false, 1, 'title text', 'class', false);"
circles = shot_js_to_html(js)
simulate_game
simulate_game(home_em: 'float', away_em: 'float', neutral: 'bool', rng: 'np.random.Generator', *, league: 'str' = 'mens') -> 'bool'
Sample one game outcome: margin ~ Normal(exp_margin, margin_sd).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
home_em | float | Home team's adjusted efficiency margin. | |
away_em | float | Away team's adjusted efficiency margin. | |
neutral | bool | True for a neutral-site game. | |
rng | Generator | A seeded numpy.random.Generator (caller owns determinism). | |
league | str | 'mens' | "mens" or "womens". |
Returns
True if the home team wins the sampled game.
Example
import numpy as np
from sportsdataverse.mbb.mbb_season_sim import simulate_game
simulate_game(20.0, 5.0, False, np.random.default_rng(0))
slow_regression
slow_regression(player_weight_matrix: 'NDArray[np.float64]', ridge_lambda: 'float', ctx: 'RapmPlayerContext') -> 'NDArray[np.float64]'
Build the Tikhonov (ridge) regression solver matrix.
Faithful port of the private RapmUtils.slowRegression
(RapmUtils.ts:756-769): (XᵀX + ridge_lambda·I)⁻¹Xᵀ, where X
is player_weight_matrix (one row per lineup, one column per player --
see calc_player_weights). See the section banner above for why
this is a plain matrix inverse (numpy.linalg.inv), not an SVD.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
player_weight_matrix | NDArray[float64] | The off/def design matrix, shape (num_lineups, ctx["num_players"]). | |
ridge_lambda | float | The Tikhonov regularization strength. | |
ctx | RapmPlayerContext | A RapmPlayerContext -- only ctx["num_players"] is read (sizes the identity matrix). |
Returns
The (num_players, num_lineups) solver matrix; apply it to a target vector via calculate_rapm.
Example
import numpy as np
from sportsdataverse.mbb.mbb_rapm import slow_regression, calculate_rapm
x = np.array([[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]])
solver = slow_regression(x, 1.0, ctx) # ctx["num_players"] == 2
rapm = calculate_rapm(solver, [1.0, 2.0, 3.0])
spearman_corr
spearman_corr(a: 'np.ndarray', b: 'np.ndarray') -> 'float'
Spearman rank correlation between two arrays.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
a | ndarray | First array of values. | |
b | ndarray | Second array of values (same length as a). |
Returns
The Spearman rank correlation coefficient.
Example
import numpy as np
from sportsdataverse._common.metrics import spearman_corr
spearman_corr(np.array([1, 2, 3]), np.array([3, 1, 2]))
start_time_from_period
start_time_from_period(period: 'int', is_women_game: 'bool') -> 'float'
The game-clock time (minutes elapsed) a period starts at
(ExtractorUtils.scala:272-281).
Women's games play four 10-minute quarters then 5-minute overtimes; men's games play two 20-minute halves then 5-minute overtimes.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
period | int | The 1-indexed period number (1/2 = halves for men, 1-4 = quarters for women, 5+ = overtimes for both). | |
is_women_game | bool | Whether to use the women's (quarters) or men's (halves) period schedule. |
Returns
The game-clock minute the period begins at.
Example
from sportsdataverse.mbb.mbb_ncaa_stints import start_time_from_period
start_time_from_period(2, is_women_game=False) # 20.0 (men's 2nd half)
start_time_from_period(1, is_women_game=True) # 0.0 (women's 1st quarter)
start_time_from_period(6, is_women_game=False) # 45.0 (men's 2nd OT)
strength_of_schedule
strength_of_schedule(results: 'pl.DataFrame', ratings: 'pl.DataFrame', *, league: 'str' = 'mens') -> 'pl.DataFrame'
Per-team SoS + Quad 1-4 record + WAB from completed games and ratings.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
results | DataFrame | Completed games with game_id, season, home_team_id, away_team_id, home_score, away_score, neutral_site. | |
ratings | DataFrame | One row per team with season, team_id, adj_em, rank (the mbb_team_ratings output). Team-id dtype must match results. | |
league | str | 'mens' | "mens" or "womens" (quad thresholds, HFA, bubble EM). |
Returns
One row per (season, team_id): season, team_id, sos, sos_rank, wab, quad1_w .. quad4_l, quality_wins. sos is the mean opponent adj_em (rank 1 = hardest schedule); quads follow the NET venue-adjusted opponent-rank thresholds; quality_wins is Quad-1 + Quad-2 wins; wab is actual wins minus a bubble-quality team's expected wins against the same schedule. Empty input returns the schema with zero rows.
Example
from sportsdataverse.mbb.mbb_strength_of_schedule import strength_of_schedule
resume = strength_of_schedule(results, ratings)
sum_event_stats
sum_event_stats(lhs: 'LineupEventStats', rhs: 'LineupEventStats') -> 'LineupEventStats'
Field-wise add two :class:`~sportsdataverse.mbb.mbb_ncaa_models
.LineupEventStats (protected def sum_event_stats, LineupUtils.scala
:1534-1622, debug-only -- the Scala's own docstring says "just used for debug"). The Scala builds this via shapeless.Generic` field-zipping;
this port is an explicit field-by-field call since Python has no
equivalent generic-programming machinery.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
lhs | LineupEventStats | The left-hand stat tree. | |
rhs | LineupEventStats | The right-hand stat tree. |
Returns
A new ~sportsdataverse.mbb.mbb_ncaa_models.LineupEventStats with every field summed (see the module's private sum_*helpers for theOptional`/nested-field summing rules).
Example
from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_event_stats
from sportsdataverse.mbb.mbb_ncaa_models import LineupEventStats
sum_event_stats(LineupEventStats.empty(), LineupEventStats.empty()).num_events
sum_shot_infos
sum_shot_infos(shot_infos: 'list[PlayerShotInfo]') -> 'Optional[PlayerShotInfo]'
Field-wise sum a list of :class:`~sportsdataverse.mbb.mbb_ncaa_models
.PlayerShotInfo\ s (sum_shot_infos, LineupUtils.scala:1625-1655`,
debug-only).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
shot_infos | list[PlayerShotInfo] | The list to combine, in order. |
Returns
None if shot_infos is empty; the single element if there's exactly one; otherwise a left-fold of pairwise field-wise sums (reduceOption).
Example
from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_shot_infos
from sportsdataverse.mbb.mbb_ncaa_models import PlayerShotInfo
sum_shot_infos([PlayerShotInfo(ast_3pm=(1, 0, 0, 0, 0)), PlayerShotInfo(ast_3pm=(0, 1, 0, 0, 0))])
talent_split_mse
talent_split_mse(scored: 'pl.DataFrame', *, k: 'float', seed: 'int' = 0) -> 'float'
Weighted MSE of the k-regressed first half predicting the raw second half.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
scored | DataFrame | mbb_shot_quality output. | |
k | float | Shrinkage pseudo-shots to evaluate. | |
seed | int | 0 | Split seed. |
Returns
sum(n_h2 * (oe_h1 * n_h1/(n_h1+k) - oe_h2)^2) / sum(n_h2).
Example
from sportsdataverse.mbb.mbb_shooter_talent import talent_split_mse
talent_split_mse(scored, k=200.0)
td_at
td_at(row: 'Tag', n: 'int') -> 'Optional[Tag]'
JSoup row >?> element("td:eq(n)"): the n-th <td> child.
Soupsieve has no :eq() positional pseudo-class (unlike JSoup), so
this is a plain 0-indexed lookup into row.find_all("td"), guarded
against an out-of-range index (JSoup's >?> returns None rather
than raising when the selector matches nothing).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
row | Tag | The row (or other container) element to search. | |
n | int | The 0-indexed <td> position. |
Returns
The n-th <td> descendant, or None if row has fewer than n + 1 of them.
Example
from sportsdataverse.mbb.mbb_ncaa_html import parse_html, td_at
soup = parse_html("<tr><td>A</td><td>B</td></tr>")
row = soup.find("tr")
td_at(row, 1).get_text() # "B"
td_at(row, 5) # None
test_positional_aware_filter
test_positional_aware_filter(sorted_to_test: 'list[dict[str, str]]', pve_frags: 'list[dict[str, Any]]', nve_frags: 'list[dict[str, Any]]') -> 'bool'
Check a positional-aware filter (from build_positional_aware_filter)
against a sorted (order_lineup-ordered) lineup array.
Faithful port of PositionUtils.testPositionalAwareFilter
(PositionUtils.ts:831-858). A fragment matches if any of its
position-restricted slots (or, when pos is empty, any slot at all)
has a code/id containing the fragment's filter text
(case-insensitive substring match). Every positive fragment must match
(vacuously true if there are none); no negative fragment may match
(vacuously true if there are none).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
sorted_to_test | list[dict[str, str]] | The ordered lineup, each a {"id": ..., "code": ...} dict (as returned by order_lineup). | |
pve_frags | list[dict[str, Any]] | Positive-filter fragments (must ALL match). | |
nve_frags | list[dict[str, Any]] | Negative-filter fragments (NONE may match). |
Returns
Whether the lineup satisfies both the positive and negative filters.
Example
::
from sportsdataverse.mbb.mbb_positions import test_positional_aware_filter
lineup = [{"code": "AnCowan", "id": "Cowan, Anthony"}]
test_positional_aware_filter(lineup, [{"filter": "cowan", "pos": []}], [])
tidy_player
tidy_player(p_in: 'str', ctx: 'TidyPlayerContext') -> 'tuple[str, TidyPlayerContext]'
Resolve a raw play-by-play name to its box-score full name, via an
ordered fallback chain (LineupErrorAnalysisUtils.tidy_player,
:76-144). Order is semantic -- ported as ordered first-non-None:
- Cache hit (see the module docstring's asymmetric-cache note).
- Exact box-score code match.
- Unique truncated-code match (
TidyPlayerContext.alt_all_players_map). - Double-barrel-surname strip retry (
"Smith-Jones"->"Jones"), recursing into this same function. - Initials (
convert_from_initials). - Jersey number (
convert_from_digits, againstctx.box_lineup.players_out). - Truncated-code + inserted-"j"-for-"junior" retry.
- Fuzzy match (
fuzzy_box_match) -- skipped (and normalized to"Team") for the four team-stat-row sentinels. - Identity fallthrough (the input is returned unresolved; a later, out-of-scope validation pass is expected to reject it).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
p_in | str | The raw play-by-play name. | |
ctx | TidyPlayerContext | The lookup context (see build_tidy_player_context). |
Returns
(resolved_name, updated_ctx) -- updated_ctx carries the new cache entry.
Example
from sportsdataverse.mbb.mbb_ncaa_names import build_tidy_player_context, tidy_player
ctx = build_tidy_player_context(box_lineup)
resolved_name, ctx = tidy_player("MITCHELL,M", ctx)
transfer_cohort
transfer_cohort(rosters: 'pl.DataFrame') -> 'pl.DataFrame'
One row per transfer: same player_id, different team_id in
consecutive seasons.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
rosters | DataFrame | Frame with player_id, team_id, season (extra columns ignored; one row per player-season-team). |
Returns
player_id: Utf8, from_team_id:Utf8, to_team_id:Utf8, from_season:Int64, to_season:Int64 -- a player transferring twice appears twice.
Example
from sportsdataverse.mbb import mbb_box_bpm, transfer_cohort
bpm = mbb_box_bpm([2025, 2026]).filter(pl.col("min") >= 150)
moves = transfer_cohort(bpm.select("player_id", "team_id", "season"))
transform_shot_location
transform_shot_location(x: 'float', y: 'float', second_half_switch: 'bool', team_shooting_left_in_first_period: 'bool', is_offensive: 'bool') -> 'tuple[float, float, float, float]'
Transforms a raw SVG pixel location into feet from the basket, always
oriented as if shooting towards the left goal (ShotEventParser .transform_shot_location, :588-620).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
x | float | Raw SVG cx pixel coordinate. | |
y | float | Raw SVG cy pixel coordinate. | |
second_half_switch | bool | Whether this shot is in the "other" half of the game from team_shooting_left_in_first_period (each False factor below flips which side is treated as "left"). | |
team_shooting_left_in_first_period | bool | Whether the team under analysis shot towards the left goal in the first period (see is_team_shooting_left_to_start). | |
is_offensive | bool | Whether the team under analysis is shooting (an opponent shot flips the expected side again). |
Returns
(x, y, alt_x, alt_y) in feet -- the believed-correct location, then the alternative (mirror-image) location, both relative to the goal the shot is (believed to be) attacking.
Example
from sportsdataverse.mbb.mbb_ncaa_shot_parser import transform_shot_location
transform_shot_location(310.2, 235, False, False, True)
update_config
update_config(**kwargs: 'object') -> 'NcaaFetchConfig'
Update the active config in place.
Returns
The (mutated) global config object.
Example
from sportsdataverse.mbb.mbb_ncaa_fetch import update_config
update_config(proxy_url="http://user:pass@1.2.3.4:8080")
using_roster_pos
using_roster_pos(pos_class: 'str', roster_pos: 'str | None') -> 'tuple[str, str | None]'
Reconcile a stats-derived position class against roster metadata.
Faithful port of PositionUtils.usingRosterPos (PositionUtils.ts:583-626).
When the classifier landed on an "unsure" bucket ("G?"/"F/C?"),
roster info narrows it (a roster "C" always wins outright); otherwise
an obviously-wrong stats classification is compromised toward the
roster-implied side, gated by pos_class_to_score thresholds.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
pos_class | str | The stats-derived position class. | |
roster_pos | str | None | The roster-reported position ("G"/"F"/"C"), or None/"" when unknown. if (rosterPos) (ts:587) is a plain JS truthiness check on a string -- "" and None behave identically (both mean "no correction"), so if not roster_pos is the faithful Python mirror, not an is None landmine. |
Returns
A (position, info) tuple. info is None when no correction/explanation applies (matches the TS undefined), else a human-readable note on why the position was adjusted.
Example
from sportsdataverse.mbb.mbb_positions import using_roster_pos
using_roster_pos("G?", "C")
validate_box_score
validate_box_score(team: 'TeamId', lineup: 'list[str]') -> 'Union[list[PlayerCodeId], ParseError]'
Checks there are no duplicates in the lineup (``BoxscoreParser
.validate_box_score, :388-404``).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
team | TeamId | The team the lineup belongs to (feeds ~sportsdataverse.mbb.mbb_ncaa_stints.build_player_code's team-scoped misspelling corrections). | |
lineup | list[str] | The raw player-name strings, in whatever order they were assembled by inject_validated_players. |
Returns
lineup, mapped to ~sportsdataverse.mbb.mbb_ncaa_models.PlayerCodeId (same order, no sort -- see the module docstring's "not sorted" note). When two teammates collide on the {first-two-letters}{Surname} scheme -- siblings, in practice -- only the colliding players are re-coded to {First}{Last} by disambiguate_sibling_codes; every other player keeps the Scala-faithful code. This is a DELIBERATE divergence from ExtractorUtils.scala, which rejects the game: since a team's roster is the same all season, one sibling pair cost the team its ENTIRE season of lineups. A ~sportsdataverse.mbb.mbb_ncaa_data_quality.ParseErroris returned only when widening cannot separate them, i.e. two players with the SAME full name -- genuinely ambiguous, so still an error. Callers must not re-derive a code from a name after this point:build_player_codewould undo the widening and silently drop one twin. Use~sportsdataverse.mbb.mbb_ncaa_names.code_from_box`, which resolves against this roster.
Example
from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import validate_box_score
from sportsdataverse.mbb.mbb_ncaa_models import TeamId
validate_box_score(TeamId("Team"), ["Player One", "Player Two"])
validate_lineup
validate_lineup(lineup_event: 'LineupEvent', box_lineup: 'LineupEvent', valid_player_codes: 'set[str]') -> 'list[ValidationError]'
Flags a lineup stint as internally inconsistent, via 3 independent
checks (LineupErrorAnalysisUtils.validate_lineup, :181-218).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
lineup_event | LineupEvent | The lineup stint to validate. | |
box_lineup | LineupEvent | The team's box-score lineup event (players is the full roster) -- used both to build the name-resolution context (see ~sportsdataverse.mbb.mbb_ncaa_names.build_tidy_player_context) and, indirectly, as the source of players_out for jersey-number resolution inside ~sportsdataverse.mbb .mbb_ncaa_names.tidy_player. | |
valid_player_codes | set[str] | Every player code that's actually on the box score / roster for this team-season. |
Returns
The failing ValidationError\ s, in declaration order (see the module docstring's "Return shape" note) -- empty if lineup_event is clean. * ValidationError.WRONG_NUMBER_OF_PLAYERS -- lineup_event doesn't have exactly 5 players on the floor. * ValidationError.UNKNOWN_PLAYERS -- some player on the floor isn't in valid_player_codes. * ValidationError.INACTIVE_PLAYERS -- some player mentioned in lineup_event's own (team-side) raw game events resolves to a code not in valid_player_codes (i.e. isn't on the floor, per the lineup being validated).
Example
from sportsdataverse.mbb.mbb_ncaa_stint_validation import validate_lineup
errors = validate_lineup(lineup_event, box_lineup, {"MiMitchell", "BbBob"})
assert not errors # a clean lineup returns []
weighted_avg
weighted_avg(mutable_acc: 'LineupStatSet', obj: 'LineupStatSet') -> 'None'
Merge obj into mutable_acc with possession weighting.
Faithful port of LineupUtils.weightedAvg (LineupUtils.ts:645).
Mutates mutable_acc in place (matching the upstream mutable-state
contract) and returns None. Each call accumulates a weighted
sum, not a weighted average -- the companion completeWeightedAvg
(upstream LineupUtils.ts:752, not yet ported) divides by the
accumulated weight totals to finish the average. The per-field weight
used at each merge step is derived from obj's own totals (e.g.
that single lineup's total_off_fga), not from any running total on
mutable_acc -- callers accumulating many lineups must call
weighted_avg once per lineup so every lineup contributes its own
weight.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
mutable_acc | LineupStatSet | The running accumulator (LineupStatSet). Mutated in place; fields absent from the accumulator are initialized to {"value": 0.0} (plus old_value / override when obj's field carries a luck-adjustment override marker) before obj's contribution is added. | |
obj | LineupStatSet | The per-lineup LineupStatSet document to merge in. |
Returns
None. mutable_acc is mutated in place.
Example
from sportsdataverse.mbb.mbb_lineup_stats import weighted_avg
acc: dict = {}
weighted_avg(acc, lineup_a)
weighted_avg(acc, lineup_b)
print(acc["off_poss"]["value"]) # plain sum (SUM_FIELDS)
# Two-lineup possession-weighted merge
acc = {}
for lineup in three_lineups:
weighted_avg(acc, lineup)
# acc now holds weighted SUMS; complete_weighted_avg (not yet
# ported) is required to turn these into rate-stat averages.
win_prob_from_margin
win_prob_from_margin(exp_margin: 'float', *, league: 'str' = 'mens') -> 'float'
Home win probability from an expected margin (normal-CDF closed form).
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
exp_margin | float | Expected home-minus-away margin in points. | |
league | str | 'mens' | "mens" or "womens" (selects the fitted margin sigma). |
Returns
Probability the home team wins, in (0, 1).
Example
from sportsdataverse.mbb.mbb_game_predict import win_prob_from_margin
win_prob_from_margin(5.0)