Skip to main content
Version: 0.1.5

WBB — additional Python functions — Analytics: ConcurrentClump–wbb_bracketology

ConcurrentClump​

ConcurrentClump(evs: 'list[RawGameEvent]' = <factory>, lineups: 'list[LineupEvent]' = <factory>) -> None

A clump of concurrent raw events, together with the lineups that end

in that clump (Concurrency.ConcurrentClump, PossessionUtils.scala :64-69).

Parameters

ParameterTypeDefaultDescription
evslist[RawGameEvent]<factory>The raw game events in this clump, in chronological order.
lineupslist[LineupEvent]<factory>The lineups (if any) whose end_min falls in this clump.

PossState​

PossState(team_stats: 'PossCalcFragment', opponent_stats: 'PossCalcFragment', prev_clump: 'ConcurrentClump') -> None

Running state threaded through calculate_possessions_by_event

(PossessionUtils.PossState, PossessionUtils.scala:39-49).

Parameters

ParameterTypeDefaultDescription
team_statsPossCalcFragmentAccumulated fragment for the team since the last lineup boundary.
opponent_statsPossCalcFragmentAccumulated fragment for the opponent since the last lineup boundary.
prev_clumpConcurrentClumpThe previously-processed merged clump (used by calculate_stats's and-one / deadball-rebound heuristics).

Methods

PossState.init​

PossState.init() -> "'PossState'"

A fresh starting state (PossState.init, :45-48) -- builds

new instances per call, matching the mutable-dataclass rationale in mbb_ncaa_models.py's .empty() factories.

apply_relative_positional_overrides​

apply_relative_positional_overrides(results: 'list[dict[str, str]]', team_season: 'str', recurse_count: 'int' = 0) -> 'list[dict[str, str]]'

Recursively re-shuffle an ordered lineup per RELATIVE_POSITION_FIXES.

Faithful port of the private PositionUtils.applyRelativePositionalOverrides (PositionUtils.ts:657-693). Finds the first rule (in table order) whose key slots all match the current results codes (a None key slot matches anything), applies that rule's rule slots (None = leave unchanged, int = 1-based back-reference into the pre-rule results, dict = literal replacement) to produce a new ordering, then recurses on the new ordering -- since one swap can expose a second rule to match (e.g. the Maryland 2019/20 Morsell/Wiggins swap can cascade into the Lindo/Smith swap). Recursion is bounded by recurse_count < len(rules) (ported verbatim from the TS bound), so it always terminates even if two rules somehow ping-ponged each other.

Parameters

ParameterTypeDefaultDescription
resultslist[dict[str, str]]The current 5-slot {"code": ..., "id": ...} ordering (PG/SG/SF/PF/C, index 0-4).
team_seasonstrKey into RELATIVE_POSITION_FIXES. A team/season absent from the table (or the recursion exhausting that team/season's rule count) returns results unchanged.
recurse_countint0Internal recursion depth counter -- callers should not pass this explicitly (mirrors the TS default parameter).

Returns

The (possibly re-shuffled) 5-slot ordering.

Example

from sportsdataverse.mbb.mbb_positions import apply_relative_positional_overrides
results = [
{"code": "AnCowan", "id": "Cowan, Anthony"},
{"code": "ErAyala", "id": "Ayala, Eric"},
{"code": "DaMorsell", "id": "Morsell, Darryl"},
{"code": "AaWiggins", "id": "Wiggins, Aaron"},
{"code": "JaSmith", "id": "Smith, Jalen"},
]
apply_relative_positional_overrides(results, "Men_Maryland_2019/20")

assign_to_right_lineup​

assign_to_right_lineup(state: 'PossState', team_stats: 'PossCalcFragment', opponent_stats: 'PossCalcFragment', clump: 'ConcurrentClump', prev_clump: 'ConcurrentClump') -> 'list[LineupEvent]'

Assign a clump's possessions to the lineup(s) ending in it

(PossessionUtils.assign_to_right_lineup, PossessionUtils.scala :418-518).

Applies the running state total (accumulated since the last lineup boundary) to the first ending lineup only, then hands off to lineup_balancer (this clump's own fragment, split across candidates if there's more than one) and finally lineup_fixer (the negative-possession clamp).

Parameters

ParameterTypeDefaultDescription
statePossStateThe running possession state since the last lineup boundary.
team_statsPossCalcFragmentThis clump's team-direction fragment.
opponent_statsPossCalcFragmentThis clump's opponent-direction fragment.
clumpConcurrentClumpThe merged clump ending one or more lineups.
prev_clumpConcurrentClumpThe previous merged clump.

Returns

The lineup(s) ending in this clump, enriched with possession counts. Empty if clump.lineups is empty (see the module docstring's landmine-index note -- unreachable via calculate_possessions_by_event).

build_3p_shot_info​

build_3p_shot_info(p: 'LineupStatSet') -> 'OffLuckShotInfo3P'

3P-only shot-decomposition wrapper.

Public port of build3PShotInfo (LuckUtils.ts:741-759) -- remaps build_shot_info's generic keys to the 3pm/ 3pa/3p suffixes used throughout the luck-adjustment engine.

Parameters

ParameterTypeDefaultDescription
pLineupStatSetThe player's LineupStatSet/IndivStatSet-shaped dict.

Returns

{"shot_info_ast_3pm", "shot_info_early_3pa", "shot_info_scramble_3pa", "shot_info_unast_3pm", "shot_info_unknown_3pM", "shot_info_total_3p"}.

Example

from sportsdataverse.mbb.mbb_luck import build_3p_shot_info

info = build_3p_shot_info(player)
print(info["shot_info_total_3p"])

build_adjusted_3p​

build_adjusted_3p(p: 'LineupStatSet', info: 'OffLuckShotInfo3P') -> 'OffLuckAdj3P'

3P-only approx-unassisted/assisted-FG% wrapper.

Public port of buildAdjusted3P (LuckUtils.ts:812-835, "retained for bwc [backwards compat]" per the upstream comment) -- a thin remap of build_adjusted_fgcalled withshot_type="3p"`.

Parameters

ParameterTypeDefaultDescription
pLineupStatSetThe (typically base-period) player dict driving off_3p/ off_3p_ast.
infoOffLuckShotInfo3PAn build_3p_shot_info-shaped dict (the "biggest sample available" per the upstream comment -- normally the base period, not the sample being luck-adjusted).

Returns

{"base3P", "unassisted3P", "assisted3P", "baseAssistPct"}.

Example

from sportsdataverse.mbb.mbb_luck import build_3p_shot_info, build_adjusted_3p

base_info = build_3p_shot_info(base_player)
adj = build_adjusted_3p(base_player, base_info)
print(adj["assisted3P"], adj["unassisted3P"])

build_efficiency_margins​

build_efficiency_margins(mutable_stat_set: 'LineupStatSet', key_override: 'str | None' = None) -> 'None'

Derive off_net / off_raw_net on a stat set, in place.

Faithful port of LineupUtils.buildEfficiencyMargins (LineupUtils.ts:145). off_net is off_adj_ppp - def_adj_ppp (adjusted efficiency margin); off_raw_net is off_ppp - def_ppp (raw/unadjusted margin). Both are only written when their two source fields are both present on mutable_stat_set.

Parameters

ParameterTypeDefaultDescription
mutable_stat_setLineupStatSetThe LineupStatSet (or team-report equivalent) to mutate in place.
key_overridestr | NoneNone"value" or "old_value" -- which sub-key to read from the source fields and write into off_net / off_raw_net. When None (the default), the upstream nonLuckKey fallback applies: use "old_value" if mutable_stat_set["off_ppp"]["old_value"] is present, otherwise "value". When given explicitly, the written field is merged onto any existing off_net / off_raw_net dict (so a second call with the other key preserves the first call's key) rather than replacing it outright.

Returns

None. mutable_stat_set is mutated in place.

Example

from sportsdataverse.mbb.mbb_lineup_stats import build_efficiency_margins

build_efficiency_margins(team_info, "value")
off_ppp = team_info.get("off_ppp")
if isinstance(off_ppp, dict) and off_ppp.get("old_value") is not None:
build_efficiency_margins(team_info, "old_value")
print(team_info["off_net"]["value"])

build_exp_3p​

build_exp_3p(info: 'OffLuckShotTypeAndAdj3P') -> 'float'

Expected made-3P count given a player's shot-type mix + shooting %s.

Public port of buildExp3P (LuckUtils.ts:838-847): (assisted 3PM * assisted3P%) + (unassisted 3PM * unassisted3P%) + (early/scramble/unknown 3PA * base3P%). Pure weighted sum -- no division, so this introduces no landmine.

Parameters

ParameterTypeDefaultDescription
infoOffLuckShotTypeAndAdj3PA dict carrying both build_3p_shot_info's shot_info_* keys and build_adjusted_3p's *3P keys (i.e. an OffLuckShotTypeAndAdj3P).

Returns

The expected number of made 3-pointers (3P% * total 3P).

Example

from sportsdataverse.mbb.mbb_luck import (
build_3p_shot_info, build_adjusted_3p, build_exp_3p,
)

base_info = build_3p_shot_info(base_player)
info = {**build_3p_shot_info(player), **build_adjusted_3p(base_player, base_info)}
expected_makes = build_exp_3p(info)

build_position​

build_position(confs: 'dict[str, float]', confs_no_height: 'dict[str, float] | None', player: 'dict[str, Any]', team_season: 'str') -> 'tuple[str, str]'

Classify a player into a position label + diagnostic trace string.

Faithful port of PositionUtils.buildPosition (PositionUtils.ts:401-580) -- the PG / s-PG / CG / WG / WF / S-PF / PF/C / C decision tree. A ABSOLUTE_POSITION_FIXES manual override short-circuits the whole tree (recursing once, with team_season="", purely to compute the diagnostic "what would this have been" string); otherwise the function walks the confidence-threshold / assist-rate / 3PT-rate branch cascade, applies the "too few effective possessions" (< 25) fallback, and reconciles the result against roster metadata via using_roster_pos.

Parameters

ParameterTypeDefaultDescription
confsdict[str, float]The 5-way positional confidence dict (TRAD_POS_LIST keys), typically the height-adjusted output of build_position_confidences.
confs_no_heightdict[str, float] | NoneThe pre-height-adjustment confidences, or None when the caller has no height data. When present, a PG <-> s-PG flip caused solely by the height adjustment is reverted (the maybeIgnoreHeight closure, ts:433-457). The check is is not None (JS object-truthiness: an empty dict is still a truthy JS object), NOT a Python-falsy if confs_no_height.
playerdict[str, Any]The player stat dict. Reads key (override lookup), off_assist / off_3pr / off_usage / off_team_poss (each {"value": N}-wrapped), and roster (a plain {"pos": ..., "role": ...} dict of un-wrapped strings).
team_seasonstr"{sport}_{team}_{season}" key into ABSOLUTE_POSITION_FIXES. Pass "" to disable override lookup for a given call (the recursive diagnostic call inside the override branch does exactly this).

Returns

A (position, diagnostic) tuple. position is one of ID_TO_POSITION's keys; diagnostic is a human-readable trace of which rule fired, byte-identical to the TS's template strings (including .toFixed(1)-style percentage formatting).

Example

from sportsdataverse.mbb.mbb_positions import build_position, TRAD_POS_LIST
confs = dict(zip(TRAD_POS_LIST, [0.9, 0.1, 0, 0, 0]))
player = {"off_assist": {"value": 0.10}, "off_3pr": {"value": 0.20},
"off_team_poss": {"value": 1000}, "off_usage": {"value": 0.20}}
build_position(confs, None, player, "Men_Boston College_2019/20")

# A manual-override short-circuit

build_position(confs, None, {"key": "Popovic, Nik",
"off_usage": {"value": 1}, "off_team_poss": {"value": 200},
"off_assist": {"value": 0.10}}, "Men_Boston College_2019/20")

build_position_confidences​

build_position_confidences(player: 'dict[str, Any]', height_in: 'float | None' = None) -> 'tuple[dict[str, float], dict[str, Any]]'

Build the 5-way positional confidence vector for a player.

Faithful port of PositionUtils.buildPositionConfidences (PositionUtils.ts:263-338). Derives the six calc_* ratios from the player's box-score fields, dot-products the resulting 17-feature vector against POSITION_FEATURE_WEIGHTS (each field regressed via regress_shot_quality and multiplied by its per-feature scale) plus the POSITION_FEATURE_INIT intercepts, applies a softmax over the five raw scores, and -- when height_in is supplied -- reweights the confidences via incorporate_height.

Parameters

ParameterTypeDefaultDescription
playerdict[str, Any]The player stat dict (ES-aggregation bucket shape); each stat field is {"value": N}. Reads total_off_assist, total_off_to, off_3p, off_efg, off_2pmid, off_2prim, total_off_fga, total_off_fta, total_off_ftm (for the calc_* ratios) plus every non-calc_ field in POSITION_FEATURE_WEIGHTS.
height_infloat | NoneNoneOptional player height in inches. When truthy, the returned confidences are height-adjusted; when None / 0, the raw softmax confidences are returned. (JS height_in ? ... : ... falsy check, ts:324 -- a 0 height is treated as "no height".)

Returns

A (confidences, diagnostics) tuple. confidences maps each TRAD_POS_LIST key (in order) to its final confidence. diagnostics carries "scores" (raw scores x 0.1, keyed by position), "confsNoHeight" (the pre-height confidences, present only when height_in is truthy, else None), and "calculated" (the six derived calc_* ratios). The upstream diag object has exactly these three fields -- no UI-only fields are dropped.

Example

from sportsdataverse.mbb.mbb_positions import build_position_confidences
confs, diags = build_position_confidences(player_bucket)
print(confs["pos_pg"], diags["calculated"]["calc_ast_tov"])

# Height-adjusted confidences

confs_h, diags_h = build_position_confidences(player_bucket, 78.0)

build_positional_aware_filter​

build_positional_aware_filter(filter_str: 'str') -> 'tuple[list[dict[str, Any]], list[dict[str, Any]], bool]'

Decompose a search-filter string into positionally-aware +ve/-ve fragments.

Faithful port of PositionUtils.buildPositionalAwareFilter (PositionUtils.ts:764-828). Picks a fragment separator by scanning [";", "/", ","] in priority order for the first one present anywhere in filter_str (a fragment separator of "!!!" -- never itself present -- is the "no separator found" fallback, which leaves the whole string as a single fragment). Splits on that separator, trims whitespace, drops empty fragments and [-prefixed ones (reserved for aggregation-key filters elsewhere in the app), then routes each fragment to the positive or negative bucket by a leading -, and parses each fragment's optional =<tokens> position spec via decomp_positional_filter_fragment`.

Parameters

ParameterTypeDefaultDescription
filter_strstrA raw filter string, e.g. "test1=pg / -test2=Pf+C / test3".

Returns

A (positive_fragments, negative_fragments, has_position) triple. Each fragment is {"filter": <lowercased name>, "pos": [indices]}. has_position is True iff any fragment (either side) carried at least one recognized position token.

Example

::

from sportsdataverse.mbb.mbb_positions import build_positional_aware_filter
build_positional_aware_filter("test1=pg / -test2=Pf+C / test3")

calc_def_player_luck_adj​

calc_def_player_luck_adj(sample: 'LineupStatSet', base: 'LineupStatSet', avg_eff: 'float') -> 'DefLuckAdjustmentDiags'

Defensive 3P-luck adjustment for a single player.

Faithful port of LuckUtils.calcDefPlayerLuckAdj (LuckUtils.ts:402-426). Unlike calc_off_player_luck_adj, this is not a pure delegation -- see the module docstring's calc_def_player_luck_adj note for the translate() remap this wraps around calc_def_team_luck_adj.

Parameters

ParameterTypeDefaultDescription
sampleLineupStatSetThe player's stat dict for the period being luck-adjusted (must carry oppo_total_def_3p_made/oppo_total_def_3p_attempts -- there is no player-level def_3p field upstream, hence the remap).
baseLineupStatSetThe player's stat dict for the baseline/reference period.
avg_efffloatLeague/context average efficiency (100 in every vendored jest call).

Returns

Same shape as calc_def_team_luck_adj, computed against the translated (oppo_* -> def_*) player stat dicts.

Example

from sportsdataverse.mbb.mbb_luck import calc_def_player_luck_adj

diags = calc_def_player_luck_adj(sample_player, base_player, 100.0)
print(diags["deltaDefAdjEff"])

calc_def_team_luck_adj​

calc_def_team_luck_adj(sample: 'LineupStatSet', base: 'LineupStatSet', avg_eff: 'float', sample_def_3pa_override: 'float | None' = None) -> 'DefLuckAdjustmentDiags'

Defensive 3P-luck adjustment for a team (or lineup).

Faithful port of LuckUtils.calcDefTeamLuckAdj (LuckUtils.ts:429-531). See the module docstring for the SoS-vs-luck-split formula (LUCK_PCT) and the shared unguarded-division landmine.

Parameters

ParameterTypeDefaultDescription
sampleLineupStatSetThe team/lineup/player stat dict for the period being luck-adjusted (e.g. an on/off split or a single lineup).
baseLineupStatSetThe team/lineup/player stat dict for the baseline/reference period.
avg_efffloatLeague/context average efficiency (100 in every vendored jest call).
sample_def_3pa_overridefloat | NoneNoneWhen given, used as sampleDef3PA instead of sample["total_def_3p_attempts"] -- see calc_off_team_luck_adj's sample_3pa_override docstring for the shared "lineup regression" rationale (LuckUtils.ts:433-434, verbatim comment).

Returns

A DefLuckAdjustmentDiags dict -- TS-verbatim keys (avgEff, luckPct, baseDef3P, baseDef3PSos, baseDef3PA, basePoss, base3PSosAdj, sampleDef3P, sampleDef3PSos, sampleDef3PA, samplePoss, sample3PSosAdj, sampleDefEfg, sampleDefPpp, sampleOffSos, sampleDef3PRate, sampleDefFGA, sampleDefOrb, avg3PSosAdj, adjDef3P, delta3P, deltaDefEfg, deltaDefPppNoOrb, deltaMissesPct, deltaDefOrbFactor, deltaPtsOffMisses, deltaDefPpp, deltaDefAdjEff).

Example

from sportsdataverse.mbb.mbb_luck import calc_def_team_luck_adj

diags = calc_def_team_luck_adj(sample_team_off, base_team, 100.0)
print(diags["deltaDefAdjEff"])

calc_off_player_luck_adj​

calc_off_player_luck_adj(sample_player: 'LineupStatSet', base_player: 'LineupStatSet', avg_eff: 'float') -> 'OffLuckAdjustmentDiags'

Offensive 3P-luck adjustment for a single player.

Faithful port of LuckUtils.calcOffPlayerLuckAdj (LuckUtils.ts:174-187). Per Task 2.1's surprise #4, this is a literal 1-player-team delegation to calc_off_team_luck_adj -- ORB effects are ignored for an individual player (the upstream comment: "the team calc basically works fine here, apart from ORBs, which we'll ignore").

Parameters

ParameterTypeDefaultDescription
sample_playerLineupStatSetThe player's stat dict for the period being luck-adjusted.
base_playerLineupStatSetThe player's stat dict for the baseline/reference period.
avg_efffloatLeague/context average efficiency (100 in every vendored jest call).

Returns

Same shape as calc_off_team_luck_adj -- identical to calling that function with sample_players=[sample_player], base_players_map={base_player["key"]: base_player}.

Example

from sportsdataverse.mbb.mbb_luck import calc_off_player_luck_adj

diags = calc_off_player_luck_adj(sample_player, base_player, 100.0)
print(diags["deltaOffAdjEff"])

calc_off_team_luck_adj​

calc_off_team_luck_adj(sample_team: 'LineupStatSet', sample_players: 'list[LineupStatSet]', base_team: 'LineupStatSet', base_players_map: 'dict[str, LineupStatSet]', avg_eff: 'float', sample_3pa_override: 'float | None' = None, manual_overrides: 'list[ManualOverride] | None' = None) -> 'OffLuckAdjustmentDiags'

Offensive 3P-luck adjustment for a team (or lineup).

Faithful port of LuckUtils.calcOffTeamLuckAdj (LuckUtils.ts:190-399). See the module docstring for the Bayesian-shrink formula, the JS-array- truthiness / object-selection landmines, and the one unguarded-division landmine this function carries.

Parameters

ParameterTypeDefaultDescription
sample_teamLineupStatSetThe team/lineup stat dict for the period being luck-adjusted (e.g. an on/off split or a single lineup).
sample_playerslist[LineupStatSet]The roster of per-player stat dicts backing sample_team (samplePlayers == players.map(on/off/baseline) per the upstream comment).
base_teamLineupStatSetThe team stat dict for the baseline/reference period (typically full-season).
base_players_mapdict[str, LineupStatSet]{player_key: base_period_player_stat_dict}.
avg_efffloatLeague/context average efficiency (100 in every vendored jest call).
sample_3pa_overridefloat | NoneNoneWhen given, used as sample3PA instead of sample_team["total_off_3p_attempts"]. Per the upstream comment (LuckUtils.ts:196-198, shared verbatim with calc_def_team_luck_adj's sample_def_3pa_override): "when calc'ing luck on lineups, each lineup gets the total sample as its regression so its average is right over the set" -- i.e. this lets every lineup in a sweep share one common 3PA denominator (the team's) for its regression target, rather than each lineup regressing against its own much smaller, noisier 3PA count. Note that calc_off_player_luck_adj itself does not pass this (its delegation call omits it entirely) -- the jest oracle's own "3P override" cross-check (LuckUtils.test.ts:100-115) instead calls calc_off_team_luck_adj directly with the player's own 3PA as this override, purely to demonstrate the parameter's effect in isolation.
manual_overrideslist[ManualOverride] | NoneNonePer-player 3P%-expectation overrides from the UI. A non-None empty list still activates the team-level override-delta branch (JS array truthiness) -- see the module docstring's landmine note. None (the default) is the "no overrides at all" case.

Returns

An OffLuckAdjustmentDiags dict -- TS-verbatim keys (avgEff, samplePoss, sample3P, sample3PA, base3PA, player3PInfo (per-player detail, sorted by descending shot_info_total_3p), sampleBase3P, regress3P, sampleOff3PRate, sampleOffFGA, sampleOffOrb, sampleOffEfg, sampleOffPpp, sampleDefSos, delta3P, deltaOffEfg, deltaMissesPct, deltaOffPppNoOrb, deltaOffOrbFactor, deltaPtsOffMisses, deltaOffPpp, deltaOffAdjEff).

Example

from sportsdataverse.mbb.mbb_luck import calc_off_team_luck_adj

diags = calc_off_team_luck_adj(
sample_team_on, sample_players_on, base_team, base_players_map, 100.0,
)
print(diags["deltaOffAdjEff"])

# With per-player manual 3P% overrides

diags = calc_off_team_luck_adj(
sample_team_on, sample_players_on, base_team, base_players_map, 100.0,
manual_overrides=[
{"rowId": "Cowan, Anthony", "statName": "off_3p", "newVal": 0.5, "use": True},
],
)

calculate_aggregated_lineup_stats​

calculate_aggregated_lineup_stats(lineups: 'list[LineupStatSet] | None') -> 'LineupStatSet'

Combine all lineups into a single team stat set.

Faithful port of LineupUtils.calculateAggregatedLineupStats (LineupUtils.ts:106). Seeds an accumulator from StatModels.emptyLineup() ({"key": "empty", "doc_count": 0}) plus an all_lineups sub-accumulator of the same shape, then merges every lineup via weighted_avg: lineups without a truthy rapmRemove key merge into the main accumulator, while rapmRemove lineups merge into all_lineups instead (their contribution is folded back in afterward). Calls complete_weighted_avg to turn the main accumulator's weighted sums into weighted averages, then -- because StatModels.emptyLineup() always carries key/doc_count and so is never considered "empty" by the upstream lodash.isEmpty check -- unconditionally re-merges the (now-averaged) team totals into all_lineups and finishes that sub-accumulator too. Finally rebuilds off_net / off_raw_net via build_efficiency_margins (value-key always; old-value-key too when the team is in luck-adjusted mode, i.e. off_ppp.old_value is present) -- but only on the top-level result, matching upstream's "don't bother for all_lineups" comment.

Parameters

ParameterTypeDefaultDescription
lineupslist[LineupStatSet] | NoneThe per-lineup LineupStatSet docs to fold together (e.g. the ES aggregation buckets under responses[0].aggregations.lineups.buckets). None or an empty list yields an all-zero/empty team stat set (mirrors the upstream lineups || [] guard).

Returns

The aggregated team-total LineupStatSet, including a nested all_lineups key holding the rapmRemove-lineups-plus-team-total composite sub-aggregate.

Example

from sportsdataverse.mbb.mbb_lineup_stats import calculate_aggregated_lineup_stats

buckets = raw_response["responses"][0]["aggregations"]["lineups"]["buckets"]
team_info = calculate_aggregated_lineup_stats(buckets)
print(team_info["off_ppp"]["value"], team_info["off_poss"]["value"])

# RAPM-exclusion flag

buckets[1]["rapmRemove"] = True # divert into all_lineups instead
team_info = calculate_aggregated_lineup_stats(buckets)

calculate_possessions​

calculate_possessions(lineup_events: 'Iterable[LineupEvent]') -> 'list[LineupEvent]'

Top-level entry point: calculate team/opponent possessions for a

sequence of lineup events (PossessionUtils.calculate_possessions, PossessionUtils.scala:371-379).

Parameters

ParameterTypeDefaultDescription
lineup_eventsIterable[LineupEvent]The lineups to enrich, in chronological order.

Returns

The lineups, each enriched with possession counts.

Example

from sportsdataverse.mbb.mbb_ncaa_possessions import calculate_possessions

enriched = calculate_possessions(lineups)
enriched[0].team_stats.num_possessions

calculate_possessions_by_event​

calculate_possessions_by_event(raw_events_as_clumps: 'Iterable[ConcurrentClump]') -> 'list[LineupEvent]'

Drive the batch loop + per-clump scoring over an already-flattened

clump stream (PossessionUtils.calculate_possessions_by_event, PossessionUtils.scala:521-573).

Parameters

ParameterTypeDefaultDescription
raw_events_as_clumpsIterable[ConcurrentClump]The unbatched clump stream, e.g. from flat-mapping lineup_as_raw_clumps over several lineups.

Returns

The lineups, each enriched with possession counts, in original order.

calculate_stats​

calculate_stats(clump: 'ConcurrentClump', prev: 'ConcurrentClump', dir: 'Direction') -> 'PossCalcFragment'

Calculate one direction's possession-fragment for one merged clump

(PossessionUtils.calculate_stats, PossessionUtils.scala:170-369).

See the upstream source's inline worked examples (and-one detection, technical/flagrant offsetting, the deadball-rebound heuristic) for the hand-annotated NCAA play-by-play snippets that motivate each step; this port reproduces every step in the same order.

Parameters

ParameterTypeDefaultDescription
clumpConcurrentClumpThe merged clump to score.
prevConcurrentClumpThe previously-processed merged clump (feeds the and-one and deadball-rebound heuristics -- see below).
dirDirectionWhich side (Direction.TEAM/Direction.OPPONENT) is "attacking" for this calculation. Named to match the Scala (shadows the dir builtin -- consistent with this port's existing precedent of naming params after their Scala originals, e.g. RawGameEvent.for_team's min).

Returns

A ~sportsdataverse.mbb.mbb_ncaa_models.PossCalcFragment for this clump/direction.

complete_weighted_avg​

complete_weighted_avg(mutable_acc: 'LineupStatSet', harmonic_weighting: 'bool' = False, regress_diffs: 'float' = 0.0) -> 'None'

Finish a weighted_avg accumulator into true weighted averages.

Faithful port of LineupUtils.completeWeightedAvg (LineupUtils.ts:752). Mutates mutable_acc in place and returns None, mirroring the upstream void + mutable-arg contract. Recomputes the per-field weight tables from mutable_acc itself (getSimpleWeights(mutableAcc, 1, regressDiffs) -- note the default_val=1, unlike weighted_avg's default_val=0), then, unless harmonic_weighting is set, calls recalculate_play_type_possto fix up the transition/scramble possession fields thatweighted_avgskipped. Finally divides every non-ignored field's accumulated weighted sum by its matching weight total (shot-type /ppp_totals/orb_totals/fta_totals/ast_totals/ generic FGA fallback);total_*andSUM_FIELDSfields are left untouched (they are already true totals, not sums to be averaged).off_ftr/def_ftr get a special non-harmonic_weightingrecompute straight from the accumulatedtotal_{off|def}_fta` rather than dividing their own weighted sum.

Parameters

ParameterTypeDefaultDescription
mutable_accLineupStatSetThe weighted_avg-accumulated LineupStatSet to finish in place. Every field with a non-total_/SUM_FIELDS key is converted from a weighted sum to a weighted average.
harmonic_weightingboolFalseWhen True, skips the recalculate_play_type_possfixup and uses a harmonic-style division foroff_ftr/def_ftr instead of the totals-based recompute. Matches the upstream default (False) used by calculate_aggregated_lineup_stats`.
regress_diffsfloat0.0Forwarded to get_simple_weights-- regression toward ~1000 possessions for on/off diff calculations. Defaults to0.0(no regression), matchingcalculate_aggregated_lineup_stats`'s call site.

Returns

None. mutable_acc is mutated in place.

Example

from sportsdataverse.mbb.mbb_lineup_stats import weighted_avg, complete_weighted_avg

acc: dict = {}
for lineup in lineups:
weighted_avg(acc, lineup)
complete_weighted_avg(acc)
print(acc["off_ppp"]["value"]) # now a true weighted average

concurrent_event_handler​

concurrent_event_handler(clumps: 'Iterable[ConcurrentClump]') -> 'list[ConcurrentClump]'

Batch a stream of singleton/boundary clumps into merged

concurrent-event clumps (Concurrency.concurrent_event_handler + StateUtils.foldLeft's clumping machinery, PossessionUtils.scala :71-111 -- see the module docstring for the full batching-predicate breakdown and the post-game-break singleton port trap).

ponytail: manual accumulate-and-flush loop replacing the generic

Clumper/StateUtils.foldLeft abstraction -- this is the ONE clumper

instantiation in the port, so a reusable abstraction buys nothing.

Lift this back into a small clumper type if a second concurrent-event

family needs the same batching later.

Parameters

ParameterTypeDefaultDescription
clumpsIterable[ConcurrentClump]An ordered stream of ConcurrentClump\ s, each either a singleton raw event (evs=[ev]) or a lineup-boundary marker (evs=[], lineups=[lineup]), e.g. from lineup_as_raw_clumps.

Returns

The merged clumps, each an in-order concatenation of one batch's evs/lineups.

count_matching​

count_matching(evs: 'Iterable[RawGameEvent]', side: 'DirFn', *parsers: 'Parser') -> 'int'

Count events on one side matching any of the given parsers.

Ports the pervasive clump.evs.collect { case side(ParseX(_)) => () } .size idiom (and its multi-arm case side(ParseX(_)) => (); case side(ParseY(_)) => () union form, when more than one parser is passed -- e.g. the and-one free-throw count, which matches either a made or a missed free throw on the same event).

Parameters

ParameterTypeDefaultDescription
evsIterable[RawGameEvent]The events to scan.
sideDirFn~sportsdataverse.mbb.mbb_ncaa_models.PossessionEvent .attacking_team or .defending_team, selecting which raw string (if any) to test per event.

Returns

The count of matching events.

get_stats_diff​

get_stats_diff(stat_set1: 'LineupStatSet', stat_set2: 'LineupStatSet', off_title: 'str', def_title: 'str | None' = None) -> 'LineupStatSet'

Straight (unweighted) field-by-field diff of two team stat sets.

Faithful port of LineupUtils.getStatsDiff (LineupUtils.ts:185). For every field on stat_set1, subtracts the matching field's value (and, when both sides carry one, old_value) from stat_set2. No possession weighting or regression -- this is a raw subtraction, unlike weighted_avg / complete_weighted_avg.

Parameters

ParameterTypeDefaultDescription
stat_set1LineupStatSetThe "from" team stat set (e.g. this team).
stat_set2LineupStatSetThe "to subtract" team stat set (e.g. the opponent, or a prior period).
off_titlestrWritten into the result's off_title field verbatim.
def_titlestr | NoneNoneWritten into the result's def_title field verbatim (None when omitted, mirroring the upstream optional arg).

Returns

A new LineupStatSet: one {"value": ..., "old_value": ..., "override": ...} dict per field present on stat_set1, plus off_title / def_title. A field becomes None (the JS undefined analog) instead of a diff dict when either side is missing a value -- e.g. because that field was never populated for one of the two stat sets.

Example

from sportsdataverse.mbb.mbb_lineup_stats import get_stats_diff

diff = get_stats_diff(team_a, team_b, "Team A", "Team B")
print(diff["off_ppp"]["value"]) # team_a.off_ppp - team_b.off_ppp

incorporate_height​

incorporate_height(height_in: 'float', confs: 'list[float]') -> 'list[float]'

Reweight positional confidences by height (Bayesian-ish height prior).

Faithful port of PositionUtils.incorporateHeight (PositionUtils.ts:346-368; see build_height_adj_probs in the linked hoop-explorer blog post). For each position i it computes a height-plausibility mass cdf(height + 1) - cdf(height - 1) under N(mean_i, sqrt2 * std_i) (the sqrt2 "height dampening" widens the variance so the effect is not too aggressive), multiplies it into the prior confidence, and renormalizes.

Parameters

ParameterTypeDefaultDescription
height_infloatPlayer height in inches.
confslist[float]The five raw (pre-height) confidences, in TRAD_POS_LIST order.

Returns

The five height-adjusted confidences, renormalized to sum to 1 (the sum_product or 1 guard makes a degenerate all-zero product a no-op rather than a divide-by-zero -- see module landmine index item 1).

Example

from sportsdataverse.mbb.mbb_positions import incorporate_height
incorporate_height(81, [0.03, 0.19, 0.49, 0.09, 0.18])

inject_luck​

inject_luck(mutable_stats: 'LineupStatSet', off_luck: 'OffLuckAdjustmentDiags | None', def_luck: 'DefLuckAdjustmentDiags | None') -> 'None'

Reversibly mutate a stat set in place with luck-adjustment deltas.

Faithful port of LuckUtils.injectLuck (LuckUtils.ts:534-650). Works on a team, lineup, or player stat dict -- only the fields already present on mutable_stats are touched (see override_mutable_val's object-presence gate), so calling this on a stat set that doesn't carry a given field (e.g. a bare {"key": ..., "doc_count": 0}placeholder) is a safe no-op for that field. Passingoff_luck=None, def_luck=Noneresets every field this function has ever touched back to its pre-luck value (see the module docstring's landmine list for the exact mechanics, including the absolute-vs-delta distinction ondef_3p/oppo_def_3p`).

Parameters

ParameterTypeDefaultDescription
mutable_statsLineupStatSetThe stat-set dict to mutate in place. May be a team/lineup stat set (carries off_net/off_raw_net/no oppo_total_def_3p_made) or a player stat set (carries oppo_total_def_3p_made, gating the extra oppo_def_3p recompute -- see the module docstring's landmine #2).
off_luckOffLuckAdjustmentDiags | NoneThe output of calc_off_team_luck_adj / calc_off_player_luck_adj, or None to omit/reset the offensive-side fields.
def_luckDefLuckAdjustmentDiags | NoneThe output of calc_def_team_luck_adj / calc_def_player_luck_adj, or None to omit/reset the defensive-side fields.

Returns

None -- this function mutates mutable_stats in place (TS injectLuck likewise returns nothing).

Example

from sportsdataverse.mbb.mbb_luck import (
calc_off_team_luck_adj, calc_def_team_luck_adj, inject_luck,
)

off_luck = calc_off_team_luck_adj(sample_team_on, sample_players_on, base_team, base_players_map, 100.0)
def_luck = calc_def_team_luck_adj(sample_team_off, base_team, 100.0)
inject_luck(sample_team_on, off_luck, def_luck)
print(sample_team_on["off_3p"])

# Reset back to the pre-luck values

inject_luck(sample_team_on, None, None)

lineup_as_raw_clumps​

lineup_as_raw_clumps(lineup: 'LineupEvent') -> 'Iterator[ConcurrentClump]'

Turn one lineup's raw events into unprocessed singleton clumps, plus a

trailing lineup-boundary marker (Concurrency.lineup_as_raw_clumps, PossessionUtils.scala:114-120).

Parameters

ParameterTypeDefaultDescription
lineupLineupEventThe lineup event to expand.

Returns

One ConcurrentClump([ev]) per raw event (in order), then a final ConcurrentClump([], [lineup]) boundary marker.

lineup_balancer​

lineup_balancer(lineups: 'list[LineupEvent]', team_stats: 'PossCalcFragment', opponent_stats: 'PossCalcFragment', clump: 'ConcurrentClump', prev_clump: 'ConcurrentClump') -> 'list[LineupEvent]'

Attribute this clump's possessions to the candidate lineup(s)

(PossessionUtils.assign_to_right_lineup.lineup_balancer, PossessionUtils.scala:429-471).

A single candidate just receives the whole clump's possessions. Multiple candidates (a lineup change landing mid-clump) are split via a greedy round-robin: for each direction, rank lineups by an "approximate" possession count computed from just that lineup's own raw events at the clump's minute, then hand out possessions one at a time to whichever lineup currently has the highest remaining approximate share.

Parameters

ParameterTypeDefaultDescription
lineupslist[LineupEvent]The candidate lineups (already updated with any running state total from assign_to_right_lineup).
team_statsPossCalcFragmentThis clump's team-direction fragment.
opponent_statsPossCalcFragmentThis clump's opponent-direction fragment.
clumpConcurrentClumpThe merged clump being assigned.
prev_clumpConcurrentClumpThe previous merged clump (only used for the first candidate's approximate stats -- see below).

Returns

New lineup copies with num_possessions incremented.

lineup_fixer​

lineup_fixer(lineups: 'list[LineupEvent]') -> 'list[LineupEvent]'

Clamp obviously-broken possession counts (``PossessionUtils

.assign_to_right_lineup.lineup_fixer, PossessionUtils.scala:490-507`).

For both team_stats and opponent_stats independently: a lineup that scored (`pts > 0``) but was attributed zero-or-fewer possessions is clamped to exactly 1 (you can't score on zero possessions); any still-negative possession count is clamped to 0.

Parameters

ParameterTypeDefaultDescription
lineupslist[LineupEvent]The lineups to fix (already balanced).

Returns

New lineup copies with clamped num_possessions.

lineup_to_team_report​

lineup_to_team_report(lineup_report: 'LineupStatSet', inc_replacement: 'bool' = False, regress_diffs: 'float' = 0.0, rep_on_off_diag_mode: 'int' = 0) -> 'LineupStatSet'

Build per-player on/off splits out of a team's lineups.

Faithful port of LineupUtils.lineupToTeamReport (LineupUtils.ts:277). For every distinct player across lineup_report["lineups"], partitions the team's lineups into ON (the player was on the floor) and OFF (they weren't) buckets, merging each bucket via weighted_avg / complete_weighted_avg. Also builds a teammates map of possession overlap with every other player, and -- when inc_replacement=True -- a "replacement" on-minus-off composite via combine_replacement_on_off`.

Lineups whose key is the empty string are skipped in the on/off-partition loop (workaround for an upstream data issue, tracked as upstream issue #53) but still contribute to the player roster. Every lineup's rapmRemove key (if present, e.g. left over from a prior calculate_aggregated_lineup_stats call sharing the same input list) is deleted as a side effect while building the roster -- lineup_to_team_report itself never consults rapmRemove.

Parameters

ParameterTypeDefaultDescription
lineup_reportLineupStatSet{"lineups": [...], "avgOff": ..., "error_code": ...} -- the per-team lineup list plus metadata (mirrors upstream's LineupStatsModel). Only lineups and error_code are consumed here.
inc_replacementboolFalseWhen True, additionally builds each player's replacement on-minus-off composite (more expensive -- scans every OFF lineup against every ON lineup for a 4-of-5-shared- players complement match).
regress_diffsfloat0.0Forwarded to combine_replacement_on_off's final complete_weighted_avgcall -- regression toward ~1000 possessions for the replacement diff (only meaningful wheninc_replacement=True`).
rep_on_off_diag_modeint0When > 0, retains diagnostic detail (myLineups on each player's replacement entry, plus lineupUsage bookkeeping) instead of discarding it after use.

Returns

{"playerMap": {code: id}, "players": [...], "error_code": ...}. Each entry in players is {"playerId", "playerCode", "teammates", "on", "off", "replacement"} -- on/off are finished LineupStatSet averages (or, for a player who's always ON, an all-zero off); replacement is None unless inc_replacement=True.

Example

from sportsdataverse.mbb.mbb_lineup_stats import lineup_to_team_report

report = lineup_to_team_report({"lineups": buckets, "error_code": None})
for player in report["players"]:
print(player["playerId"], player["on"]["off_poss"]["value"])

# With replacement (on-minus-off) splits

report = lineup_to_team_report(
{"lineups": buckets, "error_code": None},
inc_replacement=True,
regress_diffs=-500,
)

ncaa_wbb_lineups​

ncaa_wbb_lineups(pbp: 'pl.DataFrame', *, include_transition: 'bool' = False, fix_tip_in: 'bool' = True, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Aggregate WBB play-by-play into per-lineup stats (wbigballR get_lineups).

Pure delegation to sportsdataverse.mbb.mbb_ncaa_lineups.ncaa_mbb_lineups — see it for the algorithm, column contract, and the fix_tip_in vocab fix.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract (ncaa_wbb_game_pbp output).
include_transitionboolFalseAppend the trans/half split surface.
fix_tip_inboolTrueCount the real "Tip In" vocabulary (default); False reproduces R's "Tip-In" bug for oracle parity.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per lineup+team; see the MBB sibling for the column contract.

col_nametypedescription
p1character
p2character
p3character
p4character
p5character
teamcharacterTeam-side label or team identifier.
minsdouble
o_minsdouble
d_minsdouble
o_possdouble
d_possdouble
ortgdouble
drtgdouble
netrtgdouble
ptsdoublePoints scored.
d_ptsdouble
fgadoubleField goal attempts.
d_fgadouble
fgmdoubleField goals made.
d_fgmdouble
tpadouble
d_tpadouble
tpmdouble
d_tpmdouble
ftadoubleFree throw attempts.
d_ftadouble
ftmdoubleFree throws made.
d_ftmdouble
rimadouble
d_rimadouble
rimmdouble
d_rimmdouble
orbdouble
d_orbdouble
drbdouble
d_drbdouble
blkdoubleBlocks.
d_blkdouble
todoubleTo.
d_todouble
astdoubleAssists.
d_astdouble
e_possdouble
fg_pctdoubleField goal percentage (0-1).
d_fg_pctdouble
tppdouble
d_tppdouble
ftpdouble
d_ftpdouble
efg_pctdouble
d_efg_pctdouble
ts_pctdoubleTrue shooting percentage (0-1).
d_ts_pctdouble
rim_pctdouble
d_rim_pctdouble
mid_pctdouble
d_mid_pctdouble
tp_ratedouble
d_tp_ratedouble
rim_ratedouble
d_rim_ratedouble
mid_ratedouble
d_mid_ratedouble
ft_ratedoubleFt rate.
d_ft_ratedouble
ast_ratedouble
d_ast_ratedouble
to_ratedoubleTo rate.
d_to_ratedouble
blk_ratedouble
o_blk_ratedouble
orb_pctdoubleOffensive rebound percentage.
drb_pctdoubleDefensive rebound percentage.
time_per_possdouble
d_time_per_possdouble

Example

from sportsdataverse.wbb.wbb_ncaa_lineups import ncaa_wbb_lineups
lineups = ncaa_wbb_lineups(pbp)
print(lineups.shape)

ncaa_wbb_on_off​

ncaa_wbb_on_off(players: 'Union[str, Sequence[str]]', lineups: 'pl.DataFrame', *, included: 'Union[str, Sequence[str], None]' = None, excluded: 'Union[str, Sequence[str], None]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Team stats for every on/off combination of the given WBB players.

Pure delegation to sportsdataverse.mbb.mbb_ncaa_lineups.ncaa_mbb_on_off (wbigballR on_off_generator).

Parameters

ParameterTypeDefaultDescription
playersUnion[str, Sequence[str]]Player name(s) to split on (the status axis).
lineupsDataFrameLineups frame from ncaa_wbb_lineups.
includedUnion[str, Sequence[str], None]NoneOptional membership filter forwarded to ncaa_wbb_player_lineups.
excludedUnion[str, Sequence[str], None]NoneOptional membership filter forwarded to ncaa_wbb_player_lineups.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

2^k rows — status + the stat columns.

col_nametypedescription
statuscharacterStatus label.
minsdouble
o_minsdouble
d_minsdouble
o_possdouble
d_possdouble
ortgdouble
drtgdouble
netrtgdouble
ptsdoublePoints scored.
d_ptsdouble
fgadoubleField goal attempts.
d_fgadouble
fgmdoubleField goals made.
d_fgmdouble
tpadouble
d_tpadouble
tpmdouble
d_tpmdouble
ftadoubleFree throw attempts.
d_ftadouble
ftmdoubleFree throws made.
d_ftmdouble
rimadouble
d_rimadouble
rimmdouble
d_rimmdouble
orbdouble
d_orbdouble
drbdouble
d_drbdouble
blkdoubleBlocks.
d_blkdouble
todoubleTo.
d_todouble
astdoubleAssists.
d_astdouble
e_possdouble
fg_pctdoubleField goal percentage (0-1).
d_fg_pctdouble
tppdouble
d_tppdouble
ftpdouble
d_ftpdouble
efg_pctdouble
d_efg_pctdouble
ts_pctdoubleTrue shooting percentage (0-1).
d_ts_pctdouble
rim_pctdouble
d_rim_pctdouble
mid_pctdouble
d_mid_pctdouble
tp_ratedouble
d_tp_ratedouble
rim_ratedouble
d_rim_ratedouble
mid_ratedouble
d_mid_ratedouble
ft_ratedoubleFt rate.
d_ft_ratedouble
ast_ratedouble
d_ast_ratedouble
to_ratedoubleTo rate.
d_to_ratedouble
blk_ratedouble
o_blk_ratedouble
orb_pctdoubleOffensive rebound percentage.
drb_pctdoubleDefensive rebound percentage.
time_per_possdouble
d_time_per_possdouble

Example

from sportsdataverse.wbb.wbb_ncaa_lineups import ncaa_wbb_on_off
onoff = ncaa_wbb_on_off("TE-HINA.PAOPAO", lineups)
print(onoff.shape)

ncaa_wbb_player_combos​

ncaa_wbb_player_combos(lineups: 'pl.DataFrame', *, n: 'int' = 2, min_mins: 'float' = 0, included: 'Union[str, Sequence[str], None]' = None, excluded: 'Union[str, Sequence[str], None]' = None, include_transition: 'bool' = False, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Team stats for every n-player WBB combination on the court together.

Pure delegation to sportsdataverse.mbb.mbb_ncaa_lineups.ncaa_mbb_player_combos (wbigballR get_player_combos).

Parameters

ParameterTypeDefaultDescription
lineupsDataFrameLineups frame from ncaa_wbb_lineups.
nint2Combination size, 1-5.
min_minsfloat0Keep combos with total on-court minutes strictly greater than this.
includedUnion[str, Sequence[str], None]NonePlayer name(s) that must be on the court in every lineup.
excludedUnion[str, Sequence[str], None]NonePlayer name(s) that must be off the court in every lineup.
include_transitionboolFalseRe-derive the trans/half ratio surface.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per combo: team, p1..pn + the stat surface.

col_nametypedescription
teamcharacterTeam-side label or team identifier.
p1character
p2character
minsdouble
o_minsdouble
d_minsdouble
o_possdouble
d_possdouble
ortgdouble
drtgdouble
netrtgdouble
ptsdoublePoints scored.
d_ptsdouble
fgadoubleField goal attempts.
d_fgadouble
fgmdoubleField goals made.
d_fgmdouble
tpadouble
d_tpadouble
tpmdouble
d_tpmdouble
ftadoubleFree throw attempts.
d_ftadouble
ftmdoubleFree throws made.
d_ftmdouble
rimadouble
d_rimadouble
rimmdouble
d_rimmdouble
orbdouble
d_orbdouble
drbdouble
d_drbdouble
blkdoubleBlocks.
d_blkdouble
todoubleTo.
d_todouble
astdoubleAssists.
d_astdouble
e_possdouble
fg_pctdoubleField goal percentage (0-1).
d_fg_pctdouble
tppdouble
d_tppdouble
ftpdouble
d_ftpdouble
efg_pctdouble
d_efg_pctdouble
ts_pctdoubleTrue shooting percentage (0-1).
d_ts_pctdouble
rim_pctdouble
d_rim_pctdouble
mid_pctdouble
d_mid_pctdouble
tp_ratedouble
d_tp_ratedouble
rim_ratedouble
d_rim_ratedouble
mid_ratedouble
d_mid_ratedouble
ft_ratedoubleFt rate.
d_ft_ratedouble
ast_ratedouble
d_ast_ratedouble
to_ratedoubleTo rate.
d_to_ratedouble
blk_ratedouble
o_blk_ratedouble
orb_pctdoubleOffensive rebound percentage.
drb_pctdoubleDefensive rebound percentage.
time_per_possdouble
d_time_per_possdouble

Example

from sportsdataverse.wbb.wbb_ncaa_lineups import ncaa_wbb_player_combos
combos = ncaa_wbb_player_combos(lineups, n=2)
print(combos.shape)

ncaa_wbb_player_lineups​

ncaa_wbb_player_lineups(lineups: 'pl.DataFrame', *, included: 'Union[str, Sequence[str], None]' = None, excluded: 'Union[str, Sequence[str], None]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Filter a WBB lineups frame by on-court player membership.

Pure delegation to sportsdataverse.mbb.mbb_ncaa_lineups.ncaa_mbb_player_lineups (wbigballR get_player_lineups).

Parameters

ParameterTypeDefaultDescription
lineupsDataFrameLineups frame from ncaa_wbb_lineups.
includedUnion[str, Sequence[str], None]NonePlayer name(s) that must ALL be on the court.
excludedUnion[str, Sequence[str], None]NonePlayer name(s) that must NONE be on the court.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

Row-subset of lineups; schema unchanged.

col_nametypedescription
p1character
p2character
p3character
p4character
p5character
teamcharacterTeam-side label or team identifier.
minsdouble
o_minsdouble
d_minsdouble
o_possdouble
d_possdouble
ortgdouble
drtgdouble
netrtgdouble
ptsdoublePoints scored.
d_ptsdouble
fgadoubleField goal attempts.
d_fgadouble
fgmdoubleField goals made.
d_fgmdouble
tpadouble
d_tpadouble
tpmdouble
d_tpmdouble
ftadoubleFree throw attempts.
d_ftadouble
ftmdoubleFree throws made.
d_ftmdouble
rimadouble
d_rimadouble
rimmdouble
d_rimmdouble
orbdouble
d_orbdouble
drbdouble
d_drbdouble
blkdoubleBlocks.
d_blkdouble
todoubleTo.
d_todouble
astdoubleAssists.
d_astdouble
e_possdouble
fg_pctdoubleField goal percentage (0-1).
d_fg_pctdouble
tppdouble
d_tppdouble
ftpdouble
d_ftpdouble
efg_pctdouble
d_efg_pctdouble
ts_pctdoubleTrue shooting percentage (0-1).
d_ts_pctdouble
rim_pctdouble
d_rim_pctdouble
mid_pctdouble
d_mid_pctdouble
tp_ratedouble
d_tp_ratedouble
rim_ratedouble
d_rim_ratedouble
mid_ratedouble
d_mid_ratedouble
ft_ratedoubleFt rate.
d_ft_ratedouble
ast_ratedouble
d_ast_ratedouble
to_ratedoubleTo rate.
d_to_ratedouble
blk_ratedouble
o_blk_ratedouble
orb_pctdoubleOffensive rebound percentage.
drb_pctdoubleDefensive rebound percentage.
time_per_possdouble
d_time_per_possdouble

Example

from sportsdataverse.wbb.wbb_ncaa_lineups import ncaa_wbb_player_lineups
on = ncaa_wbb_player_lineups(lineups, included="TE-HINA.PAOPAO")
print(on.shape)

order_lineup​

order_lineup(player_codes_and_ids: 'list[dict[str, str]]', players_by_id: 'dict[str, dict[str, Any]]', team_season: 'str') -> 'list[dict[str, str]]'

Order a 5-man lineup X1_X2_X3_X4_X5 into PG/SG/SF/PF/C slot order.

Faithful port of PositionUtils.orderLineup (PositionUtils.ts:696-761). Greedily fits each player (in input order) to their best-scoring slot via fit_player(dominated bypos_class_to_scoreon the player'sposClass, tie-broken by their raw posConfidences), evicting and recursively re-fitting any player displaced along the way, then applies apply_relative_positional_overrides(keyed onteam_season`) as a final hand-tuned correction pass.

Parameters

ParameterTypeDefaultDescription
player_codes_and_idslist[dict[str, str]]The lineup membership, each a {"code": ..., "id": ...} dict. Order does not affect the final result (the slot-fitting algorithm is order-invariant by construction -- displaced players are always re-fit).
players_by_iddict[str, dict[str, Any]]Per-player positional info keyed by id, each a {"posConfidences": [pg, sg, sf, pf, c], "posClass": "..."} dict (the tradPosList-ordered raw confidence scores plus the classifier's ID_TO_POSITION-keyed class label).
team_seasonstrKey into RELATIVE_POSITION_FIXES for the final override pass.

Returns

A 5-element list of {"code": ..., "id": ...} dicts in PG/SG/SF/PF/C order.

Example

::

from sportsdataverse.mbb.mbb_positions import order_lineup
players_by_id = {
"Cowan, Anthony": {"posConfidences": [60, 40, 10, 0, 0], "posClass": "s-PG"},
"Ayala, Eric": {"posConfidences": [40, 60, 10, 0, 0], "posClass": "CG"},
}
order_lineup(
[{"code": "AnCowan", "id": "Cowan, Anthony"},
{"code": "ErAyala", "id": "Ayala, Eric"}],
players_by_id, "",
)

pos_class_to_score​

pos_class_to_score(pos_class: 'str') -> 'int'

Ordinal "positional weight" for a position class, PG=1000..C=8000.

Faithful port of PositionUtils.posClassToScore (PositionUtils.ts:629-654, a literal switch). Unmapped classes default to 4000 (the TS default-case comment notes "won't happen").

Parameters

ParameterTypeDefaultDescription
pos_classstrA position-class code (e.g. "PG", "WF", "C").

Returns

The class's ordinal score.

Example

::

from sportsdataverse.mbb.mbb_positions import pos_class_to_score
pos_class_to_score("WF")

project_bracket​

project_bracket(resume: 'pl.DataFrame', auto_bids: 'set[str]', *, league: 'str' = 'mens', field_size: 'int' = 68) -> 'pl.DataFrame'

Select and seed a tournament field from a per-team résumé frame.

Parameters

ParameterTypeDefaultDescription
resumeDataFrameOne row per (season, team_id) with adj_em_z, sos, wab, quad1_w (the ratings + strength-of-schedule outputs joined).
auto_bidsset[str]team_id set of conference auto-bid winners (see conference_auto_bids`); always in the field.
leaguestr'mens'"mens" or "womens" (kept for shim parity; the blend is league-agnostic).
field_sizeint68Tournament field size (68).

Returns

One row per input team: season, team_id, resume_score, projected_seed (1-16, capped for the First Four; null outside the field), at_large_prob (logistic in resume_score centred on the selection cutoff -- every selected at-large clears 0.5), auto_bid, bid (exactly field_size true).

No returns table is published for this function: no capture: it needs a resume frame (strength of schedule joined with the ratings' adj_em_z) that no package function returns.

Example

from sportsdataverse.mbb.mbb_bracketology import project_bracket
field = project_bracket(resume, auto_bids)

regress_shot_quality​

regress_shot_quality(stat: 'float', pos: 'int', feat: 'str', player: 'dict[str, Any]') -> 'float'

Shrink a small-sample shot-quality stat toward its positional average.

Faithful port of PositionUtils.regressShotQuality (PositionUtils.ts:216-258). Only the three relative shot-quality features (calc_three_relative / calc_rim_relative / calc_mid_relative) are regressed; any other feat passes stat through unchanged. A player is regressed toward the positional average whenever the relevant shot volume is below max(0.25 * total_fga, 15) (i.e. under 25% of their attempts come from that zone, floored at 15 attempts). A center (pos == 4) who took 0-2 threes and made none is left at 0 to avoid widespread changes.

Parameters

ParameterTypeDefaultDescription
statfloatThe raw (unregressed) feature value.
posintPosition index (0=pg ... 4=c).
featstrFeature field name (only the three relative shot-quality keys trigger regression; anything else is a passthrough).
playerdict[str, Any]The player stat dict; reads total_off_fga and the per-feature volume field (total_off_{3p,2pmid,2prim}_attempts), each shaped {"value": N}.

Returns

The regressed feature value (or stat unchanged when the feature is not regressed, volume is sufficient, or the center-3s carve-out fires).

Example

from sportsdataverse.mbb.mbb_positions import regress_shot_quality
player = {"total_off_fga": {"value": 25},
"total_off_3p_attempts": {"value": 1}}
regress_shot_quality(-15.5, 2, "misc_feature", player)

# Low-volume shrink toward the positional average

regress_shot_quality(100, 3, "calc_rim_relative",
{"total_off_fga": {"value": 25},
"total_off_2prim_attempts": {"value": 8}})

strength_of_schedule​

strength_of_schedule(results: 'pl.DataFrame', ratings: 'pl.DataFrame', *, league: 'str' = 'mens') -> 'pl.DataFrame'

Per-team SoS + Quad 1-4 record + WAB from completed games and ratings.

Parameters

ParameterTypeDefaultDescription
resultsDataFrameCompleted games with game_id, season, home_team_id, away_team_id, home_score, away_score, neutral_site.
ratingsDataFrameOne row per team with season, team_id, adj_em, rank (the mbb_team_ratings output). Team-id dtype must match results.
leaguestr'mens'"mens" or "womens" (quad thresholds, HFA, bubble EM).

Returns

One row per (season, team_id): season, team_id, sos, sos_rank, wab, quad1_w .. quad4_l, quality_wins. sos is the mean opponent adj_em (rank 1 = hardest schedule); quads follow the NET venue-adjusted opponent-rank thresholds; quality_wins is Quad-1 + Quad-2 wins; wab is actual wins minus a bubble-quality team's expected wins against the same schedule. Empty input returns the schema with zero rows.

No returns table is published for this function: no capture: it raises a join-key dtype mismatch on every real season (the schedule's home_team_id is Int32, the ratings' team_id is String).

Example

from sportsdataverse.mbb.mbb_strength_of_schedule import strength_of_schedule
resume = strength_of_schedule(results, ratings)

test_positional_aware_filter​

test_positional_aware_filter(sorted_to_test: 'list[dict[str, str]]', pve_frags: 'list[dict[str, Any]]', nve_frags: 'list[dict[str, Any]]') -> 'bool'

Check a positional-aware filter (from build_positional_aware_filter)

against a sorted (order_lineup-ordered) lineup array.

Faithful port of PositionUtils.testPositionalAwareFilter (PositionUtils.ts:831-858). A fragment matches if any of its position-restricted slots (or, when pos is empty, any slot at all) has a code/id containing the fragment's filter text (case-insensitive substring match). Every positive fragment must match (vacuously true if there are none); no negative fragment may match (vacuously true if there are none).

Parameters

ParameterTypeDefaultDescription
sorted_to_testlist[dict[str, str]]The ordered lineup, each a {"id": ..., "code": ...} dict (as returned by order_lineup).
pve_fragslist[dict[str, Any]]Positive-filter fragments (must ALL match).
nve_fragslist[dict[str, Any]]Negative-filter fragments (NONE may match).

Returns

Whether the lineup satisfies both the positive and negative filters.

Example

::

from sportsdataverse.mbb.mbb_positions import test_positional_aware_filter
lineup = [{"code": "AnCowan", "id": "Cowan, Anthony"}]
test_positional_aware_filter(lineup, [{"filter": "cowan", "pos": []}], [])

using_roster_pos​

using_roster_pos(pos_class: 'str', roster_pos: 'str | None') -> 'tuple[str, str | None]'

Reconcile a stats-derived position class against roster metadata.

Faithful port of PositionUtils.usingRosterPos (PositionUtils.ts:583-626). When the classifier landed on an "unsure" bucket ("G?"/"F/C?"), roster info narrows it (a roster "C" always wins outright); otherwise an obviously-wrong stats classification is compromised toward the roster-implied side, gated by pos_class_to_score thresholds.

Parameters

ParameterTypeDefaultDescription
pos_classstrThe stats-derived position class.
roster_posstr | NoneThe roster-reported position ("G"/"F"/"C"), or None/"" when unknown. if (rosterPos) (ts:587) is a plain JS truthiness check on a string -- "" and None behave identically (both mean "no correction"), so if not roster_pos is the faithful Python mirror, not an is None landmine.

Returns

A (position, info) tuple. info is None when no correction/explanation applies (matches the TS undefined), else a human-readable note on why the position was adjusted.

Example

from sportsdataverse.mbb.mbb_positions import using_roster_pos
using_roster_pos("G?", "C")

wbb_bracketology​

wbb_bracketology(season: 'int', *, as_of_date: 'Union[datetime.date, None]' = None, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Women's projected tournament field for a season.

Delegates to sportsdataverse.mbb.mbb_bracketology.mbb_bracketology with league="womens" (WBB loaders + women's constants).

Parameters

ParameterTypeDefaultDescription
seasonintSeason to project (e.g. 2024).
as_of_dateUnion[date, None]NoneOnly use games strictly before this date; None uses every completed game.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per team: season, team_id, resume_score, projected_seed, at_large_prob, auto_bid, bid -- see the mbb core for the full contract.

No returns table is published for this function: no capture: it raises a join-key dtype mismatch on every real season (the schedule's home_team_id is Int32, the ratings' team_id is String).

Example

from sportsdataverse.wbb import wbb_bracketology
field = wbb_bracketology(2024)