Skip to main content

MBB — additional Python functions — stats.ncaa.org: enrich_stats–ncaa_mbb

enrich_stats​

enrich_stats(lineup: 'LineupEvent', event_parser: 'PossessionEvent', stats: 'LineupEventStats', player_filter_coder: 'Optional[PlayerFilterCoder]' = None, player_index: 'int' = -1) -> 'LineupEventStats'

Fold a lineup's raw events into a counting-stat tree (``protected def

enrich_stats, LineupUtils.scala:115-162``). Reuses the Task 5a.3 concurrent-clump batching (~sportsdataverse.mbb.mbb_ncaa_possessions .lineup_as_raw_clumps + ~sportsdataverse.mbb.mbb_ncaa_possessions .concurrent_event_handler) rather than duplicating it -- both were already public/exported from Task 5a.3.

stats is deep-copied once up front (see the module docstring's "Scala idiom decisions"), so this function never mutates the caller's stats argument -- safe to call repeatedly against the same starting literal (e.g. a shared "empty stats" fixture).

Parameters

ParameterTypeDefaultDescription
lineupLineupEventThe lineup whose raw_game_events to fold over.
event_parserPossessionEventSelects which side (team/opponent) is "attacking".
statsLineupEventStatsThe starting stat tree (not mutated -- see above).
player_filter_coderOptional[PlayerFilterCoder]NoneOptional name -> (is_this_player, code) predicate/coder, for per-player scoping (Task 5c.4).
player_indexint-1Lineup-slot index for ~sportsdataverse.mbb .mbb_ncaa_models.PlayerShotInfo tuples (Task 5c.4; -1 for team-level calls, the only value exercised before then).

Returns

A new ~sportsdataverse.mbb.mbb_ncaa_models.LineupEventStats with every matching event folded in.

enrich_sub_error​

enrich_sub_error(location: 'str', base_id: 'str', error: 'ParseError') -> 'list[ParseError]'

Adds top-level location information to a single sub-error, returning a

list for consistency (ParseUtils.enrich_sub_error, ParseUtils.scala:91-93).

Parameters

ParameterTypeDefaultDescription
locationstrThe module in which the (now top-level) error occurred.
base_idstrAn id fragment prepended (bracket-wrapped, if non-empty) to error's existing id.
errorParseErrorThe child-parser error to enrich.

Returns

enrich_sub_errors applied to a single-element [error] list.

Example

from sportsdataverse.mbb.mbb_ncaa_data_quality import build_sub_error, enrich_sub_error

child_err = build_sub_error("game_score", error="Could not find score")
enrich_sub_error("ncaa.parse_playbyplay", "", child_err)

enrich_sub_errors​

enrich_sub_errors(location: 'str', base_id: 'str', errors: 'list[ParseError]') -> 'list[ParseError]'

Adds top-level location information to a list of sub-errors generated

by a child parser (ParseUtils.enrich_sub_errors, ParseUtils.scala:87-89).

Parameters

ParameterTypeDefaultDescription
locationstrThe module in which the (now top-level) error occurred.
base_idstrAn id fragment prepended (bracket-wrapped, if non-empty) to each error's existing id.
errorslist[ParseError]The child-parser errors to enrich (their location is replaced, not merged -- matching the Scala's ParseError(location, ..., error.messages) construction, which discards the child's own location).

Returns

A new list of ParseError, one per input error, each with location set to location and id set to build_error_id(base_id) + error.id.

Example

from sportsdataverse.mbb.mbb_ncaa_data_quality import build_sub_error, enrich_sub_errors

child_err = build_sub_error("game_time", error="Could not find time")
enrich_sub_errors("ncaa.parse_playbyplay", "", [child_err])

ensure_ev_uniqueness​

ensure_ev_uniqueness(clump: 'ConcurrentClump') -> 'ConcurrentClump'

Nudge each event's min by a tiny per-index delta so truly

concurrent (identical-min) events within a clump don't collapse under == (ensure_ev_uniqueness, LineupUtils.scala:105-111).

Parameters

ParameterTypeDefaultDescription
clumpConcurrentClumpThe clump whose events to nudge.

Returns

A new ~sportsdataverse.mbb.mbb_ncaa_possessions.ConcurrentClump with each event's min incremented by 1e-6 * index.

extract_player_from_ev​

extract_player_from_ev(shot: 'ShotEvent', pbp_event: 'MiscGameEvent', tidy_ctx: 'TidyPlayerContext') -> 'Optional[PlayerCodeId]'

Resolve the player named in pbp_event to a

~sportsdataverse.mbb.mbb_ncaa_models.PlayerCodeId (ShotEnrichmentUtils.extract_player_from_ev, PlayByPlayUtils.scala:613-635).

For a shot by the team under analysis (shot.is_off) the name is tidied against the box score before coding (so a mis-spelled play-by-play name resolves to the roster identity); for an opponent shot it is coded verbatim with no team context.

Parameters

ParameterTypeDefaultDescription
shotShotEventThe shot being enriched (only is_off is read).
pbp_eventMiscGameEventThe play-by-play event naming the player.
tidy_ctxTidyPlayerContextThe name-resolution context for this game.

Returns

The resolved PlayerCodeId, or None if the event string names no player (~sportsdataverse.mbb.mbb_ncaa_events.parse_any_play found nothing).

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import extract_player_from_ev
pc = extract_player_from_ev(shot, pbp_event, tidy_ctx)

field_keys​

field_keys(field: 'str') -> 'dict[str, str]'

Off/def stat-key names for a field (fieldKeys, ts:77-79).

Parameters

ParameterTypeDefaultDescription
fieldstrA stat field ("efg" / "3p" / "2pmid" / "2prim").

Returns

{"off": f"off_{field}", "def": f"def_{field}"}.

Example

from sportsdataverse.mbb.mbb_ncaa_strength import field_keys

keys = field_keys("3p")
print(keys["off"], keys["def"]) # off_3p def_3p

filter_matching_own​

filter_matching_own(tags: 'list[Tag]', regex: 'str') -> 'list[Tag]'

JSoup :matchesOwn(regex) applied to an already-computed candidate

list, rather than a fresh root.select(selector) call (Task 5e.2 addition; see the module docstring's note on composing this with attr_regex_filter).

Same own-text-only semantics as select_matching_own -- JSoup's Element.ownText() walks only the element's direct TextNode children, not text nested inside child elements.

Parameters

ParameterTypeDefaultDescription
tagslist[Tag]Candidate tags to filter (typically the result of an earlier .select()/attr_regex_filter call).
regexstrThe pattern each candidate's own (whitespace-collapsed) text must re.search-match.

Returns

The subset of tags whose own text contains a regex match, in the input list's order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import attr_regex_filter, filter_matching_own, parse_html
soup = parse_html('<td style="font-size:36px">92</td><td style="color:red">x</td>')
candidates = attr_regex_filter(soup.find_all("td"), "style", r"font-size:36px")
filter_matching_own(candidates, r"[0-9]+") # [<td style="font-size:36px">92</td>]

find_lineup​

find_lineup(shot: 'ShotEvent', curr_pbp: 'Optional[MiscGameEvent]', curr_lineups: 'list[LineupEvent]', lineup_it: "'PeekableIterator[LineupEvent]'") -> 'tuple[Optional[LineupEvent], list[LineupEvent]]'

Find the lineup (stint) event on the floor for shot

(ShotEnrichmentUtils.find_lineup, PlayByPlayUtils.scala:352-517).

A recursive state machine over three lists: curr_lineup (the current candidate), fallback_lineups (time-matching lineups whose raw events did not contain curr_pbp -- kept as fallbacks), and stashed_lineups (lineups pulled from the iterator but not yet stepped into, available for future shots). The branch cases (labelled 2.1-2.4 in the Scala):

  • 2.1 -- no time-matching lineup left: return the fallbacks.
  • 2.2 -- the next lineup starts after the shot: no match, stash it.
  • 2.3 -- strictly inside a lineup with no prior fallbacks: take it.
  • 2.4 -- shot is exactly at a lineup boundary (or we are already choosing among multiple fallbacks): take this lineup iff its raw game events contain curr_pbp's event string (curr_pbp is None takes it unconditionally); otherwise keep it as a fallback and recurse.

Parameters

ParameterTypeDefaultDescription
shotShotEventThe shot to place (only min / is_off are read).
curr_pbpOptional[MiscGameEvent]The already-matched play-by-play event for this shot, used to disambiguate boundary lineups; None disables that check.
curr_lineupslist[LineupEvent]Lineups pulled from the iterator on a previous call and still available (the current one first).
lineup_itPeekableIterator[LineupEvent]The shared lineup iterator (consumed in place).

Returns

(matched_lineup_or_None, lineups_to_retry_next_time) -- the second element always includes the matched lineup (so out-of-order shots sharing it still resolve) plus any leftover stash.

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import (
PeekableIterator,
find_lineup,
)
matched, retry = find_lineup(shot, None, [lineup], PeekableIterator([]))

find_missing_subs​

find_missing_subs(clump: 'BadLineupClump', box_lineup: 'LineupEvent', valid_player_codes: 'set[str]') -> 'tuple[list[LineupEvent], BadLineupClump]'

Trims a clump whose lineups carry TOO MANY players by identifying the

"ghost" player(s) a missing sub-out left behind (LineupErrorAnalysisUtils.find_missing_subs, :406-514).

Fires only when the clump's first event has >= 6 on-floor players (:415-416: candidates.size < 6 is a no-op). expected_size_diff (:419) is first_event_player_count - 5 -- the number of ghosts the trim should end up removing.

Phase 1 -- shrink the candidate pool (:437-478). Starting from the first event's players, walk the clump chronologically. At each event a candidate is confirmed present (and dropped from the pool) if it subs out (ev.players_out, skipped for the first event -- :445, literal port of clump.evs.headOption.contains(ev) as value equality ev == clump.evs[0]; for a well-formed clump of distinct events this is exactly index == 0) or is named in one of the event's team-side raw plays (parse_any_play -> ~sportsdataverse.mbb.mbb_ncaa_names .tidy_player -> ~sportsdataverse.mbb.mbb_ncaa_stints .build_player_code, :448-456; unlike validate_lineup this does NOT skip the literal "team" token -- ported verbatim). matching_index is the FIRST event index at which the pool size first equals expected_size_diff (:475); once set it freezes -- all later events are skipped in phase 1 (:439-441).

Accept gate (:479-480): the final pool must be non-empty and no larger than expected_size_diff. If matching_index never fired (the pool jumped past expected_size_diff in a single step, or never shrank to it), the gate still accepts iff the residual pool is a non-empty subset of size <= expected_size_diff -- in which case phase 3 routes every event through the "before match" branch (index > None is always false). On failure the original clump is returned unchanged.

Phase 3 -- rebuild the events (:482-503, a scanLeft ported as a manual accumulate loop that drops the seed). For events at/before matching_index the ghost pool is simply removed from players (filterNot). For events strictly after matching_index (index > matching_index -- the matched event itself is "before") players is rebuilt from the previous tidied event via ~sportsdataverse.mbb.mbb_ncaa_stints.build_new_player_list (the scanLeft threads the previously-emitted event; its seed is None, but the first event can never be an "after match" event, so the getOrElse(ev) fallback is only ever a formality -- ported faithfully all the same). The rebuilt events are partitioned by validate_lineup.

Parameters

ParameterTypeDefaultDescription
clumpBadLineupClumpThe bad-lineup clump to attempt to repair.
box_lineupLineupEventThe team's box-score lineup event (roster + name context).
valid_player_codesset[str]Every player code on the box score / roster.

Returns

(fixed_lineups, still_to_fix) -- the now-valid rebuilt events and a clump of the still-invalid ones (carrying the input's next_good); or ([], clump) on a no-op / rejected fix.

Example

from sportsdataverse.mbb.mbb_ncaa_stint_validation import (
find_missing_subs,
)
fixed, still = find_missing_subs(clump, box_lineup, valid_codes)

find_pbp_clump​

find_pbp_clump(shot_time: 'float', pbp_it: "'PeekableIterator[PlayByPlayEvent]'", curr_pbp_clump: 'list[MiscGameEvent]', maybe_next_pbp_event: 'Optional[MiscGameEvent]') -> 'tuple[list[MiscGameEvent], Optional[MiscGameEvent]]'

Gather every play-by-play shot/assist event sharing shot_time

(ShotEnrichmentUtils.find_pbp_clump, PlayByPlayUtils.scala:556-608).

If curr_pbp_clump (carried over from the previous shot) already holds events at shot_time they are returned as-is; otherwise the iterator is walked forward, discarding earlier events, accumulating the equal-time ones, and stopping (returning it as maybe_next_pbp_event) at the first later event.

Parameters

ParameterTypeDefaultDescription
shot_timefloatThe shot's game-clock minute to gather events for.
pbp_itPeekableIterator[PlayByPlayEvent]The shared play-by-play iterator (consumed in place).
curr_pbp_clumplist[MiscGameEvent]Events left over from the previous shot's clump.
maybe_next_pbp_eventOptional[MiscGameEvent]The look-ahead event stashed by the previous call, if any.

Returns

(clump, maybe_next) -- the equal-time events, plus the first strictly-later event (or None at end of stream).

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import (
PeekableIterator,
find_pbp_clump,
)
clump, nxt = find_pbp_clump(5.0, PeekableIterator([]), [], None)
# ([], None)

fix_combos​

fix_combos(first: 'str', last: 'str', code_start: 'Optional[str]' = None) -> 'list[tuple[str, Optional[str]]]'

Pair each of combos' three name variants with a shared

player-code override (DataQualityIssues.fix_combos, DataQualityIssues.scala:340-346).

Parameters

ParameterTypeDefaultDescription
firststrThe player's first name.
laststrThe player's last name.
code_startOptional[str]NoneThe forced player-code prefix for every variant, or None to leave the default build_player_code truncation behavior in place.

Returns

Three (name_variant, code_start) pairs.

fix_possible_score_swap_bug​

fix_possible_score_swap_bug(lineup: 'list[LineupEvent]', box_lineup: 'LineupEvent') -> 'list[LineupEvent]'

Undo a rare NCAA data bug where the scores get transposed

(fix_possible_score_swap_bug, LineupUtils.scala:51-90).

If the last lineup's ending score is the exact transpose of the box score's ending score, every lineup's score_info is un-transposed and pts/plus_minus are swapped/negated between team_stats and opponent_stats -- nothing else in the stat trees changes.

Parameters

ParameterTypeDefaultDescription
lineuplist[LineupEvent]The lineups to (maybe) fix, in chronological order.
box_lineupLineupEventThe trusted box-score lineup to compare the final score against.

Returns

lineup unchanged if the scores aren't transposed (or lineup is empty); otherwise a new list with every entry's score/pts/ plus_minus corrected.

get_ascending_time​

get_ascending_time(event: 'ShotEvent', period: 'int', is_women_game: 'bool') -> 'float'

Converts the descending in-period clock time to an ascending

game-elapsed time (ShotEventParser.get_ascending_time, :531-537).

Parameters

ParameterTypeDefaultDescription
eventShotEventThe shot event (only ~sportsdataverse.mbb .mbb_ncaa_models.ShotEvent.min, the raw descending clock minute, is read).
periodintThe 1-indexed period the shot was taken in.
is_women_gameboolWhether to use women's-quarters (10min) or men's- halves (20min, then 5min OTs) period lengths.

Returns

The ascending game-elapsed time, in minutes.

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import get_ascending_time
get_ascending_time(shot_with_min_4, period=1, is_women_game=False) # 16.0

get_box_lineup​

get_box_lineup(filename: 'str', in_html: 'str', team_id: 'TeamId', format_version: 'int', external_roster: 'tuple[list[str], list[RosterEntry]]' = ([], []), neutral_game_dates: 'AbstractSet[str]' = frozenset(), home_team: 'Optional[str]' = None, away_team: 'Optional[str]' = None) -> 'Union[LineupEvent, list[ParseError]]'

Gets the boxscore lineup from the HTML page (``BoxscoreParser

.get_box_lineup, :122-222``).

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name -- used both for error reporting and to extract the period via parse_period_from_filename (e.g. "test_p2.html").
in_htmlstrThe raw box-score-page HTML.
team_idTeamIdThe team this box score is being parsed for.
format_versionint0 for the legacy layout, 1 for the 2018+ layout (see the module docstring's selector-translation notes).
external_rostertuple[list[str], list[RosterEntry]]([], [])(other_players, roster_players) -- either just names, or a full roster, to validate/fuzzy-correct box names against (see inject_validated_players). Also seeds ~sportsdataverse.mbb.mbb_ncaa_models.LineupEvent.players_out on the interim lineup (roster_players, each's code replaced by its jersey number).
neutral_game_datesAbstractSet[str]frozenset()Date strings (the first whitespace-separated token of the raw date-cell text) known to be neutral-site games -- overrides the default home/away inference.
home_teamOptional[str]NoneThe game's home team when the caller already knows it, forwarded to team-name resolution so a box page that names only one side (a non-D-I opponent has no header) still resolves.
away_teamOptional[str]NoneThe game's away team, same purpose. Both are required together or neither is used.

Returns

A ~sportsdataverse.mbb.mbb_ncaa_models.LineupEvent whose players is the validated box-score lineup (natural HTML order -- see the module docstring's "not sorted" note), or a list[ParseError] if any parsing step failed.

Example

from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import get_box_lineup
from sportsdataverse.mbb.mbb_ncaa_models import TeamId

with open("tests/fixtures/ncaa/test_lineup.html", encoding="utf-8") as f:
html = f.read()
result = get_box_lineup("test_p1.html", html, TeamId("TeamA"), format_version=0)

get_config​

get_config() -> 'NcaaFetchConfig'

Return the live NcaaFetchConfig singleton.

Returns

The live singleton (cache_dir, proxy_url, the ProxyBonanza settings, timeout, impersonate, max_retries and the rotation / Terms-gate backoffs, transport).

Example

from sportsdataverse.mbb.mbb_ncaa_fetch import get_config
cfg = get_config()
print(cfg.cache_dir, cfg.timeout)

get_game_weight​

get_game_weight(opp: 'OpponentGame', field: 'str', side: 'str') -> 'float'

Weight for one game/field/side (getGameWeight, ts:119-140).

The field-specific shot volume (FGA for efg, 3PA for 3p, 2pmid_attempts / 2prim_attempts for the mid/rim fields); when that is 0 (no shots of that type), falls back to off_poss / def_poss so the game still carries weight.

Parameters

ParameterTypeDefaultDescription
oppOpponentGameOne opponent game dict.
fieldstrA stat field.
sidestr"off" or "def".

Returns

The (non-negative) game weight.

Example

from sportsdataverse.mbb.mbb_ncaa_strength import get_game_weight

game = {"off_3p_attempts": 0, "off_poss": 70}
print(get_game_weight(game, "3p", "off")) # 70.0 (poss fallback)

get_neutral_games​

get_neutral_games(filename: 'str', in_html: 'str', format_version: 'int') -> 'Union[tuple[TeamId, set[str]], list[ParseError]]'

Extracts the set of neutral/away-marked game dates from a saved NCAA

team-schedule page (TeamScheduleParser.get_neutral_games, TeamScheduleParser.scala:63-94).

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name, used only for error reporting.
in_htmlstrThe raw team-schedule-page HTML.
format_versionint0 for the legacy fieldset/legend layout, 1 for the 2018+ div.card-header/div.card-body layout.

Returns

(team, neutral_game_dates) -- the team parsed from the page's image alt attribute, and every "MM/DD/YYYY" date string found on an "@Opponent"-marked row -- or a single-element list[ParseError] if the HTML fails to parse, or the team name can't be located.

Example

from sportsdataverse.mbb.mbb_ncaa_team_parsers import get_neutral_games

with open("tests/fixtures/ncaa/test_schedule.html", encoding="utf-8") as f:
html = f.read()
result = get_neutral_games("test_schedule.html", html, format_version=0)
if isinstance(result, list):
raise RuntimeError(result) # list[ParseError]
team, neutral_dates = result

get_per_game_raw​

get_per_game_raw(opp: 'OpponentGame', field: 'str', side: 'str') -> 'Optional[float]'

Per-game raw shooting rate from one opponent row (getPerGameRaw, ts:82-116).

efg is (2pmid_made + 2prim_made + 1.5 * 3p_made) / (2pmid_att + 2prim_att + 3p_att); 3p / 2pmid / 2prim are made / attempts. Every counter read is nullish (missing -> 0); the sole guard is on total attempts.

Parameters

ParameterTypeDefaultDescription
oppOpponentGameOne opponent game dict.
fieldstrA stat field; an unknown field returns None.
sidestr"off" or "def" (selects the off_/def_ prefix).

Returns

The rate as a float, or None when the relevant attempts total is <= 0 (game skipped by the weighted means -- not a 0-rate).

Example

from sportsdataverse.mbb.mbb_ncaa_strength import get_per_game_raw

game = {"off_3p_made": 4, "off_3p_attempts": 10}
print(get_per_game_raw(game, "3p", "off")) # 0.4

get_sorted_pbp_events​

get_sorted_pbp_events(filename: 'str', in_html: 'str', box_lineup: 'LineupEvent', format_version: 'int') -> 'Union[list[PlayByPlayEvent], list[ParseError]]'

Handy util to return the play-by-play events in chronological order,

used in a few other places (PlayByPlayParser.get_sorted_pbp_events, :221-239).

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name, used only for error reporting.
in_htmlstrThe raw play-by-play-page HTML.
box_lineupLineupEventThe team's box-score lineup (supplies team/year).
format_versionint0 for the legacy layout, 1 for the 2018+ layout.

Returns

The play-by-play events in chronological (earliest-to-latest) order, or a list[ParseError] on failure. enrich=True is used internally to get the correct ascending timestamps, and its reversal is undone here (.reverse) to restore chronological order.

Example

from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import get_box_lineup
from sportsdataverse.mbb.mbb_ncaa_models import TeamId
from sportsdataverse.mbb.mbb_ncaa_pbp_parser import get_sorted_pbp_events

with open("tests/fixtures/ncaa/test_lineup.html", encoding="utf-8") as f:
box_html = f.read()
box_lineup = get_box_lineup("test_p1.html", box_html, TeamId("TeamA"), format_version=0)

with open("tests/fixtures/ncaa/test_play_by_play.html", encoding="utf-8") as f:
pbp_html = f.read()
events = get_sorted_pbp_events("test.html", pbp_html, box_lineup, format_version=0)

get_team_raw_from_per_game​

get_team_raw_from_per_game(team: 'TeamDetail', field: 'str') -> 'SideValues'

A team's field rate as the weighted mean of its per-game raws (getTeamRawFromPerGame, ts:224-250).

Same accumulation as compute_league_averages_from_per_game but scoped to one team's games; empty -> 0.

Parameters

ParameterTypeDefaultDescription
teamTeamDetailA team_details team dict.
fieldstrA stat field.

Returns

{"off": float, "def": float}.

Example

from sportsdataverse.mbb.mbb_ncaa_strength import get_team_raw_from_per_game

team = {"opponents": [{"off_3p_made": 4, "off_3p_attempts": 10}]}
print(get_team_raw_from_per_game(team, "3p")["off"]) # 0.4

get_team_triples​

get_team_triples(filename: 'str', in_html: 'str', old_format: 'bool' = False) -> 'Union[list[tuple[TeamId, str, ConferenceId]], list[ParseError]]'

Extracts (team, NCAA id, conference) triples from a saved NCAA

team-list/attendance page (TeamIdParser.get_team_triples, TeamIdParser.scala:69-91).

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name, used only for error reporting.
in_htmlstrThe raw team-list-page HTML.
old_formatboolFalseTrue for pages where the team name and conference are in separate <td>s; False (default) for pages where the conference is embedded in the team-name cell as "Team (Conf)".

Returns

One (TeamId, ncaa_id, ConferenceId) triple per row that has both a resolvable id and name/conference (rows missing either are silently skipped, matching the Scala's case _ => Nil), or a single-element list[ParseError] if the HTML itself fails to parse.

Example

from sportsdataverse.mbb.mbb_ncaa_team_parsers import get_team_triples

with open("tests/fixtures/ncaa/test_attendance_list.html", encoding="utf-8") as f:
html = f.read()
result = get_team_triples("test_attendance_list.html", html, old_format=True)

get_unified_ncaa_id​

get_unified_ncaa_id(filename: 'str', in_html: 'str') -> 'Union[Optional[str], list[ParseError]]'

Gets a player's lowest cross-season NCAA id from a saved player page

(RosterParser.get_unified_ncaa_id, RosterParser.scala:136-152).

Always uses the v1 selector table -- this bonus lookup only exists on 2018+-era pages.

Parameters

ParameterTypeDefaultDescription
filenamestrThe source file name, used only for error reporting.
in_htmlstrThe raw player-page HTML.

Returns

The numerically-lowest NCAA id found, None if the page has no tr[id^=player_season_] rows, or a single-element list[ParseError] if the HTML couldn't be parsed at all.

Example

from sportsdataverse.mbb.mbb_ncaa_roster_parser import get_unified_ncaa_id
get_unified_ncaa_id("player.html", player_page_html)

handle_common_sub_bug​

handle_common_sub_bug(clump: 'BadLineupClump', box_lineup: 'LineupEvent', valid_player_codes: 'set[str]') -> 'tuple[list[LineupEvent], BadLineupClump]'

Fixes the "2-in-1-out then a compensating 1-out" substitution bug

(LineupErrorAnalysisUtils.handle_common_sub_bug, :269-298).

Handles a single-event bad clump whose next known-good lineup carries a lone sub-out that the clump's event should have applied but didn't (e.g. IN: X, Y, Z; OUT: A, B in the bad event, then OUT: C in the good one). Fires only when all three guard conditions hold (:275-278):

  • the bad event has more players subbing IN than OUT (len(players_in) > len(players_out)),
  • the good event has no sub-ins (len(good.players_in) == 0 -- otherwise there's no way to tell which of its sub-outs to borrow), and
  • the good event has at least one sub-out (len(good.players_out) > 0).

The fix (:279-283) removes the good event's sub-outs from the bad event's on-floor players (value-equality not in -- the Scala's filterNot(good.players_out.toSet)) and appends them to the bad event's players_out (order-preserving distinct-- the Scala's(bad.players_out ++ good.players_out).distinct). The fix is **accepted only if** the result then passes validate_lineup (:284-294): on success the fixed event is returned as the sole fixedlineup and the still-to-fix clump is emptied; on failure the *fixed* event (not the original) is returned as the still-to-fix clump, keeping the samenext_good` so a later pass can try again.

Parameters

ParameterTypeDefaultDescription
clumpBadLineupClumpThe bad-lineup clump to attempt to repair.
box_lineupLineupEventThe team's box-score lineup event (roster + name context).
valid_player_codesset[str]Every player code on the box score / roster.

Returns

(fixed_lineups, still_to_fix) -- fixed_lineups is [fixed] on an accepted fix else []; still_to_fix is an empty clump on accept, the (unchanged) input clump on a guard miss, or the single-event fixed-but-still-invalid clump on a rejected fix.

Example

from sportsdataverse.mbb.mbb_ncaa_stint_validation import (
handle_common_sub_bug,
)
fixed, still = handle_common_sub_bug(clump, box_lineup, valid_codes)

inject_starting_lineup_into_box​

inject_starting_lineup_into_box(sorted_pbp_events: 'list[PlayByPlayEvent]', box_lineup: 'LineupEvent', external_roster: 'tuple[list[str], list[RosterEntry]]', format_version: 'int') -> 'LineupEvent'

Infer the starting five and reorder the box-score roster so they lead

(PlayByPlayUtils.inject_starting_lineup_into_box, PlayByPlayUtils.scala:684-845).

The v1 (2018+) NCAA box score dropped the ordered list of starters, so we reconstruct it from the play-by-play sub sequencing. A player is a starter if, walking the events forward, they are seen before their first sub-in -- either subbed out (before ever being subbed in), or named in a team-side play that isn't concurrent with a sub. Anyone subbed in before ever being seen is excluded. The reconstruction stops once five starters are found.

Parameters

ParameterTypeDefaultDescription
sorted_pbp_eventslist[PlayByPlayEvent]The full play-by-play event stream, ascending time.
box_lineupLineupEventThe box-score lineup event (its players is the full roster to reorder).
external_rostertuple[list[str], list[RosterEntry]]Unused here -- carried for signature parity with the Scala (its pipeline caller passes it). See the module note.
format_versionintUnused here -- carried for signature parity. See the module note.

Returns

A copy of box_lineup with players reordered so the inferred starters lead. If fewer than five starters could be inferred (a "40-trillion" player who was never subbed nor mentioned), the roster is ordered starters -> possible-starters -> definitely-not-starters as the best available guess.

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import inject_starting_lineup_into_box
fixed = inject_starting_lineup_into_box(pbp_events, box_lineup, ([], []), 1)

inject_validated_players​

inject_validated_players(ordered_lineup_from_box: 'list[str]', box_minus_players: 'LineupEvent', external_roster: 'tuple[list[str], list[RosterEntry]]') -> 'list[str]'

Validates box players against the roster (if available) and any

other available box scores (BoxscoreParser.inject_validated_players, :233-279).

See the module docstring's "un-threaded tidy_ctx" note -- every fuzzy-resolution call inside the loop uses the SAME original context, never the updated one a call returns (ported verbatim, including this apparent Scala oversight).

Parameters

ParameterTypeDefaultDescription
ordered_lineup_from_boxlist[str]The raw player-name strings scraped straight off the box-score page (already v0-normalized if the source was v1, by get_box_lineup's caller).
box_minus_playersLineupEventThe in-progress ~sportsdataverse.mbb.mbb_ncaa_models.LineupEvent (used only for its team field, both to scope the fuzzy-match context and to key ~sportsdataverse.mbb.mbb_ncaa_data_quality.players_missing_from_boxscore).
external_rostertuple[list[str], list[RosterEntry]](other_players, roster_players) -- extra known player names, and a full team roster (if available) to validate against / fuzzy-correct box names onto.

Returns

ordered_lineup_from_box with any name not found in roster_players fuzzy-corrected onto the closest roster name (if a roster was supplied at all), followed by any roster/other/ known-missing players not already present in that corrected list (see the module docstring's "Extra-players Set ordering" note for this trailing group's order).

Example

from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import inject_validated_players
inject_validated_players(["Player One"], box_lineup, ([], []))

is_cached​

is_cached(path: 'str', *, cache_dir: 'Optional[Path]' = None) -> 'bool'

Return whether path already has a cache file on disk.

Parameters

ParameterTypeDefaultDescription
pathstrThe stats.ncaa.org URL path (with its query string, if any).
cache_dirOptional[Path]NoneThe cache root; the active config's cache_dir when None.

Returns

True when cached_path already exists on disk.

is_end_of_game_fouling_vs_fastbreak​

is_end_of_game_fouling_vs_fastbreak(curr_clump: 'ConcurrentClump', event_parser: 'PossessionEvent') -> 'bool'

Check for intentional fouling to prolong the game, specifically so it

can be excluded from being counted as a fast break (is_end_of_game_fouling_vs_fastbreak, LineupUtils.scala:603-656).

Parameters

ParameterTypeDefaultDescription
curr_clumpConcurrentClumpThe clump to classify.
event_parserPossessionEventSelects which side of each event is "attacking".

Returns

True iff the FIRST attacking-side FT-made/FT-missed event in curr_clump.evs is both near the end of a period AND has the attacking team ahead by (0, 10] points; False if no such event exists.

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import is_end_of_game_fouling_vs_fastbreak

is_end_of_game_fouling_vs_fastbreak(curr_clump, event_parser)

is_gen2​

is_gen2(ev: 'RawGameEvent') -> 'bool'

Detect the new/"gen2" NCAA event format (EventUtils.is_gen2,

EventUtils.scala:12-14).

Parameters

ParameterTypeDefaultDescription
evRawGameEventThe raw game event to inspect.

Returns

True if ev.info contains a comma-space (", "), the gen2 format's field separator; False for the old/legacy format.

is_scramble​

is_scramble(curr_clump: 'ConcurrentClump', prev_clumps: 'list[ConcurrentClump]', event_parser: 'PossessionEvent', player_version: 'bool') -> 'tuple[Callable[[RawGameEvent], bool], str]'

Figure out if (each event of) the current clump is part of a

"scramble scenario" following an ORB (is_scramble, LineupUtils .scala:222-597).

Returns a (predicate, debug_tag) tuple -- the tuple shape is load-bearing: the oracle asserts the debug tag string directly ("N/A"/"0a"/"1aa"/"1ab"/"1b"/"2aa"/"2ab").

Parameters

ParameterTypeDefaultDescription
curr_clumpConcurrentClumpThe clump to classify.
prev_clumpslist[ConcurrentClump]Prior merged clumps, most-recent-first.
event_parserPossessionEventSelects which side of each event is "attacking".
player_versionboolUnused -- see the module docstring's is_scramble port notes (the Scala's debug-print gate this flag controls is permanently false regardless of its value).

Returns

(predicate, debug_tag) where predicate(ev) reports whether ev is part of a scramble.

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import is_scramble

predicate, tag = is_scramble(curr_clump, prev_clumps, event_parser, player_version=False)
[predicate(ev) for ev in curr_clump.evs]

is_team_shooting_left_to_start​

is_team_shooting_left_to_start(sorted_very_raw_events: 'list[tuple[int, ShotEvent]]') -> 'tuple[bool, int]'

Infers which side of the SVG court the team under analysis shoots

towards in the first period, from its own made/missed shot locations (ShotEventParser.is_team_shooting_left_to_start, :540-555).

Parameters

ParameterTypeDefaultDescription
sorted_very_raw_eventslist[tuple[int, ShotEvent]]The chronologically-sorted (period, shot) pairs, pre-geometry-transform.

Returns

(team_shooting_left_in_first_period, first_period) -- the first element of sorted_very_raw_events, if any, determines first_period; the majority side (by count) of the team's own (is_off) shots within that period determines the direction.

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import is_team_shooting_left_to_start
is_team_shooting_left_to_start([(1, shot_a), (1, shot_b)])

is_women_game​

is_women_game(sorted_very_raw_events: 'list[tuple[int, ShotEvent]]') -> 'bool'

Infers men's vs. women's game from timing evidence

(ShotEventParser.is_women_game, :558-566). Shot-parser-specific variant -- distinct from the play-by-play parser's own is_women_game (Task 5e.3), which uses PbP event timing instead of shot timing; the plan's recon flags both as "its OWN is_women_game variant" per module.

Parameters

ParameterTypeDefaultDescription
sorted_very_raw_eventslist[tuple[int, ShotEvent]]The chronologically-sorted (period, shot) pairs.

Returns

True if at least 4 periods were seen AND no shot was taken with more than 10 minutes showing on the (descending) clock in the very first event (women's quarters are 10 minutes; a shot at >10:00 remaining could only happen in a longer men's period).

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import is_women_game
is_women_game([(1, shot), (2, shot), (3, shot), (4, shot)]) # True

jsoup_text​

jsoup_text(el: 'Optional[Tag]') -> 'str'

JSoup Element.text(): all descendant text, whitespace-collapsed.

JSoup's .text() joins every text node under el (including descendants) and collapses runs of whitespace (spaces, tabs, newlines) into single spaces, trimming the ends. bs4's .get_text() does the joining but not the collapsing, so captured HTML's indentation/newlines would otherwise leak into every extracted value.

Parameters

ParameterTypeDefaultDescription
elOptional[Tag]The element to extract text from, or None.

Returns

The whitespace-collapsed text, or "" if el is None.

Example

from sportsdataverse.mbb.mbb_ncaa_html import jsoup_text, parse_html
soup = parse_html("<td>\n Akin,\tDaniel </td>")
jsoup_text(soup.find("td")) # "Akin, Daniel"

lineup_stats_bucket​

lineup_stats_bucket(ev: 'LineupEvent', *, avg_eff: 'float' = 100.0, opponent_baselines: 'Optional[dict[str, float]]' = None, doc_count: 'int' = 1) -> 'LineupStatSet'

Assemble one lineup's full 254-field {value} bucket.

lineup_stats_bucket is the Python entry point for stage 2 of the port (see the module docstring) -- the faithful composition of this module's factories in the order commonLineupAggregations.ts (572-line ES aggregation) issues them: sum ( sum_fields) -> merge the play-type pts/poss bucket_script ( play_type_pts_poss) -> mint every other rate bucket_script ( all_rate_fields) -> the SOS-adjusted-efficiency bucket_script ( adj_fields).

Parameters

ParameterTypeDefaultDescription
evLineupEventOne already-summed lineup event (team_stats/opponent_stats populated by stage 1, ~sportsdataverse.mbb.mbb_ncaa_lineup_enrich.enrich_lineup).
avg_efffloat100.0League-average efficiency passed through to adj_fields`.
opponent_baselinesOptional[dict[str, float]]NoneSOS baseline lookup passed through to adj_fields; only None` (no baselines) is implemented.
doc_countint1The ES doc_count for this bucket (number of raw events folded in).

Returns

The full bucket: every total_*/rate/adj field wrapped in {"value": <float>}, plus the structural keys key, players_array, doc_count (bare, unwrapped).

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_aggregation import lineup_stats_bucket

bucket = lineup_stats_bucket(enriched_event, doc_count=7)
bucket["off_ppp"]["value"]

lineup_stats_buckets​

lineup_stats_buckets(evs: 'list[LineupEvent]', *, avg_eff: 'float' = 100.0, opponent_baselines: 'Optional[dict[str, float]]' = None) -> 'list[LineupStatSet]'

Group events by lineup, fold each group's stats, and mint one bucket per lineup.

Python entry point for the ES terms aggregation over key (grouping by bucket_key) that feeds each lineup's docs into commonLineupAggregations.ts's sumaggs -- seecbb-on-off-analyzer/src/utils/es-queries/commonLineupAggregations.ts. This is the list-form producer the LineupStatSet consumers (mbb_lineup_stats) read from; lineup_stats_bucket` handles a single already-folded event.

Parameters

ParameterTypeDefaultDescription
evslist[LineupEvent]Raw per-possession-chunk lineup events (team_stats/opponent_stats populated by stage 1, ~sportsdataverse.mbb.mbb_ncaa_lineup_enrich .enrich_lineup), one lineup's floor time possibly split across many events.
avg_efffloat100.0League-average efficiency passed through to each bucket.
opponent_baselinesOptional[dict[str, float]]NoneSOS baseline lookup passed through to each bucket; only None (no baselines) is implemented (see adj_fields`).

Returns

One LineupStatSet per distinct lineup (bucket_key), in first-seen order, with doc_count` set to that lineup's event count.

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_aggregation import lineup_stats_buckets

buckets = lineup_stats_buckets(enriched_events)
buckets[0]["off_poss"]["value"]

load_proxybonanza_pool​

load_proxybonanza_pool(api_key: 'str', pkg: 'str', *, transport: 'Optional[PoolTransport]' = None) -> "'list[str]'"

Resolve a ProxyBonanza package into a list of http://login:pass@ip:port URLs.

Graduated from dev/ncaa_proxy.py's load_proxy_pool -- same endpoint shape, minus the .Renviron reader (creds are now explicit params, per the creds-hygiene directive).

Endpoint: GET https://api.proxybonanza.com/v1/userpackages/{pkg}.json

Parameters

ParameterTypeDefaultDescription
api_keystrProxyBonanza API key.
pkgstrProxyBonanza package id.
transportOptional[PoolTransport]NoneInjectable (url, headers) -> (status, text) callable for offline testing. Defaults to a curl_cffi GET.

Returns

One http://login:password@ip:port URL per IP in the package.

Example

def fake(url, headers):
return 200, '{"data": {"login": "u", "password": "p", "ippacks": []}}'
pool = load_proxybonanza_pool("key", "pkg", transport=fake)

matching_player​

matching_player(shot: 'ShotEvent', pbp_event: 'MiscGameEvent', tidy_ctx: 'TidyPlayerContext', code_match: 'bool') -> 'bool'

Whether the player in pbp_event matches shot's shooter

(ShotEnrichmentUtils.matching_player, PlayByPlayUtils.scala:638-652).

Parameters

ParameterTypeDefaultDescription
shotShotEventThe shot being enriched.
pbp_eventMiscGameEventThe candidate play-by-play event.
tidy_ctxTidyPlayerContextThe name-resolution context.
code_matchboolIf True, compare on player code only (looser -- lets a name that resolves to the wrong identity but the right code match); if False, require full ~sportsdataverse.mbb .mbb_ncaa_models.PlayerCodeId equality.

Returns

True if the resolved player matches shot.player under the selected comparison, else False.

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import matching_player
matching_player(shot, pbp_event, tidy_ctx, code_match=False)

misspellings​

misspellings(team: 'Optional[TeamId]') -> 'dict[str, str]'

Team-scoped misspelling map, falling back to the generic map

(DataQualityIssues.misspellings, DataQualityIssues.scala:165-322 -- see the module docstring's "fallback semantics" note for why this is a precomputed merge, not a runtime two-level lookup).

Parameters

ParameterTypeDefaultDescription
teamOptional[TeamId]The team to look up team-specific corrections for. None (like any team absent from the table) falls back to generic_misspellings.

Returns

A fresh dict -- the team's misspelling map merged with generic_misspellings, or a copy of generic_misspellings if team has no specific entries.

Example

from sportsdataverse.mbb.mbb_ncaa_data_quality import misspellings
from sportsdataverse.mbb.mbb_ncaa_models import TeamId

misspellings(TeamId("NJIT"))["Lewal, Levi"] # 'Lawal, Levi'
misspellings(TeamId("Some Unlisted Team")) # {} (generic fallback)
misspellings(None) # {} (generic fallback)

name_in_v0_box_format​

name_in_v0_box_format(v1_name: 'str') -> 'str'

Switch a v1-box-format name ("first_name names") to v0-box format

("names, first_name") (ExtractorUtils.scala:59-81).

Handles a v0-PbP-style all-caps input ("SURNAME,NAME", still seen in older files even in v1-format seasons) by first flipping it to guaranteed v1 shape, then splits on the first space to get first/last. A last starting with "(" is treated as a nickname parenthetical (e.g. "Russell (Deuce) Dean") and re-split via COMPLEX_V0_CASE_RE-- **ported verbatim including its literal quirk**: the regex's second capture group keeps the leading space before the trailing surname (e.g. yields" Dean, Russell (Deuce)", not "Dean, Russell Deuce"` as the Scala source comment's stated intent describes) and the parenthesis characters are not stripped. This is upstream behavior, not a Python-side bug -- the Scala's own pattern-match reproduces exactly this, so faithful porting keeps it.

Parameters

ParameterTypeDefaultDescription
v1_namestrThe player name as it appears in a v1 (2018+) NCAA roster or box-score row.

Returns

The name in v0 ("names, first_name") format, or v1_name (via the guaranteed-v1-format intermediate) unchanged if it has no space to split on.

Example

from sportsdataverse.mbb.mbb_ncaa_stints import name_in_v0_box_format
name_in_v0_box_format("Daniel Akin") # "Akin, Daniel"
name_in_v0_box_format("AKIN,DANIEL") # "AKIN, DANIEL" (old PbP form, flipped then re-split)

ncaa_mbb_box_scores​

ncaa_mbb_box_scores(game_ids: 'Union[str, int, Iterable[Union[str, int]]]', *, multi_games: 'bool' = False, fetcher: 'Optional[_SupportsFetchIndividualStats]' = None, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Box scores for one or more NCAA games (bigballR get_box_scores port).

Multi-game driver over parse_ncaa_bb_box (bigballR/R/all_functions.R:3603-3678): drops null ids, isolates per-game errors (failed ids are reported and skipped), binds rows, and optionally aggregates across games.

Parameters

ParameterTypeDefaultDescription
game_idsUnion[str, int, Iterable[Union[str, int]]]One id or an iterable of NCAA contest ids.
multi_gamesboolFalseWhen True, aggregate one row per (player, clean_name, team) -- counters summed, g = games played, rates recomputed from the sums (R's multi.games; grouping adapted per module docstring).
fetcherOptional[_SupportsFetchIndividualStats]NoneOptional injected fetcher exposing fetch_game_individual_stats (for offline replay/tests). Defaults to a fresh NcaaFetcher.with_browser() context per call -- stats.ncaa.org sits behind an Akamai challenge that the plain transport cannot clear.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

polars.DataFrame (or pandas with return_as_pandas=True): per-game rows in the parse_ncaa_bb_box contract, or the aggregated multi_games contract.

col_nametypedescription
game_idcharacterUnique game identifier.
box_idcharacter
playercharacterPlayer name.
clean_namecharacter
teamcharacterTeam-side label or team identifier.
mpdoubleMinutes played.
ptsdoublePoints scored.
orbdouble
drbdouble
trbdoubleCareer total rebounds.
astdoubleAssists.
todoubleTo.
stldoubleSteals.
blkdoubleBlocks.
fgadoubleField goal attempts.
fgmdoubleField goals made.
fg_pctdoubleField goal percentage (0-1).
tpadouble
tpmdouble
tp_pctdouble
ftadoubleFree throw attempts.
ftmdoubleFree throws made.
ft_pctdoubleFree throw percentage (0-1).
ts_pctdoubleTrue shooting percentage (0-1).
efg_pctdouble
foulsdoublePersonal fouls.
dqdouble
techdouble

Example

from sportsdataverse.mbb.mbb_ncaa_box_stats import ncaa_mbb_box_scores
df = ncaa_mbb_box_scores(["6470186", "6479639"])
print(df.shape)

# Season aggregate for a scraped id list

agg = ncaa_mbb_box_scores(ids, multi_games=True)

# Offline with an injected fetcher

df = ncaa_mbb_box_scores("6470186", fetcher=my_fetcher)

ncaa_mbb_date_games​

ncaa_mbb_date_games(date: 'Optional[str]' = None, *, conference: 'str' = 'All', conference_id: 'Optional[int]' = None, fetcher: "'Optional[NcaaFetcher]'" = None, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Discover every NCAA MBB game played on a date (bigballR get_date_games).

Fetches stats.ncaa.org/season_divisions/{sid}/scoreboards for the date's season and returns one row per game with the /contests/{id} game id needed by the play-by-play / box-score scrapers. Port of bigballR get_date_games (all_functions.R:1119-1427).

Parameters

ParameterTypeDefaultDescription
dateOptional[str]None"MM/DD/YYYY". Defaults to yesterday (R default).
conferencestr'All'Conference name filter (e.g. "ACC", "Big Ten", "Metro" / "MAAC"); case/punctuation-insensitive, and both the current stats.ncaa.org label and bigballR's abbreviation work. Default "All" (every conference). Unknown names raise.
conference_idOptional[int]NoneExplicit stats.ncaa.org conference id; overrides conference when given (R's conference.ID).
fetcherOptional[NcaaFetcher]NoneInjectable ~sportsdataverse.mbb.mbb_ncaa_fetch .NcaaFetcher (tests pass an offline fake). None uses NcaaFetcher.with_browser() — the page is JS-rendered behind Akamai bm-verify, so the browser transport is the live default.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per game with columns date, start_time, home, away, box_id, game_id, home_score, away_score, attendance, neutral_site, home_wins, home_losses, away_wins, away_losses (SCOREBOARD_SCHEMA). Scores stay Utf8 — they hold "Canceled" / "Ppd" for unplayed games; game_id is null for games without a box score.

col_nametypedescription
datecharacterDate in YYYY-MM-DD format.
start_timecharacter
homecharacterHome.
awaycharacterAway record.
box_idcharacter
game_idcharacterUnique game identifier.
home_scorecharacterHome team score at the time of the play.
away_scorecharacterAway team score at the time of the play.
attendancecharacterReported attendance.
neutral_sitelogicalNeutral site.
home_winsintegerHome team's wins.
home_lossesintegerHome team's losses.
away_winsintegerAway team's wins.
away_lossesintegerAway team's losses.

Example

from sportsdataverse.mbb.mbb_ncaa_scoreboard import ncaa_mbb_date_games
games = ncaa_mbb_date_games("11/11/2025")
print(games.shape)

# Useful parameter combination

acc_pd = ncaa_mbb_date_games("02/01/2025", conference="ACC",
return_as_pandas=True)

# Pipeline next step (one line)

games.filter(pl.col("game_id").is_not_null())["game_id"].to_list()

ncaa_mbb_join_pbp_shots​

ncaa_mbb_join_pbp_shots(pbp: 'pl.DataFrame', shots: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, Any]'"

Attach shot-chart coordinates to play-by-play rows (bigballR

join_pbp_shots, get_shot_locations.R:93-135).

FG attempts (shot_value 2/3) are matched to chart shots on (game_id, game_seconds, event_result == shot_result, shot_no) where shot_no is the within-second same-result sequence number on BOTH sides — free throws and non-shot rows are deliberately excluded from matching (the chart plots FGs only) and pass through NA-filled. Row count and per-game row order are preserved; the explicit shot_dist carry-through is fork-skew fix #11 (wbigballR's unexported copy drops it). Works identically for the WBB extension — feed it quarter-model pbp + shots frames.

Parameters

ParameterTypeDefaultDescription
pbpDataFrameThe 35-column snake_case pbp contract frame (see ~sportsdataverse.mbb.mbb_ncaa_game_pbp.PBP_SCHEMA).
shotsDataFrameA SHOTS_SCHEMA frame (from parse_ncaa_bb_shots / ncaa_mbb_shot_locations) covering exactly the same game ids.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

The input pbp frame + team, player, x, y, shot_dist (null on non-FG rows and unmatched FG rows), sorted by (game_id, original per-game row order) exactly as R's arrange(row, .by_group = TRUE).

col_nametypedescription
game_idcharacterUnique game identifier.
game_datecharacterGame date (YYYY-MM-DD).
homecharacterHome.
awaycharacterAway record.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
clockcharacterGame clock value.
game_timecharacterGame start time.
game_secondsinteger
home_scoreintegerHome team score at the time of the play.
away_scoreintegerAway team score at the time of the play.
event_teamcharacter
event_descriptioncharacter
player_1character
player_2character
event_typecharacterEvent / play type code (V2 PBP).
event_resultcharacter
shot_valueintegerPoint value of the shot (2 or 3).
event_lengthinteger
poss_numinteger
poss_teamcharacter
poss_lengthinteger
is_transitionlogical
home_1character
home_2character
home_3character
home_4character
home_5character
away_1character
away_2character
away_3character
away_4character
away_5character
statuscharacterStatus label.
is_garbage_timelogical
sub_deviateinteger
teamcharacterTeam-side label or team identifier.
playercharacterPlayer name.
xdoubleX.
ydoubleY.
shot_distdouble

Example

from sportsdataverse.mbb.mbb_ncaa_game_pbp import ncaa_mbb_play_by_play
from sportsdataverse.mbb.mbb_ncaa_shots import (
ncaa_mbb_join_pbp_shots,
ncaa_mbb_shot_locations,
)
pbp = ncaa_mbb_play_by_play(["6470186"])
shots = ncaa_mbb_shot_locations(["6470186"])
joined = ncaa_mbb_join_pbp_shots(pbp, shots)

# Pipeline next step (one line)

joined.filter(pl.col("x").is_not_null()).head()

ncaa_mbb_player_stats​

ncaa_mbb_player_stats(pbp: 'pl.DataFrame', *, multi_games: 'bool' = False, simple: 'bool' = False, fix_tip_in: 'bool' = True, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Aggregate bigballR-contract play-by-play into per-player box stats.

Port of bigballR get_player_stats (all_functions.R:2810-3177) + its get_mins helper (:3240-3263). Counting stats are summarised per (game, team, player), assists counted from player_2, minutes and offensive possessions derived from the ten on-court columns, and rates (FG%, TS%, eFG%, rim/mid splits, ...) computed from the counters and rounded to 3 decimals with R's round semantics. With multi_games=True the per-game rows are summed per (player, team), every rate is recomputed from the summed counters (never averaged), and GP/GS are appended.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract. May span multiple games.
multi_gamesboolFalseWhen True, aggregate across games per (player, team) — the season-stat surface. When False (default, R parity), treat each game separately and keep the game id columns.
simpleboolFalseWhen True, return the reduced 33-column (multi) / 35-column (per-game) surface without the transition / assisted / putback / block-location splits.
fix_tip_inboolTrueWhen True (default), rim and putback stats count the scrape engine's real "Tip In" vocabulary. When False, reproduce R's literal "Tip-In" test (all_functions.R:2827) — tip-ins silently excluded — for oracle parity.
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

pl.DataFrame (or pd.DataFrame): one row per player+team (+game when multi_games=False). Columns follow PLAYER_STATS_COLUMNS / PLAYER_STATS_SIMPLE_COLUMNS / PLAYER_GAME_STATS_COLUMNS / PLAYER_GAME_STATS_SIMPLE_COLUMNS. Rows sorted by the group keys (byte order, matching dplyr's C-locale group order). Empty input yields an empty frame with the documented schema.

col_nametypedescription
game_idcharacterUnique game identifier.
game_datecharacterGame date (YYYY-MM-DD).
homecharacterHome.
awaycharacterAway record.
teamcharacterTeam-side label or team identifier.
playercharacterPlayer name.
minsdouble
o_possdouble
ptsdoublePoints scored.
orbdouble
drbdouble
astdoubleAssists.
stldoubleSteals.
blkdoubleBlocks.
tovdoubleTurnovers.
pfdoublePersonal fouls.
ts_pctdoubleTrue shooting percentage (0-1).
efg_pctdouble
fgmdoubleField goals made.
fgadoubleField goal attempts.
fg_pctdoubleField goal percentage (0-1).
tpmdouble
tpadouble
tp_pctdouble
ftmdoubleFree throws made.
ftadoubleFree throw attempts.
ft_pctdoubleFree throw percentage (0-1).
rimmdouble
rimadouble
rim_pctdouble
midmdouble
midadouble
mid_pctdouble
pbackmdouble
pbackadouble
pback_pctdouble
blk_rimdouble
blk_middouble
blk_threedouble
pct_fga_transdouble
pct_tpa_transdouble
pct_rima_transdouble
pct_fgm_transdouble
pct_tpm_transdouble
pct_rimm_transdouble
pct_fgm_astdouble
pct_tpm_astdouble
pct_rimm_astdouble
pts_transdouble
orb_transdouble
drb_transdouble
ast_transdouble
stl_transdouble
blk_transdouble
tov_transdouble
ts_pct_transdouble
efg_pct_transdouble
fgm_transdouble
fga_transdouble
fg_pct_transdouble
tpm_transdouble
tpa_transdouble
tp_pct_transdouble
ftm_transdouble
fta_transdouble
ft_pct_transdouble
rimm_transdouble
rima_transdouble
rim_pct_transdouble
midm_transdouble
mida_transdouble
mid_pct_transdouble
pts_halfdouble
orb_halfdouble
drb_halfdouble
ast_halfdouble
stl_halfdouble
blk_halfdouble
tov_halfdouble
ts_pct_halfdouble
efg_pct_halfdouble
fgm_halfdouble
fga_halfdouble
fg_pct_halfdouble
tpm_halfdouble
tpa_halfdouble
tp_pct_halfdouble
ftm_halfdouble
fta_halfdouble
ft_pct_halfdouble
rimm_halfdouble
rima_halfdouble
rim_pct_halfdouble
midm_halfdouble
mida_halfdouble
mid_pct_halfdouble
pts_astdouble
fgm_astdouble
tpm_astdouble
rimm_astdouble
midm_astdouble
pts_unastdouble
efg_pct_unastdouble
fgm_unastdouble
fga_unastdouble
fg_pct_unastdouble
tpm_unastdouble
tpa_unastdouble
tp_pct_unastdouble
rimm_unastdouble
rima_unastdouble
rim_pct_unastdouble
midm_unastdouble
mida_unastdouble
mid_pct_unastdouble

Example

from sportsdataverse.mbb.mbb_ncaa_stats_agg import ncaa_mbb_player_stats
season = ncaa_mbb_player_stats(pbp, multi_games=True)
print(season.shape)

# Reduced surface, pandas out

df_pd = ncaa_mbb_player_stats(pbp, multi_games=True, simple=True, return_as_pandas=True)

# Pipeline next step (one line)

season.filter(pl.col("mins") > 50).sort("pts", descending=True).head()

ncaa_mbb_possessions​

ncaa_mbb_possessions(pbp: 'pl.DataFrame', *, simple: 'bool' = False, fix_cross_game_leak: 'bool' = True, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Aggregate bigballR-contract play-by-play into one row per possession.

Port of bigballR get_possessions (all_functions.R:3686-3745). Groups by the possession keys stamped upstream by the scrape engine (poss_num, poss_team, the ten on-court lineup columns, plus game identity), drops possessions with any missing on-court player, and — in the full variant — sorts each row's home/away lineup alphabetically so a given lineup always occupies the same columns.

Parameters

ParameterTypeDefaultDescription
pbpDataFramePlay-by-play frame in the sdv-py 35-column snake_case bigballR contract (parse_ncaa_bb_game_pbp output). May span multiple games; rows must be in scrape order.
simpleboolFalseWhen True, return only the 17-column possession/points frame (all_functions.R:3687-3694) with lineups in on-court order. When False (default), return the full 28-column frame with per-possession context columns and alpha-sorted lineups.
fix_cross_game_leakboolTrueWhen True (default, and the CORRECT behavior), window the start_event_type lag with .over("game_id") so a game's first possession has a null start event instead of inheriting the PREVIOUS game's last event. When False, reproduce R's ungrouped dplyr::lag (all_functions.R:3698) and its cross-game leak. Parity tests pass False. Ignored when simple=True (that variant emits no start_event_type).
return_as_pandasboolFalseReturn a pandas.DataFrame instead of polars.

Returns

pl.DataFrame (or pd.DataFrame) with one row per possession — 28 columns per POSSESSION_SEG_SCHEMA (full) or 17 per POSSESSIONS_SIMPLE_SCHEMA (simple). Empty input yields an empty frame carrying the documented schema.

col_nametypedescription
game_idcharacterUnique game identifier.
game_datecharacterGame date (YYYY-MM-DD).
homecharacterHome.
awaycharacterAway record.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
poss_numinteger
poss_teamcharacter
home_1character
home_2character
home_3character
home_4character
home_5character
away_1character
away_2character
away_3character
away_4character
away_5character
home_scoreintegerHome team score at the time of the play.
away_scoreintegerAway team score at the time of the play.
ptsintegerPoints scored.
is_assistedinteger
is_transitioninteger
is_garbage_timeinteger
start_event_typecharacter
first_shot_timeinteger
first_shot_typecharacter
last_event_timeinteger
last_event_typecharacter

Example

from sportsdataverse.mbb.mbb_ncaa_possession_seg import ncaa_mbb_possessions
poss = ncaa_mbb_possessions(pbp)
print(poss.shape)

# Simple points-per-possession variant

poss_pd = ncaa_mbb_possessions(pbp, simple=True, return_as_pandas=True)

# Pipeline next step (one line)

poss.group_by("poss_team").agg(pl.col("pts").mean())