Skip to main content

WBB — additional Python functions — stats.ncaa.org: playwright_transport–validate_lineup

playwright_transport​

playwright_transport(*, headless_new: 'bool' = True, challenge_wait_ms: 'int' = 8000, nav_timeout_ms: 'int' = 45000, user_agent: 'Optional[str]' = None, solve_attempts: 'int' = 3, relaunch_backoff: 'float' = 2.0) -> "'_PlaywrightTransport'"

Build the suggested stats.ncaa.org game-detail scraping transport.

Drives a real Chromium via Playwright in Chrome's new-headless mode (--headless=new) to clear the Akamai bm-verify challenge that curl_cffi cannot, then serves raw server HTML for the 5a-5e parsers. Playwright is a lazy optional import (not a hard dependency); a clear ImportError fires on first use if it is missing.

Parameters

ParameterTypeDefaultDescription
headless_newboolTrueUse --headless=new (real-GPU render, no window) -- the default and the proven-working mode. False runs old headless (headless_shell), which Akamai flags -- avoid.
challenge_wait_msint8000Milliseconds to let the bm-verify sensor run after the first navigation.
nav_timeout_msint45000Per-navigation timeout.
user_agentOptional[str]NoneOverride the Chrome UA string.
solve_attemptsint3
relaunch_backofffloat2.0

Returns

A stateful, callable FetchTransport reusing one browser for the session. Close it when done (it is a context manager, has close(), and registers an atexit safety net).

Example

from sportsdataverse.mbb.mbb_ncaa_fetch import NcaaFetcher
with NcaaFetcher.with_browser() as fetcher:
pbp = fetcher.fetch_game_pbp("1613299") # raw PBP HTML
box = fetcher.fetch_game_individual_stats("1613299") # raw box HTML
# -> feed to get_box_lineup / create_lineup_data (mbb_ncaa_*_parser)

remove_diacritics​

remove_diacritics(fragment: 'str') -> 'str'

Strip diacritical marks, e.g. "Juhász" -> "Juhasz"

(ExtractorUtils.scala:38-43: NFD normalization then removal of the combining-diacritical-marks block).

Parameters

ParameterTypeDefaultDescription
fragmentstrAny string (a full player name or a name fragment).

Returns

The string with combining marks removed.

Example

from sportsdataverse.mbb.mbb_ncaa_stints import remove_diacritics
print(remove_diacritics("Dorka Juhász")) # "Dorka Juhasz"

reorder_and_reverse​

reorder_and_reverse(reversed_partial_events: 'Iterable[PlayByPlayEvent]') -> 'list[PlayByPlayEvent]'

Orders same-minute play-by-play events so subs never enclose the plays

they logically precede/follow (ExtractorUtils.scala:435-599).

Groups consecutive events sharing the same min into a block (the input arrives in descending/reverse-chronological order, so blocks are discovered and internally accumulated in reverse too), then -- for any block containing a sub -- reorders it via inner_sort: events referencing a subbed-OUT player (or scoring no higher than the sub) land in a pre-sub group, the subs themselves come next (in ascending-score order), and events referencing a subbed-IN player (or scoring higher than the sub) land in a trailing post-sub group. Free-throw attempts sharing the sub's inferred "direction" (team vs. opponent, inferred from the nearest preceding shot/FT/foul) are pulled into the pre-sub group unless the shooter is one of the players being subbed in. Blocks with no sub are returned unchanged apart from the initial score-based sort.

Parameters

ParameterTypeDefaultDescription
reversed_partial_eventsIterable[PlayByPlayEvent]Events for one lineup event, in reverse-chronological (descending-time) order -- the natural order encountered walking play-by-play text bottom-up.

Returns

The same events, forward-chronological (ascending time), with each same-minute block internally reordered so no sub encloses a play it logically shouldn't.

Example

from sportsdataverse.mbb.mbb_ncaa_models import Score
from sportsdataverse.mbb.mbb_ncaa_stints import (
OtherTeamEvent,
SubInEvent,
reorder_and_reverse,
)
events = [
SubInEvent(0.4, Score(0, 0), "player1"),
OtherTeamEvent(0.4, Score(0, 0), "rebound"),
]
reorder_and_reverse(events)
# [OtherTeamEvent(...), SubInEvent(...)]

reset_config​

reset_config() -> 'NcaaFetchConfig'

Reset the active config to its env-var-derived defaults.

Returns

The live singleton, now holding the env-var-derived defaults again.

Example

from sportsdataverse.mbb.mbb_ncaa_fetch import update_config, reset_config
update_config(timeout=5)
reset_config()

right_kind_of_shot​

right_kind_of_shot(shot: 'ShotEvent', pbp_event: 'MiscGameEvent', strict: 'bool') -> 'bool'

Whether pbp_event's shot type is compatible with shot's

distance and make/miss (ShotEnrichmentUtils.right_kind_of_shot, PlayByPlayUtils.scala:659-679).

The distance-in-the-data is approximate, so exact 2-vs-3 discrimination is impossible; this only rules out the obvious mismatches (a clearly-short shot matched to a 3, or vice versa) and always requires make/miss agreement.

Parameters

ParameterTypeDefaultDescription
shotShotEventThe shot being enriched (pts/dist read).
pbp_eventMiscGameEventThe candidate play-by-play event.
strictboolIf True, also apply the distance gate; if False, only the make/miss agreement is required.

Returns

True if the event could plausibly be this shot.

Example

from sportsdataverse.mbb.mbb_ncaa_pbp_glue import right_kind_of_shot
right_kind_of_shot(shot, pbp_event, strict=True)

run_iterative_adjustment_with_hca​

run_iterative_adjustment_with_hca(teams: 'Sequence[TeamDetail]', team_by_name: 'dict[str, TeamDetail]', fields: 'Sequence[str]', league_averages: 'LeagueAverages', poss_splits: 'dict[str, PossessionSplits]', *, max_iterations: 'int' = 100, tolerance: 'float' = 1e-06) -> 'IterationResult'

KenPom-style SoS + HCA fixed-point solver (runIterativeAdjustmentWithHCA, ts:306-527).

Each iteration (Jacobi -- all teams read the previous iteration's adjustments, then commit together):

  1. Per team/field, adjust every game adj_game = raw_game * (league / (opp_adj +/- hca)) and take the weighted mean; a field with no valid games keeps its current value.
  2. Re-estimate per-field HCA from home/away possession-imbalance residuals hca = sum((raw - pred) * |imbalance|) / sum(|imbalance|) over teams with |imbalance| >= IMBALANCE_MIN.

Stops when the max per-team/field change drops below tolerance or after max_iterations sweeps (the HCA re-estimate still runs on the final sweep). The cross-guard on the per-game branch, the asymmetric residual prediction, and the cross-named opponent strengths are all preserved -- see the module docstring's landmine list.

Parameters

ParameterTypeDefaultDescription
teamsSequence[TeamDetail]The teams to solve over.
team_by_namedict[str, TeamDetail]{team_name: team_detail} for opponent lookup.
fieldsSequence[str]The stat fields to solve.
league_averagesLeagueAveragesOutput of compute_league_averages_from_per_game.
poss_splitsdict[str, PossessionSplits]{team_name: PossessionSplits }.
max_iterationsint100Iteration cap (default MAX_ITERATIONS; pin to 1 to inspect a single sweep).
tolerancefloat1e-06Convergence tolerance (default TOLERANCE).

Returns

An IterationResult (adj_values, hca_per_field).

Example

from sportsdataverse.mbb.mbb_ncaa_strength import (
STRENGTH_ADJUSTED_FIELDS,
compute_league_averages_from_per_game,
compute_possession_splits,
run_iterative_adjustment_with_hca,
)

by_name = {t["team_name"]: t for t in teams}
league = compute_league_averages_from_per_game(teams)
splits = {t["team_name"]: compute_possession_splits(t) for t in teams}
result = run_iterative_adjustment_with_hca(
teams, by_name, STRENGTH_ADJUSTED_FIELDS, league, splits,
)
print(result.hca_per_field["3p"]["hca_off"])

select_contains​

select_contains(root: 'Tag', selector: 'str', text: 'str') -> 'list[Tag]'

JSoup root.select(sel + ":contains(text)"): candidates whose full

text (own + every descendant's) case-insensitively CONTAINS text as a plain substring -- not a regex (Task 5e.2 addition; see the module docstring's "Critical divergence" note).

JSoup's :contains() is documented case-insensitive substring containment; soupsieve's :-soup-contains() (the non-deprecated spelling of its :contains()) is case-SENSITIVE, with no case-insensitive variant of its own. Reproducing JSoup's actual semantics therefore needs this helper rather than :-soup-contains().

Parameters

ParameterTypeDefaultDescription
rootTagThe element to search within.
selectorstrA plain (soupsieve-legal) CSS selector for the structural part of the match (everything before :contains).
textstrThe plain substring each candidate's collapsed text must case-insensitively contain.

Returns

Every selector match whose jsoup_text case-insensitively contains text, in document order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_contains
soup = parse_html("<td>game date:</td><td>Location:</td>")
select_contains(soup, "td", "Game Date:") # [<td>game date:</td>]

select_matching​

select_matching(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'

JSoup root.select(sel + ":matches(regex)"): candidates whose full

text (own + every descendant's) matches regex.

Soupsieve has no :matches() pseudo-class equivalent, so this runs the plain structural selector first, then filters by re.search over each candidate's jsoup_text (own text plus descendants', matching JSoup's :matches() semantics -- as opposed to select_matching_own's own-text-only :matchesOwn()).

Parameters

ParameterTypeDefaultDescription
rootTagThe element to search within.
selectorstrA plain (soupsieve-legal) CSS selector.
regexstrThe pattern each candidate's collapsed text must re.search-match.

Returns

Every selector match whose jsoup_text contains a regex match, in document order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching
soup = parse_html("<div><p>Home Team</p><p>Away Team</p></div>")
select_matching(soup, "p", r"^Home") # [<p>Home Team</p>]

select_matching_own​

select_matching_own(root: 'Tag', selector: 'str', regex: 'str') -> 'list[Tag]'

JSoup root.select(sel + ":matchesOwn(regex)"): candidates whose

OWN text only (excluding descendant elements' text) matches regex.

JSoup's Element.ownText() walks only the element's direct TextNode children, not text nested inside child elements -- the same distinction bs4 draws between a tag's direct bs4.NavigableString children and its full .get_text().

Parameters

ParameterTypeDefaultDescription
rootTagThe element to search within.
selectorstrA plain (soupsieve-legal) CSS selector.
regexstrThe pattern each candidate's own (whitespace-collapsed) text must re.search-match.

Returns

Every selector match whose own text contains a regex match, in document order.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, select_matching_own
soup = parse_html('<div class="card-header">Coach <b>Info</b></div>')
select_matching_own(soup, "div.card-header", r"^Coach")
# [<div class="card-header">Coach <b>Info</b></div>]

shot_js_to_html​

shot_js_to_html(js: 'str') -> 'list[Tag]'

Converts client-side addShot(...) JS calls into parseable

circle.shot HTML, for pages where the shot map is built on the fly rather than baked into the initial HTML (ShotEventParser .shot_js_to_html, :266-283). See the module docstring's "Scala idiom decision" note -- the Scala's builders/browser parameters are dropped here since the Scala body never actually uses them.

Parameters

ParameterTypeDefaultDescription
jsstrThe concatenated <script> text containing one or more addShot(x, y, ..., 'title', ...) calls, one per line.

Returns

The circle.shot elements reconstructed from every matching line (non-matching lines, e.g. the addShot function definition line itself, are silently skipped).

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import shot_js_to_html
js = "addShot(27.0, 77.0, 392, false, 1, 'title text', 'class', false);"
circles = shot_js_to_html(js)

start_time_from_period​

start_time_from_period(period: 'int', is_women_game: 'bool') -> 'float'

The game-clock time (minutes elapsed) a period starts at

(ExtractorUtils.scala:272-281).

Women's games play four 10-minute quarters then 5-minute overtimes; men's games play two 20-minute halves then 5-minute overtimes.

Parameters

ParameterTypeDefaultDescription
periodintThe 1-indexed period number (1/2 = halves for men, 1-4 = quarters for women, 5+ = overtimes for both).
is_women_gameboolWhether to use the women's (quarters) or men's (halves) period schedule.

Returns

The game-clock minute the period begins at.

Example

from sportsdataverse.mbb.mbb_ncaa_stints import start_time_from_period
start_time_from_period(2, is_women_game=False) # 20.0 (men's 2nd half)
start_time_from_period(1, is_women_game=True) # 0.0 (women's 1st quarter)
start_time_from_period(6, is_women_game=False) # 45.0 (men's 2nd OT)

sum_event_stats​

sum_event_stats(lhs: 'LineupEventStats', rhs: 'LineupEventStats') -> 'LineupEventStats'

Field-wise add two :class:`~sportsdataverse.mbb.mbb_ncaa_models

.LineupEventStats (protected def sum_event_stats, LineupUtils.scala :1534-1622, debug-only -- the Scala's own docstring says "just used for debug"). The Scala builds this via shapeless.Generic` field-zipping; this port is an explicit field-by-field call since Python has no equivalent generic-programming machinery.

Parameters

ParameterTypeDefaultDescription
lhsLineupEventStatsThe left-hand stat tree.
rhsLineupEventStatsThe right-hand stat tree.

Returns

A new ~sportsdataverse.mbb.mbb_ncaa_models.LineupEventStats with every field summed (see the module's private sum_*helpers for theOptional`/nested-field summing rules).

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_event_stats
from sportsdataverse.mbb.mbb_ncaa_models import LineupEventStats

sum_event_stats(LineupEventStats.empty(), LineupEventStats.empty()).num_events

sum_shot_infos​

sum_shot_infos(shot_infos: 'list[PlayerShotInfo]') -> 'Optional[PlayerShotInfo]'

Field-wise sum a list of :class:`~sportsdataverse.mbb.mbb_ncaa_models

.PlayerShotInfo\ s (sum_shot_infos, LineupUtils.scala:1625-1655`, debug-only).

Parameters

ParameterTypeDefaultDescription
shot_infoslist[PlayerShotInfo]The list to combine, in order.

Returns

None if shot_infos is empty; the single element if there's exactly one; otherwise a left-fold of pairwise field-wise sums (reduceOption).

Example

from sportsdataverse.mbb.mbb_ncaa_lineup_enrich import sum_shot_infos
from sportsdataverse.mbb.mbb_ncaa_models import PlayerShotInfo

sum_shot_infos([PlayerShotInfo(ast_3pm=(1, 0, 0, 0, 0)), PlayerShotInfo(ast_3pm=(0, 1, 0, 0, 0))])

td_at​

td_at(row: 'Tag', n: 'int') -> 'Optional[Tag]'

JSoup row >?> element("td:eq(n)"): the n-th <td> child.

Soupsieve has no :eq() positional pseudo-class (unlike JSoup), so this is a plain 0-indexed lookup into row.find_all("td"), guarded against an out-of-range index (JSoup's >?> returns None rather than raising when the selector matches nothing).

Parameters

ParameterTypeDefaultDescription
rowTagThe row (or other container) element to search.
nintThe 0-indexed <td> position.

Returns

The n-th <td> descendant, or None if row has fewer than n + 1 of them.

Example

from sportsdataverse.mbb.mbb_ncaa_html import parse_html, td_at
soup = parse_html("<tr><td>A</td><td>B</td></tr>")
row = soup.find("tr")
td_at(row, 1).get_text() # "B"
td_at(row, 5) # None

transform_shot_location​

transform_shot_location(x: 'float', y: 'float', second_half_switch: 'bool', team_shooting_left_in_first_period: 'bool', is_offensive: 'bool') -> 'tuple[float, float, float, float]'

Transforms a raw SVG pixel location into feet from the basket, always

oriented as if shooting towards the left goal (ShotEventParser .transform_shot_location, :588-620).

Parameters

ParameterTypeDefaultDescription
xfloatRaw SVG cx pixel coordinate.
yfloatRaw SVG cy pixel coordinate.
second_half_switchboolWhether this shot is in the "other" half of the game from team_shooting_left_in_first_period (each False factor below flips which side is treated as "left").
team_shooting_left_in_first_periodboolWhether the team under analysis shot towards the left goal in the first period (see is_team_shooting_left_to_start).
is_offensiveboolWhether the team under analysis is shooting (an opponent shot flips the expected side again).

Returns

(x, y, alt_x, alt_y) in feet -- the believed-correct location, then the alternative (mirror-image) location, both relative to the goal the shot is (believed to be) attacking.

Example

from sportsdataverse.mbb.mbb_ncaa_shot_parser import transform_shot_location
transform_shot_location(310.2, 235, False, False, True)

update_config​

update_config(**kwargs: 'object') -> 'NcaaFetchConfig'

Update the active config in place.

Returns

The (mutated) global config object.

Example

from sportsdataverse.mbb.mbb_ncaa_fetch import update_config
update_config(proxy_url="http://user:pass@1.2.3.4:8080")

validate_box_score​

validate_box_score(team: 'TeamId', lineup: 'list[str]') -> 'Union[list[PlayerCodeId], ParseError]'

Checks there are no duplicates in the lineup (``BoxscoreParser

.validate_box_score, :388-404``).

Parameters

ParameterTypeDefaultDescription
teamTeamIdThe team the lineup belongs to (feeds ~sportsdataverse.mbb.mbb_ncaa_stints.build_player_code's team-scoped misspelling corrections).
lineuplist[str]The raw player-name strings, in whatever order they were assembled by inject_validated_players.

Returns

lineup, mapped to ~sportsdataverse.mbb.mbb_ncaa_models.PlayerCodeId (same order, no sort -- see the module docstring's "not sorted" note). When two teammates collide on the {first-two-letters}{Surname} scheme -- siblings, in practice -- only the colliding players are re-coded to {First}{Last} by disambiguate_sibling_codes; every other player keeps the Scala-faithful code. This is a DELIBERATE divergence from ExtractorUtils.scala, which rejects the game: since a team's roster is the same all season, one sibling pair cost the team its ENTIRE season of lineups. A ~sportsdataverse.mbb.mbb_ncaa_data_quality.ParseErroris returned only when widening cannot separate them, i.e. two players with the SAME full name -- genuinely ambiguous, so still an error. Callers must not re-derive a code from a name after this point:build_player_codewould undo the widening and silently drop one twin. Use~sportsdataverse.mbb.mbb_ncaa_names.code_from_box`, which resolves against this roster.

Example

from sportsdataverse.mbb.mbb_ncaa_boxscore_parser import validate_box_score
from sportsdataverse.mbb.mbb_ncaa_models import TeamId
validate_box_score(TeamId("Team"), ["Player One", "Player Two"])

validate_lineup​

validate_lineup(lineup_event: 'LineupEvent', box_lineup: 'LineupEvent', valid_player_codes: 'set[str]') -> 'list[ValidationError]'

Flags a lineup stint as internally inconsistent, via 3 independent

checks (LineupErrorAnalysisUtils.validate_lineup, :181-218).

Parameters

ParameterTypeDefaultDescription
lineup_eventLineupEventThe lineup stint to validate.
box_lineupLineupEventThe team's box-score lineup event (players is the full roster) -- used both to build the name-resolution context (see ~sportsdataverse.mbb.mbb_ncaa_names.build_tidy_player_context) and, indirectly, as the source of players_out for jersey-number resolution inside ~sportsdataverse.mbb .mbb_ncaa_names.tidy_player.
valid_player_codesset[str]Every player code that's actually on the box score / roster for this team-season.

Returns

The failing ValidationError\ s, in declaration order (see the module docstring's "Return shape" note) -- empty if lineup_event is clean. * ValidationError.WRONG_NUMBER_OF_PLAYERS -- lineup_event doesn't have exactly 5 players on the floor. * ValidationError.UNKNOWN_PLAYERS -- some player on the floor isn't in valid_player_codes. * ValidationError.INACTIVE_PLAYERS -- some player mentioned in lineup_event's own (team-side) raw game events resolves to a code not in valid_player_codes (i.e. isn't on the floor, per the lineup being validated).

Example

from sportsdataverse.mbb.mbb_ncaa_stint_validation import validate_lineup
errors = validate_lineup(lineup_event, box_lineup, {"MiMitchell", "BbBob"})
assert not errors # a clean lineup returns []