Skip to main content
Version: 0.1.5

NBA — additional Python functions — Analytics: add_ctg–nba_referee

add_ctg_shot_zones​

add_ctg_shot_zones(enhanced_pbp: 'pl.DataFrame') -> 'pl.DataFrame'

Append CTG's shot-location zone (ctg_shot_zone) to an enhanced PBP frame.

CTG's zones differ from the official NBA zones emitted by ~sportsdataverse.nba.nba_shot_zones.add_shot_zones: CTG splits the midrange at the free-throw-line distance rather than at the paint boundary.

  • at_rim — shot distance < 4 ft ("Shots within 4 feet of the basket").
  • short_mid — 4 ft <= distance < 14 ft ("outside of 4 feet, but inside of ~14 feet (the free throw line distance)").
  • long_mid — >= 14 ft, inside the arc.
  • corner_3 — a three "below the break" (|x_legacy| >= 220 and y_legacy <= 87.5).
  • arc_3 — any other three (CTG's "non-corner three").

Parameters

ParameterTypeDefaultDescription
enhanced_pbpDataFrameFrame from ~sportsdataverse.nba.nba_enhanced_pbp.enhanced_pbp_from_payload.

Returns

The input frame with a ctg_shot_zone Utf8 column appended (null on non-field-goal rows). Empty input returns a zero-row frame carrying the column — never raises.

col_nametypedescription
order_indexinteger
action_numberintegerSequential action number within a game (V3 PBP).
clockcharacterGame clock value.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
team_idintegerUnique team identifier.
team_tricodecharacterThree-letter team code (e.g. 'LAS' / 'NYL').
person_idintegerUnique player identifier (V3 endpoints).
player_namecharacterPlayer name.
player_name_icharacterPlayer name i.
x_legacyintegerV2-format X coordinate (preserved for V3-to-V2 compatibility).
y_legacyintegerV2-format Y coordinate (preserved for V3-to-V2 compatibility).
shot_distanceintegerShot distance from the basket, in feet.
shot_resultcharacterShot result ('Made' / 'Missed').
is_field_goalinteger1 if the action was a field goal; 0 otherwise.
score_homecharacterScore home.
score_awaycharacterScore away.
points_totalintegerRunning total of points scored.
locationcharacterLocation.
descriptioncharacterLong-form description text.
action_typecharacterAction type label (e.g. 'Made Shot', 'Substitution').
sub_typecharacterAction sub-type label.
video_availableintegerVideo available.
shot_valueintegerPoint value of the shot (2 or 3).
action_idintegerUnique action identifier within a game (V3 PBP).
game_idcharacterUnique game identifier.
seconds_remainingdoubleSeconds remaining in the period.
event_typecharacterEvent / play type code (V2 PBP).
is_made_shotlogical
is_missed_shotlogical
is_free_throwlogical
is_reboundlogical
is_turnoverlogical
is_foullogical
is_substitutionlogical
is_jump_balllogical
is_timeoutlogical
is_periodlogical
ctg_shot_zonecharacter

Example

from sportsdataverse.nba.nba_enhanced_pbp import enhanced_pbp_from_payload
from sportsdataverse.nba.nba_play_context import add_ctg_shot_zones
pbp = add_ctg_shot_zones(enhanced_pbp_from_payload(payload))
print(pbp.filter(pl.col("ctg_shot_zone").is_not_null())["ctg_shot_zone"].value_counts())

add_play_context​

add_play_context(enhanced_pbp: 'pl.DataFrame', *, transition_seconds: 'float' = 6.0, transition_variant: 'str' = 'hoop_math', starters_on_court: 'Optional[dict[int, int]]' = None) -> 'pl.DataFrame'

Build possessions and enrich them with the full CTG play-context surface.

One call: ~sportsdataverse.nba.nba_possessions.build_possessions -> add_start_type_detail -> add_transition -> flag_heave_possessions -> flag_garbage_time.

The CTG filter columns are flags, not filters — nothing is dropped. Apply

Parameters

ParameterTypeDefaultDescription
enhanced_pbpDataFrameFrame from ~sportsdataverse.nba.nba_enhanced_pbp.enhanced_pbp_from_payload.
transition_secondsfloat6.0Transition initial-play cutoff (default 10.0).
transition_variantstr'hoop_math'See add_transition.
starters_on_courtOptional[dict[int, int]]NoneOptional starters-on-floor counts; see flag_garbage_time.

Returns

The possession frame (POSSESSIONS_SCHEMA) plus every column in PLAY_CONTEXT_POSSESSIONS_SCHEMA.

col_nametypedescription
game_idcharacterUnique game identifier.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
possession_numberintegerPossession number.
offense_team_idintegerUnique identifier for offense team.
defense_team_idinteger
start_order_indexinteger
end_order_indexinteger
start_seconds_remainingdouble
end_seconds_remainingdouble
pointsintegerPoints scored.
is_second_chancelogical
number_in_periodinteger
possession_start_typecharacter
count_as_possessionlogical
fg2ainteger
fg2minteger
fg3aintegerThree-point field goal attempts.
fg3mintegerThree-point field goals made.
ftaintegerFree throw attempts.
ftmintegerFree throws made.
orebintegerOffensive rebounds.
drebintegerDefensive rebounds.
tovintegerTurnovers.
possession_start_type_detailcharacter
possession_start_type_ctgcharacter
seconds_to_first_playdouble
is_transitionlogical
transition_sourcecharacter
possession_contextcharacter
is_heave_possessionlogical
is_garbage_timelogical
garbage_time_basischaracter

Example

from sportsdataverse.nba.nba_enhanced_pbp import enhanced_pbp_from_payload
from sportsdataverse.nba.nba_play_context import add_play_context
poss = add_play_context(enhanced_pbp_from_payload(payload))
print(poss["possession_start_type_ctg"].value_counts())

add_start_type_detail​

add_start_type_detail(possessions: 'pl.DataFrame', enhanced_pbp: 'pl.DataFrame') -> 'pl.DataFrame'

Append the full pbpstats start-type taxonomy to a possession frame.

Upgrades the engine's coarse 5-value possession_start_type into:

  • possession_start_type_detail — the zone-split pbpstats vocabulary: Off{AtRim|ShortMidRange|LongMidRange|Corner3|Arc3}{Make|Miss|Block}, OffFTMake / OffFTMiss, OffLiveBallTurnover, OffTimeout, OffDeadball.
  • possession_start_type_ctg — the coarse bucket CTG reports on: off_made / off_live_rebound / off_steal / off_deadball / off_timeout (see ~nba_play_context_constants.CTG_START_BUCKETS).

Precedence (pbpstats): period start > timeout > previous boundary event. A team rebound (person_id == 0) is a dead-ball start even though a rebound row exists; a timeout beats a made basket (an after-timeout possession is OffTimeout, not OffMadeShot).

Parameters

ParameterTypeDefaultDescription
possessionsDataFrameFrame from ~sportsdataverse.nba.nba_possessions.build_possessions.
enhanced_pbpDataFrameThe enhanced PBP frame those possessions were built from.

Returns

possessions with the two columns appended. Empty input returns a zero-row frame carrying them — never raises.

col_nametypedescription
game_idcharacterUnique game identifier.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
possession_numberintegerPossession number.
offense_team_idintegerUnique identifier for offense team.
defense_team_idinteger
start_order_indexinteger
end_order_indexinteger
start_seconds_remainingdouble
end_seconds_remainingdouble
pointsintegerPoints scored.
is_second_chancelogical
number_in_periodinteger
possession_start_typecharacter
count_as_possessionlogical
fg2ainteger
fg2minteger
fg3aintegerThree-point field goal attempts.
fg3mintegerThree-point field goals made.
ftaintegerFree throw attempts.
ftmintegerFree throws made.
orebintegerOffensive rebounds.
drebintegerDefensive rebounds.
tovintegerTurnovers.
possession_start_type_detailcharacter
possession_start_type_ctgcharacter

Example

from sportsdataverse.nba.nba_possessions import build_possessions
from sportsdataverse.nba.nba_play_context import add_start_type_detail
poss = add_start_type_detail(build_possessions(pbp), pbp)
print(poss["possession_start_type_ctg"].value_counts())

add_transition​

add_transition(possessions: 'pl.DataFrame', enhanced_pbp: 'pl.DataFrame', *, transition_seconds: 'float' = 6.0, variant: 'str' = 'hoop_math') -> 'pl.DataFrame'

Flag possessions that started in transition, and time their initial play.

CTG defines transition as beginning at the possession start and ending "once the defense is set", without publishing a seconds threshold. We therefore time the possession's initial play — its first shot attempt, trip to the line, or turnover (CTG's own definition of a "play") — and call the possession transition when that play lands within transition_seconds.

Variants (~nba_play_context_constants.TRANSITION_VARIANTS):

  • hoop_math (default) — any non-timeout start type qualifies.
  • haslametrics — steal starts only (conservative).
  • bigballr — the previous possession must have ended live (off_made / off_live_rebound / off_steal); a dead-ball start can never be transition.

The first possession of a period is never transition. After a timeout the defense is set by construction, so off_timeout never qualifies under any variant.

Parameters

ParameterTypeDefaultDescription
possessionsDataFrameFrame carrying possession_start_type_ctg (i.e. the output of add_start_type_detail).
enhanced_pbpDataFrameThe enhanced PBP frame the possessions were built from.
transition_secondsfloat6.0Initial-play cutoff. Default 10.0 (hoop-math). Calibrate against Synergy transition frequency (league mean ~15-16%).
variantstr'hoop_math'One of ~nba_play_context_constants.TRANSITION_VARIANTS.

Returns

possessions with seconds_to_first_play (Float64, null when the possession had no play), is_transition (Boolean), transition_source (Utf8: steal / live_rebound / made / deadball; null when not transition) and possession_context (Utf8: transition / halfcourt / misc) appended.

col_nametypedescription
game_idcharacterUnique game identifier.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
possession_numberintegerPossession number.
offense_team_idintegerUnique identifier for offense team.
defense_team_idinteger
start_order_indexinteger
end_order_indexinteger
start_seconds_remainingdouble
end_seconds_remainingdouble
pointsintegerPoints scored.
is_second_chancelogical
number_in_periodinteger
possession_start_typecharacter
count_as_possessionlogical
fg2ainteger
fg2minteger
fg3aintegerThree-point field goal attempts.
fg3mintegerThree-point field goals made.
ftaintegerFree throw attempts.
ftmintegerFree throws made.
orebintegerOffensive rebounds.
drebintegerDefensive rebounds.
tovintegerTurnovers.
possession_start_type_detailcharacter
possession_start_type_ctgcharacter
seconds_to_first_playdouble
is_transitionlogical
transition_sourcecharacter
possession_contextcharacter

Example

poss = add_transition(add_start_type_detail(poss, pbp), pbp)
print(poss["is_transition"].mean()) # transition frequency

# Tune the knob against the Synergy oracle

poss8 = add_transition(poss, pbp, transition_seconds=8.0)

box_features​

box_features(player_logs: 'pl.DataFrame', team_logs: 'pl.DataFrame', *, game_ids: 'Optional[List[str]]' = None) -> 'pl.DataFrame'

Aggregate per-player per-100-possession box features over a set of games.

Restricting game_ids to a fold's games is the harness leakage guard.

Per-100 possessions are computed per game (so mid-window trades use each game's own team pace), then summed — the result is fully deterministic.

Parameters

ParameterTypeDefaultDescription
player_logsDataFramePer-player-per-game box lines (game_id, team_id, player_id, min, and the counting stats in STATS`).
team_logsDataFramePer-team-per-game lines (game_id, team_id, min, fga, oreb, tov, fta) for the possession estimate.
game_idsOptional[List[str]]NoneOptional subset of game_id to include (default: all).

Returns

One row per player: player_id, the STATSper-100 rates,min(total),gp` (games). Empty frame with that schema on empty input.

col_nametypedescription
player_idintegerUnique player identifier.
ptsdoublePoints scored.
fg3mdoubleThree-point field goals made.
fgadoubleField goal attempts.
ftadoubleFree throw attempts.
astdoubleAssists.
orebdoubleOffensive rebounds.
drebdoubleDefensive rebounds.
stldoubleSteals.
blkdoubleBlocks.
tovdoubleTurnovers.
pfdoublePersonal fouls.
mindoubleMinutes played.
gpintegerGames played.

build_play_context_shots​

build_play_context_shots(possessions: 'pl.DataFrame', enhanced_pbp: 'pl.DataFrame', *, putback_seconds: 'float' = 2.0) -> 'pl.DataFrame'

Build the per-shot frame carrying CTG's play context.

CTG assigns context per play, not per possession: one possession can contain a transition miss, a halfcourt reset and a putback. This frame is the play-level view — one row per field-goal attempt.

  • is_putback — pbpstats field_goal.py:112-144: an unassisted 2-point attempt whose preceding event is a real offensive rebound by the same player, within putback_seconds. A three is never a putback.
  • is_second_chance_shot — the shot follows an offensive rebound earlier in the same possession.
  • shot_context — transition / putback / halfcourt. Transition wins over putback, reproducing CTG exactly: "if a team comes down in transition and misses a shot but gets a putback, that putback is classified as part of the overall transition event."

Parameters

ParameterTypeDefaultDescription
possessionsDataFrameFrame from add_transition (needs is_transition).
enhanced_pbpDataFrameThe enhanced PBP frame the possessions were built from.
putback_secondsfloat2.0Rebound-to-shot window. Default 2.0 (pbpstats).

Returns

Polars DataFrame with schema PLAY_CONTEXT_SHOTS_SCHEMA — one row per field-goal attempt. Empty input returns the zero-row schema.

col_nametypedescription
game_idcharacterUnique game identifier.
possession_numberintegerPossession number.
order_indexinteger
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
team_idintegerUnique team identifier.
person_idintegerUnique player identifier (V3 endpoints).
shot_valueintegerPoint value of the shot (2 or 3).
shot_madelogical
ctg_shot_zonecharacter
is_assistedlogical
is_putbacklogical
is_second_chance_shotlogical
shot_contextcharacter

Example

shots = build_play_context_shots(poss, pbp)
print(shots.group_by("shot_context").len())
print(shots.filter(pl.col("is_putback") == True).height)

build_possession_shooting​

build_possession_shooting(enhanced_pbp: 'pl.DataFrame') -> 'pl.DataFrame'

Build the per-shooter companion frame from an enhanced play-by-play DataFrame.

Companion to build_possessions: instead of one team-level row per possession, emits one row per distinct shooter (player_id) per possession, with their own fg2a/fg2m/fg3a/fg3m/fta/ftm counts. Shares the same possession-group traversal as build_possessions via assemble` — the two frames are always built from a single consistent pass over the play-by-play. Consumed by WP2's luck-adjusted shooting response.

Parameters

ParameterTypeDefaultDescription
enhanced_pbpDataFramePolars DataFrame with schema ENHANCED_PBP_SCHEMA (from ~sportsdataverse.nba.nba_enhanced_pbp.enhanced_pbp_from_payload). An empty or malformed frame returns a zero-row frame with POSSESSION_SHOOTING_SCHEMA — never raises.

Returns

Polars DataFrame with schema POSSESSION_SHOOTING_SCHEMA. One row per (possession_number, player_id) pair. Events with person_id == 0 are skipped (unattributable to a shooter — they still count toward build_possessions' team-level totals). Per-possession sums of the six shooting columns match the corresponding build_possessions columns exactly.

col_nametypedescription
game_idcharacterUnique game identifier.
possession_numberintegerPossession number.
player_idintegerUnique player identifier.
team_idintegerUnique team identifier.
fg2ainteger
fg2minteger
fg3aintegerThree-point field goal attempts.
fg3mintegerThree-point field goals made.
ftaintegerFree throw attempts.
ftmintegerFree throws made.

Example

import json, pathlib
from sportsdataverse.nba.nba_enhanced_pbp import enhanced_pbp_from_payload
from sportsdataverse.nba.nba_possessions import build_possession_shooting

payload = json.loads(pathlib.Path("playbyplayv3.json").read_text())
pbp = enhanced_pbp_from_payload(payload)
sh = build_possession_shooting(pbp)
print(sh.shape, sh.schema["player_id"])

# Per-player shooting totals

import polars as pl
totals = sh.group_by("player_id").agg(
pl.col("fg3m").sum(), pl.col("ftm").sum()
)
print(totals.head())

clutch_delta​

clutch_delta(clutch: 'pl.DataFrame', ratings: 'pl.DataFrame') -> 'pl.DataFrame'

Clutch net-rating delta vs a full-game baseline, per (season, team_id).

clutch_delta = clutch_net_rating - adj_net_rtg. Joins clutch to the baseline ratings frame on (season, team_id) (asserting dtype agreement first).

Parameters

ParameterTypeDefaultDescription
clutchDataFrameFrame with season, team_id, clutch_net_rating, clutch_poss.
ratingsDataFrameFull-game baseline with season, team_id, adj_net_rtg (the stats full-game net, or any per-team baseline).

Returns

One row per matched (season, team_id): season, team_id, clutch_net_rating, adj_net_rtg, clutch_delta, clutch_poss. Empty input returns that schema with zero rows.

No returns table is published for this function: no capture: its clutch frame is built from stats.nba.com leaguedashteamclutch, which answers HTTP 403 to the datacenter IP the docs are built on.

Example

from sportsdataverse.nba.nba_clutch import clutch_delta
d = clutch_delta(clutch_frame, baseline_frame)

flag_garbage_time​

flag_garbage_time(possessions: 'pl.DataFrame', enhanced_pbp: 'pl.DataFrame', *, starters_on_court: 'Optional[dict[int, int]]' = None) -> 'pl.DataFrame'

Flag CTG garbage time (excluded from CTG stats by default).

CTG (exact): "the game has to be in the 4th quarter, the score differential has to be **>= 25 for minutes 12-9, >= 20 for minutes 9-6, and

= 10 for the remainder of the quarter**. Additionally, there have to be two or fewer starters on the floor combined between the two teams. Importantly, the game can never go back to being non-garbage time, or this clock resets."

The margin x minutes bands are reproduced exactly, evaluated on the score at each possession's start. The reset semantics fall out of that per-possession evaluation: if the trailing team claws back inside the band's threshold, the condition stops holding and those possessions are NOT garbage time (CTG's own "comeback is not counted as garbage time" example); if the lead re-expands, the flag turns back on.

The starters clause is applied only when starters_on_court is supplied — it needs lineup + box START_POSITION data this frame does not carry. Without it the flag is the margin-only superset of CTG's definition (it can flag a blowout stretch in which the starters are still on the floor), and garbage_time_basis records which rule was actually used. This is a deliberate, documented divergence — do not read a margin_only flag as CTG-exact.

Parameters

ParameterTypeDefaultDescription
possessionsDataFrameFrame with period, start_seconds_remaining and start_order_index.
enhanced_pbpDataFrameThe enhanced PBP frame (supplies the running score).
starters_on_courtOptional[dict[int, int]]NoneOptional map possession_number -> number of starters on the floor across BOTH teams. When given, a possession is garbage time only if that count is <= ~nba_play_context_constants.GARBAGE_TIME_MAX_STARTERS.

Returns

possessions with Boolean is_garbage_time and Utf8 garbage_time_basis ("margin_and_starters" or "margin_only") appended.

col_nametypedescription
game_idcharacterUnique game identifier.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
possession_numberintegerPossession number.
offense_team_idintegerUnique identifier for offense team.
defense_team_idinteger
start_order_indexinteger
end_order_indexinteger
start_seconds_remainingdouble
end_seconds_remainingdouble
pointsintegerPoints scored.
is_second_chancelogical
number_in_periodinteger
possession_start_typecharacter
count_as_possessionlogical
fg2ainteger
fg2minteger
fg3aintegerThree-point field goal attempts.
fg3mintegerThree-point field goals made.
ftaintegerFree throw attempts.
ftmintegerFree throws made.
orebintegerOffensive rebounds.
drebintegerDefensive rebounds.
tovintegerTurnovers.
is_garbage_timelogical
garbage_time_basischaracter

Example

poss = flag_garbage_time(poss, pbp)
print(poss.filter(pl.col("is_garbage_time") == True).height)

# CTG-exact, with the starters clause

poss = flag_garbage_time(poss, pbp, starters_on_court=starters_by_possession)

flag_heave_possessions​

flag_heave_possessions(possessions: 'pl.DataFrame') -> 'pl.DataFrame'

Flag CTG's "projected heave possessions" (excluded from CTG stats by default).

CTG (exact): "possessions that start with 4 or fewer seconds on the game clock at the end of one of the first three quarters." Q4/OT are exempt — a late Q4 possession is a real possession.

Parameters

ParameterTypeDefaultDescription
possessionsDataFrameAny frame with period and start_seconds_remaining.

Returns

possessions with a Boolean is_heave_possession column appended.

col_nametypedescription
game_idcharacterUnique game identifier.
periodintegerPeriod of the game (1-4 quarters; 5+ for OT).
possession_numberintegerPossession number.
offense_team_idintegerUnique identifier for offense team.
defense_team_idinteger
start_order_indexinteger
end_order_indexinteger
start_seconds_remainingdouble
end_seconds_remainingdouble
pointsintegerPoints scored.
is_second_chancelogical
number_in_periodinteger
possession_start_typecharacter
count_as_possessionlogical
fg2ainteger
fg2minteger
fg3aintegerThree-point field goal attempts.
fg3mintegerThree-point field goals made.
ftaintegerFree throw attempts.
ftmintegerFree throws made.
orebintegerOffensive rebounds.
drebintegerDefensive rebounds.
tovintegerTurnovers.
off_player_1integer
off_player_2integer
off_player_3integer
off_player_4integer
off_player_5integer
def_player_1integer
def_player_2integer
def_player_3integer
def_player_4integer
def_player_5integer
lineup_sourcecharacter
is_heave_possessionlogical

Example

poss = flag_heave_possessions(poss)
clean = poss.filter(pl.col("is_heave_possession") == False)

hoopshype_salaries​

hoopshype_salaries(*, proxy: 'Any' = None, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

League-wide NBA player salaries from HoopsHype.

One row per player per contract season (current plus the future seasons HoopsHype lists), for the whole league (~600 players).

HoopsHype is a Next.js app: the single /salaries/players/ page paginates client-side and only ~20 rows survive a static fetch, but each team page embeds that team's complete roster in <script id="__NEXT_DATA__">. This walks the 30 slugs in HOOPSHYPE_TEAMS serially -- ~30 requests per call -- and parses that JSON. A team page that fails is warned about and skipped rather than aborting the league.

Pacing is environment-tunable, never hardcoded in the fetch path:

============================ ================================================== SDV_PY_HOOPSHYPE_DELAY seconds slept between team pages (default 0.5) ============================ ==================================================

Parameters

ParameterTypeDefaultDescription
proxyAnyNoneProxy configuration forwarded to ~sportsdataverse.dl_utils.download (requests proxies= shape).
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per player-season with player_id, player, first_name, last_name, team_id, team, season, salary, cap_allocation, team_option, player_option, two_way and qualifying_offer. Ids are Utf8, money is Float64, options are Boolean. All 30 pages failing yields a zero-row frame with that schema.

col_nametypedescription
player_idcharacterUnique player identifier.
playercharacterPlayer name.
first_namecharacterPlayer's first name.
last_namecharacterPlayer's last name.
team_idcharacterUnique team identifier.
teamcharacterTeam-side label or team identifier.
seasonintegerSeason year.
salarydoubleTotal cap-counting salary for the season ($).
cap_allocationdoubleCap allocation for the season (USD).
team_optionlogicalWhether the season is a team option.
player_optionlogicalWhether the season is a player option.
two_waylogicalWhether it is a two-way contract.
qualifying_offerlogicalWhether it is a qualifying offer.

Example

from sportsdataverse.nba import hoopshype_salaries

salaries = hoopshype_salaries()
print(salaries.shape)

# As pandas

salaries_pd = hoopshype_salaries(return_as_pandas=True)

# Pipeline next step (this season's top-paid)

salaries.filter(pl.col("season") == 2026).sort("salary", descending=True).head()

lineup_play_context​

lineup_play_context(possessions: 'pl.DataFrame', *, min_poss: 'int' = 0, league_non_transition_ppp: 'Optional[float]' = None, apply_ctg_filters: 'bool' = True, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Roll possessions up into a per-5-man-lineup Play-Context table.

The lineup analogue of team_play_context: same metric columns, grouped by the five players on the floor for the offense. Lineups are identified by lineup_id — the five player ids sorted ascending and hyphen-joined — so the same five players always land in the same bucket regardless of slot order.

Requires the off_player_1..5 columns from ~sportsdataverse.nba.nba_possessions.attach_possession_lineups (which passes the play-context columns through, so the two compose in either order).

Parameters

ParameterTypeDefaultDescription
possessionsDataFrameFrame from add_play_context with lineups attached.
min_possint0Drop lineups below this possession count (CTG's tables carry a minimum; 0 keeps everything, which is what the partition identity needs).
league_non_transition_pppOptional[float]NonePts+/Poss baseline; see team_play_context.
apply_ctg_filtersboolTrueDrop garbage-time / heave / non-counting possessions first.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per (team, lineup) with LINEUP_PLAY_CONTEXT_SCHEMA. Empty input returns a zero-row frame with that schema.

col_nametypedescription
offense_team_idintegerUnique identifier for offense team.
lineup_idcharacter
off_player_1integer
off_player_2integer
off_player_3integer
off_player_4integer
off_player_5integer
possintegerPoss.
pointsintegerPoints scored.
pts_per_100double
transition_possinteger
transition_pointsinteger
transition_freqdouble
transition_pts_per_100double
non_transition_pts_per_100double
transition_pts_added_per_100double
halfcourt_possinteger
halfcourt_pts_per_100double
freq_off_stealdouble
freq_off_live_rebounddouble

Example

poss = attach_possession_lineups(add_play_context(enh), oncourt, enh, home_team_id=home)
lu = lineup_play_context(poss, min_poss=25)
print(lu.sort("pts_per_100", descending=True).head())

make_prob_by_context​

make_prob_by_context(ptshots: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "'dict[str, Union[pl.DataFrame, pd.DataFrame]]'"

Marginal FG% tables by defender distance and by shot clock.

The public API exposes defender-distance and shot-clock only as aggregate bucket tables (playerdashptshots), not per-shot fields, so this aggregates Σfgm/Σfga across players within each bucket.

Parameters

ParameterTypeDefaultDescription
ptshotsDataFrameThe stacked playerdashptshots fixture — one frame with a result_set tag (ClosestDefenderShooting / ShotClockShooting) plus bucket, fga, fgm.
return_as_pandasboolFalseReturn pandas DataFrames instead of polars.

Returns

{"defender": frame, "shot_clock": frame} each with rows per bucket (bucket, fga, fgm, fg_pct). Missing result sets return the zero-row schema.

Example

from sportsdataverse.nba.nba_shot_value import make_prob_by_context
tables = make_prob_by_context(ptshots)
tables["defender"].sort("fg_pct")

make_prob_joint​

make_prob_joint(defender: 'pl.DataFrame', shot_clock: 'pl.DataFrame', overall_fg_pct: 'float', *, return_as_pandas: 'bool' = False) -> "'Union[pl.DataFrame, pd.DataFrame]'"

Independence-combined defender x shot-clock make probability.

Combines the two marginal FG% tables under a conditional-independence assumption via odds multipliers: odds(p) = p/(1-p); odds_joint = odds_overall * (odds_def/odds_overall) * (odds_clock/odds_overall); joint = odds_joint/(1+odds_joint). This assumes defender distance and shot-clock effects are independent given the league baseline — a simplification (a late clock correlates with tighter defense), documented here so callers weigh it.

Parameters

ParameterTypeDefaultDescription
defenderDataFrameThe "defender" marginal table from make_prob_by_context (bucket, fg_pct).
shot_clockDataFrameThe "shot_clock" marginal table (bucket, fg_pct).
overall_fg_pctfloatThe league overall FG% baseline.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per (close_def_dist_range, shot_clock_range): close_def_dist_range:Utf8, shot_clock_range:Utf8, joint_fg_pct:Float64. Empty inputs return the zero-row schema.

No returns table is published for this function: no capture: its inputs come from make_prob_by_context on stats.nba.com tracking data, which answers HTTP 403 to the datacenter IP the docs are built on.

Example

from sportsdataverse.nba.nba_shot_value import make_prob_by_context, make_prob_joint
t = make_prob_by_context(ptshots)
joint = make_prob_joint(t["defender"], t["shot_clock"], 0.47)

nba_availability​

nba_availability(seasons: "'int | list[int]'", *, league: 'str' = 'nba', gleague_bridge: 'bool' = False, return_as_pandas: 'bool' = False) -> "'pl.DataFrame | pd.DataFrame'"

Project games-available % for a season (or seasons) from career GP history.

avail_pct is availability, not skill -- it is the only output of this function and is never combined into a value/rating column by this module (sportsdataverse.nba.nba_rookie_projection.nba_rookie_projection reports it as a separate column too).

Parameters

ParameterTypeDefaultDescription
seasonsint | list[int]A season (end year, e.g. 2020) or list of seasons.
leaguestr'nba'"nba", "wnba", or "gleague".
gleague_bridgeboolFalseWhen True (and league != "gleague"), also pulls each season's G-League (league_id="20") bulk GP as a development-outcome bridge feature before scoring. Best-effort: gracefully absent (never raises) when the G-League bulk call returns no rows for a season; the returned schema is unaffected either way since the bridge column isn't part of the bundled artifact's scored features.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

Frame player_id:Utf8, season:Int64, avail_pct:Float64 (clipped to [0, 1]). Empty seasons -> zero-row schema.

No returns table is published for this function: no capture: it reads stats.nba.com, which answers HTTP 403 to the datacenter IP the docs are built on; the function works from a residential IP.

Example

from sportsdataverse.nba import nba_availability
proj = nba_availability(2019)
print(proj.sort("avail_pct").head())

nba_box_logs​

nba_box_logs(season: 'str', *, league_id: 'str' = '00', season_type: 'str' = 'Regular Season', fetch: 'Optional[Callable[..., pl.DataFrame]]' = None) -> 'Dict[str, pl.DataFrame]'

Fetch per-player and per-team game logs for a season (bulk, one call each).

Parameters

ParameterTypeDefaultDescription
seasonstrNBA season in "2023-24" form.
league_idstr'00'LeagueID ("00" NBA).
season_typestr'Regular Season'SeasonType ("Regular Season").
fetchOptional[Callable[..., DataFrame]]NoneInjectable nba_stats_leaguegamelog replacement for offline tests.

Returns

{"player": <per-player-game logs>, "team": <per-team-game logs>} as snake-cased polars frames.

Example

from sportsdataverse.nba.nba_box_logs import nba_box_logs
logs = nba_box_logs("2023-24")
print(logs["player"].shape)

nba_foul_drawing​

nba_foul_drawing(season: 'str', *, league_id: 'str' = '00', base: 'Optional[pl.DataFrame]' = None, advanced: 'Optional[pl.DataFrame]' = None, player_mix: 'Optional[pl.DataFrame]' = None, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Expected FTA + residual foul-drawing skill from Synergy play-type mix.

lg_ft_rate_t = poss-weighted league mean of ft_freq for play type t (Σ_players(ft_freq_t·poss_t) / Σ_players(poss_t)). expected_fta = scale · Σ_t poss_t · lg_ft_rate_t; foul_draw_skill = 100·(fta − expected_fta)/poss.

ft_freq (Synergy's ft_poss_pct) counts FT-drawing trips, not individual free-throw attempts (a shooting foul is typically a 2-shot trip) -- scale = Σ actual fta / Σ raw trip-based estimate is derived from the fetched season itself (never hard-coded) so the trip-to-attempt conversion self-normalizes and Σ expected_fta ≡ Σ fta holds exactly.

Parameters

ParameterTypeDefaultDescription
seasonstrSeason string, e.g. "2023-24".
league_idstr'00'"00" NBA (default), "10" WNBA, "20" G-League.
baseOptional[DataFrame]NoneInjected nba_stats_leaguedashplayerstats (Base measure) frame: player_id, fta, poss (bypasses the live fetch).
advancedOptional[DataFrame]NoneInjected Advanced-measure frame with player_id, pfd (personal fouls drawn); optional -- pfd is null when omitted (fta is the always-present proxy).
player_mixOptional[DataFrame]NoneInjected Synergy player-level offensive mix: player_id, play_type, poss, ft_freq.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per player: player_id (Int64), poss/fta/ expected_fta/foul_draw_skill (Float64), pfd (Float64, null when advanced has no data for that player or is omitted). Zero-row frame with this schema when the inputs are empty (sparse-coverage leagues never raise).

No returns table is published for this function: no capture: it reads stats.nba.com, which answers HTTP 403 to the datacenter IP the docs are built on; the function works from a residential IP.

Example

from sportsdataverse.nba import nba_foul_drawing
f = nba_foul_drawing("2023-24")
print(f.sort("foul_draw_skill", descending=True).head(10))

# Injected offline (oracle / test) path

f = nba_foul_drawing("2023-24", base=base_df, player_mix=mix_df)

# Pipeline next step

f.filter(pl.col("poss") >= 200).sort("foul_draw_skill", descending=True)

nba_l2m​

nba_l2m(game_id: 'str | int', *, raw: 'bool' = False, return_as_pandas: 'bool' = False, proxy: 'dict | None' = None) -> 'dict[str, Any]'

Fetch and parse an NBA Last Two Minute report from official.nba.com.

Retrieves the L2M report for a given game and returns parsed tables of calls, game metadata, and error statistics. The game_id is zero-padded to 10 digits (e.g., 42500405 becomes "0042500405"). A report is published for any game that is within 3 points (5 points before the 2017-18 season) at any point during the last two minutes of the fourth quarter or overtime -- not only playoff games.

Parameters

ParameterTypeDefaultDescription
game_idstr | intNBA game ID (can be int or str). Automatically zero-padded to 10 digits.
rawboolFalseIf True, return the raw JSON payload (dict) instead of parsed DataFrames.
return_as_pandasboolFalseIf True, return pandas DataFrames instead of polars.
proxydict | NoneNoneOptional proxy dict passed through to the HTTP layer.

Returns

If raw=True, the raw JSON dict. Otherwise, a dict with keys "calls", "game", "stats" mapping to DataFrames as documented in parse_nba_l2m.

col_nametypedescription
calls.game_idcharacter10-digit NBA game id (zero-padded), the join key back to the L2M game and stats tables.
calls.periodintegerPeriod number extracted from period_name's digits: 4 for the fourth quarter, 5+ for overtime periods (Q5 = OT1 through Q8 = OT4).
calls.period_namecharacterRaw NBA.com period label for the graded play, e.g. Q4 for the fourth quarter or Q5 for the first overtime.
calls.pc_timecharacterRaw game clock string from the L2M report before normalization, in MM:SS or MM:SS.t tenths-of-a-second format.
calls.seconds_remainingdoubleSeconds remaining in the period at the graded play, parsed out of pc_time.
calls.call_typecharacterRaw "Call: Type" label from the report, e.g. "Foul: Shooting" or "Turnover: 24 Second Violation"; whitespace (including doubled spaces) is NOT cleaned here -- the cleaned, upper-cased parts are call and type.
calls.callcharacterUpper-cased category before the colon in call_type, e.g. FOUL, TURNOVER, or STOPPAGE.
calls.typecharacterUpper-cased detail after the colon in call_type, e.g. SHOOTING or 24 SECOND VIOLATION.
calls.committingcharacterPlayer, team nickname, or coach responsible for the graded action.
calls.disadvantagedcharacterPlayer or team nickname disadvantaged by the graded action, blank when not applicable.
calls.decisioncharacterNormalized grading decision: CC (correct call), CNC (correct non-call), IC (incorrect call), or INC (incorrect non-call); null when the report marks the play undetectable at game speed.
calls.decision_rawcharacterRaw grading code from the report before normalization, including the rare NCC/NCI aliases folded into decision.
calls.commentcharacterGrader's free-text explanation of the ruling, sometimes naming a player as "Last (TEAM)".
calls.difficultycharacterGrader's difficulty rating for the play, e.g. Observable, Difficult, Blatant/Hard, or Undetectable.
calls.video_event_idcharacterL2M video event id from the report (source field VideolLink, sic); not a stats.nba.com play-by-play EVENTNUM, so it cannot be joined to pbp data.
calls.pos_idintegerPossession id assigned by the report, rising monotonically across the game; rows sharing pos_start/pos_end belong to the same possession.
calls.pos_startcharacterGame clock at the start of the possession containing this graded play.
calls.pos_endcharacterGame clock at the end of the possession containing this graded play.
calls.pos_team_idinteger10-digit NBA team id of the team in possession during the graded play.
game.game_idcharacter10-digit NBA game id (zero-padded), the join key to the calls and stats tables.
game.game_datedateGame date parsed from the report's local tip-off timestamp (no timezone offset is applied).
game.season_typecharacterSeason type inferred from the third digit of game_id: preseason, regular, all-star, playoffs, play-in, or nba-cup-final.
game.home_team_idinteger10-digit NBA team id (1610612xxx) of the home team.
game.away_team_idinteger10-digit NBA team id (1610612xxx) of the away team.
game.home_team_abbrcharacterThree-letter abbreviation of the home team.
game.away_team_abbrcharacterThree-letter abbreviation of the away team.
game.home_team_namecharacterHome team nickname as published in the report, e.g. Thunder rather than Oklahoma City Thunder.
game.away_team_namecharacterAway team nickname as published in the report.
game.home_scoreintegerHome team's final score for the game.
game.away_scoreintegerAway team's final score for the game.
game.l2m_commentscharacterReport-level note from the league, e.g. a technical-issue disclaimer about missing video; null for almost every game.
stats.game_idcharacter10-digit NBA game id (zero-padded), the join key to the calls and game tables.
stats.stat_namecharacterName of the report-level error-rate statistic: Calls, Errors in Favor, or Possessions in Favor.
stats.homeintegerValue of stat_name attributed to the home team.
stats.awayintegerValue of stat_name attributed to the away team.

Example

from sportsdataverse.nba.nba_officiating import nba_l2m
result = nba_l2m("0042500405")
calls = result["calls"]
print(f"Game had {calls.height} tracked plays in the L2M window")

# Parse as pandas instead

result = nba_l2m("0042500405", return_as_pandas=True)
calls_pd = result["calls"]

# Access raw JSON

payload = nba_l2m("0042500405", raw=True)
print(payload["game"])

nba_l2m_games​

nba_l2m_games(season: 'int | str', *, return_as_pandas: 'bool' = False, proxy: 'dict | None' = None) -> 'Union[pl.DataFrame, pd.DataFrame]'

Fetch the list of games with Last Two Minute reports for an NBA season.

Retrieves and parses the L2M season index page from official.nba.com, returning a table of all games for which L2M reports exist. JSON reports exist only from 2019-01-01 onward; earlier seasons' index pages list PDFs, which this function ignores. A release loader for historical (PDF-era) reports is planned but does not exist yet.

Parameters

ParameterTypeDefaultDescription
seasonint | strThe NBA season (end year) as a 4-digit int or string, e.g. 2026 for the 2025-26 season.
return_as_pandasboolFalseIf True, return a pandas DataFrame instead of polars.
proxydict | NoneNoneOptional proxy dict passed through to the HTTP layer.

Returns

A polars.DataFrame (or pandas.DataFrame when return_as_pandas=True) with schema {"game_id": Utf8, "season": Int32, "season_type": Utf8, "label": Utf8}, one row per unique game ID in page order.

col_nametypedescription
game_idcharacter10-digit NBA game id (zero-padded), parsed from the season listing page's report link.
seasonintegerNBA season end year passed to the function (e.g. 2026 for the 2025-26 season), stamped onto every row.
season_typecharacterSeason type inferred from the third digit of game_id: preseason, regular, all-star, playoffs, play-in, or nba-cup-final.
labelcharacterMatchup label text scraped from the listing page link, e.g. "Thunder 125, Rockets 124 (2OT)".

Example

from sportsdataverse.nba.nba_officiating import nba_l2m_games
df = nba_l2m_games(2026)
print(f"Season had {df.height} games with L2M reports")

# Get playoff games only

df = nba_l2m_games(2026)
playoffs = df.filter(df["season_type"] == "playoffs")

# Convert to pandas

df = nba_l2m_games(2026, return_as_pandas=True)

nba_play_context​

nba_play_context(game_id: 'str', league_id: 'str' = '00', *, transition_seconds: 'float' = 6.0, transition_variant: 'str' = 'hoop_math', return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Fetch one game and return its possessions with the full CTG play-context surface.

Single live call to nba_stats_playbyplayv3, then add_play_context. Works for NBA (league_id="00"), WNBA and the G-League — stats.wnba.com ships the same play-by-play shapes, and every threshold here is league-agnostic.

Parameters

ParameterTypeDefaultDescription
game_idstrTen-character game identifier (e.g. "0022200001").
league_idstr'00'League identifier ("00" NBA, "10" WNBA, "20" G-League).
transition_secondsfloat6.0Transition initial-play cutoff (default 10.0).
transition_variantstr'hoop_math'See add_transition.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

Possession frame with the play-context columns. Empty/malformed payloads return a zero-row frame — never raises.

No returns table is published for this function: no capture: it reads stats.nba.com, which answers HTTP 403 to the datacenter IP the docs are built on; the function works from a residential IP.

Example

from sportsdataverse.nba.nba_play_context import nba_play_context
poss = nba_play_context("0022200001")
print(poss["possession_start_type_ctg"].value_counts())

# Transition rate for the game

import polars as pl
clean = poss.filter(
(pl.col("is_garbage_time") == False) & (pl.col("is_heave_possession") == False)
)
print(clean["is_transition"].mean())

nba_player_ages​

nba_player_ages(season: 'str', *, league_id: 'str' = '00', fetch: 'Optional[Callable[..., pl.DataFrame]]' = None) -> 'pl.DataFrame'

Per-player age for a season (bulk), for the DARKO aging curve.

Parameters

ParameterTypeDefaultDescription
seasonstrNBA season, e.g. "2023-24".
league_idstr'00'LeagueID ("00" NBA).
fetchOptional[Callable[..., DataFrame]]NoneInjectable nba_stats_leaguedashplayerbiostats replacement for offline tests.

Returns

Frame player_id:Int64, age:Float64.

No returns table is published for this function: no capture: it reads stats.nba.com, which answers HTTP 403 to the datacenter IP the docs are built on; the function works from a residential IP.

Example

from sportsdataverse.nba import nba_player_ages
ages = nba_player_ages("2023-24")
print(ages.head())

nba_player_identity​

nba_player_identity(player_logs: 'pl.DataFrame') -> 'pl.DataFrame'

Human-readable identity for every player in a season's box logs.

Model outputs key on player_id alone, which makes them unusable without a second lookup -- a leaderboard reads 1628983 instead of Shai Gilgeous-Alexander. This derives the display columns from the season's own game logs, so they are season-accurate: a player's team is what he actually played for that year, not his current one (which is what a player directory would give and would silently mislabel every historical season).

A traded player has rows for several teams. team_* is his primary team by minutes -- the one a reader means when they say "his team that season" -- and teams lists every abbreviation he appeared for, in descending minutes, so a trade is visible rather than silently collapsed.

Parameters

ParameterTypeDefaultDescription
player_logsDataFramePer-player-per-game rows from leaguegamelog (the player_or_team="P" variant), carrying player_id, player_name, team_id, team_abbreviation, team_name and min.

Returns

One row per player_id with PLAYER_IDENTITY_SCHEMA. An empty input -- or one missing any required column, min included -- gives the zero-row frame with that schema, so callers can join unconditionally. min is required rather than optional: without it every team totals zero minutes and "primary team" quietly degrades to whichever team_id sorts first, which looks like an answer but is not one.

col_nametypedescription
player_idintegerUnique player identifier.
player_namecharacterPlayer name.
team_idintegerUnique team identifier.
team_abbreviationcharacterShort team abbreviation (e.g. 'LAS').
team_namecharacterFull team display name (e.g. 'Las Vegas Aces').
teamscharacterNested list of member-team membership spans.

Example

import polars as pl
from sportsdataverse.nba import nba_player_identity

logs = pl.DataFrame({
"player_id": [1628983],
"player_name": ["Shai Gilgeous-Alexander"],
"team_id": [1610612760],
"team_abbreviation": ["OKC"],
"team_name": ["Oklahoma City Thunder"],
"min": [34.0],
})
ratings = pl.DataFrame({"player_id": [1628983], "war": [21.9]})
named = ratings.join(nba_player_identity(logs), on="player_id", how="left")
print(named.select("player_name", "team_name", "war"))

nba_player_positions​

nba_player_positions(season: 'str', *, league_id: 'str' = '00', fetch: 'Optional[Callable[..., pl.DataFrame]]' = None) -> 'pl.DataFrame'

Fetch league-wide listed positions for a season as numeric 1-5.

Parameters

ParameterTypeDefaultDescription
seasonstrNBA season, e.g. "2023-24".
league_idstr'00'LeagueID ("00" NBA, "10" WNBA, "20" G-League).
fetchOptional[Callable[..., DataFrame]]NoneInjectable nba_stats_playerindex replacement for offline tests.

Returns

Frame with columns player_id:Int64, position_num:Float64.

No returns table is published for this function: no capture: it reads stats.nba.com, which answers HTTP 403 to the datacenter IP the docs are built on; the function works from a residential IP.

Example

from sportsdataverse.nba import nba_player_positions
pos = nba_player_positions("2023-24")
print(pos.head())

# Offline / injectable fetch for testing

import polars as pl
stub = lambda **kw: pl.DataFrame({"person_id": [1], "position": ["PG"]})
pos = nba_player_positions("2023-24", fetch=stub)

nba_player_props​

nba_player_props(season: 'int', game_id: 'str', home_team_id: 'str', away_team_id: 'str', *, league_id: 'str' = '00', return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Per-player expected prop lines + team pace projection for a matchup.

Loads the season's player box logs + team ratings, computes per-minute player_rates, and projects each player's line onto their mean minutes and the matchup's pace factor (exp_poss / avg_pace). Only the two teams in the matchup are returned.

Parameters

ParameterTypeDefaultDescription
seasonintEnd year of the season (e.g. 2024).
game_idstrThe game id (passed through for the caller's join; not used to filter historical rates).
home_team_idstrHome team id.
away_team_idstrAway team id.
league_idstr'00'"00" NBA / "10" WNBA / "20" G-League.
return_as_pandasboolFalseReturn a pandas frame instead of polars.

Returns

One row per player on either team: player_id, team_id, stat_pts_exp, stat_reb_exp, stat_ast_exp, stat_fg3m_exp, pace_proj. Empty input returns that schema with zero rows.

col_nametypedescription
player_idcharacterUnique player identifier.
team_idcharacterUnique team identifier.
stat_pts_expdouble
stat_reb_expdouble
stat_ast_expdouble
stat_fg3m_expdouble
pace_projdouble

Example

from sportsdataverse.nba.nba_player_props import nba_player_props
props = nba_player_props(2024, "401585828", "2", "6")

nba_playtype_ratings​

nba_playtype_ratings(season: 'str', *, league_id: 'str' = '00', off_team: 'Optional[pl.DataFrame]' = None, def_team: 'Optional[pl.DataFrame]' = None, schedule: 'Optional[pl.DataFrame]' = None, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Season play-type-adjusted offensive/defensive team ratings.

Fetches (or uses injected) Synergy offensive/defensive team frames plus the league schedule, computes raw per-type efficiency (raw_playtype_efficiency), opponent-adjusts it (adjust_playtype_efficiency), then rolls up to one row per team.

Parameters

ParameterTypeDefaultDescription
seasonstrSeason string, e.g. "2023-24".
league_idstr'00'"00" NBA (default), "10" WNBA, "20" G-League.
off_teamOptional[DataFrame]NoneInjected Synergy offensive team frame (bypasses the live fetch -- used for tests / oracle fixtures).
def_teamOptional[DataFrame]NoneInjected Synergy defensive team frame.
scheduleOptional[DataFrame]NoneInjected team_id/opp_team_id schedule frame.
return_as_pandasboolFalseReturn a pandas DataFrame instead of polars.

Returns

One row per team: team_id (Int64), adj_off/adj_def/adj_net (Float64) roll-ups, plus per-type wide columns adj_off_ppp_<playtype>/adj_def_ppp_<playtype>/off_freq_<playtype> (Float64) for each play type present in the data. adj_off = Σ_t off_freq_t · adj_off_ppp_t · 100 (symmetric for adj_def off def_freq_t); adj_net = adj_off - adj_def. Returns a zero-row frame with the base roll-up schema when the upstream fetch is empty (sparse-coverage leagues never raise).

No returns table is published for this function: no capture: it reads stats.nba.com, which answers HTTP 403 to the datacenter IP the docs are built on; the function works from a residential IP.

Example

from sportsdataverse.nba import nba_playtype_ratings
r = nba_playtype_ratings("2023-24")
print(r.sort("adj_off", descending=True).head(10))

# Injected offline (oracle / test) path

r = nba_playtype_ratings("2023-24", off_team=off_df, def_team=def_df, schedule=sched_df)

# Pipeline next step

r.filter(pl.col("adj_net") > 0).sort("adj_net", descending=True)

nba_ratings_panel​

nba_ratings_panel(model: 'AnyModel', possessions: 'pl.DataFrame', dates: 'Optional[Sequence[datetime.date]]' = None, *, return_as_pandas: 'bool' = False) -> 'Union[pl.DataFrame, pd.DataFrame]'

Player-ratings-through-date long panel: one row per (player_id, date).

Refit-per-checkpoint (v1; no warm-start incrementality) — each date's row calls ratings_as_of independently, so the panel is leakage-free by construction: a possession dated after a given checkpoint can never affect that checkpoint's row, no matter what other dates are also being computed or what future rows exist in possessions. Cost is a full refit per checkpoint date; for a season's sparse RAPM-family design this is seconds per date, not minutes — acceptable for a nightly/daily cadence but not for live in-game updating (out of scope; see spec non-goals).

Parameters

ParameterTypeDefaultDescription
modelAnyModelA harness model conforming to nba_model_validation.AnyModel.
possessionsDataFrameA possession+lineup frame with a game_date (pl.Date) column (as emitted by compile_nba_season).
datesOptional[Sequence[date]]NoneCheckpoint dates to compute. None (default) uses every distinct game_date present in possessions, sorted ascending — a rating for every game day, matching what EPM/LEBRON publish nightly. Duplicates are deduped; input order does not matter (the output is always sorted by date).
return_as_pandasboolFalseReturn pandas instead of polars.

Returns

Long frame with RATINGS_PANEL_SCHEMA columns (player_id, date, o_rating, d_rating, rating). Zero-row (that schema) when possessions is empty or no date yields any players.

No returns table is published for this function: no capture: its possessions input needs a game_date column, and neither nba_possessions nor the released possessions carry one.

Example

import datetime
from sportsdataverse.nba.nba_model_validation import RidgeRapmModel
from sportsdataverse.nba.nba_ratings_panel import nba_ratings_panel

checkpoints = [datetime.date(2023, 11, 1), datetime.date(2023, 12, 1)]
panel = nba_ratings_panel(RidgeRapmModel(), season_poss, dates=checkpoints)
print(panel.filter(pl.col("player_id") == 201939).sort("date"))

# Every game day, no explicit grid

panel = nba_ratings_panel(RidgeRapmModel(), season_poss)

nba_raw_store_season_frame​

nba_raw_store_season_frame(endpoint: 'str', season: 'int', variant: 'Optional[str]' = None, *, result_set: 'Optional[str]' = None, raw_store_dir: 'RawStoreDir' = None) -> "Optional['pl.DataFrame']"

Read a committed SEASON-LEVEL capture from the raw store, parsed to a frame.

The per-game half of the store is served by the read-through per-game path; this is the season-keyed half (leaguegamelog, playerindex, leaguedashplayerbiostats, ...) that the -raw scraper writes as

  • {endpoint}/{season}/{variant}.json -- parameterized captures, where variant is the slugified parameter sweep (e.g. "regular-season", "regular-season_totals"), and
  • {endpoint}/{season}.json -- unparameterized captures (e.g. playerindex), i.e. variant=None.

Roots may be a local checkout or an http(s):// base (the raw repo served over raw.githubusercontent / a CDN), so a consumer -- notably the hoopR-nba-stats-data model producer -- runs clone-free in CI against the same committed tree the per-game compile reads.

Parameters

ParameterTypeDefaultDescription
endpointstrstats.nba.com endpoint slug (the store subdirectory).
seasonintSeason END year (2024 = 2023-24) -- the key both halves of the store are filed under.
variantOptional[str]NoneCapture variant slug, or None for an unparameterized capture.
result_setOptional[str]NoneNamed result set to return when the payload carries several; defaults to the first frame that parses.
raw_store_dirRawStoreDirNoneStore root spec (dir or URL base) or per-endpoint mapping; None falls back to the env vars, "" disables.

Returns

The parsed polars.DataFrame, or None when the store is unset, the capture is absent, or the payload carries no usable frame -- so a caller can cleanly fall back to a live fetch.

col_nametypedescription
season_idcharacterUnique season identifier.
team_idintegerUnique team identifier.
team_abbreviationcharacterShort team abbreviation (e.g. 'LAS').
team_namecharacterFull team display name (e.g. 'Las Vegas Aces').
game_idcharacterUnique game identifier.
game_datecharacterGame date (YYYY-MM-DD).
matchupcharacterMatchup.
wlcharacterWl.
minintegerMinutes played.
fgmintegerField goals made.
fgaintegerField goal attempts.
fg_pctdoubleField goal percentage (0-1).
fg3_mintegerThree-point field goals made.
fg3_aintegerThree-point field goal attempts.
fg3_pctdoubleThree-point field goal percentage (0-1).
ftmintegerFree throws made.
ftaintegerFree throw attempts.
ft_pctdoubleFree throw percentage (0-1).
orebintegerOffensive rebounds.
drebintegerDefensive rebounds.
rebintegerRebounds per game.
astintegerAssists.
stlintegerSteals.
blkintegerBlocks.
tovintegerTurnovers.
pfintegerPersonal fouls.
ptsintegerPoints scored.
plus_minusintegerPlus/minus point differential while on court.
video_availableintegerVideo available.

Example

from sportsdataverse.nba import nba_raw_store_season_frame
base = "https://raw.githubusercontent.com/sportsdataverse/hoopR-nba-stats-raw/main/nba_stats/json"
logs = nba_raw_store_season_frame("leaguegamelog", 2024, "regular-season", raw_store_dir=base)

# Fall back to a live fetch when the capture is absent

frame = nba_raw_store_season_frame("playerindex", 2024, raw_store_dir=base)
positions = frame if frame is not None else nba_stats_playerindex(season="2023-24")

nba_referee_assignments​

nba_referee_assignments(date: 'str | _dt.date', *, league: 'str' = 'nba', raw: 'bool' = False, return_as_pandas: 'bool' = False, proxy: 'dict | None' = None) -> 'dict[str, Any]'

Fetch and parse NBA referee assignments for a given date from official.nba.com.

Retrieves one league's referee crew assignments and replay-center officials for a date: NBA, G-League or WNBA, picked by league (raw=True returns all three leagues' blocks). The crew_position column (1–4) represents the feed's slot order; slot 1 is inferred to be the crew chief. The season column converts from the feed's format to an END year: START+1 for NBA/G-League (two-calendar-year seasons) and START unchanged for WNBA.

Parameters

ParameterTypeDefaultDescription
datestr | dateThe date to fetch assignments for (str in "YYYY-MM-DD" format or datetime.date).
leaguestr'nba'The league to extract ("nba", "gl", or "wnba"). Defaults to "nba".
rawboolFalseIf True, return the raw JSON payload (dict) with all three leagues instead of parsed DataFrames.
return_as_pandasboolFalseIf True, return pandas DataFrames instead of polars.
proxydict | NoneNoneOptional proxy dict passed through to the HTTP layer.

Returns

If raw=True, the raw JSON dict with keys "nba", "gl", "wnba". Otherwise, a dict with keys "officials" and "replay_center" mapping to DataFrames as documented in parse_nba_referee_assignments.

col_nametypedescription
officials.leaguecharacterLeague the assignment belongs to: nba, gl (G League), or wnba.
officials.game_idcharacter10-digit game id (zero-padded) for the assigned game.
officials.game_datedateGame date parsed from the feed's MM/DD/YYYY format.
officials.seasonintegerSeason end year, converted from the feed's code: start year + 1 for NBA/G League's two-calendar-year seasons, start year unchanged for WNBA's single-year seasons.
officials.season_typecharacterSeason type decoded from the feed's season code first digit: preseason, regular, all-star, playoffs, play-in, or nba-cup-final.
officials.game_codecharacterLeague game code in YYYYMMDD/AWYHOM format, matching the away and home team abbreviations.
officials.home_team_idinteger10-digit team id of the home team.
officials.home_team_abbrcharacterThree-letter abbreviation of the home team.
officials.away_team_idinteger10-digit team id of the away team.
officials.away_team_abbrcharacterThree-letter abbreviation of the away team.
officials.crew_positionintegerFeed's official slot order (1-4); slot 1 is inferred to be the crew chief since the API does not label roles.
officials.official_idintegerNumeric official id from the feed (source field official{n}_code); expected to match stats.nba.com's OFFICIAL_ID.
officials.official_namecharacterOfficial's display name for this crew slot.
officials.jersey_numcharacterOfficial's jersey number as a string, from the feed's official{n}_JNum field.
replay_center.leaguecharacterLeague the replay-center staffing belongs to: nba, gl, or wnba.
replay_center.game_datedateDate the replay-center official worked; a date-level staffing record, not tied to one game.
replay_center.official_idintegerNumeric replay-center official id from the feed.
replay_center.official_namecharacterReplay-center official's display name for that date.

Example

from sportsdataverse.nba.nba_officiating import nba_referee_assignments
result = nba_referee_assignments("2026-06-13")
officials = result["officials"]
print(f"Found {officials.height} official slots")

# Fetch WNBA assignments for the same date

result = nba_referee_assignments("2026-06-13", league="wnba")
wnba_officials = result["officials"]