Skip to main content

NFL — additional Python functions — Other

NflConfig​

NflConfig(cache_mode: 'CacheMode' = 'memory', cache_dir: 'Optional[Path]' = None, cache_duration: 'int' = 86400, verbose: 'bool' = True, timeout: 'int' = 30, user_agent: 'str' = 'sportsdataverse-py-nfl') -> None

Runtime configuration for sdv-py NFL loaders.

Fields mirror nflreadpy's NflreadpyConfig so users can swap engines without changing call sites. The defaults are conservative: in-memory caching with a 24-hour TTL, verbose progress bars on, 30-second HTTP timeout.

Parameters

ParameterTypeDefaultDescription
cache_modeCacheMode'memory'
cache_dirOptional[Path]None
cache_durationint86400
verboseboolTrue
timeoutint30
user_agentstr'sportsdataverse-py-nfl'

Example

from sportsdataverse.nfl import get_config
cfg = get_config() # NflConfig instance
cfg.cache_mode # "memory"
cfg.cache_duration # 86400 (24h)
cfg.timeout # 30 (seconds)

# Construct a fresh instance directly (rarely needed -- prefer ``update_config``)

from sportsdataverse.nfl import NflConfig
cfg = NflConfig(cache_mode="off", timeout=10)

adjust_pressure_pairs​

adjust_pressure_pairs(pairs: 'pl.DataFrame', *, max_iter: 'int' = 50, tol: 'float' = 0.0001) -> 'pl.DataFrame'

Opponent-adjust matchup pressure rates via an additive fixed point.

Fits rate(off, def) ~ mu + alpha_off + beta_def per season by alternating dropback-weighted residual means (league-mean-centered); league-agnostic (no NFL constant inside).

Parameters

ParameterTypeDefaultDescription
pairsDataFrameOutput of pressure_pairs (or any frame with season, off_team, def_team, dropbacks, pressures).
max_iterint50Fixed-point iteration cap.
tolfloat0.0001Max-abs-change convergence tolerance.

Returns

Per (season, team): raw allowed/generated rates + counts and adj_pressure_rate_allowed (mu + alpha) / adj_pressure_rate_generated (mu + beta).

col_nametypedescription
seasonintegerSeason of the aggregate.
teamcharacterTeam abbreviation.
dropbacks_offintegerOffensive dropbacks (qb_dropback plays).
pressures_allowedintegerSacks plus QB hits allowed on the team's own dropbacks.
pressure_rate_alloweddoublepressures_allowed / dropbacks_off (raw).
dropbacks_defintegerOpponent dropbacks faced on defense.
pressures_generatedintegerSacks plus QB hits generated against opponent dropbacks.
pressure_rate_generateddoublepressures_generated / dropbacks_def (raw).
adj_pressure_rate_alloweddoubleOpponent-adjusted allowed pressure rate (mu + team offense effect from the additive fixed point).
adj_pressure_rate_generateddoubleOpponent-adjusted generated pressure rate (mu + team defense effect from the additive fixed point).

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.nfl_line_grades import (
adjust_pressure_pairs, pressure_pairs,
)
adj = adjust_pressure_pairs(pressure_pairs(load_nfl_pbp([2023])))
print(adj.sort("adj_pressure_rate_generated", descending=True).head())

cached_loader​

cached_loader(func: 'F') -> 'F'

Decorator that adds caching to a load_nfl_* function.

Honors the active NflConfig.cache_mode:

  • memory: dict-based per-process cache.
  • filesystem: parquet-based cross-process cache under cache_dir.
  • off: no caching, function runs every time.

The cache key is the hash of (qualified_name, args, kwargs) with return_as_pandas excluded so memory / disk hits work regardless of which return shape the caller asked for. The cache always stores the polars frame internally and converts to pandas on read when requested.

Parameters

ParameterTypeDefaultDescription
funcF

Example

import polars as pl
from sportsdataverse.nfl.cache import cached_loader

@cached_loader
def load_my_thing(season: int, return_as_pandas: bool = False):
# ... fetch parquet, build a polars frame ...
return pl.DataFrame({"season": [season]})

df1 = load_my_thing(2024) # network hit, populates cache
df2 = load_my_thing(2024) # served from cache
df_pd = load_my_thing(2024, return_as_pandas=True)
# `return_as_pandas` is excluded from the cache key, so the
# polars hit is reused and converted to pandas on the way out.

# Switch caching modes at runtime

from sportsdataverse.nfl import clear_cache, update_config

update_config(cache_mode="filesystem") # parquet-on-disk reuse
df3 = load_my_thing(2024) # writes parquet under cache_dir
clear_cache() # wipe both memory + filesystem
update_config(cache_mode="off") # bypass cache entirely

clean_nfl_pbp​

clean_nfl_pbp(df: 'pl.DataFrame', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Canonicalize names/ids/teams on a play-by-play frame (nflfastR clean_pbp port).

See the module docstring for the full column set added, the compute-if-absent scope note on pass/rush, and the lookaround -> capture-group regex rewrites.

Parameters

ParameterTypeDefaultDescription
dfDataFrameAn nflverse-shape (or ESPN/native) play-by-play polars.DataFrame. Required columns: desc, epa, game_id, play_id, season, posteam. See the module docstring for the full optional-column-with-default list.
return_as_pandasboolFalseIf True, return a pandas.DataFrame; otherwise a polars.DataFrame (default).

Returns

The input frame with every §6 column added/overwritten (idempotent -- pre-existing values of those columns, except pass/rush, are dropped and recomputed). A zero-row input yields a zero-row frame carrying the full documented schema rather than raising.

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.nfl_clean import clean_nfl_pbp

pbp = load_nfl_pbp([2023])
cleaned = clean_nfl_pbp(pbp)
print(cleaned.select("name", "id", "fantasy").head())

# Pandas output

cleaned_pd = clean_nfl_pbp(pbp, return_as_pandas=True)

# Pipeline next step (one line)

import polars as pl
cleaned.filter(pl.col("play") == 1).group_by("passer").len()

clear_cache​

clear_cache() -> 'None'

Clear both memory and filesystem caches.

Memory: empties the in-process dict. Filesystem: removes all entries under config.cache_dir. The directory itself is preserved so subsequent writes succeed without needing mkdir.

The models/ subdirectory is deliberately preserved — it holds download-on-demand model artifacts (e.g. the ~34 MB xyac_model.ubj) that are expensive to re-fetch. Clearing the data cache should not force a model re-download; delete <cache_dir>/models/ by hand to drop those.

Example

from sportsdataverse.nfl import clear_cache, load_nfl_pbp
clear_cache()
pbp = load_nfl_pbp(seasons=[2024])

# Pair with a cache-mode switch

from sportsdataverse.nfl import clear_cache, update_config
update_config(cache_mode="filesystem")
# ... lots of cached calls accumulate parquet files on disk ...
clear_cache() # wipe disk + memory together

compose_counting_projection​

compose_counting_projection(rate_proj: 'pl.DataFrame', avail_proj: 'pl.DataFrame', *, rate_col: 'str' = 'proj_rate', volume_col: 'str' = 'proj_volume') -> 'pl.DataFrame'

Compose skill and availability into a counting projection.

The only place skill (rate x volume) and availability meet: proj_counting = rate * volume * proj_availability, joined on player_id (dtype-asserted).

Parameters

ParameterTypeDefaultDescription
rate_projpl.DataFrameSkill projection carrying player_id + rate_col + volume_col.
avail_projpl.DataFrameAvailability projection carrying player_id + proj_availability.
rate_colstr'proj_rate'Rate column name in rate_proj.
volume_colstr'proj_volume'Volume column name in rate_proj.

Returns

rate_proj columns plus proj_availability and proj_counting:Float64.

col_nametypedescription
player_idcharacternflverse gsis player id (character join key; asserted Utf8 on both sides of the join).
proj_ratedoubleProjected per-opportunity rate carried through from the skill projection (rate_col).
proj_volumedoubleProjected opportunity volume carried through from the skill projection (volume_col).
proj_availabilitydoubleProjected availability rate in [0, 1] from nfl_availability_projection.
proj_countingdoubleComposed counting projection - proj_rate x proj_volume x proj_availability (the only place skill and availability meet).

Example

import polars as pl
from sportsdataverse.nfl.nfl_availability import compose_counting_projection
out = compose_counting_projection(rate_frame, avail_frame)

efficiency_ratings​

efficiency_ratings(plays: 'pl.DataFrame', *, config: 'RatingsConfig | None' = None) -> 'pl.DataFrame'

One row per team: opponent-adjusted offense/defense EPA per play.

Filters plays to competitive non-special-teams scrimmage plays (special != 1, qb_kneel != 1, qb_spike != 1, min_competitive_wp <= wp <= max_competitive_wp, non-null epa/posteam/defteam) and fits opponent_adjusted_ridge on epa. Callers pass an already as-of-date-filtered frame (the public nfl_ratings entry point does the date filter) -- this function is pure.

Parameters

ParameterTypeDefaultDescription
playsDataFrameAn load_nfl_pbp-schema frame carrying game_id, posteam, defteam, home_team, epa, wp, special, qb_kneel, qb_spike.
configRatingsConfig | NoneNoneTuning knobs (ridge_lambda + the competitive-wp window); defaults to RatingsConfig.

Returns

One row per team_id (Utf8) with adj_off_epa / adj_def_epa / adj_net (Float64, adj_net = adj_off_epa - adj_def_epa) and games (Int64). Zero-row, correctly-typed on empty/fully-filtered input.

Example

from sportsdataverse.nfl.nfl_ratings import efficiency_ratings
ratings = efficiency_ratings(pbp)
ratings.sort("adj_net", descending=True).head()

env_adjusted_make_prob​

env_adjusted_make_prob(pbp: 'pl.DataFrame') -> 'pl.DataFrame'

Add base_make_prob + environment-adjusted exp_make_prob.

`exp_make_prob = sigmoid(logit(base) + b_windwind + b_temp(temp-baseline)

  • b_alt*altitude_kft)with coefficients fromsportsdataverse.nfl.nfl_scheme_constants.ENVIRONMENT_FG_COEFand altitude fromSTADIUM_ALTITUDE[home_team]`. Dome / closed-roof kicks (and missing readings) are treated as neutral (wind 0, temp = baseline).

Parameters

ParameterTypeDefaultDescription
pbpDataFrameFG-attempt rows with yardline_100 / roof / temp / wind / home_team (+ season or era0..era4 / fg_roof).

Returns

The input plus base_make_prob and exp_make_prob (Float64).

col_nametypedescription
base_make_probdoubleShipped fg_model make probability (with nfl4th long-kick clamps applied).
exp_make_probdoubleEnvironment-adjusted make probability (logit shift for long-kick clamp correction, wind, temperature and altitude).

Example

import polars as pl
from sportsdataverse.nfl.nfl_kicker_rating import env_adjusted_make_prob
fg = pl.read_parquet("tests/fixtures/nfl_scheme/fg_attempts_2019_2023.parquet")
out = env_adjusted_make_prob(fg)
print(out.select("base_make_prob", "exp_make_prob").describe())

espn_nfl_teams​

espn_nfl_teams(return_as_pandas=False, **kwargs) -> 'pl.DataFrame'

espn_nfl_teams - look up NFL teams

Parameters

ParameterTypeDefaultDescription
return_as_pandasboolFalseIf True, returns a pandas dataframe. If False, returns a polars dataframe.

Returns

Polars dataframe containing teams for the requested league. This function caches by default, so if you want to refresh the data, use the command sportsdataverse.nfl.espn_nfl_teams.clear_cache().

col_nametypedescription
team_abbreviationcharacterTeam abbreviation; team_detail = TRUE only.
team_alternate_colorcharacterAlternate team color; team_detail = TRUE only.
team_colorcharacterPrimary team color; team_detail = TRUE only.
team_display_namecharacterFull team display name; team_detail = TRUE only.
team_idcharacterESPN team id.
team_is_activelogicalTRUE if the team is currently active.
team_is_all_starlogicalTRUE if the row represents an All-Star team.
team_locationcharacterTeam location / school name; team_detail = TRUE only.
team_logosintegerTeam logo metadata.
team_namecharacterTeam nickname; team_detail = TRUE only.
team_nicknamecharacterTeam nickname label; team_detail = TRUE only.
team_short_display_namecharacterShort team display name; team_detail = TRUE only.
team_slugcharacterTeam slug for the stat row.
team_uidcharacterESPN universal team identifier (UID format 's:40~l:...~t:...').

Example

from sportsdataverse.nfl import espn_nfl_teams
teams = espn_nfl_teams()
teams.shape

# Pandas round-trip

teams_pd = espn_nfl_teams(return_as_pandas=True)
teams_pd[["team_abbreviation", "team_display_name"]].head()

# Force a refresh after upstream ESPN updates

espn_nfl_teams.cache_clear() # underlying lru_cache
teams = espn_nfl_teams()

fg_make_probability​

fg_make_probability(yardline_100: 'np.ndarray', fg_roof: 'np.ndarray', era: 'np.ndarray') -> 'Optional[np.ndarray]'

Predict FG make probability from the bundled fg_model (public wrapper).

Thin supported alias over the private underscore-prefixed helper so downstream consumers (e.g. the kicker-rating spine) reuse the shipped model through a public import instead of a private reach.

Parameters

ParameterTypeDefaultDescription
yardline_100ndarrayKick spot (yards from the opponent end zone); the attempt distance is yardline_100 + 18.
fg_roofndarray1.0 when roof == "outdoors" else 0.0, per kick.
erandarray(n, 5) one-hot era matrix (era0..era4, season cuts 2001/2005/2013/2017).

Returns

Make probabilities (with nfl4th's long-kick clamps), or None when the bundled model is unavailable.

Example

import numpy as np
from sportsdataverse.nfl.nfl_fourth_down import fg_make_probability
p = fg_make_probability(
np.array([30.0]), np.array([1.0]),
np.array([[0.0, 0.0, 0.0, 0.0, 1.0]]),
)
print(p)

fit_nfl_field_position_ep​

fit_nfl_field_position_ep(pbp: 'pl.DataFrame', *, exclude_garbage: 'bool' = True) -> 'pl.DataFrame'

Fit the NFL EP-by-starting-yardline curve from released espn_nfl_pbp plays.

Extracts one row per drive (starting yard line from the offense's own goal, realized drive points) and fits the monotone curve with sportsdataverse.cfb.cfb_field_position.fit_field_position_ep -- the same estimator and target the college curve uses. This is how the bundled artifact was produced; re-run it on newer seasons to refresh it.

Parameters

ParameterTypeDefaultDescription
pbpDataFrameplays in the released espn_nfl_pbp shape (any number of seasons concatenated). Needs the drive fields (drive.id, drive.result, drive.start.yardLine), homeTeamId, period and start.pos_team.id / start.def_pos_team.id.
exclude_garbageboolTruedrop drives that start in garbage time.

Returns

yardline_own: Int64 (1..99), ep: Float64 -- monotone non-decreasing. Empty input returns a zero-row frame.

Example

import polars as pl
from sportsdataverse.nfl import fit_nfl_field_position_ep
pbp = pl.concat([pl.read_parquet(f) for f in files], how="diagonal_relaxed")
curve = fit_nfl_field_position_ep(pbp)
curve.write_parquet("nfl_field_position_ep.parquet")

get_2pt_probs​

get_2pt_probs(pbp_df: "Union[pl.DataFrame, 'pd.DataFrame']") -> 'pd.DataFrame'

The PAT-vs-2pt decision surface for post-touchdown states (CFB-shaped).

The NFL twin of sportsdataverse.cfb.cfb_two_point.get_2pt_probs. It runs the same three-outcome enumeration get_2pt_wp uses, but returns the decision columns instead of folding them into wp_td. The two option

Parameters

ParameterTypeDefaultDescription
pbp_dfUnion[DataFrame, 'DataFrame']Post-touchdown states in nflverse column space (the same inputs get_4th_down_probs takes; score_differential is the scoring team's lead after the six points). Prepared frames are accepted as-is.

Returns

A pandas frame with go_index plus two_pt_wp, xp_wp, prob_2pt, two_pt_recommendation ("go_for_2" iff two_pt_wp > xp_wp else "kick_xp") and two_pt_wp_diff (two_pt_wp - xp_wp). All NaN / null when the models are unavailable.

Example

from sportsdataverse.nfl.nfl_fourth_down import get_2pt_probs
out = get_2pt_probs(touchdown_states)
print(out[["two_pt_wp", "xp_wp", "two_pt_recommendation"]].head())

get_2pt_wp​

get_2pt_wp(pbp_df: "Union[pl.DataFrame, 'pd.DataFrame']") -> 'pd.DataFrame'

Win probability of the PAT-vs-2pt choice after a touchdown (nfl4th get_2pt_wp).

For each row, scores the post-touchdown state under three scoring outcomes (0 / 1 / 2 added points) from the kicking-off team's ensuing-drive WP, and combines them with the 2-pt conversion probability (two_pt_model) and the PAT make probability (the FG model at yardline_100 = 15) into wp_td — the better of go-for-2 and kick-the-PAT.

Parameters

ParameterTypeDefaultDescription
pbp_dfUnion[DataFrame, 'DataFrame']Play-by-play frame (polars or pandas) of post-touchdown states, already carrying the prepared state columns (see module docstring).

Returns

A pandas frame with go_index, yardline_100 (always 0) and wp_td. wp_td is NaN when the WP / 2-pt models are unavailable.

Example

from sportsdataverse.nfl.nfl_fourth_down import get_2pt_wp
out = get_2pt_wp(touchdown_states)
print(out[["go_index", "wp_td"]].head())

get_4th_down_probs​

get_4th_down_probs(pbp_df: "Union[pl.DataFrame, 'pd.DataFrame']") -> 'pd.DataFrame'

Full 4th-down decision surface (nfl4th add_4th_probs) + recommendation.

Runs get_go_wp, get_fg_wp, get_punt_wp on the fourth-down rows and adds the combined option columns plus:

  • go_boost -- nfl4th's headline number: 100 * (go_wp - max(fg_wp, punt_wp)) in percentage points (a NaN punt_wp is treated as 0).
  • fourth_down_recommendation -- the max-WP choice among {go, punt, field_goal} (NaN options are excluded).
  • go_wp_diff / punt_wp_diff / fg_wp_diff -- each option's WP minus the recommended option's WP (the recommended option's diff is 0, the others <= 0). NaN where the option WP is NaN.

Parameters

ParameterTypeDefaultDescription
pbp_dfUnion[DataFrame, 'DataFrame']Play-by-play frame (polars or pandas) of fourth-down situations (the nflverse-shape output of load_nfl_pbp; see module docstring for required columns).

Returns

A pandas copy of pbp_df with the decision columns added. Empty input returns the input plus empty decision columns.

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.nfl_fourth_down import get_4th_down_probs

pbp = load_nfl_pbp([2023])
fourth = pbp.filter((pl.col("down") == 4) & pl.col("yardline_100").is_not_null())
out = get_4th_down_probs(fourth)
print(out[["go_wp", "punt_wp", "fg_wp", "go_boost", "fourth_down_recommendation"]].head())

get_config​

get_config() -> 'NflConfig'

Return the live NflConfig singleton.

The same object is returned on every call; mutate via update_config rather than reassigning fields directly so future hooks (e.g. logging on config change) have a single choke point.

Example

from sportsdataverse.nfl import get_config
cfg = get_config()
print(cfg.cache_mode, cfg.cache_duration, cfg.cache_dir)

# Pair with ``update_config`` to verify a change took effect

from sportsdataverse.nfl import update_config, get_config
update_config(cache_mode="off")
assert get_config().cache_mode == "off"

get_fg_wp​

get_fg_wp(pbp_df: "Union[pl.DataFrame, 'pd.DataFrame']") -> 'pd.DataFrame'

Expected win probability of attempting a field goal (nfl4th get_fg_wp).

The make probability comes from the self-trained fg_model (a binary:logistic XGBoost re-train of the original mgcv GAM, features [yardline_100, fg_roof, fg_era]), shrunk by 0.9 for kicks at/beyond yardline_100 = 38 and zeroed at/beyond yardline_100 = 45 (>= ~63-yard kicks). The made-FG state (opponent receives a touchback kickoff at the 25, kicking team +3) and the missed-FG state (opponent takes over 8 yards back of the spot, capped at the 80) are each scored with win probability; fg_wp = make_prob * make_wp + (1 - make_prob) * miss_wp.

Parameters

ParameterTypeDefaultDescription
pbp_dfUnion[DataFrame, 'DataFrame']Play-by-play frame (polars or pandas) of fourth-down situations.

Returns

A pandas copy of pbp_df plus fg_make_prob, make_fg_wp, miss_fg_wp and fg_wp (from the kicking team's perspective). All four are NaN when the FG model or WP model is unavailable.

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.nfl_fourth_down import get_fg_wp

pbp = load_nfl_pbp([2023])
fourth = pbp.filter((pl.col("down") == 4) & pl.col("yardline_100").is_not_null())
out = get_fg_wp(fourth)
print(out[["fg_make_prob", "fg_wp"]].head())

get_go_wp​

get_go_wp(pbp_df: "Union[pl.DataFrame, 'pd.DataFrame']") -> 'pd.DataFrame'

Expected win probability of going for it on 4th down (nfl4th get_go_wp).

The fd_model 76-class yards-gained distribution is expanded per play; each outcome's hypothetical post-play game state (turnover-on-downs flip, +6 touchdown with the PAT/2-pt branch routed through get_2pt_wp, 6-second runoff, goal-to-go distance shrink) is scored with win probability and the end-of-game kneel-out clamps are applied; the option value is the prob-weighted WP.

Parameters

ParameterTypeDefaultDescription
pbp_dfUnion[DataFrame, 'DataFrame']Play-by-play frame (polars or pandas) of fourth-down situations carrying the prepared state columns (see module docstring). The frame is prepared internally if it lacks the derived columns.

Returns

A pandas copy of pbp_df plus go_wp (prob-weighted WP of going for it), first_down_prob (P(conversion)), wp_succeed (mean WP over conversion outcomes) and wp_fail (mean WP over failure outcomes). All are NaN when the fourth-down / WP models are unavailable (FD_MODEL_AVAILABLE / WP_MODEL_AVAILABLE).

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.nfl_fourth_down import get_go_wp

pbp = load_nfl_pbp([2023])
fourth = pbp.filter((pl.col("down") == 4) & pl.col("yardline_100").is_not_null())
out = get_go_wp(fourth)
print(out[["go_wp", "first_down_prob"]].head())

get_punt_wp​

get_punt_wp(pbp_df: "Union[pl.DataFrame, 'pd.DataFrame']") -> 'pd.DataFrame'

Expected win probability of punting on 4th down (nfl4th get_punt_wp).

The punt landing distribution (punt_data: yardline_after / pct / muff per yardline_100) is joined per play; possession is flipped to the receiving team, with return-touchdown (yardline_after == 100) and muff (muff == 1) recoveries flipping the ball back to the punting team; each landing spot's ensuing-drive WP is scored and the option value is the prob-weighted WP from the punting team's perspective.

Parameters

ParameterTypeDefaultDescription
pbp_dfUnion[DataFrame, 'DataFrame']Play-by-play frame (polars or pandas) of fourth-down situations.

Returns

A pandas copy of pbp_df plus punt_wp. punt_wp is NaN where the punt distribution has no support for the play's yardline_100 (inside the punting team's own 31, where the table is empty — matching the R reference's left-join NA behavior) or when the WP model is unavailable.

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.nfl_fourth_down import get_punt_wp

pbp = load_nfl_pbp([2023])
fourth = pbp.filter((pl.col("down") == 4) & pl.col("yardline_100").is_not_null())
out = get_punt_wp(fourth)
print(out[["punt_wp"]].head())

opponent_adjusted_ridge​

opponent_adjusted_ridge(plays: 'pl.DataFrame', *, off_col: 'str', def_col: 'str', home_col: 'str', resp_col: 'str', lam: 'float', penalize_home: 'bool' = False, hfa_col: 'str | None' = None) -> 'tuple[pl.DataFrame, float, float]'

Ridge-regress resp_col on offense + defense team indicators + HFA.

League-agnostic (column names are arguments): builds the full offense/defense-indicator + intercept + home design and solves the ridge normal equations beta = (X'X + lam*R)^-1 X'y. Only team coefficients are penalised; the intercept (and, unless penalize_home, the home term) is free. Moved (T7.2) from sportsdataverse.nfl.nfl_ratings. Callers: NFL ratings, and CFB adjusted EPA (cfb_adjusted_epa._fit_team_strengths, via hfa_col); cfb_ratings uses the different dropped_level_ridge.

Parameters

ParameterTypeDefaultDescription
playsDataFrameOne row per play. Rows with a null off_col / def_col / resp_col must be filtered by the caller.
off_colstrColumn naming the offense (possession) team.
def_colstrColumn naming the defense team.
home_colstrColumn naming the home team (HFA indicator is off_col == home_col).
resp_colstrNumeric response column (e.g. epa).
lamfloatRidge penalty applied to the team coefficients.
penalize_homeboolFalseAlso penalise the home-field coefficient (default False).
hfa_colstr | NoneNoneNumeric column used as the home regressor as-is (e.g. CFB's +1 home offense / 0 neutral site / -1 away), in place of the off_col == home_col indicator, which then goes unused. Must not contain nulls (raises ValueError).

Returns

A (frame, intercept, home_coef) tuple: frame has one row per team (team_id Utf8, off_coef / def_coef Float64); intercept is the league baseline; home_coef the fitted HFA in response units. Zero-row frame + (0.0, 0.0) on empty input.

Example

from sportsdataverse.nfl.nfl_ratings import opponent_adjusted_ridge
frame, intercept, hfa = opponent_adjusted_ridge(
plays, off_col="posteam", def_col="defteam",
home_col="home_team", resp_col="epa", lam=200.0,
)
frame.sort("off_coef", descending=True).head()

playcall_features​

playcall_features(pbp: 'pl.DataFrame', participation: 'Optional[pl.DataFrame]' = None) -> 'pl.DataFrame'

Build the play-call feature frame (one row per offensive run/pass play).

Filters to plays with pass == 1 or rush == 1, derives the 5-class family label (scramble > deep/short pass > inside/outside run), and left-joins the optional participation frame for personnel counts.

Parameters

ParameterTypeDefaultDescription
pbpDataFramenflverse-format pbp with the pre-snap feature columns + pass / rush / qb_scramble / pass_length / run_location / run_gap and xpass.
participationOptional[DataFrame]NoneOptional nflverse participation frame with game_id / play_id / offense_personnel.

Returns

Keys + PLAYCALL_FEATURE_ORDER columns + family + is_pass. Personnel columns are null (has_participation=0) when no participation row matches.

col_nametypedescription
game_idcharacternflverse game identifier (Utf8 join key).
play_idintegernflverse play identifier within the game (Int64 join key).
seasonintegerSeason of the play.
weekintegerWeek of the play.
posteamcharacterOffense (possession) team abbreviation.
downdoubleDown (1-4) at the snap.
ydstogodoubleYards to go for a first down.
yardline_100doubleYards from the opponent end zone at the snap.
score_differentialdoubleOffense score minus defense score at the snap.
half_seconds_remainingdoubleSeconds remaining in the half.
game_seconds_remainingdoubleSeconds remaining in the game.
wpdoubleStart-of-play win probability for the offense.
shotgundouble1 when the offense lined up in shotgun.
no_huddledouble1 when the play was run without a huddle.
xpassdoubleShipped nflfastR-parity expected-dropback probability for the play.
n_rbdoubleRunning backs in the offensive personnel grouping (null without participation data).
n_tedoubleTight ends in the offensive personnel grouping (null without participation data).
n_wrdoubleWide receivers in the offensive personnel grouping (null without participation data).
has_participationinteger1 when a participation row matched the play, else 0.
familycharacter5-class play-call label (inside_run, outside_run, short_pass, deep_pass, scramble).
is_passinteger1 when the play was a pass (including scrambles), else 0.

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.ep_wp import calculate_xpass
from sportsdataverse.nfl.nfl_playcall import playcall_features
feat = playcall_features(calculate_xpass(load_nfl_pbp([2023])))
print(feat["family"].value_counts())

player_usage_efficiency​

player_usage_efficiency(player_stats: 'pl.DataFrame', *, as_of_week: 'int', era: 'str' = 'modern') -> 'pl.DataFrame'

Per-player as-of usage + efficiency with empirical-Bayes shrinkage.

Aggregates one season of week-level player stats over weeks strictly before as_of_week (the leakage boundary), then shrinks every usage (per-game attempts / carries / targets) and efficiency (yards + TDs per opportunity) stat toward its position prior: (n * player_value + kappa * prior) / (n + kappa) with n = games played and kappa the stat family's fitted shrinkage.

Parameters

ParameterTypeDefaultDescription
player_statsDataFrameOne season of load_nfl_player_stats() rows (columns player_id, position, recent_team, week, attempts, passing_yards, passing_tds, carries, rushing_yards, rushing_tds, targets, receiving_yards, receiving_tds).
as_of_weekintOnly weeks < as_of_week are used.
erastr'modern'Constants era key (supplies kappas + position priors).

Returns

One row per player_id (Utf8) whose position has a prior table: position / team_id (Utf8, latest team), games (Int64), exp_attempts / exp_carries / exp_targets (Float64, shrunk per-game usage), ypa / ypc / ypt / pass_td_rate / rush_td_rate / rec_td_rate (Float64, shrunk per-opportunity efficiency). Zero-row, correctly-typed on empty input.

Example

import polars as pl
import sportsdataverse.nfl as nfl
stats = nfl.load_nfl_player_stats().filter(pl.col("season") == 2023)
usage = nfl.player_usage_efficiency(stats, as_of_week=10)
usage.sort("exp_attempts", descending=True).head()

predict_margin​

predict_margin(home_adj_net: 'float', away_adj_net: 'float', neutral: 'bool', *, era: 'str' = 'modern') -> 'float'

Expected home scoring margin from two net ratings.

points_per_net * (home_adj_net - away_adj_net) plus the era HFA on non-neutral fields.

Parameters

ParameterTypeDefaultDescription
home_adj_netfloatHome team's adj_net (EPA/play units).
away_adj_netfloatAway team's adj_net.
neutralboolTrue drops the home-field advantage.
erastr'modern'Constants era key (default "modern").

Returns

Expected home margin in points (positive = home favored).

Example

from sportsdataverse.nfl.nfl_market import predict_margin
predict_margin(0.10, -0.05, False)

predict_total​

predict_total(home_adj_off: 'float', home_adj_def: 'float', away_adj_off: 'float', away_adj_def: 'float', *, era: 'str' = 'modern') -> 'float'

Expected combined point total from the four efficiency components.

avg_total + total_scale * (home_adj_off + away_adj_def + away_adj_off + home_adj_def). The four ratings are summed because each side's scoring rises with its own offense and with the opponent's EPA-allowed (adj_def is lower = better defense) -- same semantics as the shipped CFB analog. (The plan text wrote this with a minus; that sign flips a good defense into raising the total, so the analog's sum is used.)

Parameters

ParameterTypeDefaultDescription
home_adj_offfloatHome adj_off_epa.
home_adj_deffloatHome adj_def_epa (lower = better defense).
away_adj_offfloatAway adj_off_epa.
away_adj_deffloatAway adj_def_epa.
erastr'modern'Constants era key.

Returns

Expected combined total in points.

Example

from sportsdataverse.nfl.nfl_market import predict_total
predict_total(0.10, -0.02, 0.05, 0.01)

pressure_pairs​

pressure_pairs(pbp: 'pl.DataFrame') -> 'pl.DataFrame'

Per (season, off_team, def_team) dropbacks + pressures (matchup grid).

Parameters

ParameterTypeDefaultDescription
pbpDataFrame

Returns

col_nametypedescription
seasonintegerSeason of the matchup aggregate.
off_teamcharacterOffense team abbreviation.
def_teamcharacterDefense team abbreviation.
dropbacksintegerOffense dropbacks in the matchup.
pressuresintegerSacks plus QB hits in the matchup.

reset_config​

reset_config() -> 'NflConfig'

Reset the active config to its env-var-derived defaults.

Convenience for tests / interactive sessions that want to undo a chain of update_config() calls without restarting the interpreter.

Example

from sportsdataverse.nfl import update_config, reset_config
update_config(cache_mode="off", timeout=5)
# ... do work ...
reset_config() # back to env-derived defaults

scoreboard_event_parsing​

scoreboard_event_parsing(event)

Normalize one ESPN scoreboard event into a flatter shape.

Splits the competitors list into home / away siblings, hoists notes / broadcast metadata onto the competition root, and drops the fields the schedule helper does not need (odds, leaders, geoBroadcasts, etc.). Used internally by espn_nfl_schedule.

Parameters

ParameterTypeDefaultDescription
eventDictA single events[i] dict from the ESPN scoreboard endpoint.

Returns

The mutated event dict with normalized home / away / broadcast keys.

Example

from sportsdataverse.dl_utils import download
from sportsdataverse.nfl.nfl_schedule import scoreboard_event_parsing
url = "http://site.api.espn.com/apis/site/v2/sports/football/nfl/scoreboard"
payload = download(url=url).json()
for ev in payload.get("events", []):
scoreboard_event_parsing(ev)
ev["competitions"][0]["home"]["abbreviation"]

shield_nfl_pbp​

shield_nfl_pbp(game_detail: 'Optional[Dict[str, Any]]' = None, shield_game_id: 'Optional[str]' = None, *, enrich: 'bool' = True, context: 'Optional[Dict[str, Any]]' = None, game_id: 'Optional[str]' = None) -> 'pl.DataFrame'

Build one NFL game's nflverse-shape play-by-play from Shield, at ANY game phase.

The live entry point: the same parser build_pbp runs on the archived nfl/raw finals, plus the four things a game still being played needs — the in-progress drive's possession, game-outcome columns held null until the feed says FINAL, a next-snap row from summary, and provisional rows flagged (see sportsdataverse.nfl.shield_pbp.live). Safe to poll: pass the payload you already have via game_detail (no network), or a shield_game_id to fetch it.

Parameters

ParameterTypeDefaultDescription
game_detailOptional[Dict[str, Any]]NoneA Shield experience/v2/gamedetails payload (the raw body, or a {"data": ...} envelope). Takes precedence over shield_game_id, so tests and pollers that already hold a payload never touch the network.
shield_game_idOptional[str]NoneShield game uuid, fetched via sportsdataverse.nfl.nfl_game_details_v2 with include_drive_chart=True, return_parsed=False when game_detail is None.
enrichboolTrueRun sportsdataverse.nfl.ep_wp.enrich_nfl_pbp on the result (default True) for the nfl_model_pbp EP/EPA/WP/WPA/CP/CPOE columns. Pass False for the base frame only (no model loads).
contextOptional[Dict[str, Any]]NoneGame context {"roof": ..., "spread_line": ..., "total_line": ...} the Shield feed omits. Unset fields fall back to the nflverse schedule row for this game, then to live.DEFAULT_CONTEXT (outdoors / 2.5 / 55.5, the same default the ESPN processor uses).
game_idOptional[str]NoneOverride the nflverse game_id (computed from the payload when None).

Returns

A polars DataFrame, one row per play (plus, while summary.phase is INGAME, one current-situation row), carrying the nfl_model_pbp columns — the build_pbp base frame, the EP/WP enrichment when enrich is True, and: | col_name | type | description | |----------|------|-------------| | live_phase | str | The payload's summary.phase: PREGAME, INGAME, HALFTIME, FINAL or FINAL_OVERTIME. | | is_play | int | 1 for a real play; 0 for the feed's GAME_START / END_QUARTER / END_GAME markers and the current-situation row. | | provisional | int | 1 when the feed has not closed the play (playEndTime null) and it is in the trailing run of such plays of a non-final game — its text, yardage and stats may still change. Always 0 on a final game. | home_score / away_score / result are null until the game is final. The current-situation row is not inert once enrich is True: it is the next state, so it also completes the previous play's lead-diff columns (epa, qb_epa, wpa, vegas_wpa, the total_* running sums). That play is usually still provisional, so those values can move on the next poll. A payload Shield has not populated a drive chart for (every scheduled game before kickoff) returns a zero-row frame carrying only the three live columns — check df.is_empty() before selecting anything else.

Example

import polars as pl
from sportsdataverse.nfl import shield_nfl_pbp

df = shield_nfl_pbp(shield_game_id="a9a8944e-4feb-11f1-abca-2c54536568a9")
df.filter(pl.col("is_play") == 0).select("posteam", "down", "ydstogo", "wp")

shield_to_espn_summary​

shield_to_espn_summary(game_detail: 'Mapping[str, Any]', idmap_row: 'Mapping[str, Any]', *, parsed: 'Optional[pl.DataFrame]' = None, odds: 'Optional[Mapping[str, Any]]' = None, player_stats: 'Optional[Mapping[str, Any]]' = None, team_stats: 'Optional[Mapping[str, Any]]' = None) -> 'Tuple[Dict[str, Any], List[str]]'

Project one Shield game (any phase) onto an ESPN-summary-shaped dict.

Parameters

ParameterTypeDefaultDescription
game_detailMapping[str, Any]A Shield experience/v2/gamedetails payload (raw body or a {"data": ...} envelope) -- the same object sportsdataverse.nfl.shield_pbp.build.shield_nfl_pbp consumes.
idmap_rowMapping[str, Any]The game's pre-kickoff id-map row (sportsdataverse.football.sources.idmap.GAME_SCHEMA): espn_event_id, home_espn_team_id and away_espn_team_id are required; the optional home_team / away_team sub-dicts supply the era-correct espn_abbr.
parsedOptional[DataFrame]NoneThe frame shield_nfl_pbp(game_detail, enrich=False) already produced. Built here when None -- pass it to parse the payload once for both projections.
oddsOptional[Mapping[str, Any]]None{gameSpread, overUnder, homeFavorite, gameSpreadAvailable} (the stored closing line, sportsdataverse.football.sources.idmap._odds_override_from_row). Becomes the summary's one-provider pickcenter.
player_statsOptional[Mapping[str, Any]]NoneA Shield /football/v2/stats/live/player-statistics/{gameId} body. Becomes boxscore.players in ESPN's exact shape (ten categories, athletes carrying ESPN ids from the players crosswalk). Omitted -> the box stays empty and no ESPN athlete id is attached to any play.
team_statsOptional[Mapping[str, Any]]NoneA Shield /football/v2/stats/live/team-statistics/{gameId} body. Becomes boxscore.teams -- the authoritative countable team totals NFLPlayProcess.create_box_score prefers over its play-by-play derivation.

Returns

(summary, notes). | item | type | description | |---|---|---| | summary | dict | An ESPN-summary-shaped payload: header (season/week/competitions/competitors/status), drives.previous (+ drives.current while the game is live), gameInfo, pickcenter, boxscore (filled when player_stats/team_stats are given) and passthrough arrays. Feed it to espn_nfl_pbp(summary=). | | notes | list[str] | Adapter-side degradations worth surfacing in provenance: a missing summary.timeouts block, a missing summary.homeTeam/awayTeam team id, a PAT with no touchdown to fold into, plays outside the drive chart, and (pre-2014) play ids that do not join ESPN's own. |

Example

import json
from sportsdataverse.nfl import NFLPlayProcess, shield_to_espn_summary

# any Shield gamedetails body -- here the copy nfl-raw keeps
with open("nfl/raw/2025/2025_07_LA_JAX.json") as fh:
game = json.load(fh)
row = {"espn_event_id": "401772635", "home_espn_team_id": "30", "away_espn_team_id": "14"}
summary, notes = shield_to_espn_summary(game, row)
proc = NFLPlayProcess(gameId=401772635, join_participants=False)
proc.espn_nfl_pbp(summary=summary)
result = proc.run_processing_pipeline()

special_teams_ratings​

special_teams_ratings(plays: 'pl.DataFrame', *, config: 'RatingsConfig | None' = None) -> 'pl.DataFrame'

One row per team: opponent-adjusted special-teams EPA per play.

Reuses opponent_adjusted_ridge (no forked solver) restricted to special == 1 plays with resp_col="epa"; adj_st_epa is the off_coef (the special-teams unit acting as "offense" on the play). Teams appearing anywhere in plays but on no special-teams play get the documented neutral fill adj_st_epa = 0.0.

Parameters

ParameterTypeDefaultDescription
playsDataFrameAn load_nfl_pbp-schema frame carrying posteam, defteam, home_team, epa, special. Not pre-filtered -- this function selects the ST plays itself.
configRatingsConfig | NoneNoneTuning knobs (only ridge_lambda is consulted); defaults to RatingsConfig.

Returns

One row per team_id (Utf8) with adj_st_epa (Float64). Zero-row, correctly-typed when plays is empty.

Example

from sportsdataverse.nfl.nfl_ratings import special_teams_ratings
st = special_teams_ratings(pbp)
st.sort("adj_st_epa", descending=True).head()

team_game_pace​

team_game_pace(pbp: 'pl.DataFrame') -> 'pl.DataFrame'

Per team-game pace + pass-rate-over-expected.

sec_per_play is the per-drive elapsed game_seconds_remaining divided by drive plays, averaged over the team's offensive drives (kneels / spikes / no_plays excluded). Neutral = wp in [0.2, 0.8] and half_seconds_remaining > 120. proe is the mean pass_oe over dropbacks.

Parameters

ParameterTypeDefaultDescription
pbpDataFramenflverse-format pbp with game_id / season / week / posteam / drive / play_type / qb_dropback / pass_oe / game_seconds_remaining / wp / half_seconds_remaining.

Returns

One row per (game_id, season, week, posteam) with off_plays, sec_per_play, neutral_plays, neutral_sec_per_play, proe. Empty input yields a zero-row frame with this schema.

col_nametypedescription
game_idcharacternflverse game identifier (Utf8 join key).
seasonintegerSeason of the game.
weekintegerWeek of the game.
posteamcharacterOffense team abbreviation.
off_playsintegerOffensive plays in the game (kneels, spikes and no_plays excluded).
sec_per_playdoubleMean over the team's drives of elapsed game clock divided by drive plays.
neutral_playsintegerOffensive plays in neutral situations (wp in [0.2, 0.8], over 2 minutes left in the half).
neutral_sec_per_playdoublesec_per_play computed on neutral-situation plays only.
proedoubleMean pass_oe over the team's dropbacks in the game (percentage points).
dropbacksintegerDropbacks with a non-null pass_oe (the proe denominator).

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.nfl_gamescript import team_game_pace
pace = team_game_pace(load_nfl_pbp([2023]))
print(pace.sort("sec_per_play").head())

team_name_fn​

team_name_fn(expr: 'pl.Expr') -> 'pl.Expr'

Fold historical/relocated team codes onto their current abbreviation.

Verbatim port of nflfastR's team_name_fn (a plain stringr::str_replace_all over a 10-entry named vector). Operates as a substring replace (not a full-value lookup) so it also fixes embedded codes like "SD 49" -> "LAC 49" on yard-line columns. The 10 from-codes are disjoint from all of their to-values, so the order of the 10 sequential replacements does not matter (verified in tests.nfl.test_nfl_clean).

Parameters

ParameterTypeDefaultDescription
exprExprA polars.Expr over a Utf8 column (e.g. pl.col("posteam")).

Returns

The same expression with every occurrence of the 10 historical codes replaced by their current-franchise code.

team_pressure_rates​

team_pressure_rates(pbp: 'pl.DataFrame') -> 'pl.DataFrame'

Per (season, team) raw pressure rates, both sides of the ball.

Parameters

ParameterTypeDefaultDescription
pbpDataFramenflverse-format pbp with season / posteam / defteam / qb_dropback / sack / qb_hit.

Returns

Per (season, team): dropbacks_off, pressures_allowed, pressure_rate_allowed, dropbacks_def, pressures_generated, pressure_rate_generated. Empty input yields a zero-row frame.

col_nametypedescription
seasonintegerSeason of the aggregate.
teamcharacterTeam abbreviation.
dropbacks_offintegerOffensive dropbacks (qb_dropback plays).
pressures_allowedintegerSacks plus QB hits allowed on the team's own dropbacks.
pressure_rate_alloweddoublepressures_allowed / dropbacks_off.
dropbacks_defintegerOpponent dropbacks faced on defense.
pressures_generatedintegerSacks plus QB hits generated against opponent dropbacks.
pressure_rate_generateddoublepressures_generated / dropbacks_def.

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.nfl_line_grades import team_pressure_rates
rates = team_pressure_rates(load_nfl_pbp([2023]))
print(rates.sort("pressure_rate_generated", descending=True).head())

update_config​

update_config(**kwargs: 'object') -> 'NflConfig'

Update the active config in place.

Returns

The (mutated) global config object, for chaining or inspection.

Example

from sportsdataverse.nfl import update_config
update_config(cache_mode="filesystem", cache_duration=3600)

# Disable caching for development

update_config(cache_mode="off")

# Point cache at a custom directory

update_config(cache_dir="~/sdv-cache")

win_prob_from_margin​

win_prob_from_margin(exp_margin: 'float', *, era: 'str' = 'modern') -> 'float'

Home win probability from an expected margin (Gaussian margin model).

Parameters

ParameterTypeDefaultDescription
exp_marginfloatExpected home margin in points.
erastr'modern'Constants era key (supplies margin_sd).

Returns

Phi(exp_margin / margin_sd) in [0, 1].

Example

from sportsdataverse.nfl.nfl_market import win_prob_from_margin
win_prob_from_margin(3.0)