Skip to main content
Version: 0.1.5

NFL — additional Python functions — Analytics

calculate_nfl_standings​

calculate_nfl_standings(games: 'pl.DataFrame', *, teams: 'pl.DataFrame | None' = None, tiebreaker_depth: 'int' = 3, playoff_seeds: 'int | None' = None, return_as_pandas: 'bool' = False) -> "pl.DataFrame | 'pd.DataFrame'"

Compute NFL division standings + conference playoff seeds.

A reduced port of the tiebreaker ladder nflfastR delegates to the external nflseedR package (see the module docstring for the exact scope). Games are doubled into one row per team per game, regular-season win/loss/tie records are computed per team, and ties are broken win_pct -> head-to-head -> division record -> conference record, to the depth configured by tiebreaker_depth.

Parameters

ParameterTypeDefaultDescription
gamesDataFrameA load_nfl_schedule-shaped frame: game_id, season, game_type, week, home_team, away_team, home_score, away_score. Only game_type == "REG" rows with both scores present are used.
teamsDataFrame | NoneNoneA load_nfl_teams-shaped frame (team_abbr, team_conf, team_division). When None (default), calls sportsdataverse.nfl.load_nfl_teams. Must cover every team abbreviation appearing in games -- a team absent from teams gets null conf/division and is silently pooled into the (season, None) division/conference group rather than raising.
tiebreaker_depthint31 (win_pct only), 2 (adds head-to-head + division record), or 3 (default; adds conference record too).
playoff_seedsint | NoneNoneNumber of teams per conference that receive a non-null seed. When None (default), uses the 2020 playoff -format cutover: 6 for seasons <= 2019, 7 for 2020+.
return_as_pandasboolFalseIf True return a pandas DataFrame; else polars.

Returns

A polars (or pandas) DataFrame with one row per (season, team): conf, division, div_rank, seed (null past playoff_seeds), team, games, wins, losses, ties, win_pct (ties count as 0.5 win), div_pct, conf_pct. Sorted by (season, division, div_rank, seed).

col_nametypedescription
seasoninteger4 digit number indicating to which season(s) the specified timeframe belongs to.
confcharacter
divisioncharacter
div_rankinteger
seedinteger
teamcharacterNFL team. Uses official abbreviations as per NFL.com
gamesintegerGames played in career
winsinteger
lossesinteger
tiesinteger
win_pctdouble
div_pctdouble
conf_pctdouble

Example

from sportsdataverse.nfl import calculate_nfl_standings, load_nfl_schedule
games = load_nfl_schedule(seasons=[2023])
standings = calculate_nfl_standings(games)
standings.filter(standings["div_rank"] == 1)

# Injected teams frame (offline)

standings = calculate_nfl_standings(games, teams=my_teams_df)

# Pipeline next step (one line)

standings.sort(["conf", "seed"]).select("team", "seed", "win_pct")

compose_counting_projection​

compose_counting_projection(rate_proj: 'pl.DataFrame', avail_proj: 'pl.DataFrame', *, rate_col: 'str' = 'proj_rate', volume_col: 'str' = 'proj_volume') -> 'pl.DataFrame'

Compose skill and availability into a counting projection.

The only place skill (rate x volume) and availability meet: proj_counting = rate * volume * proj_availability, joined on player_id (dtype-asserted).

Parameters

ParameterTypeDefaultDescription
rate_projpl.DataFrameSkill projection carrying player_id + rate_col + volume_col.
avail_projpl.DataFrameAvailability projection carrying player_id + proj_availability.
rate_colstr'proj_rate'Rate column name in rate_proj.
volume_colstr'proj_volume'Volume column name in rate_proj.

Returns

rate_proj columns plus proj_availability and proj_counting:Float64.

col_nametypedescription
player_idcharacternflverse gsis player id (character join key; asserted Utf8 on both sides of the join).
proj_ratedoubleProjected per-opportunity rate carried through from the skill projection (rate_col).
proj_volumedoubleProjected opportunity volume carried through from the skill projection (volume_col).
proj_availabilitydoubleProjected availability rate in [0, 1] from nfl_availability_projection.
proj_countingdoubleComposed counting projection - proj_rate x proj_volume x proj_availability (the only place skill and availability meet).

Example

import polars as pl
from sportsdataverse.nfl.nfl_availability import compose_counting_projection
out = compose_counting_projection(rate_frame, avail_frame)

nfl_availability_projection​

nfl_availability_projection(seasons: 'List[int]', target_season: 'int', *, team_games: 'int' = 17, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Empirical-Bayes availability projection: expected fraction of team games.

Shrinks each player's historical availability toward the fitted position

Parameters

ParameterTypeDefaultDescription
seasonsList[int]History seasons to load snap counts/rosters for.
target_seasonintThe season being projected.
team_gamesint17Regular-season team games.
return_as_pandasboolFalseIf True, returns a pandas dataframe.

Returns

player_id:Utf8, target_season:Int64, position:Utf8, proj_availability:Float64, proj_games:Float64, proj_games_missed:Float64. Empty history returns a zero-row frame.

col_nametypedescription
player_idcharacternflverse gsis player id (character join key).
target_seasonintegerThe season being projected (features use strictly earlier seasons only).
positioncharacterRoster position from the most recent visible season.
proj_availabilitydoubleProjected availability rate in [0, 1] - empirical-Bayes shrinkage of historical snap-based availability toward the fitted position base rate, then the fold-fit linear recalibration.
proj_gamesdoubleExpected games available - proj_availability x team_games (17).
proj_games_misseddoubleExpected games missed - team_games minus proj_games.

Example

from sportsdataverse.nfl.nfl_availability import nfl_availability_projection
avail = nfl_availability_projection([2021, 2022, 2023], 2024)
avail.sort("proj_games").head()

nfl_game_script​

nfl_game_script(seasons: 'Union[int, List[int]]', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Team-season pace / PROE / expected-plays engine.

Loads pbp + schedules for seasons, aggregates per-game pace to the team-season level, and computes expected plays per game from the fitted sportsdataverse.nfl.nfl_scheme_constants.PACE_CONSTANTS.

Parameters

ParameterTypeDefaultDescription
seasonsUnion[int, List[int]]Season or list of seasons (nflverse pbp coverage).
return_as_pandasboolFalseWhen True, return a pandas.DataFrame.

Returns

Per (season, team): games, off_plays_pg, sec_per_play, neutral_sec_per_play, proe, exp_plays_pg, plays_oe, pace_rank (1 = fastest neutral pace). Empty seasons yield a zero-row frame with this schema.

col_nametypedescription
seasonintegerSeason of the aggregate.
teamcharacterTeam abbreviation.
gamesintegerGames included in the aggregate.
off_plays_pgdoubleRealized offensive plays per game.
sec_per_playdoubleSeason mean of the per-game sec_per_play.
neutral_sec_per_playdoubleSeason mean neutral-situation seconds per play (lower = faster).
proedoubleSeason pass-rate over expected, dropback-weighted so it reconciles exactly with the pbp pass_oe aggregate.
exp_plays_pgdoubleExpected plays per game from the fitted PACE_CONSTANTS OLS (own pace, opponent pace, market total).
plays_oedoubleRealized minus expected plays per game.
pace_rankintegerRank of neutral pace within the season (1 = fastest).

Example

from sportsdataverse.nfl.nfl_gamescript import nfl_game_script
gs = nfl_game_script([2023])
print(gs.sort("proe", descending=True).head())

# Pipeline next step

gs.filter(pl.col("plays_oe") > 0).sort("plays_oe", descending=True).head()

nfl_play_call_probabilities​

nfl_play_call_probabilities(pbp: 'pl.DataFrame', participation: 'Optional[pl.DataFrame]' = None, *, models_dir: 'Optional[str]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Score the bundled play-call classifier over offensive plays.

Parameters

ParameterTypeDefaultDescription
pbpDataFramenflverse-format pbp (must carry xpass; run sportsdataverse.nfl.ep_wp.calculate_xpass first if not).
participationOptional[DataFrame]NoneOptional participation frame for personnel features.
models_dirOptional[str]NoneOptional directory holding nfl_playcall.ubj (defaults to the bundled package artifact; no first-use download).
return_as_pandasboolFalseWhen True, return a pandas.DataFrame.

Returns

Keys + per-family probabilities p_inside_run / p_outside_run / p_short_pass / p_deep_pass / p_scramble, p_pass (pass-family sum), pred_family (argmax) and pass_oe_model (100 * (is_pass - p_pass)). Empty input yields a zero-row frame with this schema.

col_nametypedescription
game_idcharacternflverse game identifier (Utf8 join key).
play_idintegernflverse play identifier within the game (Int64 join key).
seasonintegerSeason of the play.
weekintegerWeek of the play.
posteamcharacterOffense (possession) team abbreviation.
p_inside_rundoublePredicted probability of an inside run (guard/center gap or middle).
p_outside_rundoublePredicted probability of an outside run (end/tackle or off-middle).
p_short_passdoublePredicted probability of a short pass.
p_deep_passdoublePredicted probability of a deep pass.
p_scrambledoublePredicted probability of a QB scramble.
p_passdoublePredicted pass probability (short + deep + scramble family sum).
pred_familycharacterArgmax family among the five class probabilities.
pass_oe_modeldoublePass-rate over model expectation for the play, 100 * (is_pass - p_pass).

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.ep_wp import calculate_xpass
from sportsdataverse.nfl.nfl_playcall import nfl_play_call_probabilities
out = nfl_play_call_probabilities(calculate_xpass(load_nfl_pbp([2023])))
print(out.select("p_pass", "pred_family").head())

# Pipeline next step

out.group_by("posteam").agg(pl.col("p_pass").mean()).sort("p_pass")

nfl_play_call_tendencies​

nfl_play_call_tendencies(pbp: 'pl.DataFrame', participation: 'Optional[pl.DataFrame]' = None, *, models_dir: 'Optional[str]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Aggregate scored play-call probabilities to team-season tendencies.

Parameters

ParameterTypeDefaultDescription
pbpDataFramenflverse-format pbp (with xpass).
participationOptional[DataFrame]NoneOptional participation frame.
models_dirOptional[str]NoneOptional directory holding nfl_playcall.ubj.
return_as_pandasboolFalseWhen True, return a pandas.DataFrame.

Returns

Per (season, posteam): plays, mean_p_pass, pass_rate, proe (100 * (pass_rate - mean_p_pass)) and the family mix shares share_<family>. Empty input yields a zero-row frame.

col_nametypedescription
seasonintegerSeason of the aggregate.
posteamcharacterOffense team abbreviation.
playsintegerOffensive run/pass plays scored.
mean_p_passdoubleMean model pass probability across the team's plays.
pass_ratedoubleActual pass rate (scrambles count as passes).
proedoublePass rate over expected, 100 * (pass_rate - mean_p_pass).
share_inside_rundoubleShare of plays labeled inside_run.
share_outside_rundoubleShare of plays labeled outside_run.
share_short_passdoubleShare of plays labeled short_pass.
share_deep_passdoubleShare of plays labeled deep_pass.
share_scrambledoubleShare of plays labeled scramble.

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.ep_wp import calculate_xpass
from sportsdataverse.nfl.nfl_playcall import nfl_play_call_tendencies
t = nfl_play_call_tendencies(calculate_xpass(load_nfl_pbp([2023])))
print(t.sort("proe", descending=True).head())

nfl_player_props​

nfl_player_props(seasons: 'int | list[int]', *, as_of_date: 'datetime.date | None' = None, era: 'str' = 'modern', lines: 'pl.DataFrame | None' = None, return_as_pandas: 'bool' = False) -> 'pl.DataFrame | pd.DataFrame'

Empirical-Bayes player-prop projections, leakage-safe per week.

For every game in the requested season(s) (or, with as_of_date, every game on/after that date), projects each rostered QB/RB/WR/TE's stat-family mean as usage x efficiency x matchup x game-script:

  • usage + efficiency from player_usage_efficiency built as-of that game's week (weeks strictly before it),
  • the matchup multiplier from the opponent's adj_def_epa in sportsdataverse.nfl.nfl_ratings.nfl_ratings (as-of the week's first game date),
  • game script from the native expected margin (sportsdataverse.nfl.nfl_market.nfl_predict_games) -- the market line is never read (binding non-market boundary).

Parameters

ParameterTypeDefaultDescription
seasonsint | list[int]Season (e.g. 2023) or list of seasons.
as_of_datedate | NoneNoneWhen given, only games with gameday >= as_of_date are projected (history before each game's week still feeds the projections). None projects every week of the season(s).
erastr'modern'Constants era key.
linesDataFrame | NoneNoneOptional market lines to score p_over against -- columns game_id / player_id / stat (Utf8) + line (Float64), e.g. built from espn_nfl_game_propbets (ESPN only serves propbets for upcoming games). None leaves line / p_over null.
return_as_pandasboolFalseIf True, returns a pandas DataFrame.

Returns

One row per (player-game, stat): season / week (Int64), game_id / player_id / position / team_id / opp_team_id / stat (Utf8), proj_mean / proj_sd / line / p_over (Float64; p_over = 1 - Phi((line - proj_mean) / proj_sd) when a line is joined, else null). Stats are passing_yards (QB), rushing_yards (RB), receiving_yards (WR/TE). Zero-row, correctly-typed when there is nothing to project.

col_nametypedescription
seasonintegerSeason the projection belongs to.
weekintegerWeek of the projected game; only weeks strictly before it feed the projection.
game_idcharacterGame identifier from the schedule (nflverse id, e.g. "2023_06_DET_TB").
player_idcharacternflverse GSIS player identifier (character join key).
positioncharacterPlayer position (QB, RB, WR, TE) selecting the projected stat family.
team_idcharacterPlayer's team nflverse abbreviation as of the projection week.
opp_team_idcharacterOpponent team nflverse abbreviation (drives the matchup multiplier).
statcharacterProjected stat name (passing_yards for QB, rushing_yards for RB, receiving_yards for WR/TE).
proj_meandoubleProjected stat mean - EB-shrunk usage x efficiency x opponent matchup x game-script.
proj_sddoubleResidual standard deviation for the stat family (fitted on the 2023 as-of backtest).
linedoubleMarket prop line joined from the caller-supplied lines frame (e.g. espn_nfl_game_propbets); null when no line is available.
p_overdoubleProbability the player exceeds line, 1 - Phi((line - proj_mean) / proj_sd); null without a line.

Example

from sportsdataverse.nfl import nfl_player_props
props = nfl_player_props(2023)
props.filter(props["stat"] == "passing_yards").head()

# Upcoming-only, as-of a date

import datetime as dt
props = nfl_player_props(2024, as_of_date=dt.date(2024, 11, 1))

nfl_predict_games​

nfl_predict_games(games: 'pl.DataFrame', ratings: 'pl.DataFrame', *, era: 'str' = 'modern', odds: 'pl.DataFrame | None' = None, return_as_pandas: 'bool' = False) -> 'pl.DataFrame | pd.DataFrame'

Vectorized pregame predictions (+ display-only market edge) per game.

Joins ratings twice (home/away) onto the schedule and computes the three closed-form predictions. odds is display-only: it feeds market_edge = exp_margin - close_spread_home and never the predictions themselves (the binding non-market boundary).

Parameters

ParameterTypeDefaultDescription
gamesDataFrameOne row per game: game_id (Utf8), home_team_id / away_team_id (Utf8 team abbreviations), neutral_site (Boolean).
ratingsDataFrameThe sportsdataverse.nfl.nfl_ratings.nfl_ratings output (needs team_id, adj_off_epa, adj_def_epa, adj_net).
erastr'modern'Constants era key.
oddsDataFrame | NoneNoneOptional market frame (game_id, close_spread_home -- the market's expected home margin, positive = home favored). Games absent from odds get a null market_edge.
return_as_pandasboolFalseIf True, returns a pandas DataFrame.

Returns

One row per input game: game_id / home_team_id / away_team_id (Utf8), neutral_site (Boolean), exp_margin / home_win_prob / exp_total / market_edge (Float64; market_edge null without odds). Zero-row, correctly-typed on empty input.

col_nametypedescription
game_idcharacterGame identifier carried through from the input schedule.
home_team_idcharacterHome team nflverse abbreviation (character; the ratings team_id join key).
away_team_idcharacterAway team nflverse abbreviation (character; the ratings team_id join key).
neutral_sitelogicalWhether the game is at a neutral site (home-field advantage is dropped when true).
exp_margindoubleExpected home scoring margin in points (points_per_net * net rating differential + the fitted home-field advantage on non-neutral fields).
home_win_probdoubleHome win probability, Phi(exp_margin / margin_sd) under a Gaussian margin model.
exp_totaldoubleExpected combined point total (avg_total + total_scale * the four-way efficiency matchup sum).
market_edgedoubleDisplay-only native-minus-market spread edge (exp_margin - close_spread_home); null when no odds frame is supplied.

Example

from sportsdataverse.nfl import nfl_ratings
from sportsdataverse.nfl.nfl_market import nfl_predict_games
ratings = nfl_ratings(2023)
preds = nfl_predict_games(games, ratings)
preds.sort("home_win_prob", descending=True).head()

# With a market edge (display only)

preds = nfl_predict_games(games, ratings, odds=odds)

nfl_punter_value​

nfl_punter_value(seasons: 'Union[int, List[int]]', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Punter net-field-position value over expected.

Expected net comes from the shipped punt landing distribution (nfl_fourth_down._load_punt_data) evaluated at each punt's line of scrimmage; realized net is kick_distance - return_yards - 20*touchback.

Parameters

ParameterTypeDefaultDescription
seasonsUnion[int, List[int]]Season or list of seasons.
return_as_pandasboolFalseWhen True, return a pandas.DataFrame.

Returns

Per (season, punter_player_id): punts, gross_avg, net_avg, exp_net_avg, net_over_expected, epa. Empty seasons yield a zero-row frame with this schema.

col_nametypedescription
seasonintegerSeason of the aggregate.
punter_player_idcharacternflverse punter GSIS id (Utf8 join key).
puntsintegerPunts with a recorded kick distance.
gross_avgdoubleMean gross punt distance (yards).
net_avgdoubleMean net distance, kick_distance - return_yards - 20 * touchback.
exp_net_avgdoubleMean expected net from the shipped punt landing distribution at each punt's line of scrimmage.
net_over_expecteddoublenet_avg minus exp_net_avg (yards of field position per punt over expectation).
epadoubleTotal EPA on the punter's punt plays (kicking-team perspective).

Example

from sportsdataverse.nfl.nfl_special_teams import nfl_punter_value
pv = nfl_punter_value([2023])
print(pv.head())

nfl_season_standings​

nfl_season_standings(games: 'pl.DataFrame', *, ranks: 'str' = 'CONF', tiebreaker_depth: 'str' = 'SOS', playoff_seeds: 'Optional[int]' = None, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Compute NFL standings with the real NFL tiebreaking procedures.

Faithful polars port of nflseedR::nfl_standings() (v2 engine, R/standings.R L82-155): initializes records, points, win percentages, SOV and SOS from a games frame, then resolves division ranks, conference ranks (playoff seeds) and draft order through the full NFL tiebreaker cascades.

Parameters

ParameterTypeDefaultDescription
gamesDataFrameGames frame with one row per game. Required columns: sim or season (identifier), game_type ('REG', 'WC', 'DIV', 'CON', 'SB'), week, away_team, home_team, and result (home score minus away score; no missing values allowed). away_score / home_score are additionally required for tiebreaker_depth='POINTS' and enable the pf/pa/pd output columns.
ranksstr'CONF'One of 'DIV', 'CONF' (default), 'DRAFT', or 'NONE' — which rank columns (and thus tiebreakers) to compute. 'DRAFT' implies 'CONF' implies 'DIV'.
tiebreaker_depthstr'SOS'One of 'SOS' (default), 'PRE-SOV', 'POINTS', or 'RANDOM'. Controls how deep the tiebreaker cascade goes before falling back to a coin toss.
playoff_seedsOptional[int]NoneIf not None, only conference ranks up to this value are resolved with tiebreakers; deeper ranks are returned as null. Must be in 1-16.
return_as_pandasboolFalseIf True, return a pandas DataFrame.

Returns

A standings frame with one row per (sim/season, team) including records, win_pct/div_pct/conf_pct, sov, sos, and the requested div_rank/conf_rank/draft_rank columns plus *_tie_broken_by bookkeeping. conf_rank is the playoff seed.

col_nametypedescription
seasonintegerSeason identifier from the input games frame (named sim instead when the input used a sim column).
confcharacterConference of the team (AFC or NFC).
divisioncharacterDivision of the team (e.g. "AFC East").
teamcharacterTeam abbreviation.
gamesintegerNumber of regular season games played.
winsdoubleRegular season wins with ties counted as half a win.
true_winsintegerRegular season wins excluding ties (outright wins only).
lossesintegerRegular season losses.
tiesintegerRegular season ties.
pfintegerPoints scored across regular season games (points for); present only when the input carries home_score and away_score.
paintegerPoints allowed across regular season games (points against); present only when the input carries scores.
pdintegerRegular season point differential (pf minus pa); present only when the input carries scores.
win_pctdoubleRegular season win percentage with ties counted as half a win.
div_pctdoubleWin percentage in games against division opponents (0 when the team played no division games).
conf_pctdoubleWin percentage in games against conference opponents (0 when the team played no conference games).
sovdoubleStrength of victory - combined win percentage of all opponents the team defeated (0 for winless teams).
sosdoubleStrength of schedule - combined win percentage of all opponents the team faced.
div_rankintegerRank within the division (1-4) after applying the NFL division tiebreaking procedures.
div_tie_broken_bycharacterTiebreaker step that resolved the team's division rank (e.g. "Head-To-Head Win PCT (2)" or "Coin Toss"); null when the rank needed no tiebreaker.
conf_rankintegerConference rank, i.e. the playoff seed, after applying the NFL conference tiebreaking procedures; null beyond playoff_seeds when that argument is set.
conf_tie_broken_bycharacterTiebreaker step that resolved the team's conference rank; null when the rank needed no tiebreaker.
exitcharacterRound of the team's final game - REG, WC, DIV, CON, SB, or SB_WIN for the Super Bowl winner (returned with ranks="DRAFT").
draft_rankintegerDraft pick position (1 = first overall pick) derived from postseason exit, win percentage, SOS and the draft tiebreaking procedures (returned with ranks="DRAFT").
draft_tie_broken_bycharacterTiebreaker step that resolved the team's draft rank; null when the rank needed no tiebreaker.

Example

import sportsdataverse.nfl as nfl
games = nfl.load_schedules([2024])
standings = nfl.nfl_season_standings(games, ranks="DRAFT")
print(standings.shape)

# Playoff seeds only, pandas output

df = nfl.nfl_season_standings(
games, ranks="CONF", playoff_seeds=7, return_as_pandas=True
)

# Pipeline next step (one line)

standings.filter(pl.col("conf_rank") <= 7).sort("conf", "conf_rank")

nfl_special_teams_epa​

nfl_special_teams_epa(seasons: 'Union[int, List[int]]', *, return_as_pandas: 'bool' = False) -> "Union[pl.DataFrame, 'pd.DataFrame']"

Special-teams EPA by team-unit.

Units: punt / punt_return / kickoff / kickoff_return / field_goal / extra_point. On each punt/kickoff the kicking team's unit carries the play EPA signed to the kicking team and the return team's unit its negation, so a team's units sum to its total ST-play EPA.

Parameters

ParameterTypeDefaultDescription
seasonsUnion[int, List[int]]Season or list of seasons.
return_as_pandasboolFalseWhen True, return a pandas.DataFrame.

Returns

Per (season, team, unit): plays, epa, epa_per_play. Empty seasons yield a zero-row frame with this schema.

col_nametypedescription
seasonintegerSeason of the aggregate.
teamcharacterTeam abbreviation.
unitcharacterSpecial-teams unit (punt, punt_return, kickoff, kickoff_return, field_goal, extra_point).
playsintegerPlays credited to the unit.
epadoubleTotal EPA credited to the unit (kicking team carries the play EPA signed to it; the return team carries its negation).
epa_per_playdoubleEPA per play for the unit.

Example

from sportsdataverse.nfl.nfl_special_teams import nfl_special_teams_epa
st = nfl_special_teams_epa([2023])
print(st.filter(pl.col("unit") == "punt").sort("epa", descending=True).head())

playcall_features​

playcall_features(pbp: 'pl.DataFrame', participation: 'Optional[pl.DataFrame]' = None) -> 'pl.DataFrame'

Build the play-call feature frame (one row per offensive run/pass play).

Filters to plays with pass == 1 or rush == 1, derives the 5-class family label (scramble > deep/short pass > inside/outside run), and left-joins the optional participation frame for personnel counts.

Parameters

ParameterTypeDefaultDescription
pbpDataFramenflverse-format pbp with the pre-snap feature columns + pass / rush / qb_scramble / pass_length / run_location / run_gap and xpass.
participationOptional[DataFrame]NoneOptional nflverse participation frame with game_id / play_id / offense_personnel.

Returns

Keys + PLAYCALL_FEATURE_ORDER columns + family + is_pass. Personnel columns are null (has_participation=0) when no participation row matches.

col_nametypedescription
game_idcharacternflverse game identifier (Utf8 join key).
play_idintegernflverse play identifier within the game (Int64 join key).
seasonintegerSeason of the play.
weekintegerWeek of the play.
posteamcharacterOffense (possession) team abbreviation.
downdoubleDown (1-4) at the snap.
ydstogodoubleYards to go for a first down.
yardline_100doubleYards from the opponent end zone at the snap.
score_differentialdoubleOffense score minus defense score at the snap.
half_seconds_remainingdoubleSeconds remaining in the half.
game_seconds_remainingdoubleSeconds remaining in the game.
wpdoubleStart-of-play win probability for the offense.
shotgundouble1 when the offense lined up in shotgun.
no_huddledouble1 when the play was run without a huddle.
xpassdoubleShipped nflfastR-parity expected-dropback probability for the play.
n_rbdoubleRunning backs in the offensive personnel grouping (null without participation data).
n_tedoubleTight ends in the offensive personnel grouping (null without participation data).
n_wrdoubleWide receivers in the offensive personnel grouping (null without participation data).
has_participationinteger1 when a participation row matched the play, else 0.
familycharacter5-class play-call label (inside_run, outside_run, short_pass, deep_pass, scramble).
is_passinteger1 when the play was a pass (including scrambles), else 0.

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.ep_wp import calculate_xpass
from sportsdataverse.nfl.nfl_playcall import playcall_features
feat = playcall_features(calculate_xpass(load_nfl_pbp([2023])))
print(feat["family"].value_counts())

player_usage_efficiency​

player_usage_efficiency(player_stats: 'pl.DataFrame', *, as_of_week: 'int', era: 'str' = 'modern') -> 'pl.DataFrame'

Per-player as-of usage + efficiency with empirical-Bayes shrinkage.

Aggregates one season of week-level player stats over weeks strictly before as_of_week (the leakage boundary), then shrinks every usage (per-game attempts / carries / targets) and efficiency (yards + TDs per opportunity) stat toward its position prior: (n * player_value + kappa * prior) / (n + kappa) with n = games played and kappa the stat family's fitted shrinkage.

Parameters

ParameterTypeDefaultDescription
player_statsDataFrameOne season of load_nfl_player_stats() rows (columns player_id, position, recent_team, week, attempts, passing_yards, passing_tds, carries, rushing_yards, rushing_tds, targets, receiving_yards, receiving_tds).
as_of_weekintOnly weeks < as_of_week are used.
erastr'modern'Constants era key (supplies kappas + position priors).

Returns

One row per player_id (Utf8) whose position has a prior table: position / team_id (Utf8, latest team), games (Int64), exp_attempts / exp_carries / exp_targets (Float64, shrunk per-game usage), ypa / ypc / ypt / pass_td_rate / rush_td_rate / rec_td_rate (Float64, shrunk per-opportunity efficiency). Zero-row, correctly-typed on empty input.

col_nametypedescription
player_idcharacterPlayer ID (aka GSIS ID) as defined by nflreadr::load_rosters
positioncharacterPrimary position as reported by NFL.com
team_idcharacter
gamesintegerGames played in career
exp_attemptsdouble
exp_carriesdouble
exp_targetsdouble
ypadouble
ypcdouble
yptdouble
pass_td_ratedouble
rush_td_ratedouble
rec_td_ratedouble

Example

import polars as pl
import sportsdataverse.nfl as nfl
stats = nfl.load_nfl_player_stats().filter(pl.col("season") == 2023)
usage = nfl.player_usage_efficiency(stats, as_of_week=10)
usage.sort("exp_attempts", descending=True).head()

predict_margin​

predict_margin(home_adj_net: 'float', away_adj_net: 'float', neutral: 'bool', *, era: 'str' = 'modern') -> 'float'

Expected home scoring margin from two net ratings.

points_per_net * (home_adj_net - away_adj_net) plus the era HFA on non-neutral fields.

Parameters

ParameterTypeDefaultDescription
home_adj_netfloatHome team's adj_net (EPA/play units).
away_adj_netfloatAway team's adj_net.
neutralboolTrue drops the home-field advantage.
erastr'modern'Constants era key (default "modern").

Returns

Expected home margin in points (positive = home favored).

Example

from sportsdataverse.nfl.nfl_market import predict_margin
predict_margin(0.10, -0.05, False)

predict_total​

predict_total(home_adj_off: 'float', home_adj_def: 'float', away_adj_off: 'float', away_adj_def: 'float', *, era: 'str' = 'modern') -> 'float'

Expected combined point total from the four efficiency components.

avg_total + total_scale * (home_adj_off + away_adj_def + away_adj_off + home_adj_def). The four ratings are summed because each side's scoring rises with its own offense and with the opponent's EPA-allowed (adj_def is lower = better defense) -- same semantics as the shipped CFB analog. (The plan text wrote this with a minus; that sign flips a good defense into raising the total, so the analog's sum is used.)

Parameters

ParameterTypeDefaultDescription
home_adj_offfloatHome adj_off_epa.
home_adj_deffloatHome adj_def_epa (lower = better defense).
away_adj_offfloatAway adj_off_epa.
away_adj_deffloatAway adj_def_epa.
erastr'modern'Constants era key.

Returns

Expected combined total in points.

Example

from sportsdataverse.nfl.nfl_market import predict_total
predict_total(0.10, -0.02, 0.05, 0.01)

team_game_pace​

team_game_pace(pbp: 'pl.DataFrame') -> 'pl.DataFrame'

Per team-game pace + pass-rate-over-expected.

sec_per_play is the per-drive elapsed game_seconds_remaining divided by drive plays, averaged over the team's offensive drives (kneels / spikes / no_plays excluded). Neutral = wp in [0.2, 0.8] and half_seconds_remaining > 120. proe is the mean pass_oe over dropbacks.

Parameters

ParameterTypeDefaultDescription
pbpDataFramenflverse-format pbp with game_id / season / week / posteam / drive / play_type / qb_dropback / pass_oe / game_seconds_remaining / wp / half_seconds_remaining.

Returns

One row per (game_id, season, week, posteam) with off_plays, sec_per_play, neutral_plays, neutral_sec_per_play, proe. Empty input yields a zero-row frame with this schema.

col_nametypedescription
game_idcharacternflverse game identifier (Utf8 join key).
seasonintegerSeason of the game.
weekintegerWeek of the game.
posteamcharacterOffense team abbreviation.
off_playsintegerOffensive plays in the game (kneels, spikes and no_plays excluded).
sec_per_playdoubleMean over the team's drives of elapsed game clock divided by drive plays.
neutral_playsintegerOffensive plays in neutral situations (wp in [0.2, 0.8], over 2 minutes left in the half).
neutral_sec_per_playdoublesec_per_play computed on neutral-situation plays only.
proedoubleMean pass_oe over the team's dropbacks in the game (percentage points).
dropbacksintegerDropbacks with a non-null pass_oe (the proe denominator).

Example

from sportsdataverse.nfl import load_nfl_pbp
from sportsdataverse.nfl.nfl_gamescript import team_game_pace
pace = team_game_pace(load_nfl_pbp([2023]))
print(pace.sort("sec_per_play").head())

win_prob_from_margin​

win_prob_from_margin(exp_margin: 'float', *, era: 'str' = 'modern') -> 'float'

Home win probability from an expected margin (Gaussian margin model).

Parameters

ParameterTypeDefaultDescription
exp_marginfloatExpected home margin in points.
erastr'modern'Constants era key (supplies margin_sd).

Returns

Phi(exp_margin / margin_sd) in [0, 1].

Example

from sportsdataverse.nfl.nfl_market import win_prob_from_margin
win_prob_from_margin(3.0)