Skip to contents

cfbfastR 3.0.0.9000 (development version)

The development version is numbered 3.0.0.9000 so it can be told apart from the CRAN 3.0.0 release, which ships the previous model generation. Until now both builds reported 3.0.0, so packageVersion("cfbfastR") could not distinguish them — while producing different EPA/WPA for the same play. A four-component version means you are on the GitHub build. See the model note below before mixing outputs from the two.

Air yards side the catch spot by the game’s own text abbreviations

The 2025+ ESPN vendor text spots the catch with each school’s own abbreviation (UHM, GSO, USC for South Carolina, Sac St, BC.), which is frequently not the payload’s home_team_abbreviation / away_team_abbreviation. The helper behind air_yards, air_yardsToEndzone and yards_after_catch now learns each game’s abbreviations from its own "... to the ABC nn" end spots against yards_to_goal_end (majority vote per abbreviation) before falling back to the payload abbreviations, reads every observed token shape (UA 10, BC.41, Sac St10, NC ST19) and resolves a spot at the 50 without one. On 2025 new-template games this lifts in-game air-yards coverage from 72% to 90% of pass plays; every spot-phrase play now resolves. Same logic as sdv-py #418; the parity oracle was re-captured from sdv-py 9efee9f1 and extended with five affected 2025/26 games (tests/testthat/fixtures/parity/README.md).

Expected Points model now comes from the shared cfb_model_artifacts bundle

The EP model is now the XGBoost artifact published in cfb_model_artifactsthe same artifact sportsdataverse-py scores with, so both libraries agree on EPA for a given play and a retrain updates both from one publish (#138).

  • Fixes #5epa_wpa = TRUE no longer aborts with predict.nnet(): missing values in 'x' on mid-era CFBD data (seasons ~2006–2013), in either engine.
  • EP scoring is consolidated behind one internal helper, so the seven next-score probability columns keep their historical names and order and no downstream code changed. The bundle’s class order differs from the retired model’s with no fixed point between them; the permutation is read from the bundle’s MANIFEST.json and asserted in tests, because getting it wrong yields EP that is wrong yet plausible-looking.
  • Model artifacts are cached under the package cache dir and refreshed on the cfbfastR.cache_duration TTL (default 24h), so a republished model is picked up without a package update. An expired cached copy is still used if the release is unreachable.
  • xgboost (Suggests) is now required to score EP and its floor moved to >= 1.7 for .ubj support. Without it — or offline with no cached copy — the retired nnet model is loaded as a fallback and still works.
  • The Win Probability model moved to the bundle’s wp_naive.ubj on the same terms. Eleven of its twelve features already existed on the frame; only is_home is derived.
  • The Field Goal model moved to the bundle’s era-aware fg_model.ubj (yards_to_goal + one-hot era0..era3). season is now threaded through .run_epa_wpa() and the exported create_epa() / epa_fg_probs() gain a season argument (defaulting to NULL). Scoring the era-aware model without a season is an error rather than a silent all-zero one-hot.
  • New: completion probability. cp and cpoe columns are added on pass plays, cpoe on the percentage-point scale 100 * (completion - cp), matching sportsdataverse-py. The model loads lazily on first use and the stage degrades to NA columns rather than failing.
  • New: spread-aware win probability. A vegas_wp column is added alongside the unchanged naive wp_before — the same split nflfastR draws between wp and vegas_wp. It is NA wherever the game has no pre-game line (the ESPN path, and CFBD before 2013). Sign convention verified over 83 games: CFBD’s spread is negative when the home team is favoured, and the model reads positive spread_time as the team in possession being favoured. vegas_wpa and vegas_wp_after are derived alongside it, using the same turnover and half/period-end overlays as the naive wpa.
  • New: expected pass rate. xpass and pass_oe columns are added on scrimmage plays (nflfastR’s xpass / pass_oe), pass_oe on the percentage-point scale 100 * (pass - xpass). Note this model uses an ordinal rule-era feature cutting at 2006/2013/2017, which is a different encoding from the FG model’s one-hot era0..era3 (2006/2013/2020).
  • Existing EPA/WPA values will change: this is a different model generation. Rebuild rather than mixing old and new outputs in one dataset.

Decision surfaces and QBR from the same bundle

The three remaining cfb_model_artifacts artifacts are now used: the two-point, fourth-down and QBR models (#140). Unlike the model columns above, these are analytic surfaces — they build the game state that WOULD follow a choice, score it, and compare. They are ports of cfb4th and degrade to NA columns rather than failing when a model, the punt table, arrow, or a pre-game line is missing.

  • New: the two-point decision. On offensive touchdowns, two_pt_wp / xp_wp / two_pt_wp_diff / two_pt_recommendation compare going for two against kicking, using the conversion probability prob_2pt and the opponent’s ensuing-drive win probability for each outcome. Verified bit-identical to sportsdataverse-py to eight decimals. Read two_pt_wp_diff rather than the bare recommendation: cfb4th’s rule has no margin, and the bundled model is optimistic about college conversion rates.
  • New: the fourth-down decision. On fourth downs, go_wp / punt_wp / fg_wp and the comparison columns go_boost, go_wp_diff, fg_wp_diff, punt_wp_diff and fourth_down_recommendation. The go branch expands the fourth-down model’s 76-class yards-gained distribution to one state per possible outcome; the punt branch joins the empirical end-yardline distribution (NA inside the 31, where cfb4th’s table is empty); the field-goal branch carries cfb4th’s policy clamps (zero beyond 42 yards-to-goal, 0.9x from 35 out). first_down_prob, wp_succeed, wp_fail, make_fg_wp, miss_fg_wp and fourth_down_fg_make_prob are exposed alongside. Note the last of those is namespaced: fg_make_prob already means the make probability of the field goal that was actually attempted.
  • New: create_qbr(). Per-quarterback, per-game leverage-weighted EPA components scored through the bundled QBR model. It emits its own table rather than columns on the play-by-play frame, so it is an entry point you call on a modeled frame, not a pipeline stage.
  • These three surfaces have no cross-language oracle for the fourth-down branches — sportsdataverse-py’s get_go_wp() raises on pandas 3 — so they were ported from the cfb4th R source and gated behaviourally against known college conversion and field-goal rates. Three sign conventions were found inverted in the Python port while doing so; each produced a confident wrong recommendation rather than an error, and each is pinned by test here.
  • Every decision column is NA on the ESPN engine, which carries no pre-game spread.

cfbfastR v3.0.0

CRAN release: 2026-08-24

New release-dataset loaders (48 functions)

cfbfastR now loads every published CFB dataset on the sportsdataverse-data release repo, closing the gap with sportsdataverse-py’s loader surface. All loaders return cfbfastR_data-tagged tibbles, accept seasons = TRUE for the full published range, and support the dbConnection/tablename database write-through.

Which play-by-play loader do I want? Nothing is deprecated — the classic functions are unchanged; the new families add sources:

  • load_cfb_pbp() (unchanged) — the classic cfbfastR EPA/WPA play-by-play, FBS 2014+.
  • load_espn_cfb_pbp() (new) — the ESPN-derived play-by-play (469 columns, EPA/WPA + participant ids), 2004+.
  • load_ncaa_mfb_pbp() (new) — stats.ncaa.org play-by-play incl. FCS and lower divisions, 2013+; load_ncaa_mfb_pbp_cfbfastr() is the same data reshaped onto cfbfastR pbp column conventions for cross-source binds.

This release adds a 65-function ESPN college-football API layer, expanding cfbfastR’s ESPN surface from 8 wrappers to 73. The new wrappers expose ESPN’s core-v2 endpoints in ESPN’s own ID space — complementary to the CollegeFootballData (cfbd_*) wrappers, and the natural join partners for espn_cfb_pbp() / espn_cfb_scoreboard(). Every wrapper was verified live against the 2023, 2024, and 2025 seasons.

Naming alignment with the sportsdataverse convention (this dev cycle, never on CRAN): espn_cfb_player_statistics() is renamed to espn_cfb_player_career_stats() (the core-v2 /athletes/{id}/statistics career view, matching hoopR/wehoop/sportsdataverse-py). New espn_cfb_player_stats_v3() wraps the comprehensive web-common-v3 /athletes/{id}/stats payload (all categories, long format) — the _v3 companion to espn_cfb_player_stats() (core-v2 season statistics).

New Fox Sports API wrappers (fox_cfb_*)

A read-only Fox Sports “Bifrost” college-football layer (api.foxsports.com/bifrost/v1/cfb/*), complementary to the espn_cfb_* and cfbd_* families. Eight wrappers flatten Fox’s layout-oriented JSON (sections → tables → rows → cells) into tidy cfbfastR-tagged tibbles. Reverse-engineering notes and an OpenAPI 3.1 spec live in the sdv-internal-refs repo. Verified live against the 2025 season.

New ESPN wrappers — football-specific metrics

  • espn_cfb_powerindex() — ESPN’s College Football Power Index (FPI): every predictive metric and efficiency component, in long format.
  • espn_cfb_qbr() — Total Quarterback Rating (QBR) and the full set of clutch-weighted EPA components, one row per qualified passer.
  • espn_cfb_futures() — the season betting-futures board (national championship, conference, and award markets) with each sportsbook’s American odds.
  • espn_cfb_recruits() — ESPN’s recruiting board for a class, one row per recruit with grade, position/state/region rank, committed school, and hometown.

New ESPN wrappers — players

New ESPN wrappers — teams

New ESPN wrappers — game detail

New ESPN wrappers — catalogs and season metadata

New Yahoo Sports wrappers

New CollegeFootballData wrappers

  • cfbd_betting_ats() — season against-the-spread (ATS) summary records by team, wrapping the CollegeFootballData /teams/ats endpoint.
  • cfbd_stats_game_havoc() — per-game havoc statistics (total / front-seven / defensive-back havoc events and rates, offense and defense), wrapping the CollegeFootballData /stats/game/havoc endpoint.
  • cfbd_pbp_data_v2() is a new public function: a modular successor to cfbd_pbp_data() that runs the same EPA/WPA pipeline through a single shared engine (.run_epa_wpa()) and a canonical play-type taxonomy (.pbp_play_types()). The legacy cfbd_pbp_data() is unchanged.
  • espn_cfb_pbp_v2() now sources play-by-play and meta through the shared engine, requests participants = "wide" and team_participants = "wide" from espn_cfb_game_drives(), and adds the meta columns home_team_name, home_team_color, home_team_alternate_color, home_team_rank (and away_*) via the new .espn_pbp_game_meta() bridge. Output is a strict superset of legacy espn_cfb_pbp() on the meta columns.

Play-by-play engine — v2 is now the default

  • cfbd_pbp_data() and espn_cfb_pbp() now run the v2 engine by default. Both gain penalty enforcement resolution, ESPN-resolved player names, the *_player_id columns and the output tier selector without a code change. The previous behaviour is one argument away — engine = "legacy" per call, or options(cfbfastR.pbp_engine = "legacy") for a session — and a once-per-session message says so. engine = "auto" continues to mean “whatever this release considers current”. tests/testthat/test-pbp_equivalence.R asserts v2 reproduces the legacy frames column-for-column, with an explicit allow-list of intentional deltas.

  • Play-by-play now overwrites the regex-extracted *_player_name values with ESPN’s own participants[] names (2014 onward), so a capture that trailed narration ("Rod Smith 3 Yd"), abbreviated, or carried a team code becomes the real name. Ported from sportsdataverse’s CFBPlayProcess.__join_participants and verified against a 60-game offline oracle (5 games from each of 2004, 2006, 2008, 2010, 2013, 2014, 2017, 2019, 2020, 2021, 2023, 2025, including the postseason): 1,236 of 9,545 × 11 name cells change, with zero divergence from the Python. The stage runs before id resolution, so the roster matcher gets a clean key rather than narration.

  • espn_cfb_pbp_v2() gains resolve_names (default TRUE). When epa_wpa = TRUE it spends one memoised request per game on ESPN’s play-by-play sidecar, which supplies (a) full athlete names — "Jalen Mitchell" rather than the core-v2 roster’s "J. Mitchell" — and (b) the per-player box score as a second identity source. The box score is the only identity source on the large share of games where ESPN 404s the roster resource; adding it cuts *_player_id divergence from sdv-py from 52/163,656 to 26/166,482 — fewer mismatches over more plays. Set resolve_names = FALSE for a bulk sweep that would rather have the short names than the requests.

  • Play-by-play now resolves which team each event belongs to, adding 30 columns cfbfastR could not previously produce: the special-teams flip (kicking_team, return_team, punt_return_team, kick_return_team, fg_team, punt_team), the event-credit columns (sack_team, interception_team, pass_breakup_team, forced_fumble_team, fumble_recovery_team), the fumble/recovery chain (fumble_or_muff, fumbling_team, recovery_team, recovery_team_2), the per-side turnover model (is_turnover, turnover_team, int_turnover, pos_fumble_lost, def_fumble_lost, is_pos_team_turnover, is_def_pos_team_turnover, is_st_turnover, is_blocked_punt_turnover, is_blocked_fg_turnover), penalty attribution (penalized_team, penalty_team_id, penalty_yards_signed) and the id-keyed pos_team_id / def_pos_team_id. Ported from sportsdataverse’s CFBPlayProcess.__add_attribution_cols and verified against the 60-game offline oracle at zero divergence over 267,260 cells.

    The turnover flags are framed per side because one play can lose the ball twice — the offense fumbles, the defense recovers and fumbles back — and a single boolean cannot say that both teams turned it over. Blocked punts and blocked field goals deliberately stay out of is_turnover: ESPN’s official box counts only giveaways, so folding them in would break the reconciliation against it. They get their own flags instead.

    On the ESPN path, ESPN’s own per-play turnover flag is preserved as espn_is_turnover rather than being silently overwritten. The two legitimately differ — ESPN’s also fires on blocked kicks.

  • Play-by-play gains air_yards, air_yardsToEndzone and yards_after_catch, splitting a completed pass into the yards thrown and the yards run after the catch. ESPN states the catch point as "caught at OU35"; the stated yard line belongs to whichever team owns that side of the field, so the abbreviation is sided against the possessing and defending teams with the same prefix-tolerant matcher the recovery and penalty teams use. yards_after_catch is computed for completions only. Ported from sportsdataverse’s CFBPlayProcess.__add_air_yards_cols.

    One deliberate divergence from the Python: sdv-py’s pattern has no article, so it resolves "caught at OU35" and silently drops "thrown to the ARK30". Both forms occur in ESPN’s text — and cfbfastR’s core-v2 feed uses the article form almost exclusively, where a verbatim port would have matched nothing at all. This implementation accepts both. The parity test is partitioned accordingly: exact agreement on the 137 oracle rows sdv-py resolves, plus an explicit assertion that the 7 article-form rows it drops are recovered here.

  • Play-by-play gains pass_depth, pass_direction, rush_direction and qb_hurry, read from the play text ("short"/"deep", "left"/"middle"/"right", and ESPN’s "QB hurried by" annotation). Null where ESPN omits the phrase — sacks, screens, and older seasons that never annotated depth or direction.

  • espn_cfb_pbp_v2() now refuses to run the EPA/WPA models on an obviously malformed feed — no plays, or implausibly few or many for a game that has finished — and returns the unmodeled frame with a warning instead. A truncated game models perfectly cleanly and produces EPA, drive results and a box score that all look reasonable and are all wrong, and nothing downstream can distinguish that from a real blowout. The count rules apply only to completed games so a live feed is never rejected. Ported from CFBPlayProcess.corrupt_pbp_check.

  • A pbp-to-boxscore parity gate now guards the test suite. The per-play parity tests check a play against itself; this aggregates play-by-play by team and compares it against ESPN’s own team box, which is the only cheap end-to-end judge of whether parsing put the right events on the right team. An attribution bug leaves every per-play assertion green and shows up here immediately. Ported from sportsdataverse’s tools/validation/checks/boxscore_parity, adapted from a per-season harness check to a per-game library gate.

    It encodes three conventions proven against the box, two of which are the opposite of the NFL’s: NCAA charges a sack to rushing (attempt and yardage both), pass attempts exclude sacks, and a penalty belongs to the team that committed it — a positive penalty_yards_signed means the offence gained, so the defence was flagged. Floors are measured from the shipped 60-game corpus, never guessed, and are per-stat because parity is strongly era-dependent (interceptions reconcile at 97%, 2004-inclusive rushing yardage at 31%).

CFBD API coverage

Audited against the CollegeFootballData OpenAPI spec (5.24.1, 74 endpoints).

15 endpoints that had no wrapper now have one: cfbd_playoffs_cfp(), cfbd_playoffs_cfp_games(), cfbd_playoffs_cfp_participants(), cfbd_conference_affiliations(), cfbd_conference_changes(), cfbd_coaches_profile(), cfbd_coaches_seasons(), cfbd_coaches_tenures(), cfbd_ratings_core(), cfbd_ratings_srs_expanded(), cfbd_teams_fbs(), cfbd_stats_player_success(), cfbd_stats_player_success_game(), cfbd_player_season_overview() and cfbd_info_usage(). Every one was exercised against the live API before being committed.

26 parameters added to existing wrappers — most importantly division on ten more functions, plus defense / offense_conference / defense_conference / conference / division on cfbd_pbp_data(), competition and round on cfbd_game_info() (College Football Playoff filtering), line_provider on cfbd_betting_lines(), conference on cfbd_play_stats_player() and recruit_type on cfbd_recruiting_position(). cfbd_conferences() previously took no arguments at all and now accepts year and division.

New validate_division() covering fbs / fcs / ii / ii/iii / iii. This validates locally because CFBD ignores an unrecognised filter value rather than rejecting it — so without it a typo silently returns every division.

Two spec parameters were deliberately not exposed after testing them: /rankings declares latest and final as booleans, but the API returns HTTP 400 for every form of both. poll is validated to "cfp", the only value it accepts.

Bug fixes

  • espn_cfb_team_coaches() — the year argument is deprecated. ESPN’s core-v2 coaches endpoint returns the current coach whatever season is requested, echoing the requested year back in the response, so historical calls silently returned today’s coach labelled with the old season (#125). year now defaults to most_recent_cfb_season(); passing any other season warns and is coerced, rather than returning misattributed data. Existing calls keep working.

  • espn_cfb_teams() returned zero rows, because site.api.espn.com now answers HTTP 403 to a spoofed browser User-Agent. The failure was silent — the wrapper caught it and returned an empty frame — and every consumer degraded to NA, which took home, away, pos_team, def_pos_team, offense_play, defense_play and every team abbreviation on the ESPN play-by-play path down with it. Measured 2026-08-19: the endpoint answers 200 with httr2’s default UA, with curl/8.5.0, or with Accept/Origin/Referer and no UA at all, and 403 with the Chrome string. The User-Agent header is dropped.

    Probing every ESPN host the package uses narrowed the blast radius to exactly two callersespn_cfb_teams() and espn_cfb_team_schedule(), the only two that combine site.api.espn.com with the spoofed header. espn_cfb_team_schedule() was returning zero rows for the same reason and now returns data. The other ~65 occurrences of the header sit on sports.core.api.espn.com and site.web.api.espn.com, which answer 200 either way, so they are left alone.

    Worth knowing for anyone adding an ESPN call: cdn.espn.com does something worse than a 403 under the browser UA — it answers 200 with a zero-byte body, so nothing raises and the parse silently yields nothing. tests/testthat/test-espn_http_headers.R is a source-level guard that fails if the header is re-added to a site.api caller.

  • defense_play was a copy-paste duplicate of offense_play in the ESPN adapter — both case_when() branches returned the home team — so it named the team with the ball on every ESPN play.

  • Play-by-play gains pos_team_id / def_pos_team_id / offense_play_id / defense_play_id. pos_team, def_pos_team, offense_play and defense_play are team NAMES resolved through the ESPN teams catalog, and when that catalog is unavailable they all go NA together — which silently disabled team-aware roster matching, dropping every player-id lookup to the global-unique fallback. The ids come straight off the play and are always present. (espn_cfb_teams() currently returns zero rows, so this is the live condition, not a hypothetical.)

  • .espn_cfb_participant_roster() is now memoised alongside the ESPN catalog helpers. espn_cfb_pbp_v2() needs one game’s roster twice — once to name participants, once to resolve player ids — and a season sweep asked for it once per game; both now cost a single request. The memoised-helper list is a single constant shared by .onLoad() and espn_cfb_clear_cache(), which had been a second hand-maintained copy that could silently drift into caching a helper it never cleared.

  • .run_epa_wpa_by_game() had no roster argument, so cfbd_pbp_data_v2() could never resolve *_player_id columns however the roster was supplied. It now threads rosters and participants through, sliced per game_id.

  • espn_cfb_pbp() now builds its request URL with the ?event= query separator (previously concatenated as summaryevent=, which returned HTTP 404 for every game) and initializes its return frame before the tryCatch so an upstream failure no longer throws object 'plays_df' not found.

  • cfbd_pbp_data_v2() and espn_cfb_pbp_v2() preserve character id_play precision through the EPA/WPA pipeline. The legacy shared helper used unquoted numeric literals in two ifelse calls (a historical id_play swap for one game), which silently coerced character id_play to numeric and then lost precision past 2^53 — breaking the play-id join-back in espn_cfb_pbp_v2(). The modular .pbp_clean_pbp_dat() quotes those literals so id_play stays character; the legacy clean_pbp_dat() is unchanged.

Internal changes

  • httr -> httr2 migration. cfbfastR’s HTTP layer now uses the modern httr2 package (>= 1.0.0) instead of the legacy httr. End users running existing wrapper calls (cfbd_*, espn_cfb_*) should see no behavioural change – the migration is internal. Custom code that calls get_req() or check_status() directly must update from httr::content(res, as = "text") to httr2::resp_body_string(res) and from httr::status_code(res) to httr2::resp_status(res).
  • Proxy support. get_req() now resolves a proxy in the order: explicit proxy argument -> getOption("cfbfastR.proxy") -> http_proxy / https_proxy env vars. The proxy value accepts either a URL string or a named list with url / port / username / password / auth for authenticated proxies.
  • Dependency footprint trimmed. lubridate, progressr, memoise, cachem, and magrittr have moved out of Imports (21 -> 16). lubridate is gone entirely – its two ymd_hm() |> with_tz() calls in espn_cfb_schedule.R are now base-R as.POSIXct(format = "%Y-%m-%dT%H:%M", tz = "UTC") + attr(., "tzone"). progressr, memoise, and cachem moved to Suggests and the helpers degrade gracefully when missing: load_cfb_pbp() / cfbd_pbp_data() / pbp_epa_wpa_engine() run without a progress bar when progressr is absent; ESPN catalog wrappers run uncached when memoise / cachem are absent (espn_cfb_clear_cache() becomes a no-op). Drops the Imports count below the >20 R CMD check NOTE threshold.
  • Native pipe migration. All 1,419 %>% chains in R/, plus 137 across vignettes/ and tests/, were converted to the base-R native pipe |>. magrittr is no longer an Imports; downstream consumers that load cfbfastR purely for its functions don’t get %>% re-exported anymore. User-visible impact is minimal – the public API is unchanged and dplyr (which is in Imports) still re-exports %>% for users who want to keep writing it. Two non-mechanical fixes were needed during the sweep: three |> [[("url") chains in cfbd_betting.R and cfbd_coaches.R (rejected as RHS in R 4.1’s |>) became |> purrr::pluck("url"); seven |> tibble::tibble(col = .data$.) constructs were a magrittr quirk that silently duplicated the LHS into both a . and the named column – rewritten to tibble::tibble(col = <lhs>), which drops the redundant . column.
  • Test-time CFBD throttle. A new tests/testthat/setup-cfbd-throttle.R adds a 1-second sleep before every CFBD request made by devtools::test() / R CMD check. It works by monkey-patching cfbfastR:::get_req() for the duration of the test session (restored via withr::defer(., teardown_env())) – the package code is unchanged, so interactive and production calls pay no penalty. Tunable via options(cfbfastR.test_request_delay = N) (default 1; set to 0 for unthrottled local runs). Resolves the cascading HTTP 429 skip-if-empty results that were turning otherwise-green test runs into “all green, mostly skipped.” withr joins Suggests to declare the test-side dependency cleanly (it was already a transitive dep of testthat).

cfbfastR v2.2.0

  • Fixes a bug in validate_week() utility function where some inputs were not being handled correctly (i.e. week 16). Fixes trickle down to cfbd_pbp_data() and other functions.
  • Default value for season_type parameter in cfbd_game_info() and cfbd_play_stats_player() function changed from “regular” to “both” to align with other functions in the package.

cfbfastR v2.1.0

  • Fixes a bug in cfbd_pbp_data() where play-by-play data for some games were not as expected.
  • Improves add_yardage() where plays with missing yardage values were not being handled correctly.

cfbfastR v2.0.0

CRAN release: 2025-09-09

Breaking Changes to Loading Functions

  • All load_cfb_*() functions now use sportsdataverse-data releases or the CollegeFootballData.com API as their underlying data source to remain in compliance with CFBD API terms and conditions (See Note below).
  • Updated load_cfb_pbp() dataset to include various team- and game-level ID’s and flags that were not being included, like home_team_id, away_team_id, season_type, venue_id, some drive_* columns, a half-dozen player stat columns, etc. Essentially, all the leg-work users have undoubtedly had to do while using these datasets is mostly just included now. The downside: this means end users need to check their pipelines which build off these datasets to ensure behavior is as expected and all your joins are doing what is intended.

Now upgraded to the CFBD v2 API

Special thanks are in order for our newest contributor, Brad Hill (@bradisbrad) for providing most of the v2 upgrade via his first PR to cfbfastR!! 🙌🏽 👑 🥇 Your contributions are most appreciated by the community.

Note: The free-tier API key for the CFBD v2 API has a strict 1k calls/month limit, so plan your workflows accordingly! If you receive errors mentioning r Request failed [429], you have most likely run out of API calls for the month in your membership tier.

cfbfastR v1.9.5

  • fixed breaking bug related to stringi v1.8 update in cfbd_play_pbp_data() EPA and WPA processing
  • Minor documentation and test updates

cfbfastR v1.9.4

cfbfastR v1.9.3

cfbfastR v1.9.2

cfbfastR v1.9.1

  • Improved drive_pts logic in play-by-play data.
  • Fixed an issue that occasionally made the cfbd_game_team_stats() function return data in a long format
  • Minor documentation and test updates

cfbfastR v1.9.0

CRAN release: 2022-06-13

Added functions to access ESPN API:
Added functions to pull data from the data repo:

cfbfastR v1.8.0

  • All functions now default to return tibbles.
  • Added S3 method to print outputs with data info and retrieval timestamps. Thank you to Tan Ho (@tanho36) for the idea.

cfbfastR v1.7.1

cfbfastR v1.7.0

cfbfastR v1.6.7

cfbfastR v1.6.6

  • Updated function cfbd_pbp_data() to account for additional timeout cases (namely, kickoffs/extra point attempts).

cfbfastR v1.6.5

cfbfastR v1.6.4

CRAN release: 2021-10-27

  • Changed options to revert to old options on exit of function.
  • Removed check_github functions.

cfbfastR v1.6.3

  • Switched package urls in DESCRIPTION again.

cfbfastR v1.6.2

  • Switched package urls in README and DESCRIPTION files to https://

cfbfastR v1.6.1

  • Removed source urls from many package documentation entries.
  • Updated a test to skip on CRAN

cfbfastR v1.6.0

  • Added cfbd_ratings_elo() function
  • Fixed a bug in update_cfb_db() where the function failed when trying to load recent games from the data repo. (#35)
  • Added the option cfbfastR.dbdirectory that allows to set the database directory in update_cfb_db() globally.

cfbfastR v1.5.2

  • Remove verbose parameter

cfbfastR v1.5.1

Minor release
  • Removed calculated columns from cfbd_stats_season_team() that were not behaving correctly
  • Fixed bug where only_fbs input in cfbd_team_info() was ignored. It is now possible to get the team info for all the colleges in the API instead of only FBS schools.
  • Removed default year from cfbd_metrics_ppa_teams. cfbd_metrics_ppa_teams and cfbd_metrics_ppa_players_season now require one of team or year to be specified

cfbfastR v1.5.0

cfbfastR v1.4.0

cfbfastR v1.3.3

cfbfastR v1.3.2

Added ID linking to cfbd_recruiting_players()

cfbfastR v1.3.0-1

Added three NFL draft functions:

cfbfastR v1.2.1

Minor release
  • Added headshot_url to outputs of cfbd_team_roster()

  • Renamed returns in cfbd_game_box_advanced():

    • rushing_line_yd_avg to plural rushing_line_yds_avg
    • rushing_second_lvl_yd_avg to plural rushing_second_lvl_yds_avg
    • rushing_open_field_yd_avg to plural rushing_open_field_yds_avg
  • Completed documentation for all returns except cfbd_pbp_data()

  • Continued work on intro vignette

cfbfastR v1.2.0-1

Add significant documentation to the package
ESPN/CFBD metrics function variable return standardization

cfbfastR v1.1.0

Add loading from Data Repository functionality
Add support for parallel processing and progress updates
  • Added furrr, future, and progressr dependencies to the package to allow for parallel processing of the play-by-play data with progress updates if desired.

cfbfastR v1.0.0

Function Naming Convention Change
  • All functions sourced from the College Football Data API will start with cfbd_ as opposed to cfb_ (as in cfbscrapR). One additional cfbd_ function has been added that corresponds to the result when cfbd_pbp_data() has the parameter epa_wpa=FALSE. It has now been separated into its own function for clarity cfbd_plays(). The parameter and functionality still exists in cfbd_pbp_data() but we expect this function will still exist but made obsolete in favor of a function more closely matching nflfastR’s naming conventions.

  • Similarly, data and metrics sourced from ESPN will begin with espn_ as opposed to cfb_. In particular, the two functions are now espn_ratings_fpi() and espn_metrics_wp()

  • Data generated from any of the cfbfastR methods will use cfb_

College Football Data API Keys

The CollegeFootballData API now requires an API key, here’s a quick run-down:

CFBD_API_KEY = XXXX-YOUR-API-KEY-HERE-XXXXX

Save the script and restart your RStudio session, by clicking Session (in between Plots and Build) and click Restart R (n.b. there also exists the shortcut Ctrl + Shift + F10 to restart your session). If set correctly, from then on you should be able to use any of the cfbd_ functions without any other changes.

  • For less consistent usage: At the beginning of every session or within an R environment, save your API key as the environment variable CFBD_API_KEY (with quotations) using a command like the following.

{r} Sys.setenv(CFBD_API_KEY = "XXXX-YOUR-API-KEY-HERE-XXXXX")

  • Added API Key methods. If you forget to set your environment variable, functions will give you a warning and ask for one.