
CFB Data Stats Examples
Saiem Gilani,
Akshay Easwaran, Jared Lee
Source: vignettes/cfbd_stats.Rmd
cfbd_stats.RmdLoad and Install Packages
if (!requireNamespace('pak', quietly = TRUE)){
install.packages('pak')
}
pak::pak(c("dplyr", "tidyr", "gt"))## ℹ Loading metadata database
## ✔ Loading metadata database ... done
##
##
## → Package library at /home/runner/work/_temp/Library.
## ✔ All system requirements are already installed.
##
## ℹ No downloads are needed
## ℹ Installing system requirements
## ℹ Executing `sudo sh -c apt-get -y update`
## Get:1 file:/etc/apt/apt-mirrors.txt Mirrorlist [144 B]
## Hit:6 https://packages.microsoft.com/ubuntu/22.04/prod jammy InRelease
## Hit:2 http://azure.archive.ubuntu.com/ubuntu jammy InRelease
## Hit:3 http://azure.archive.ubuntu.com/ubuntu jammy-updates InRelease
## Hit:7 https://dl.google.com/linux/chrome-stable/deb stable InRelease
## Hit:4 http://azure.archive.ubuntu.com/ubuntu jammy-backports InRelease
## Hit:5 http://azure.archive.ubuntu.com/ubuntu jammy-security InRelease
## Reading package lists...
## ℹ Executing `sudo sh -c apt-get -y install libicu-dev libcurl4-openssl-dev libssl-dev cmake make libuv1-dev pandoc libnode-dev libxml2-dev`
## Reading package lists...
## Building dependency tree...
## Reading state information...
## libicu-dev is already the newest version (70.1-2).
## make is already the newest version (4.3-4.1build1).
## pandoc is already the newest version (2.9.2.1-3ubuntu2).
## cmake is already the newest version (3.22.1-1ubuntu1.22.04.2).
## libcurl4-openssl-dev is already the newest version (7.81.0-1ubuntu1.27).
## libssl-dev is already the newest version (3.0.2-0ubuntu1.29).
## libuv1-dev is already the newest version (1.43.0-1ubuntu0.1).
## libxml2-dev is already the newest version (2.9.13+dfsg-1ubuntu0.12).
## libnode-dev is already the newest version (12.22.9~dfsg-1ubuntu3.6).
## 0 upgraded, 0 newly installed, 0 to remove and 86 not upgraded.
## ✔ 3 pkgs + 56 deps: kept 59 [9.1s]
##
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
##
## filter, lag
## The following objects are masked from 'package:base':
##
## intersect, setdiff, setequal, union
library(tidyr)
library(gt)
# cfbfastR is deliberately NOT in the pak() list above. pkgdown and
# R CMD check render this vignette against the package being built, and
# installing from CRAN here would overwrite that dev build with the last
# release -- so any function added since it would vanish mid-render.
library(cfbfastR)
# pak::pak("sportsdataverse/cfbfastR")Settling 2019 LSU and 2013 Florida State offense debates
Get Season Statistics by Team
team_season_stats <- dplyr::bind_rows(
cfbd_stats_season_team(year=2019, team = "LSU"),
cfbd_stats_season_team(year=2013, team = "Florida State")
)
logos <- read.csv("https://raw.githubusercontent.com/sportsdataverse/cfbfastR-data/main/themes/logos.csv")
logos<- logos |> dplyr::select(-"conference")
df_team_season <- team_season_stats |>
dplyr::left_join(logos, by=c("team"="school"))
df_team_season_long <- as.data.frame(t(as.matrix(df_team_season)))
colnames(df_team_season_long) <- df_team_season$teamGet Season Advanced Statistics by Team
df_team_season_adv <- dplyr::bind_rows(
cfbd_stats_season_advanced(2019, team = "LSU"),
cfbd_stats_season_advanced(2013, team = "Florida State")
)
df_team_season_adv <- df_team_season_adv |>
dplyr::left_join(logos, by=c("team"="school"))Get Game Advanced Stats
df_team_game_adv <- dplyr::bind_rows(
cfbd_stats_game_advanced(2019, team = "LSU"),
cfbd_stats_game_advanced(2013, team = "Florida State")
)
df_team_game_adv <- df_team_game_adv |>
dplyr::left_join(logos, by=c("team"="school"))Get Season Statistics by Player
source("https://raw.githubusercontent.com/sportsdataverse/cfbfastR-data/main/themes/gt_theme_code_SG.R")
passing_df <- dplyr::bind_rows(
cfbd_stats_season_player(2019, team = "LSU", category = "passing"),
cfbd_stats_season_player(2013, team = "Florida State", category = "passing")) |>
dplyr::left_join(logos, by=c("team"="school")) |>
dplyr::group_by(team) |>
dplyr::mutate(logo = sprintf('<img src="%s" height="30" alt="%s logo">', logo, team)) |>
dplyr::select(logo,
player,
passing_completions,
passing_att,
passing_yds,
passing_td,
passing_int,
passing_ypa) |>
arrange( desc(passing_yds), team)## Adding missing grouping variables: `team`
passing_df |> gt() |>
tab_header(title = "Passing Summary") |>
cols_label(logo="",
player = "Player",
passing_completions = "C",
passing_att = "Att",
passing_yds = "Yds",
passing_td = "TDs",
passing_int = "INTs",
passing_ypa = "YPA") |>
data_color(
columns = c("passing_yds"),
colors = scales::col_numeric(
palette = "RdBu",
domain = c(-6000,6000)
)
) |>
data_color(
columns = c("passing_td"),
colors = scales::col_numeric(
palette = "RdBu",
domain = c(-60,60)
)
) |>
data_color(
columns = c("passing_td"),
colors = scales::col_numeric(
palette = "RdBu",
domain = c(-60,60)
)
) |>
# Render the team logos from pre-built <img> HTML that carries alt text
# (gt::web_image() cannot set alt; important for screen-reader accessibility).
fmt_markdown(columns = "logo") |>
tab_source_note(source_note = md("**Table:** @SaiemGilani | **Data:** @CFB_Data with @cfbfastR v2.0.0")) |>
gt_theme_538(table.width = px(550))## Warning: Since gt v0.9.0, the `colors` argument has been deprecated.
## • Please use the `fn` argument instead.
## This warning is displayed once every 8 hours.
| Passing Summary | |||||||
| Player | C | Att | Yds | TDs | INTs | YPA | |
|---|---|---|---|---|---|---|---|
| LSU | |||||||
| Joe Burrow | 402 | 527 | 5671 | 60 | 6 | 10.8 | |
| Myles Brennan | 24 | 40 | 353 | 1 | 1 | 8.8 | |
| Florida State | |||||||
| Jameis Winston | 257 | 384 | 4057 | 40 | 10 | 10.6 | |
| Jake Coker | 18 | 36 | 250 | 0 | 1 | 6.9 | |
| Sean Maguire | 13 | 21 | 116 | 2 | 2 | 5.5 | |
| Table: @SaiemGilani | Data: @CFB_Data with @cfbfastR v2.0.0 | |||||||
Passing and rushing splits (2025 onward)
The cfbd_passing_*() and cfbd_rushing_*()
families carry charting detail the season-stats endpoints above do not:
passing production split across seven pass locations, rushing across
four run directions, and play-level frames with air yards, yards after
catch and carrier attribution.
These frames are wide by construction — one
23-column passing production block repeats for every location, and the
team endpoints repeat the whole thing for offense and defense, so
cfbd_passing_teams_season() is 371 columns. Select what you
need rather than printing them whole.
tex_pass <- cfbd_passing_players_season(year = 2025, team = "Texas")
tex_pass |>
dplyr::filter(attempts > 20) |>
dplyr::select(player, attempts, completion_rate, average_depth_of_target, ppa) |>
dplyr::arrange(dplyr::desc(attempts))## ── Player season passing data from CollegeFootballData.com ─────────────────────
## ℹ Data updated: 2026-09-19 04:19:02 UTC
## # A tibble: 1 × 5
## player attempts completion_rate average_depth_of_target ppa
## <chr> <int> <dbl> <dbl> <dbl>
## 1 Arch Manning 399 0.607 8.5 0.274
Where the splits earn their keep is comparing a passer by area of the
field. The bucket columns are named
locations_<bucket>_<stat>:
buckets <- c("short_left", "short_middle", "short_right",
"deep_left", "deep_middle", "deep_right", "unknown")
tex_pass |>
dplyr::filter(attempts == max(attempts)) |>
# name the seven buckets explicitly: a wildcard on `_attempts` would also
# sweep up `locations_<bucket>_successful_attempts`, which is a different
# statistic, and silently double the rows.
dplyr::select(player, paste0("locations_", buckets, "_attempts")) |>
tidyr::pivot_longer(
-player,
names_to = "location", values_to = "attempts",
names_pattern = "locations_(.*)_attempts"
) |>
dplyr::arrange(dplyr::desc(attempts))## # A tibble: 7 × 3
## player location attempts
## <chr> <chr> <int>
## 1 Arch Manning unknown 189
## 2 Arch Manning short_right 74
## 3 Arch Manning short_left 71
## 4 Arch Manning short_middle 41
## 5 Arch Manning deep_left 11
## 6 Arch Manning deep_right 7
## 7 Arch Manning deep_middle 5
Note how large the unknown bucket is. Location is parsed
from play text, so a sizeable share of attempts never get one — which is
exactly why the next point matters.
A caveat that matters for every mean you compute from these: the
*_available columns are denominators, not
statistics. CFBD parses air yards, YAC and location out of play
text, so an attempt can be counted while those stay unknown. Divide by
the matching *_available column, not by
attempts:
tex_pass |>
dplyr::filter(attempts > 20) |>
dplyr::transmute(
player,
attempts,
air_yards_parsed = air_yards_attempts_available,
# right: the denominator CFBD actually measured
adot_correct = total_air_yards / air_yards_attempts_available,
# wrong: silently understated wherever parsing was incomplete
adot_naive = total_air_yards / attempts
)## ── Player season passing data from CollegeFootballData.com ─────────────────────
## ℹ Data updated: 2026-09-19 04:19:02 UTC
## # A tibble: 1 × 5
## player attempts air_yards_parsed adot_correct adot_naive
## <chr> <int> <int> <dbl> <dbl>
## 1 Arch Manning 399 211 8.54 4.52
Play-level frames expose the same idea as parse flags. Filter on
location_analysis_eligible (passing) or
direction_analysis_eligible (rushing) before trusting a
split column:
cfbd_passing_plays(year = 2025, week = 5, outcome = "completion") |>
dplyr::filter(location_analysis_eligible, !is.na(target)) |>
dplyr::select(offense, defense, passer, target, total_yards, ppa, success) |>
head(10)## ── Passing plays data from CollegeFootballData.com ────── cfbfastR 3.0.0.9000 ──
## ℹ Data updated: 2026-09-19 04:19:04 UTC
## # A tibble: 10 × 7
## offense defense passer target total_yards ppa success
## <chr> <chr> <chr> <chr> <int> <dbl> <lgl>
## 1 Notre Dame Arkansas C.J. Carr Eli Raridon 18 1.66 TRUE
## 2 Arkansas Notre Dame Taylen Green O'Mega Blake 8 0.910 TRUE
## 3 Notre Dame Arkansas C.J. Carr Will Pauling 22 1.89 TRUE
## 4 Arkansas Notre Dame Taylen Green O'Mega Blake 33 2.09 TRUE
## 5 Notre Dame Arkansas C.J. Carr Malachi Fields 21 2.30 TRUE
## 6 Arkansas Notre Dame Taylen Green Andreas Paaske 8 0.752 TRUE
## 7 Notre Dame Arkansas C.J. Carr Jeremiyah Love 25 2.44 TRUE
## 8 Notre Dame Arkansas C.J. Carr Jeremiyah Love 7 2.61 TRUE
## 9 Arkansas Notre Dame Taylen Green Jalen Brown 11 1.46 TRUE
## 10 Notre Dame Arkansas C.J. Carr Jordan Faison 10 0.504 TRUE
Check a column before you build on it. As of this writing CFBD
populates target on the play frame (1,976 of 2,003 week-5
completions) but has not populated play-level
air_yards or yards_after_catch at all — R
types those logical because every value is NA.
The season aggregates do carry air yards, which is why
total_air_yards above is non-zero while the play column is
empty. The same applies to rush_direction on
cfbd_rushing_plays().
Earlier seasons are not an error — these endpoints begin in 2025, and CFBD answers a 2024 request with an empty array, so the functions return a zero-row data frame rather than failing.
College Football Mapping for Stats Categories
## ── Stat categories for CollegeFootballData.com ────────── cfbfastR 3.0.0.9000 ──
## ℹ Data updated: 2026-09-19 04:19:04 UTC
## # A tibble: 38 × 1
## category
## <chr>
## 1 completionAttempts
## 2 defensiveTDs
## 3 extraPoints
## 4 fieldGoalPct
## 5 fieldGoals
## 6 firstDowns
## 7 fourthDownEff
## 8 fumblesLost
## 9 fumblesRecovered
## 10 interceptions
## # ℹ 28 more rows
Our Authors
Related SportsDataverse packages
- cfbfastR - college football
- hoopR - men’s basketball
- wehoop - women’s basketball
- baseballr - baseball
- fastRhockey - hockey
- oddsapiR - betting odds
- sportyR - playing surfaces
- sportsdataverse-py - the Python package
- sportsdataverse-R - the R meta-package