How the Gridpex model works
We built this to be defensible, not to impress. Every number on the site comes from standard, named football analytics computed on real data — never an invented formula. Here's exactly how it works, and what it can and can't do.
Model as of Sep 1 · through week 1 · drive simulation
See the graded record →The data
Eleven seasons of college football (2014–2024) from CollegeFootballData: every play, drive, game, and box score, plus ratings (SP+, FPI, SRS, Elo), betting lines, weather, travel, talent and returning production. Roughly 8,300 FBS games.
Opponent-adjusted metrics, not raw stats
Raw stats lie — 6 yards a play against a bad defense isn't 6 yards against a good one. We use EPA (expected points added) and success rate as the efficiency base, opponent-adjust them via ridge regression as of each week (so a Week 8 number only knows Weeks 1–7), and strip garbage time. Player value uses WEPA — the same EPA, opponent-adjusted — and a PAAR-style value over a replacement-level backup. These are the field's standards, not ours.
The prediction model
Every projected score, spread and win probability on the site comes from a possession-level Monte-Carlo. We fit each team's offensive and defensive efficiency from the drives it has actually run, opponent-adjusted, then simulate the game drive by drive several thousand times. The projected margin is the average outcome of those simulations and the win probability is simply how often each side wins them — not a formula applied to a rating gap.
Before a team has played, there are no drives to fit, so a preseason prior stands in: returning production, recruiting and roster talent, and last season's form carried forward. That prior is weaker than the in-season model and we say so — a projection built on it is labelled preseason wherever it appears, and it hands over to the drive model as soon as a team has banked enough possessions. FBS-vs-FCS games, where one side has no comparable data at all, are priced by a separate calibrated tier model.
drive ratings = opponent-adjusted points per drive, fit on every drive played so far
margin, total = mean of 8,000 simulated games
win_prob = share of simulations won, recalibrated against realised outcomes
live constants: HFA=4.5 WP_CAL_A=0.0476 WP_CAL_B=0.1074 COLD_GAIN=2.1 MARGIN_GAIN=2.6
A blend of Elo and SP+ ratings remains in the code as a fallback, used only for a game the simulation has not priced. Every surface labels which engine produced the number it is showing, and the public record is broken out by engine so a change of engine cannot quietly flatter the total.
CORE — the efficiency rating
CORE answers a narrower question than a power rating: how good is a team per play, once you remove the situations it happened to face and the opponents it happened to draw? Two stages, both stated so you can argue with either.
Remove the situation. For each play we predict its EPA from the game state alone — period, seconds left, down, distance, field position, home/away/neutral, pass or rush. No team identity is anywhere in those features, so what the model learns is what an average offense does from that spot. The residual is the part the situation does not explain.
Not the score. Score margin used to be one of those features and is deliberately gone. Margin is not a neutral situation: it is caused by team quality, and much of it is caused by the defense. Conditioning on it taught the model to expect more from a leading team’s plays and then charged the shortfall to that team’s offense — so a side carried by its defense was scored as though it could not move the ball. Across all eleven seasons the correlation between a team’s offense and defense ratings ran the wrong way, and the adjustment agreed with SP+ less than doing nothing at all did. Removing it flips that correlation to the sign talent implies and raises the rating’s agreement with itself, split half against half, from 0.56 to 0.74.
Remove the opponent. Offenses and defenses form one connected network across a season. We fit both sides at once — residual ~ offense(team) + defense(opponent) — by ridge regression over team dummies. Ridge rather than plain least squares because in September the schedule graph is barely connected, and an unregularised fit hands an enormous rating to whoever beat up an overmatched opener.
Reported in points per 100 plays above average. Overtime, garbage time, FBS-vs-FCS games, special teams and non-scrimmage plays are excluded. It is a measurement, not a forecast — the model that predicts games is the drive-sim above, and the two are deliberately separate objects.
And how sure it is. Ridge gives an analytic standard error, so every team carries one — and /ratings prints the range of places a team could plausibly occupy rather than a rank pretending to be exact. It is usually wide. On a normal board the top five overlap each other, and a mid-table team spans thirty places; in a season with a dominant leader the top of the range collapses to two or three. That contrast is the honest content of the number.
The bars are checked rather than asserted: fitting a season’s odd and even weeks separately and standardising each team’s difference by its own two errors should give a spread of 1.00 if they are right-sized, and it gives 0.93 and 0.94. ⚠ It is sampling error only — it does not include the shrinkage the ridge applies deliberately, and it cannot know a team changed in November.
Pressure signature — and what it deliberately does not claim
College football has no public blitz rate, no public pressure rate and no public coverage charting. We checked directly rather than assuming: the play-level QB Hurry field in the data we buy is empty in every season, and the season-level version is a scorekeeper’s judgement that a quarter of schools never record at all. The companies that do chart it price that data for professional teams.
Sacks and tackles for loss are credited to named players, though, and rosters give positions. A safety recording a sack is a blitz almost by definition; a defensive end recording one usually is not. So the position mix of a defense’s sacks is a scheme signature, and we compute it for every FBS team from 2016. It is stable year to year (r = 0.51, higher than havoc’s own 0.36), and it agrees at r = 0.93 with the same measure built from a completely independent feed.
It is not a blitz rate. A five-man rush that never gets home leaves no trace, so two defenses that blitz identically will differ here if one has better rushers. It measures where pressure lands, not where it is sent from.
And it is a style, not a grade. Across 1,193 team-seasons, the off-ball share of a defense’s sacks correlates with its opponent-adjusted quality at r = −0.021, 95% interval [−0.078, +0.036]. That is a measured null rather than an untested one: a full standard-deviation swing moves a defense by at most a point a game. Getting home with four is not better than sending five, and we never rank teams on it. What does predict defensive quality is how much a defense disrupts at all — havoc rate, at r = +0.56 — which is why that number sits beside the split rather than inside it.
Drives, and why not yards per game
A team gets ten to fourteen possessions in a football game, and the only question the scoreboard asks is what it does with each one. Yards per game is a rate on the wrong denominator: it rewards pace, which is a choice, not a quality. Drive efficiency reports points per drive, starting field position, share of the available yards gained, scoring opportunities (drives reaching the opponent's 40), what a team does once it gets there, and three-and-out rate — for both sides of the ball.
The hard part is deciding what counts. A drive that starts with 24 seconds left and ends in a kneel is a real row in the data and a fake possession; left in, it drags every team's points per drive down in proportion to how often their games were already decided. We drop end-of-half and end-of-game possessions, anything starting inside the last 40 seconds of a half, and kickoff/uncategorised rows.
Play-calling, after the situation is controlled for
“The most pass-happy offense in the country” is nearly always a statement about a schedule and a scoreboard. Teams that trail throw, teams with a lead run, and teams facing third-and-nine throw because it is third-and-nine. Pass rate over expected fits the league's pass rate on the game state alone — with no team identity in the features — and reports the difference. That is a statement about the staff.
The same construction on fourth down gives two numbers that disagree constantly, and both ship: how often a staff goes for it relative to its peers, and how often relative to what the expected-points model says is correct. Across 117,697 fourth downs, coaches go 21.4% of the time and the model says 34.3%.
The fourth-down curves are measured, not borrowed
Our fourth-down model prices go, kick and punt in expected points. Until 2026-09-01 all four curves behind it were published-shape approximations with anchor points typed in by hand, and the module said so. We have 1.7M plays; the curves are measurable, so we measured them — an expected-points model fit on 1.4M scrimmage plays from 2015-2025, and the conversion, field-goal and punt curves read off the play that follows each attempt, so returns, touchbacks and penalties are inside the number rather than modelled.
The comparison is worth publishing because the old curves were wrong in a direction that mattered. Expected points sat ~0.3–0.46 too high through the middle of the field, and the field-goal curve read 80% from 40 yards against a measured 71.6%, and 58% from 50 against a measured 52.3%. Eight points of make probability at 40–50 yards is exactly the band where go-versus-kick is a live decision. (The conversion curve, for what it is worth, was already right: 4th-and-1 at 68% against a measured 68.8%.)
Comparing teams, and comparing leagues
Every metric on the team explorer comes from one frame with one definition per metric, so “explosiveness” means the same thing on both axes of a scatter. Where a metric exists both raw and opponent-adjusted, both ship and both are labeled.
⛔ Raw metrics cannot be compared across conferences. A league that plays itself twelve times has its own schedule strength baked into every raw number, which is the single most common way conference comparisons go wrong. That is what CORE is for, and the comparison page says so on the page. It also draws every team rather than a bar per league: a conference mean is dragged by four teams, and the question “compare the conferences” is really about how far two distributions overlap.
How we validate it
Everything is walk-forward backtested: the model is only ever scored on games it was trained before, never on data it has seen. A Week 8 projection knows Weeks 1–7 and nothing else — no future results, no closing line, no hindsight.
That discipline is the whole point. It is easy to build a model that looks brilliant on games it has already seen; the only test that matters is the one it takes cold. When an internal check turns up a number inflated by leakage, we fix the leak rather than publish the number.
What we don't claim
No model is a crystal ball, and anyone selling you 70%+ winners is lying. Our projections are one well-built input, not a guarantee — treat them that way. We never call an EPA-based value a “grade” (we can't source charted grades, snaps, routes or air yards), and we label usage as play share, not snaps.
For player availability we serve our own stored copy of the feed and ingest the official SEC / Big Ten / ACC / Big 12 / CFP availability reports as they publish each game week (ESPN's aggregation in the meantime). We treat it as context, not a betting signal — the market prices known absences efficiently, so we show it to inform you, not to sell you an edge.
Questions about the method, or want to see a number defended? Read how Gridpex works and makes money.
