Skip to content

CORE Ratings

Analysis

How good each team is per play, once you account for who they played and the situations they were in. Zero is an average team; higher is better. points per 100 qualifying plays, above average. Built from 94,257 plays across 773 games.

Offense against defense

Offense across, defense up. Best teams are top right. Both numbers are adjusted for who each team played, so a soft schedule does not flatter anyone.

-20+0+20-20+0+20CORE offense (points per 100 plays)← worsebetter →CORE defense (vs average)← worsebetter →
130 teams · dashed lines mark the league medianr = -0.43(r² = 0.18)

Where efficiency and the poll disagree

The gap between a team’s CORE rank and its AP rank. This is the column worth arguing about — it is where one of the two measures is about to be shown up.

#TeamCOREOffenseDefense
6613Oklahoma+27.0+25.6#4-1.4#57
201633Baylor+14.8+0.8#58-14.0#12
221633Texas+14.7+12.1#10-2.6#50
231736Iowa State+14.2+9.2#20-5.0#41
332049TCU+9.0-2.7#82-11.8#16
362454Oklahoma State+7.6+5.0#35-2.6#49
463169Kansas State+4.0+1.8#49-2.2#52
745499Texas Tech-3.7+0.8#57+4.6#88
8056104West Virginia-4.6-5.2#95-0.6#62
10071111Kansas-9.6-2.2#79+7.5#103

Four ways to measure a team, and where each fails

This site publishes all four. That is not indecision — each answers a different question, and where they disagree is the information, not the noise.

Results · Elo, SRS

Measures: who beat whom, and by how much

Fails when: It cannot tell a good team from a lucky one. Elo moves on outcomes, and a season is twelve of them — a team that won three one-score games is rated as though it earned all three.

Efficiency · CORE, SP+

Measures: how good each play was, adjusted for situation and opponent

Fails when: It is indifferent to winning. A team can be the most efficient in the country and lose to the second-most because a punt was blocked, and the rating will not notice, because the rating is not about that.

The market · closing spreads

Measures: everything anyone knows, priced

Fails when: It is not published as a rating and it has no memory. It is the sharpest number available on any single game and it will not tell you who the fourth-best team in the country is.

Human polls · AP, coaches

Measures: reputation, weighted by results

Fails when: Inertia and brand. A preseason ranking survives contact with reality for weeks, and a team that starts unranked has to be much better than one that started at 12 to pass it.

How CORE is computed

  1. Remove the situation. For each play, predict its EPA from the game state alone — period, seconds left, score margin, down, distance, field position, home/away/neutral, pass or rush. No team identity is in those features, so what the model learns is what an average offense does from that spot. The residual is the part the situation does not explain.
  2. Remove the opponent. Offenses and defenses form one connected network across a season. Fit both sides at once — residual ~ offense(team) + defense(opponent) — by ridge regression over team dummies. Ridge rather than plain least squares because in September the schedule graph is barely connected, and an unregularised fit hands a huge rating to whoever beat up an overmatched opener.
  3. Report in points per 100 plays. Offense is points created above average; defense is points allowed above average, so lower is better; overall is the difference.

Excluded: overtime, garbage time, FBS vs FCS, special teams, non-scrimmage plays.

How sure is any of this? The small range under each rank is where a team would sit if its rating were one standard error better, or worse, than the estimate — about a two-in-three chance of covering the truth. It is wide on purpose: a season is a dozen games, and most of this table is one good afternoon away from a different order. Those bars are checked rather than claimed: fitting 2025’s odd and even weeks separately and standardising every team’s difference by its own two errors gives a spread of 0.93 on offense and 0.94 on defense, where 1.00 is exactly right-sized. It covers sampling error only — not the shrinkage the ridge applies on purpose, and not a team that changed in November.

What this is not. Not a prediction, not a forecast, not a projection of next week. It is a measurement of what happened per play, adjusted. The model that forecasts games is a different object and lives on the model record, where its actual edge is stated with the same bluntness.