Skip to content

College Football Fourth-Down Study: 22,218 Decisions, 2015–2025 — The Goal-Line Mistake Costing 272 Points a Year

The fourth down nobody argues about

Twenty-two thousand decisions across eleven seasons. College football spent a decade learning to go for it — and the most expensive mistake left on the field is a twenty-one-yard field goal that splits the uprights.

The Numbers Room
Ratings & power ·
16 min read

There is a version of every football argument that happens on Sunday and a version that happens in a spreadsheet, and they are almost never about the same play. The Sunday argument is about the fourth-and-short that failed — the one with the sideline reaction shot, the one that gets a coordinator's name trending. The spreadsheet argument is about a field goal that went straight through the uprights and drew no reaction from anybody.

We built an expected-points model on 1.8 million plays and graded every fourth-down decision in every FBS game between 2015 and 2025 — 148,495 of them across 9,471 games. Coaches matched the model on 81.1%. What follows is about the other 18.7%: how the sport's thinking changed over a decade, which coaches changed with it, where the remaining money is, and — last, and least comfortably — whether any of it is worth as much as we all assume.

The shape of 148,495 decisions

Every fourth down in all 9,471 FBS games, 2015–2025, by what the coach chose.

Puntedand the model agreed 93% of the time71%
Went for itthe model agreed 47% of the time24%
Kickedthe model agreed 75% of the time20%

The model disagreed with 18.7% of them. That disagreement is what this piece is about.

The decade college football changed its mind

Sort every disagreement into two kinds. Either a coach went for it when kicking or punting was worth more — call that over-aggression — or he did the reverse. In 2015 those two errors were close to evenly split. They are not close now.

Share of fourth-down mistakes that were over-aggression

50% would mean a coach was equally likely to err bold as timid.

201551.6%
201657.6%
201754.9%
201855.8%
201956.8%
202060.9%
202163.3%
202265.4%
202364.8%
202464.0%
202567.0%

+1.60 percentage points a year. ~2,600 graded mistakes per season, full corpus.

That is the analytics movement, visible in the error term. But the aggregate hides the more interesting story, which is *who moved*. In 2015 the Group of Five went for it more often than the Power conferences — smaller programs, less to lose, less scrutiny. A decade later that has reversed.

How often teams went for it on fourth down
Power conferencesGroup of Five
2015
20.1%
23.7%
2019
20.7%
23.1%
2022
23.9%
25.8%
2025
28.5%
26.9%

Power teams started three points behind and finished nearly two points ahead. The gap closed in 2022 and inverted after.

The programs with the most to lose were the last to move and are now the furthest along (+1.16 percentage points a year against the Group of Five's +0.92). That is what adoption of an idea looks like when the idea is correct and the early adopters are the people who could least afford a bad Saturday: the resource-rich hold out, then overshoot.

Kirk Ferentz and Jeff Monken are playing different sports

Aggregates are comfortable. Names are not. Here is the same measurement at the level of the man with the headset — every coach with at least 120 graded decisions, scored on how much more often he went for it than the model said to.

Went for it more often than the model said to

Percentage points. Positive = more aggressive than expected points. Min. 120 decisions.

Jeff MonkenArmy+14.6
Dave ArandaBaylor+14.1
Lane KiffinFAU / Ole Miss+13.3
Scot LoefflerBowling Green+12.5
Lance LeipoldBuffalo / Kansas+11.5
Sonny DykesSMU / TCU+11.3
Gary PattersonTCU−1.4
Brian KellyNotre Dame / LSU−4.0
Paul ChrystWisconsin−6.3
Kirk FerentzIowa−8.8

48 qualifying coaches. The spread from Ferentz to Monken is 24 points of go-rate — four and a half standard deviations of coaching style.

Two things are worth pulling out of that chart, and neither is the one it looks like it is saying.

The first is that **aggression and error are not the same measurement.** Jeff Monken is the most aggressive coach in the data and loses *less* expected value per game than the median coach, because Army's triple option converts fourth-and-2 at a rate the model's league-wide curve does not credit it for. A leaderboard that reads aggression as error is measuring style and calling it skill. Lane Kiffin is third in aggression and among the ten most expensive; Monken is first in aggression and comfortably in the better half. The correlation people assume is there is not.

The second is the sheer width of the distribution. Across 48 coaches the standard deviation of this gap is 5.2 percentage points, and the range is 24 — from Ferentz kicking and punting nine points more often than the model wants to Monken going nearly fifteen points more. These men are coaching the same game under the same rules with access to the same public research, and on this one axis they behave as differently as it is possible to behave.

The most misplayed spot on a football field is the opponent's four-yard line

Now to where the money actually is. Our model says something specific and slightly counterintuitive about short field position: it wants you to **kick almost everywhere inside the 12**. A 22-yard field goal is close to free, going for it from the 8 gains little, and the expected-points arithmetic is not close. With one exception.

Fourth down inside the 6 — expected points of each choice
Go for itKick the field goal
4th & goal at the 2
4.80
3.36
4th & 1 at the 2
3.19
3.36
4th & goal at the 3
4.46
3.34
4th & 2 at the 3
2.99
3.34
4th & goal at the 4
4.26
3.31
4th & 3 at the 4
2.76
3.31

The only situations the model wants you to go are fourth-and-GOAL. One yard further back and it flips.

Fourth-and-goal from the two is worth **1.44 expected points more** if you go than if you kick. Fourth-and-1 from the two — the same field position, a yard of cushion behind the goal line — flips to the kick. The model is not asking for bravery. It is asking for a distinction, and the distinction is whether converting scores a touchdown or merely continues a drive.

Coaches do not make that distinction. Here is how often they actually kicked from each yard line on fourth-and-five-or-less, against how often the model wanted them to.

Kicked the field goal — what coaches did vs what the model wanted
Coaches kickedModel wanted a kick
At the 1
7.30%
0%
At the 2
33.3%
13.5%
At the 3
53.5%
23.6%
At the 4
67.5%
29.2%
At the 5
68.6%
40.1%

The gap peaks at the four-yard line: coaches kick 67.5% of the time where the model wants 29.2%.

**At the four-yard line the disagreement is 38 percentage points**, the widest anywhere on the field. Between the two and the five, coaches kick roughly twice as often as the arithmetic supports — and the reason is not hard to guess. Three points from the four-yard line feels like a drive that worked. Nobody in the stadium boos it. No panel discusses it.

The arithmetic underneath is almost embarrassingly simple. From inside the ten, on fourth-and-five or less:

Inside the 10, fourth-and-5-or-less: what each choice returns

Expected points, from what actually happened in 1,180 such decisions.

Go for it56.2% convert x ~7 points · n=6073.93 pts
Kick92.8% make x 3 points · n=5732.78 pts

And failing to convert leaves the opponent pinned inside their own 10 — which the raw comparison does not even credit.

Ninety-three percent of those kicks are good. Fifty-six percent of the conversions succeed. And the 56% is still worth more, because seven is a lot more than three and because a failed conversion hands the opponent the ball on their own goal line rather than at the 25.

Where the 3,179 points went

Across the full corpus, 6,373 decisions were kicks the model wanted to be conversion attempts, and they cost 3,179 expected points. That damage is not spread evenly — it is almost entirely a goal-line phenomenon.

Expected points lost to kicking instead of going, by field position

912 decisions, 2015–2025.

Inside the 52,081 kicks · 95.1% good−1,895
Beyond the 54,292 kicks−1,284

2,081 decisions inside the 5 account for 60% of the damage from a third of the calls.

Across the eleven seasons this one error — kicking a short field goal that should have been a conversion attempt — cost the sport 3,179 expected points — about 290 a season. Not one team. The sport.

The obvious explanation is lead protection: a coach with a lead takes the sure three. We tested it, and it is not what is happening. Teams leading by one score kick 56.6% of the time on short yardage inside the 20 where the model wants 53.1% — a gap of three and a half points. Teams leading by nine or more kick 55.6% where the model wants 55.2%: no gap at all. The error is not coaches protecting leads. It is coaches taking points that are on the table in front of them, regardless of the situation.

There is a genuinely encouraging footnote, though. This particular error is being fixed. In 2015 coaches kicked 3.6 percentage points more often than the model wanted in short-yardage situations inside the 20. By 2025 they kick 10.5 points *less* often than it wants — a swing of fourteen points in a decade, moving at 1.26 points a year. The goal-line error is the last redoubt of a habit that is otherwise in retreat.

The most expensive chip shot in the sport belongs to Nick Saban

Rank coaches by expected points lost specifically to kicking when the model wanted a conversion attempt, per game, and the name at the top is the most successful coach of the modern era.

Expected points lost per game to kicking when the model wanted a go

This bucket only — not overall decision quality. Minimum 12 such kicks.

Nick SabanAlabama · 14 kicks−0.45
Mike GundyOklahoma State · 19 kicks−0.38
Dave DoerenNC State · 13 kicks−0.32
Matt CampbellIowa State · 12 kicks−0.25

Alabama was a heavy favorite in nearly all of these, which is exactly the situation the model understands least.

Read that carefully, because the obvious reading is wrong. This is not evidence that Nick Saban was bad at fourth downs. It is the clearest illustration in the dataset of what an expected-points model cannot see.

Alabama was a three-touchdown favorite in most of these situations. A three-touchdown favorite taking three guaranteed points instead of a 56% chance at seven is trading expected points for **certainty**, and certainty is exactly what a team with a large edge should be buying. Variance is the underdog's friend and the favorite's enemy. The model maximizes average points and has no opinion about which team wants the average.

An expected-points model does not know what a team is trying to do.

It maximizes the mean. A heavy favorite wants certainty and a heavy underdog wants the tail, and neither of those is the mean. A great deal of what registers as coaching error is that gap.

This is also why the aggression gap in the first chart sits almost entirely with underdogs. A team favoured by 14 goes for it 21.9% of the time against a model that wants 22.7% — within a point. A team getting 7 to 14 goes for it 24.1% against a model that wants 16.6%. The underdogs are not being reckless. They are buying variance, which is the correct purchase, and being marked down for it by an instrument that prices only the average.

Two calls, two verdicts

Put a real decision under each end of that idea.

**Penn State at Ohio State, 2017. Fourth-and-15 from their own 36, down one, 1:22 on the clock. James Franklin went for it.** Ninety-five percent of coaches in that exact state punt, and our model agrees emphatically — it scores punting 3.16 expected points better, the harshest verdict it renders on any high-leverage call in eleven seasons. Penn State gained nothing. Ohio State ran out the clock and won 39–38.

**Our verdict: nobody can tell you, and that is the point.** The case for Franklin is real — punting there does not produce a punt and a subsequent possession, it produces Ohio State kneeling. Expected points assumes the game continues and that a possession carries generic future value. With 82 seconds left and one point down, a possession is worth precisely the chance to score with it, and an expected-points model does not know the game is about to end. But the case against him is also real: with two timeouts, a punt buys a plausible path to the ball back. This is exactly the kind of call the models cannot separate — see the last section — and the confident verdicts delivered on it that Monday, in both directions, were claiming a precision nobody had.

**Kansas vs TCU, 2016. Fourth-and-4 at the TCU 4, third quarter, Kansas ahead by two. David Beaty kicked.** It was good. Kansas led by five. No segment was produced about it and no reasonable person would have objected — a made field goal from the four looks like a drive that worked.

**Our verdict: the model is right, and this is the one that mattered.** Kansas lost 24–23. The play that would have filled the postgame — a hopeless fourth-and-22 heave from their own 31 with forty seconds left — is the call our model punishes hardest and understands least, for exactly the reason it misjudged Franklin. The call that actually decided the game drew no comment at all.

That pattern is not a coincidence. Controversy concentrates in the fourth quarter of one-score games, which is precisely where score and clock overwhelm the assumptions an expected-points model rests on. The decisions where such a model is most reliable — second quarter, ordinary field position, no clock pressure — generate no argument whatsoever, because nothing dramatic happens immediately afterward. The sport has built its entire fourth-down discourse on the subset of decisions its instruments handle worst.

Do the aggressive teams win? The answer looks obvious and is wrong

The natural next question, and the one every coach's critics reach for: do the teams that go for it too often actually lose more? Take every team-season in the sample, measure how much more often it went for it than the model wanted, and line that up against how often it won.

Adjusted win rate by quartile of over-aggression

Win% after removing what the betting market expected of that team. 105 team-seasons.

Least aggressivewent −7.0% vs model · won 63.0%+7.6 pts
Third quartilewent +2.0% · won 50.9%+2.0 pts
Second quartilewent +7.5% · won 53.7%+0.1 pts
Most aggressivewent +15.6% · won 40.4%−9.9 pts

A 17.5-point spread, cleanly monotone. Correlation −0.31. It is also almost entirely an illusion.

Seventeen and a half points of adjusted win rate between the calmest quartile and the boldest, moving monotonically. If you wanted a chart to prove that analytics-brained aggression is losing football games, that is the chart, and you will find versions of it in circulation.

It is backwards. **Losing causes aggression.** A team down three scores in the fourth quarter faces fourth-and-6 from midfield and goes, because punting is conceding. A team having a bad season faces that situation ten times. The aggression is not producing the losing; the losing is producing the aggression, and the correlation is picking up the arrow pointing the wrong way.

The test is straightforward. Restrict the aggression measurement to decisions where the scoreboard cannot yet have driven the behavior — first-half calls, or calls taken while the game is within one score — and see whether the relationship survives.

Correlation between over-aggression and adjusted win rate

The same measurement, restricted to decisions the score cannot have caused.

All decisionsthe headline correlation−0.31
Within one scoreany quarter−0.12
First half onlyscore has not diverged−0.09
First half, one scorethe cleanest read−0.07

Once desperation is excluded, four-fifths of the effect disappears. The raw number was measuring which teams were behind.

From −0.31 to −0.07. There is no meaningful relationship between how aggressive a coach is on fourth down and whether his team wins — not once you stop letting the scoreboard contaminate the measurement of the behavior. Anyone showing you the first chart without the second is showing you a scoreboard, not a coaching philosophy.

Why this one habit survived when the others died

Everything above leaves a puzzle. The sport corrected its punting instincts in nine years. It did not correct the goal-line field goal. Same coaches, same research, same decade — so why did one lesson take and the other one not?

The answer is in what each mistake feels like the moment after you make it. Take every call the model disliked and ask how often it nonetheless outperformed its own expectation — how often, in other words, the coach walked off the field feeling vindicated.

How often each kind of mistake still felt like the right call

Share of wrong decisions that beat their own expected value. Punts excluded — their outcome is already an average.

Kicked when the model said gon=6,618 · CI 74.2–76.375.3%
Went when the model said kickn=16,269 · CI 47.5–49.148.3%

A wrong kick is rewarded three times as often as a wrong gamble. That asymmetry is the whole story.

Seventy-five percent against forty-eight. And it gets starker the closer you get to the end zone.

The reinforcement schedule, by where the kick was taken

Share of these kicks that were good.

Inside the 5n=2,081 · CI 94.1–96.095.1%
Beyond the 5n=4,29266.9%

Inside the 5, a decision the model rates as a mistake succeeds 93.6% of the time.

**A wrong decision that succeeds 95% of the time is not a habit anybody was ever going to break.** Every failed fourth-down gamble produces a stadium groan, a talk-radio segment and a coach explaining himself; roughly half of them fail, so the feedback arrives constantly and the profession adjusted. The goal-line field goal produces three points and a helmet tap nineteen times out of twenty (95.1%, on 2,081 of them). There is no correction signal, because from the sideline there is nothing to correct.

Compare the gambles by distance and the same logic runs the other way: going for it on fourth-and-1 or 2 works 66.9% of the time, on fourth-and-6-to-10 it works 42.0%, and on fourth-and-11-plus it works 25.1%. The aggressive mistakes are punished often enough, and publicly enough, to teach. The conservative one is rewarded almost every time and teaches nothing.

Going for it: how often it worked, by distance

Conversion rate on fourth downs where the model preferred a kick or a punt.

4th-and-1 or 2n=73566.9%
4th-and-3 to 5n=50850.8%
4th-and-6 to 10n=83342.0%
4th-and-11+n=38625.1%

The gambles fail often and loudly. That is why the profession learned from them.

This is, in the end, the most useful thing in the data. It is not that coaches are irrational. It is that the two mistakes carry wildly different feedback, and human beings — including very good ones, with analytics departments — learn from the feedback they get rather than the arithmetic they could look up. The errors that hurt visibly got fixed. The error that pays out nineteen times in twenty is still sitting there, at the opponent's four-yard line, worth roughly 290 points a season to whoever notices first.

How much is any of this actually worth?

A piece like this one is under an obligation to answer the question it has been dodging: does fourth-down decision-making win football games? We took every team-season in the sample, measured its fourth-down expected points per game, and compared it against how often that team won.

Win rate by quartile of fourth-down decision quality

Adjusted win rate — win% after removing what the betting market expected of that team.

Worst quartile−1.81 EP per game−5.4 pts
Third−1.08 EP per game−1.0 pts
Second−0.76 EP per game+2.6 pts
Best quartile−0.42 EP per game+4.0 pts

105 team-seasons. Correlation between decision quality and win rate: +0.08.

The gradient is real and it is in the right direction — about nine points of adjusted win rate between the best and worst quartiles, which over a twelve-game season is roughly one win. But the correlation is **+0.08**, which means fourth-down decision quality explains well under one percent of who wins. It is a real effect and a small one, and anybody selling it as larger is selling something.

That is the honest scale of the thing. Fourth downs are worth about a tenth of a win a season to an average team, and up to a full win at the extremes. They are not why Georgia beats Vanderbilt. They are the margin available to a coach who has already done everything else — which is exactly why the coaches at the top of the sport have spent a decade chasing it.

How certain is any of this? Less than anyone admits

One last result, and it is the one that should change how you read every fourth-down take you see this season — including ours.

We built a second model to check this work: a win-probability model rather than an expected-points one, trained on score, clock, field position, down, distance and timeouts across the same corpus. Held out on 2024–25 it is well calibrated — AUC 0.884, Brier 0.139, worst-decile error 0.015. Then we refit it on the same data several times and asked it the same goal-line questions.

Size of the recommended 'go zone' — same model, same data, four refits

Goal-line grid, trailing by 14. 30 squares in total.

Refit 126 of 30
Refit 25 of 30
Refit 35 of 30
Refit 420 of 30

Nothing changed between these four runs except ordinary fitting noise.

The recommendation flips because the underlying edge is tiny — on the order of two hundredths of a win probability between going and kicking — so ordinary fit noise decides the answer. A single model reports “GO” with total confidence on a decision it cannot actually resolve.

Running an ensemble of ten models on bootstrap resamples of the games, and refusing to call any situation where they agree less than 80% of the time, **twenty of the thirty goal-line squares for a trailing team are coin flips.** This is the argument Brill, Yurko and Wyner make in *Analytics, have some humility* (2023), and we reproduced it accidentally while trying to check our own work.

A model can be beautifully calibrated and still be unable to tell you what to do.

Calibration is a statement about averages across many situations. A decision is a comparison between two specific ones, and a model can get the first right while having no resolving power on the second.

Which is why this piece grades decisions in expected points and says so, rather than dressing the numbers up as verdicts. Expected points is a transparent accounting of what a choice is worth on the scoreboard. It is not a claim about who should have won, and anyone — including us — telling you a fourth-down call was definitively wrong is claiming a precision these models do not have.

What to take from this

  • The sport learned. Fourth-down errors went from an even split to two-thirds over-aggression in nine years, and the Power conferences — the last to move — are now the most aggressive of all.
  • Coaching styles remain enormously wide. Twenty-four percentage points of go-rate separate Kirk Ferentz from Jeff Monken, and aggression tells you almost nothing about whether a coach is good at this.
  • The unfixed money is on the goal line. Coaches kick at the four-yard line 67.5% of the time against a model that wants 29.2%, and 59% of the entire cost of this error sits inside the five.
  • And the whole thing is worth about one win a season at the extremes, which is both less than the argument suggests and more than nothing.

The last thought is the one we would keep. An expected-points model is a measuring instrument, not a referee. It is least reliable in the closing minutes of a one-score game — which is where every argument happens — and most reliable on a second-quarter fourth-and-goal from the three, where nobody is arguing at all. The best available use of a model like this is not to settle the fight everyone is already having. It is to point at the play nobody noticed.

Discussion

Weigh in on the analysis — the best takes rise to the top.

0 Replies

Sign in to join the discussion.