Skip to content

How the MLB numbers work

MLB

What powers the live game page, the power ratings and the postseason odds — and, at each step, what the number cannot see. Everything here is built from free data: the MLB Stats API and pitch-level Statcast.

Run expectancy (RE24)

Baseball has 24 base-out states — eight arrangements of runners times zero, one or two outs. For each one we measured, from our own play-by-play, how many runs a team goes on to score before the inning ends. That table is the backbone of everything else: it prices every play in runs. An inning starts at 0.51 expected runs; loaded with nobody out it is 2.41. The league averaged 0.51 runs an inning.

Runners0 out1 out2 out
Empty0.510.270.10
1st0.900.530.23
2nd1.130.680.33
1st & 2nd1.510.930.46
3rd1.400.950.36
1st & 3rd1.811.180.50
2nd & 3rd2.041.400.56
Loaded2.411.620.78

What it cannot see. It is a league average. It does not know who is batting next — the simulator below does.

The game simulator

The game page projects the rest of a game by playing it out thousands of times, one plate appearance at a time. Each plate appearance draws an outcome — strikeout, walk, hit by pitch, single, double, triple, home run, or an out in play — from a blend of the actual batter's and the actual pitcher's true-talent rates (Bill James's log5, applied per outcome). Those rates are each player's own history pulled toward the league average, harder for rare, luck-driven events like home runs than for strikeouts. Runners then advance on a probabilistic base-out model, and the real lineups bat in order.

Layered on top: park factors from our own home/road data (COL is the biggest run park at 1.11×), the platoon split, each team's bullpen, a small home-field term, and the pitcher's fatigue.

What it cannot see. Weather — the live simulation does not use it — or a roof's state, injuries announced after lineups post, and a manager's bullpen plans. It needs the real lineups: until they are in the feed, the game page falls back to a league-average projection.

Pitcher fatigue

From every pitch, we measure how a pitcher's fastball velocity falls as his pitch count climbs, against his own fresh-arm baseline — the league loses about 0.6 mph per 100 pitches — and how hitters do better each time they see him. Hitters' expected wOBA rises from .323 the first time through the order to .342 the third; in runs, the third time through costs about 0.54 runs per nine innings over the first. The live page reads each pitcher's velocity pitch by pitch and says how far he is below his season norm.

What it cannot see. Velocity is a proxy. A pitcher can lose command without losing a tick, and the curve cannot see that.

Pitch quality (Stuff+)

Our own pitch-quality grade, so we are not renting anyone else's. A gradient-boosted model predicts each pitch's run value from its physical shape only: velocity, spin, movement, extension, release point and arm angle — with breaking balls measured against the pitcher's own fastball, and left-handers mirrored so both hands share one frame. A pitcher's grade is the average over his pitches, scaled so 100 is league average and 10 points is one standard deviation. 1,145 pitchers are graded; each pitch type needs 25 thrown to get its own grade. In our fit the grade tracks a pitcher's strikeout rate at r = 0.59.

What it cannot see. It is deliberately location-free: it does not see where a pitch was thrown, the count, sequencing, or what happened to the ball. Great stuff with poor command still gives up runs — which is why the leaderboard puts ERA beside it.

Live win probability

For every combination of inning, half, outs, runners and score margin, we counted how often the home team went on to win in our play-by-play. Sparse states fall back to inning and margin, then to margin alone. Before the first pitch the home side won 51.7% of the time. A separate leverage table says how much the current state can swing the game.

What it cannot see. The table is a league average — it does not know the two teams. The simulator above is what knows the matchup.

Power ratings

Power ratings are Pythagenpat: the win percentage a team's runs scored and allowed should have produced, RSx / (RSx + RAx), with the exponent x = ((RS + RA) / games)0.287 rising with the run environment. The gap between that and the real record is shown as exactly that — a gap, not a prediction that it closes.

What it cannot see. Raw runs: no adjustment for schedule, park, or roster changes during the season. Blowouts count in full.

Postseason odds

The postseason odds play the bracket out 20,000 times. A team's strength is its Pythagorean win% pulled toward .500 by 69 games of average play; one game is log5 of the two strengths plus home field (51.7% for the home side between equal teams). Each series is played game by game in its real format — best of 3 all at the higher seed, best of 5 in a 2-2-1, best of 7 in a 2-3-2 — and a series already under way starts from its real score. The random seed is fixed per season, so the same inputs print the same numbers.

What it cannot see. This is a run-differential model, not the game simulator. It does not know who starts each game, and October rotations shrink to three or four starters; it treats a team's whole season as one strength.

Data: MLB Stats API (schedule, live feed, standings, stats) and Baseball Savant Statcast (pitch-level). Model code: model/mlb in the Gridpex repository.