Reference · MLB

MLB: what the numbers mean

A starting pitcher, a park and a lineup card, every day for six months.

What a cell is

Every cell on the for/allowed grids is the same kind of number: fantasy points per game, at one position, under one scoring system . Nothing on the grid is a raw box-score statistic, and nothing on it is a projection. It is what actually happened, converted into the points a MLB lineup would have scored for it.

The left half of the grid is what a team's own players did ( For ). The right half is what the players opposite them did ( Allowed ). A defence that gives up a lot to tight ends has a high Allowed figure at TE — that is the number everyone is actually looking for, and it is why the grids exist.

Why the number is not a plain average

A plain season average has two problems. Early in a season it is built on three games and jumps around; late in a season it counts a game from October the same as a game from last night.

Both are fixed the same way. The window decides how much each game counts: the season (every game equal), the last five or ten, or an exponentially weighted average where a game's weight halves every few games. The control bar above every grid chooses it, and the n column tells you how many games of information the window actually left you with.

Shrinkage handles the other half. A figure built on four games is pulled toward a prior — last season's figure for the same team and position, blended with the league average and penalised when the coaching staff turned over — in proportion to how little evidence there is. With k games of prior weight and n games observed, the number you see is `(n × observed + k × prior) / (n + k)`. By midseason the prior has almost no effect. In week two it is most of the number, which is correct: in week two you do not know anything yet.

Turn shrinkage off in the control bar to see the raw window mean. The difference between the two is the honest measure of how much of the figure is evidence.

What the colours and the ranks mean

Colour is per column, on that column's own range, red through white to green. Green is always the good end for the team in the row — which means the polarity flips between the For half and the Allowed half, and between a points column and a points-conceded column. The header says which way each one runs.

A rank is that team's place among the league's teams on that exact figure, on the same window, with the same shrinkage. Ranks are recomputed with the grid, so changing the window changes the ranks. Rank 1 is the largest number: first in For is the best offence at that position; first in Allowed is the softest defence.

What is specific to MLB

The probable pitcher is the matchup. An opposing starter changes a lineup's expected output more than any defensive quality does, and the grid's Allowed figure by position is a team-season average that does not know who is pitching tonight. Read it beside the probable, not instead of him.

Park factors are real and are not in the grid. Run environment varies by 20% or more between parks. The model applies park factors; the for/allowed grids are raw.

The lineup card lands two to four hours out. Until it does, the eighth hitter and whether your man is in at all are guesses. Lineup freshness is tracked on the data-health page because a stale card is worse than no card.

There is no key-player cohort for pitchers. When the probable changes, the change is shown and nothing is adjusted, because the with/without evidence for a specific pitcher is too thin to shrink honestly. Every other league adjusts; MLB deliberately does not.

Positions are the fantasy positions — P, C, 1B, 2B, 3B, SS and OF — and a player qualifies where the feed lists him for that game, not where he has played most.

Where the numbers come from

Box scores are fetched from the public feed when a game goes final, and then fetched again later and rescored, because first reports are corrected — a reclassified assist or a corrected rushing attempt changes a fantasy line. A game is only counted as coverage once it has been scored from a final box; a live, provisional fetch is not coverage.

Positions come from the feed and are corrected by hand where the feed is wrong, because a player listed at the wrong position poisons both his own team's For and the opponent's Allowed. Lines, salaries and prices are recorded as observed, with the time they were observed, and never back-filled.

What this cannot tell you

These are descriptive statistics. A team that has allowed a lot to a position has allowed a lot to a position: it has not been proven soft, and the schedule it faced is not controlled for. The splits on each team page (home and away, against the top and bottom of the league) exist because the raw figure hides exactly that.

Not the casual single-entry player who wants a picks feed, and not the professional with a Python stack and their own data. Both are better served elsewhere, and writing for either makes the grids look wrong.

Nothing here is betting advice. See the disclaimer in the footer, and the responsible-gambling page beside it.

Where to go next

The grid itself is at `/mlb/offense` (For) and `/mlb/defense` (Allowed); the two sides beside tonight's slate are at `/mlb/matchups`. Every cell opens the games behind it. The glossary of every other column in the product is on the help page.

Columns, in one line each

For
Fantasy points this team's players at this position scored per game, under the scoring system in the control bar.
Allowed
Fantasy points this team's opponents' players at this position scored against it, per game.
Rank
Where this team sits among the league's teams on that figure. Rank 1 is the highest number, which is good on For and bad on Allowed.
Lg avg
The mean across every team in the league, on the same window.
GP
Games in the window that contributed to the figure.
n
Effective games: a weighted window counts a recent game as more than one old one, so n is usually smaller than GP.
Probable
The announced starting pitcher. Shown, never used to adjust the grid.
Park
The venue's run environment factor. Applied by the model, not by the grid.