What Counts as an Average League?

What Counts as an Average League?

Every number on this site is measured against the league a player was in — so how long a memory should that league average have?

← White Papers · Published 5 August 2026 · Affects every z-score, and therefore every comp on the site

Update · 9 August 2026. Since this paper was published, the Spectral Index has been redefined as a prospect’s expected career value — the probability he reaches the majors multiplied by how valuable he projects to be if he does — shown as four numbers: Confidence, Impact, Bust risk, and Boom chance. The league-baseline windows examined here feed the same underlying league-relative statistics regardless of how the final index is assembled; the finding stands. See Punished for Making Contact and The Glove That Kept Disappearing.

The question

In 2026, the Double-A Southern League started hitting for power. Isolated power across the league landed at .1673. The year before it was .1097 — a 52% jump in one season, the kind of move that usually means a new baseball, a new park, or a new set of hitters.

The site did not notice. Asked what an average Southern League hitter did in 2026, it answered .1282.

That gap matters because nothing on this site is an absolute number. A prospect’s power is not “he slugged .190 of isolated power” — it is “he was this far above his league.” Every comparison, every comp, every Spectral Index runs through the league baseline. If the baseline is wrong by .039, then every Southern League hitter in 2026 is credited with .039 of power he did not actually have relative to his peers, and he gets matched against the wrong historical players.

So: why did the site think an environment hitting .167 was hitting .128?

Why there is a window at all

Because it was averaging three years together.

The reason for doing that is real, and it is worth stating in its strongest form before tearing it down. A single league-season is a limited sample. The Appalachian League might hit .140 one year and .120 the next purely because a few different teenagers showed up, not because anything about the environment changed. If you set the baseline from one year alone, you inherit all of that randomness, and you end up measuring players against a target that jitters for no reason.

Averaging three years suppresses the jitter. The cost is that it also suppresses real movement — it is, by construction, partly a memory of a league that no longer exists.

That is exactly what happened here, and the arithmetic is worth seeing because it leaves nothing to interpret:

Southern League season Isolated power Plate appearances
2024 .1174 39,945
2025 .1097 39,659
2026 .1673 30,234

Pool all three and weight by plate appearances and you get .1284 — which is, to within rounding, the .1282 the site was using. The baseline was not broken or miscomputed. It was correctly reporting the average of a league that had spent two of the last three years being a different league.

This is the standard bias-variance tradeoff, and there is no arguing it from first principles. Either the noise you remove is worth more than the staleness you add, or it isn’t. It has to be measured.

How to measure it without cheating

There is an obvious trap. Ask “which window best predicts what the league hit this year” and a one-year window wins automatically, because a one-year window is what the league hit this year. It would be grading a number against itself.

So every player is permanently assigned to one of two halves, by a fixed hash of his player ID that never changes across seasons. Then:

  • The baseline is built from half A only.
  • Truth is what half B actually did that season — a real, observed number, computed from players whose data was never used to build the estimate.

Now a one-year window has no unfair advantage. It is being asked to predict a disjoint group of players, and if last year’s league average is a poor description of this year’s environment, it will be punished for that just like any other candidate.

Every number in the rest of this paper comes from that setup, across roughly 800 league-seasons running from the late 1970s to 2025.

The league averages want a one-year memory

Counting it the simple way

Out of 767 league-seasons, how often was the one-year baseline closer to the held-out truth than the three-year baseline? Fifty percent would mean the window makes no difference at all.

What is being measured One-year window was closer Median miss, 1-year Median miss, 3-year
Pitcher home-run rate 60.6% 0.00091 0.00130
Hitter walk rate 55.4% 0.00378 0.00436
Hitter isolated power 53.8% 0.00573 0.00679
Pitcher walk rate 53.8% 0.00370 0.00422
Pitcher strikeout rate 52.9% 0.00524 0.00574
Hitter strikeout rate 50.5% 0.00665 0.00688

This is not a landslide and should not be sold as one. Hitter strikeout rate is a coin flip. The typical miss shrinks on all six, but on a straight count of league-seasons the one-year window wins 50-61% of the time.

Where the gain actually lives

The win rate is modest because most league-seasons barely move, and when a league doesn’t move, the window can’t matter — every candidate gets the same answer right.

Splitting league-seasons by how much the league actually moved tells a much sharper story:

Mean miss, 3-year Mean miss, 1-year 1-year closer
The 10% of leagues that moved most 0.0160 0.0087 81%
The quietest 25% of leagues 0.0055 0.0061 49%

In the leagues that actually moved, the three-year window’s error is 84% larger and it loses four times out of five. In the quiet leagues the three-year window is very slightly better — exactly what the smoothing argument predicts, and it is real.

That is the tradeoff, measured instead of assumed. And it reframes the finding: it is not that one year beats three years generally. It is that three years fails specifically whenever there is something to notice.

The number that settles it

Win rates depend on how you choose to score a miss. This one doesn’t.

An estimator that is merely noisy misses in both directions at random. An estimator that is stale misses in the direction the league moved — when power goes up it reads low, when power goes down it reads high. That’s a testable difference. Correlate each window’s error against how much the league actually moved year over year, across 691 full-season league-seasons:

correlation between the miss and the league’s actual movement
Three-year window (what the site used) −0.717
One-year window −0.292

−0.717 is not noise. It is lag. More than half the variance in the three-year window’s error is explained by the thing it is supposed to be tracking. Switching to a one-year window cuts that correlation by 59%.

This is the strongest single result in this paper, because it requires no choice of loss function and no judgment call. It is a direct measurement that the old baseline was biased, not merely imprecise — and bias is the failure mode that averaging more years cannot fix, because averaging more years is what causes it.

The motivating league, worked end to end

2025 Double-A Southern League, with the baseline built from half the league and checked against the other half:

Isolated power Miss
What half B actually hit (truth) .1017
One-year baseline from half A .1173 +.0156
Three-year baseline from half A .1304 +.0287

The three-year window calls the Southern League a .130 environment while the held-out half of that same league is hitting .102.

2024 is starker: truth .1178, one-year says .1172 — a miss of .0006, essentially exact — while three-year says .1477, a miss of .0299.

2022 goes the other way, with the three-year window closer by 38%. That is included deliberately. The one-year window does lose league-seasons, and the Southern League is not a clean sweep even in the case that motivated the whole investigation.

Widening out: of the ten largest three-year misses on isolated power among full-season leagues since 1980, the one-year window was closer in ten out of ten — and two of those ten are the 2024 and 2025 Southern League.

Where the evidence points the other way

Two honest counterweights, both of which cut against the change:

The 1970s prefer the three-year window. On a straight count of league-seasons, one year was closer in only 37.5% of 1970s hitter isolated-power cells, 39.6% for walk rate, 43.8% for strikeouts. This disagrees with the typical-miss measurement over the same years, and the disagreement is informative: in the 1970s the three-year window wins more league-seasons but the one-year window wins the big ones. With only 48 league-seasons it is also the thinnest decade in the data. From 1980 onward, both ways of counting agree on the one-year window.

Rookie ball genuinely wants smoothing — which is the next section.

Rookie ball is the exception, and keeps three years

Whether a league’s year-to-year movement is worth tracking depends on whether that movement is real. Split-half reliability answers that directly: of the swing you see in a league’s isolated power from one year to the next, how much is a genuine change in environment versus a different set of players showing up?

Level Reliability (bootstrap median, 90% interval) Reading
Rookie 0.08 [−0.76, 0.44] indistinguishable from zero
Low-A 0.53 [0.15, 0.73] moderate
Double-A 0.46 [0.09, 0.65] about half real, half noise
Triple-A 0.80 [0.70, 0.87] mostly real

At Triple-A, year-to-year movement is mostly signal, and smoothing it away destroys information. At Rookie level, it is indistinguishable from noise — complex-league rosters turn over almost completely every year, so a “change in the league” is mostly just a change in who is standing there.

So the league averages use one year at full-season levels and three years at Rookie level. That is not a hedge; it is the only place in the data where the smoothing argument survives contact with a measurement.

One caveat recorded in full: by the count of league-seasons, Rookie ball mildly prefers one year too. The two ways of scoring disagree there, and three years was kept as the conservative choice — it leaves the noisiest level exactly as it was.

The spread wants a longer memory, not a shorter one

Here the investigation turned up something nobody was looking for.

A league baseline is two numbers, not one. There is the average, and there is the spread — how far apart the league’s hitters are from each other. Spread is what converts a raw gap into a z-score: being .050 above average means something very different in a league where everyone is bunched together than in one where they are scattered.

Both numbers were being smoothed over the same three years, for no reason other than that they were computed in the same block of code.

Swept separately across windows of one through five years:

Quantity 1 yr 2 yr 3 yr 4 yr 5 yr Best
Spread of hitter isolated power 0.038227 0.035862 0.034546 0.034416 0.034391 5 yr
Spread of hitter walk rate 0.025971 0.023088 0.022036 0.021708 0.021668 5 yr, +1.7%
Spread of hitter strikeout rate 0.028801 0.026441 0.025683 0.025789 0.026164 3 yr
Spread of pitcher home-run rate 0.012561 0.010182 0.0093433 0.0091086 0.0090575 5 yr, +3.1%
Spread of pitcher walk rate 0.017997 0.016071 0.0152904 0.0150504 0.0147654 5 yr, +3.4%
Spread of pitcher strikeout rate 0.016205 0.013833 0.0131073 0.0125558 0.0123077 5 yr, +6.1%

Five of six improve monotonically all the way out to five years. The spread wants to be smoothed harder than it was, in the same build where the average wants to be smoothed less.

That sounds contradictory and isn’t. The average of a league is a fast-moving physical fact — a new baseball changes it in one offseason. The dispersion of talent around that average is a slow structural property of how a level is stocked, and it is measured far less precisely from one season, because a standard deviation needs more data than a mean does. Fast-moving quantity, short memory. Slow-moving, noisily-measured quantity, long memory.

A second test agrees. An entirely separate check — asking whether the same player, in the same season, who appeared in two different leagues gets the same z-score from both — independently lands on a five-year window for hitter power and strikeouts, at about +2.0% each. The two tests share no methodology and point the same way.

League age wants two years, and this one had been invisible

The third quantity riding on that single three-year constant was the average age of each league — which is not a diagnostic. It feeds the age adjustment that decides how much credit a 19-year-old gets for holding his own against 23-year-olds.

Quantity 1 yr 2 yr 3 yr 4 yr 5 yr Best
Hitter league age 0.589126 0.552076 0.609820 0.685794 0.769139 2 yr, +9.5%
Pitcher league age 0.660380 0.643716 0.700177 0.798247 0.899284 2 yr, +8.1%

Two years, on both sides, by the largest margin of any of the three quantities — and it is a genuine interior optimum, degrading in both directions.

The reason this went unnoticed for so long is worth recording, because it is a general lesson about how a constant hides a question. Every earlier comparison in this project had asked “one year or three?” — and age loses that particular comparison. It is better at three than at one, so at every previous look it appeared to be a quantity that wanted smoothing and was already getting the right amount. Only sweeping the full curve reveals that both endpoints of the old question were wrong.

Structurally, league age was being computed inside the same rolling pass as the batting rates, which is precisely why it inherited their window. It now gets its own pass.

What the single constant was really doing

Three quantities, three answers: the average wants one year, the spread wants five, the age wants two.

The old setting was three for all of them. Which means the number 3 was not a compromise anybody reached — it was the best value for none of the three quantities it controlled. The problem was never that 3 was the wrong number. It was that one number was being asked three different questions.

What it does to actual players

Fixing the baseline moves individual players, and it moves them in the direction the mechanism predicts. Change in power z-score for the 2026 Milwaukee prospects used as the anchor case throughout this work:

Player Change in power z-score
Brady Ebel −0.219
Andrew Fischer −0.204
Blake Burke −0.174
Jesus Made −0.170
Josh Adamczewski −0.146
Eric Bitonti −0.112
Jett Williams −0.089
Leanders Matos −0.072
Diego Frontado −0.054
Luke Adams +0.011
Braylon Payne +0.011

They go down, and they should: these hitters were in a league that was hitting for far more power than the site believed, so their power was being overstated relative to their peers. Correcting the baseline takes back credit that belonged to the environment.

The two exceptions are the useful part. Luke Adams moved up to Triple-A, so he is graded against a different 2026 environment than the Double-A group — and he moves the other way. A change that pushed every Milwaukee prospect in the same direction would be a red flag for a global rescaling rather than a genuine per-league correction. This one separates players by which league they were actually in.

Across all 34,712 players with a profile, the average absolute move is 0.029 z-score units, with the 90th percentile at 0.068, and it is nearly symmetric — 50% down, 47% up. Some players are helped, some hurt, depending entirely on whether their league was hotter or colder than its three-year memory.

And the case that started it: the 2026 Southern League baseline went from .1282 to .1670, against an actual .1673. A 24% error became a 0.2% error.

What it does not do

Here is the result that keeps this from being a straightforward success story.

All of the above establishes that the league baselines are more accurate. It does not establish that the site’s headline number — the Spectral Index — got better at its actual job.

Testing that means asking a question with a real answer: for players whose careers are already over, does the Spectral Index identify who reached the major leagues? That can be scored on four decades of finished careers. The comparison is run on the same players in both versions, resampled together, because the two versions produce nearly identical scores and comparing them independently would be far too forgiving.

Cohort Old New Change
Hitters 1980s 0.726 0.729 +0.004 no real change
Hitters 1990s 0.779 0.768 −0.011 worse
Hitters 2000s 0.829 0.824 −0.005 no real change
Hitters 2015-19 0.864 0.865 +0.001 no real change
Pitchers 1980s 0.863 0.869 +0.006 better
Pitchers 1990s 0.872 0.869 −0.002 no real change
Pitchers 2000s 0.892 0.893 +0.001 no real change
Pitchers 2015-19 0.895 0.892 −0.003 no real change

Six of eight cohorts show no real change. The average across all eight is −0.001. One cohort is genuinely better, one genuinely worse.

Two things have to be said about that honestly. First, eight tests at 95% confidence produce about 0.4 false alarms by chance, and this produced two, one in each direction — roughly what chance alone delivers. Second, this instrument was built to answer “is the product useful,” not “did change X help”: it reads each player’s current profile rather than what was known at the time, so it is descriptive rather than a clean forecast.

But it is the best available evidence on the question, and it does not say the change helped. More accurate league baselines did not produce a better Spectral Index.

That is less strange than it first sounds. The Spectral Index is driven by how often statistically similar players reached the majors, and comps are chosen by relative position within a league. Shifting a whole league’s baseline shifts everyone in that league together, which reshuffles comps far less than it reshuffles the underlying z-scores. The correction is real, and it largely cancels at the point where the index is computed.

Which of the three changes did that?

Three windows moved at once, so the table above cannot say which one is responsible. Answering that took three more complete rebuilds of the database, each changing exactly one window away from the old three-year setting and leaving the other two alone, every one scored against the same original baseline.

Averages
3 → 1
Spreads
3 → 5
Age
3 → 2
All three
Hitters 1980s +0.0044 +0.0005 −0.0046 +0.0036
Hitters 1990s −0.0102 −0.0004 −0.0042 −0.0110
Hitters 2000s −0.0030 −0.0064 −0.0028 −0.0051
Hitters 2015-19 +0.0010 −0.0020 −0.0019 +0.0009
Pitchers 1980s +0.0065 +0.0024 +0.0004 +0.0063
Pitchers 1990s −0.0027 +0.0019 +0.0003 −0.0023
Pitchers 2000s +0.0011 −0.0005 +0.0001 +0.0009
Pitchers 2015-19 −0.0018 −0.0022 −0.0012 −0.0032

Bold = the change is real at 95% confidence rather than inside the noise.

The averages change owns both of the big movements. The 1990s hitter degradation is −0.0102 on its own against −0.0110 for all three together; the 1980s pitcher improvement is +0.0065 against +0.0063. Both headline effects are the one-year league average, and the other two windows are passengers.

That is worth stating flatly because the prediction going in was the opposite. The expectation was that stretching the spread window to five years would be the culprit — importing stale dispersion into the 1990s, a decade when the offensive environment was moving fast. That hypothesis is dead. The spread change does essentially nothing to 1990s hitters (−0.0004), and its only real damage is in the 2000s.

The genuinely uncomfortable result is the age window. Look down that column: it is negative in seven of eight cohorts and significantly negative in four, with no cohort it significantly helps. Averaged across all eight it is the worst of the three changes.

And it is the change with the strongest upstream evidence. Two years beat three years at predicting held-out league age by 9.5% for hitters and 8.1% for pitchers — the largest margin any of the three windows produced, a clean interior optimum degrading in both directions. A more accurate league age makes the Spectral Index slightly worse, consistently, across four decades and both sides.

There is no comfortable way to reconcile those two facts, and this paper does not pretend to. The most likely explanation is that the age adjustment downstream was fit and calibrated while league age carried a three-year window, so its constants silently absorbed some of that smoothing; handing it a sharper input moves it off the point it was tuned for. That is a hypothesis. It has not been tested, and testing it would mean re-fitting the age adjustment against the new age input.

One more caveat, stated because it limits how hard the table above can be read. The three isolated effects do not quite add up to the combined effect — they miss by 0.003 on average and by as much as 0.007, which is meaningful next to effects of this size. The three windows interact rather than acting independently, and for hitters the combination is consistently less harmful than its parts would suggest. So the column labels are the right first-order attribution, not a clean decomposition.

What is still unknown

  • There is no test that can properly see this class of change. A league-wide baseline error moves every player in that league in the same direction, so any measurement comparing players within a league is blind to it by construction — the error cancels inside the comparison. The tests that can see it (held-out league truth) are upstream of the Spectral Index; the test that measures the Spectral Index cannot see it. Closing that gap is the honest next piece of work, and shipping this did not close it.
  • The age window is the loose end. It has the best upstream evidence of the three changes and the worst downstream behaviour. The leading explanation — that the age adjustment’s own constants were fit against a three-year league age and quietly absorbed that smoothing — is untested, and the test is to re-fit the age adjustment against the new input.
  • Why the one-year average specifically hurts 1990s hitters is not established. It is now known to be the responsible change, which is a different thing from knowing the mechanism.
  • The Rookie-level window rests on two ways of counting that disagree. Three years was kept as the conservative choice, not because the evidence was clean.

A note on method

Every number here was produced by the site’s own pipeline functions, not by a reimplementation. The window sweeps call the same rolling_prior_years() and rolling_prior_means() that build the shipped database, applied to half the league.

That is a standing rule on this project, learned the hard way: an independent reimplementation of the Spectral Index, written to make research easier, silently drifted out of date and produced an entire thread of confidently wrong findings before anyone caught it. Before being used here, the replica used for the Spectral Index comparison was re-checked line by line against the live site’s own code.

One control is worth naming because it caught a real problem. The analysis rebuilds the league context itself, so it was first required to reproduce the shipped league context exactly. It matched to zero on every column that the database stores — but the check also revealed that several computed columns are never written back to the database at all and therefore could not be verified. That is reported rather than glossed: the control confirmed what it could confirm, and named what it could not.

← Back to White Papers

Prospect Wavelength is not affiliated with MLB, nor with any of their digital properties. Email any questions, comments, or concerns to stealofhome@prospectwavelength.com