Update · 9 August 2026. Since this paper was published, the Spectral Index has been redefined as a prospect’s expected career value — the probability he reaches the majors multiplied by how valuable he projects to be if he does — shown as four numbers: Confidence, Impact, Bust risk, and Boom chance. The league-baseline windows examined here feed the same underlying league-relative statistics regardless of how the final index is assembled; the finding stands. See Punished for Making Contact and The Glove That Kept Disappearing.
The question
In 2026, the Double-A Southern League started hitting for power. Isolated power across the league landed at .1673. The year before it was .1097 — a 52% jump in one season, the kind of move that usually means a new baseball, a new park, or a new set of hitters.
The site did not notice. Asked what an average Southern League hitter did in 2026, it answered .1282.
That gap matters because nothing on this site is an absolute number. A prospect’s power is not “he slugged .190 of isolated power” — it is “he was this far above his league.” Every comparison, every comp, every Spectral Index runs through the league baseline. If the baseline is wrong by .039, then every Southern League hitter in 2026 is credited with .039 of power he did not actually have relative to his peers, and he gets matched against the wrong historical players.
So: why did the site think an environment hitting .167 was hitting .128?
Why there is a window at all
Because it was averaging three years together.
The reason for doing that is real, and it is worth stating in its strongest form before tearing it down. A single league-season is a limited sample. The Appalachian League might hit .140 one year and .120 the next purely because a few different teenagers showed up, not because anything about the environment changed. If you set the baseline from one year alone, you inherit all of that randomness, and you end up measuring players against a target that jitters for no reason.
Averaging three years suppresses the jitter. The cost is that it also suppresses real movement — it is, by construction, partly a memory of a league that no longer exists.
That is exactly what happened here, and the arithmetic is worth seeing because it leaves nothing to interpret:
| Southern League season | Isolated power | Plate appearances |
|---|---|---|
| 2024 | .1174 | 39,945 |
| 2025 | .1097 | 39,659 |
| 2026 | .1673 | 30,234 |
Pool all three and weight by plate appearances and you get .1284 — which is, to within rounding, the .1282 the site was using. The baseline was not broken or miscomputed. It was correctly reporting the average of a league that had spent two of the last three years being a different league.
This is the standard bias-variance tradeoff, and there is no arguing it from first principles. Either the noise you remove is worth more than the staleness you add, or it isn’t. It has to be measured.
How to measure it without cheating
There is an obvious trap. Ask “which window best predicts what the league hit this year” and a one-year window wins automatically, because a one-year window is what the league hit this year. It would be grading a number against itself.
So every player is permanently assigned to one of two halves, by a fixed hash of his player ID that never changes across seasons. Then:
- The baseline is built from half A only.
- Truth is what half B actually did that season — a real, observed number, computed from players whose data was never used to build the estimate.
Now a one-year window has no unfair advantage. It is being asked to predict a disjoint group of players, and if last year’s league average is a poor description of this year’s environment, it will be punished for that just like any other candidate.
Every number in the rest of this paper comes from that setup, across roughly 800 league-seasons running from the late 1970s to 2025.
The league averages want a one-year memory
Counting it the simple way
Out of 767 league-seasons, how often was the one-year baseline closer to the held-out truth than the three-year baseline? Fifty percent would mean the window makes no difference at all.
| What is being measured | One-year window was closer | Median miss, 1-year | Median miss, 3-year |
|---|---|---|---|
| Pitcher home-run rate | 60.6% | 0.00091 | 0.00130 |
| Hitter walk rate | 55.4% | 0.00378 | 0.00436 |
| Hitter isolated power | 53.8% | 0.00573 | 0.00679 |
| Pitcher walk rate | 53.8% | 0.00370 | 0.00422 |
| Pitcher strikeout rate | 52.9% | 0.00524 | 0.00574 |
| Hitter strikeout rate | 50.5% | 0.00665 | 0.00688 |
This is not a landslide and should not be sold as one. Hitter strikeout rate is a coin flip. The typical miss shrinks on all six, but on a straight count of league-seasons the one-year window wins 50-61% of the time.
Where the gain actually lives
The win rate is modest because most league-seasons barely move, and when a league doesn’t move, the window can’t matter — every candidate gets the same answer right.
Splitting league-seasons by how much the league actually moved tells a much sharper story:
| Mean miss, 3-year | Mean miss, 1-year | 1-year closer | |
|---|---|---|---|
| The 10% of leagues that moved most | 0.0160 | 0.0087 | 81% |
| The quietest 25% of leagues | 0.0055 | 0.0061 | 49% |
In the leagues that actually moved, the three-year window’s error is 84% larger and it loses four times out of five. In the quiet leagues the three-year window is very slightly better — exactly what the smoothing argument predicts, and it is real.
That is the tradeoff, measured instead of assumed. And it reframes the finding: it is not that one year beats three years generally. It is that three years fails specifically whenever there is something to notice.
The number that settles it
Win rates depend on how you choose to score a miss. This one doesn’t.
An estimator that is merely noisy misses in both directions at random. An estimator that is stale misses in the direction the league moved — when power goes up it reads low, when power goes down it reads high. That’s a testable difference. Correlate each window’s error against how much the league actually moved year over year, across 691 full-season league-seasons:
| correlation between the miss and the league’s actual movement | |
|---|---|
| Three-year window (what the site used) | −0.717 |
| One-year window | −0.292 |
−0.717 is not noise. It is lag. More than half the variance in the three-year window’s error is explained by the thing it is supposed to be tracking. Switching to a one-year window cuts that correlation by 59%.
This is the strongest single result in this paper, because it requires no choice of loss function and no judgment call. It is a direct measurement that the old baseline was biased, not merely imprecise — and bias is the failure mode that averaging more years cannot fix, because averaging more years is what causes it.
The motivating league, worked end to end
2025 Double-A Southern League, with the baseline built from half the league and checked against the other half:
| Isolated power | Miss | |
|---|---|---|
| What half B actually hit (truth) | .1017 | — |
| One-year baseline from half A | .1173 | +.0156 |
| Three-year baseline from half A | .1304 | +.0287 |
The three-year window calls the Southern League a .130 environment while the held-out half of that same league is hitting .102.
2024 is starker: truth .1178, one-year says .1172 — a miss of .0006, essentially exact — while three-year says .1477, a miss of .0299.
2022 goes the other way, with the three-year window closer by 38%. That is included deliberately. The one-year window does lose league-seasons, and the Southern League is not a clean sweep even in the case that motivated the whole investigation.
Widening out: of the ten largest three-year misses on isolated power among full-season leagues since 1980, the one-year window was closer in ten out of ten — and two of those ten are the 2024 and 2025 Southern League.
Where the evidence points the other way
Two honest counterweights, both of which cut against the change:
The 1970s prefer the three-year window. On a straight count of league-seasons, one year was closer in only 37.5% of 1970s hitter isolated-power cells, 39.6% for walk rate, 43.8% for strikeouts. This disagrees with the typical-miss measurement over the same years, and the disagreement is informative: in the 1970s the three-year window wins more league-seasons but the one-year window wins the big ones. With only 48 league-seasons it is also the thinnest decade in the data. From 1980 onward, both ways of counting agree on the one-year window.
Rookie ball genuinely wants smoothing — which is the next section.
Rookie ball is the exception, and keeps three years
Whether a league’s year-to-year movement is worth tracking depends on whether that movement is real. Split-half reliability answers that directly: of the swing you see in a league’s isolated power from one year to the next, how much is a genuine change in environment versus a different set of players showing up?
| Level | Reliability (bootstrap median, 90% interval) | Reading |
|---|---|---|
| Rookie | 0.08 [−0.76, 0.44] | indistinguishable from zero |
| Low-A | 0.53 [0.15, 0.73] | moderate |
| Double-A | 0.46 [0.09, 0.65] | about half real, half noise |
| Triple-A | 0.80 [0.70, 0.87] | mostly real |
At Triple-A, year-to-year movement is mostly signal, and smoothing it away destroys information. At Rookie level, it is indistinguishable from noise — complex-league rosters turn over almost completely every year, so a “change in the league” is mostly just a change in who is standing there.
So the league averages use one year at full-season levels and three years at Rookie level. That is not a hedge; it is the only place in the data where the smoothing argument survives contact with a measurement.
One caveat recorded in full: by the count of league-seasons, Rookie ball mildly prefers one year too. The two ways of scoring disagree there, and three years was kept as the conservative choice — it leaves the noisiest level exactly as it was.
The spread wants a longer memory, not a shorter one
Here the investigation turned up something nobody was looking for.
A league baseline is two numbers, not one. There is the average, and there is the spread — how far apart the league’s hitters are from each other. Spread is what converts a raw gap into a z-score: being .050 above average means something very different in a league where everyone is bunched together than in one where they are scattered.
Both numbers were being smoothed over the same three years, for no reason other than that they were computed in the same block of code.
Swept separately across windows of one through five years:
| Quantity | 1 yr | 2 yr | 3 yr | 4 yr | 5 yr | Best |
|---|---|---|---|---|---|---|
| Spread of hitter isolated power | 0.038227 | 0.035862 | 0.034546 | 0.034416 | 0.034391 | 5 yr |
| Spread of hitter walk rate | 0.025971 | 0.023088 | 0.022036 | 0.021708 | 0.021668 | 5 yr, +1.7% |
| Spread of hitter strikeout rate | 0.028801 | 0.026441 | 0.025683 | 0.025789 | 0.026164 | 3 yr |
| Spread of pitcher home-run rate | 0.012561 | 0.010182 | 0.0093433 | 0.0091086 | 0.0090575 | 5 yr, +3.1% |
| Spread of pitcher walk rate | 0.017997 | 0.016071 | 0.0152904 | 0.0150504 | 0.0147654 | 5 yr, +3.4% |
| Spread of pitcher strikeout rate | 0.016205 | 0.013833 | 0.0131073 | 0.0125558 | 0.0123077 | 5 yr, +6.1% |
Five of six improve monotonically all the way out to five years. The spread wants to be smoothed harder than it was, in the same build where the average wants to be smoothed less.
That sounds contradictory and isn’t. The average of a league is a fast-moving physical fact — a new baseball changes it in one offseason. The dispersion of talent around that average is a slow structural property of how a level is stocked, and it is measured far less precisely from one season, because a standard deviation needs more data than a mean does. Fast-moving quantity, short memory. Slow-moving, noisily-measured quantity, long memory.
A second test agrees. An entirely separate check — asking whether the same player, in the same season, who appeared in two different leagues gets the same z-score from both — independently lands on a five-year window for hitter power and strikeouts, at about +2.0% each. The two tests share no methodology and point the same way.
League age wants two years, and this one had been invisible
The third quantity riding on that single three-year constant was the average age of each league — which is not a diagnostic. It feeds the age adjustment that decides how much credit a 19-year-old gets for holding his own against 23-year-olds.
| Quantity | 1 yr | 2 yr | 3 yr | 4 yr | 5 yr | Best |
|---|---|---|---|---|---|---|
| Hitter league age | 0.589126 | 0.552076 | 0.609820 | 0.685794 | 0.769139 | 2 yr, +9.5% |
| Pitcher league age | 0.660380 | 0.643716 | 0.700177 | 0.798247 | 0.899284 | 2 yr, +8.1% |
Two years, on both sides, by the largest margin of any of the three quantities — and it is a genuine interior optimum, degrading in both directions.
The reason this went unnoticed for so long is worth recording, because it is a general lesson about how a constant hides a question. Every earlier comparison in this project had asked “one year or three?” — and age loses that particular comparison. It is better at three than at one, so at every previous look it appeared to be a quantity that wanted smoothing and was already getting the right amount. Only sweeping the full curve reveals that both endpoints of the old question were wrong.
Structurally, league age was being computed inside the same rolling pass as the batting rates, which is precisely why it inherited their window. It now gets its own pass.
What the single constant was really doing
Three quantities, three answers: the average wants one year, the spread wants five, the age wants two.
The old setting was three for all of them. Which means the number 3 was not a compromise anybody reached — it was the best value for none of the three quantities it controlled. The problem was never that 3 was the wrong number. It was that one number was being asked three different questions.
What it does to actual players
Fixing the baseline moves individual players, and it moves them in the direction the mechanism predicts. Change in power z-score for the 2026 Milwaukee prospects used as the anchor case throughout this work:
| Player | Change in power z-score |
|---|---|
| Brady Ebel | −0.219 |
| Andrew Fischer | −0.204 |
| Blake Burke | −0.174 |
| Jesus Made | −0.170 |
| Josh Adamczewski | −0.146 |
| Eric Bitonti | −0.112 |
| Jett Williams | −0.089 |
| Leanders Matos | −0.072 |
| Diego Frontado | −0.054 |
| Luke Adams | +0.011 |
| Braylon Payne | +0.011 |
They go down, and they should: these hitters were in a league that was hitting for far more power than the site believed, so their power was being overstated relative to their peers. Correcting the baseline takes back credit that belonged to the environment.
The two exceptions are the useful part. Luke Adams moved up to Triple-A, so he is graded against a different 2026 environment than the Double-A group — and he moves the other way. A change that pushed every Milwaukee prospect in the same direction would be a red flag for a global rescaling rather than a genuine per-league correction. This one separates players by which league they were actually in.
Across all 34,712 players with a profile, the average absolute move is 0.029 z-score units, with the 90th percentile at 0.068, and it is nearly symmetric — 50% down, 47% up. Some players are helped, some hurt, depending entirely on whether their league was hotter or colder than its three-year memory.
And the case that started it: the 2026 Southern League baseline went from .1282 to .1670, against an actual .1673. A 24% error became a 0.2% error.
What it does not do
Here is the result that keeps this from being a straightforward success story.
All of the above establishes that the league baselines are more accurate. It does not establish that the site’s headline number — the Spectral Index — got better at its actual job.
Testing that means asking a question with a real answer: for players whose careers are already over, does the Spectral Index identify who reached the major leagues? That can be scored on four decades of finished careers. The comparison is run on the same players in both versions, resampled together, because the two versions produce nearly identical scores and comparing them independently would be far too forgiving.
| Cohort | Old | New | Change | ||
|---|---|---|---|---|---|
| Hitters | 1980s | 0.726 | 0.729 | +0.004 | no real change |
| Hitters | 1990s | 0.779 | 0.768 | −0.011 | worse |
| Hitters | 2000s | 0.829 | 0.824 | −0.005 | no real change |
| Hitters | 2015-19 | 0.864 | 0.865 | +0.001 | no real change |
| Pitchers | 1980s | 0.863 | 0.869 | +0.006 | better |
| Pitchers | 1990s | 0.872 | 0.869 | −0.002 | no real change |
| Pitchers | 2000s | 0.892 | 0.893 | +0.001 | no real change |
| Pitchers | 2015-19 | 0.895 | 0.892 | −0.003 | no real change |
Six of eight cohorts show no real change. The average across all eight is −0.001. One cohort is genuinely better, one genuinely worse.
Two things have to be said about that honestly. First, eight tests at 95% confidence produce about 0.4 false alarms by chance, and this produced two, one in each direction — roughly what chance alone delivers. Second, this instrument was built to answer “is the product useful,” not “did change X help”: it reads each player’s current profile rather than what was known at the time, so it is descriptive rather than a clean forecast.
But it is the best available evidence on the question, and it does not say the change helped. More accurate league baselines did not produce a better Spectral Index.
That is less strange than it first sounds. The Spectral Index is driven by how often statistically similar players reached the majors, and comps are chosen by relative position within a league. Shifting a whole league’s baseline shifts everyone in that league together, which reshuffles comps far less than it reshuffles the underlying z-scores. The correction is real, and it largely cancels at the point where the index is computed.
Which of the three changes did that?
Three windows moved at once, so the table above cannot say which one is responsible. Answering that took three more complete rebuilds of the database, each changing exactly one window away from the old three-year setting and leaving the other two alone, every one scored against the same original baseline.
| Averages 3 → 1 |
Spreads 3 → 5 |
Age 3 → 2 |
All three | ||
|---|---|---|---|---|---|
| Hitters | 1980s | +0.0044 | +0.0005 | −0.0046 | +0.0036 |
| Hitters | 1990s | −0.0102 | −0.0004 | −0.0042 | −0.0110 |
| Hitters | 2000s | −0.0030 | −0.0064 | −0.0028 | −0.0051 |
| Hitters | 2015-19 | +0.0010 | −0.0020 | −0.0019 | +0.0009 |
| Pitchers | 1980s | +0.0065 | +0.0024 | +0.0004 | +0.0063 |
| Pitchers | 1990s | −0.0027 | +0.0019 | +0.0003 | −0.0023 |
| Pitchers | 2000s | +0.0011 | −0.0005 | +0.0001 | +0.0009 |
| Pitchers | 2015-19 | −0.0018 | −0.0022 | −0.0012 | −0.0032 |
Bold = the change is real at 95% confidence rather than inside the noise.
The averages change owns both of the big movements. The 1990s hitter degradation is −0.0102 on its own against −0.0110 for all three together; the 1980s pitcher improvement is +0.0065 against +0.0063. Both headline effects are the one-year league average, and the other two windows are passengers.
That is worth stating flatly because the prediction going in was the opposite. The expectation was that stretching the spread window to five years would be the culprit — importing stale dispersion into the 1990s, a decade when the offensive environment was moving fast. That hypothesis is dead. The spread change does essentially nothing to 1990s hitters (−0.0004), and its only real damage is in the 2000s.
The genuinely uncomfortable result is the age window. Look down that column: it is negative in seven of eight cohorts and significantly negative in four, with no cohort it significantly helps. Averaged across all eight it is the worst of the three changes.
And it is the change with the strongest upstream evidence. Two years beat three years at predicting held-out league age by 9.5% for hitters and 8.1% for pitchers — the largest margin any of the three windows produced, a clean interior optimum degrading in both directions. A more accurate league age makes the Spectral Index slightly worse, consistently, across four decades and both sides.
There is no comfortable way to reconcile those two facts, and this paper does not pretend to. The most likely explanation is that the age adjustment downstream was fit and calibrated while league age carried a three-year window, so its constants silently absorbed some of that smoothing; handing it a sharper input moves it off the point it was tuned for. That is a hypothesis. It has not been tested, and testing it would mean re-fitting the age adjustment against the new age input.
One more caveat, stated because it limits how hard the table above can be read. The three isolated effects do not quite add up to the combined effect — they miss by 0.003 on average and by as much as 0.007, which is meaningful next to effects of this size. The three windows interact rather than acting independently, and for hitters the combination is consistently less harmful than its parts would suggest. So the column labels are the right first-order attribution, not a clean decomposition.
What is still unknown
- There is no test that can properly see this class of change. A league-wide baseline error moves every player in that league in the same direction, so any measurement comparing players within a league is blind to it by construction — the error cancels inside the comparison. The tests that can see it (held-out league truth) are upstream of the Spectral Index; the test that measures the Spectral Index cannot see it. Closing that gap is the honest next piece of work, and shipping this did not close it.
- The age window is the loose end. It has the best upstream evidence of the three changes and the worst downstream behaviour. The leading explanation — that the age adjustment’s own constants were fit against a three-year league age and quietly absorbed that smoothing — is untested, and the test is to re-fit the age adjustment against the new input.
- Why the one-year average specifically hurts 1990s hitters is not established. It is now known to be the responsible change, which is a different thing from knowing the mechanism.
- The Rookie-level window rests on two ways of counting that disagree. Three years was kept as the conservative choice, not because the evidence was clean.
A note on method
Every number here was produced by the site’s own pipeline functions, not by a reimplementation. The window sweeps call the same rolling_prior_years() and rolling_prior_means() that build the shipped database, applied to half the league.
That is a standing rule on this project, learned the hard way: an independent reimplementation of the Spectral Index, written to make research easier, silently drifted out of date and produced an entire thread of confidently wrong findings before anyone caught it. Before being used here, the replica used for the Spectral Index comparison was re-checked line by line against the live site’s own code.
One control is worth naming because it caught a real problem. The analysis rebuilds the league context itself, so it was first required to reproduce the shipped league context exactly. It matched to zero on every column that the database stores — but the check also revealed that several computed columns are never written back to the database at all and therefore could not be verified. That is reported rather than glossed: the control confirmed what it could confirm, and named what it could not.