Status. Nothing on this page shipped. This is a record of a change that was measured, gated, and declined — including the one version of it that genuinely worked, and the specific reason it was still not good enough.
Update · 16 August 2026 — re-measured, and the verdict holds. Everything below was originally measured against the model as it stood on 15 August. That model has since been corrected: minor-league seasons a player played after he had already established himself in the majors were being counted as part of his own prospect record, which is a peek at the answer the model is trying to predict, and they are now excluded. Because that correction changes exactly what the model knows about a player’s own record, this page’s whole argument had to be re-run rather than assumed — the question here is how much a player’s own power and speed add on top of what his comps already say, and both halves of that moved.
It came back essentially unchanged. The era split is still there and still points in opposite directions, at −3.1 standard deviations one way (0 of 80 resamples) and +3.6 the other (80 of 80) — within noise of what it was before the correction. It still puts the wrong players on the board: Karim Garcia now enters the top 25 at 10th, with two more sub-5-win careers alongside him, and Wade Boggs falls out of the top 450, to 773rd. The decision to decline stands, on evidence that has now survived a serious change to the model underneath it.
That is worth stating for a reason beyond this page. A companion investigation running the same day did have its verdict overturned by the same correction — an era disagreement it was built on turned out to be an artifact of the contaminated features. This one wasn’t. Two findings that looked alike were not alike, and the only way to know which was which was to re-run both.
What this page is
The Spectral Index is built in two stages. First the site finds a prospect’s twenty closest historical comparables. Then four separate models take that comp pool and turn it into the numbers you actually see: the chance he reaches the majors, how good he’d be if he did, his bust risk, and his boom chance.
The four statistics used to find the comps are the prospect’s own walk rate, strikeout rate, power, and speed. Once the comps are chosen, those four numbers are discarded. They never reach the three prediction models. Those models see what happened to the twenty comps, plus the prospect’s batting average, his walk-to-strikeout ratio, and his position — but nothing about his own power or his own legs.
Stated that way it sounds like an obvious oversight, and an earlier measurement suggested fixing it was worth more than anything else left on the list. This page is about why that measurement did not survive contact with the current model.
The idea that should have worked
Take two real players the site scores almost identically on the comp pool.
Adrian Beltre and Carlos Quentin both come out at +2.27 on the comp-outcome number — the single summary of “what became of players like this.” That is the only thing about their comparables the prediction models are told.
What the models are not told is that Beltre’s own minor-league power was +2.44 in league-adjusted terms and Quentin’s was +1.10. On power, the models cannot tell these two apart at all.
Beltre finished with 93.7 career wins above replacement. Quentin finished with 10.5.
The same setup, further down: Alex Rodriguez and Don LeJohn share a comp score of +0.75. Rodriguez’s own power was +2.77, LeJohn’s −1.45. Rodriguez finished at 117.4 career WAR; LeJohn at 0.3.
Those examples are honest about the setup — they show exactly what the model can and cannot see. They are not evidence that adding power would help, and it is worth saying so plainly, because the aggregate answer turns out to run the other way.
Why it doesn’t work: the comps were built out of those numbers
Here is the part that is easy to miss. The comps were selected using power and speed in the first place. A prospect with big power gets comparables with big power. The comp-outcome number is therefore not independent of the prospect’s own bat — it is partly a restatement of it.
Measured across every settled hitter who reached the majors, the comp term already accounts for 41.2% of a player’s own power and 8.2% of his own speed. Nearly half the power information was already in the model before anyone proposed adding it.
So the real question is not “does power predict a good career.” It plainly does. On its own, own power correlates +0.20 with career WAR. The real question is whether power says anything after the comps have spoken. And there the answer reverses:
| correlation with career WAR | |
|---|---|
| Own power, by itself | +0.20 |
| Own power, against what the comp term missed | −0.22 |
| Own speed, against what the comp term missed | −0.12 |
Among players who reached the majors, once you know what their comparables did, the ones whose own power ran ahead of what their comp profile implied had worse careers, not better. Both halves of the database agree on the sign (−0.25 for players finishing in the minors by 2000, −0.20 for 2001–2010).
There is a plausible baseball reason. If a hitter’s raw power outruns his comp pool, some of that excess is context rather than talent — a hitter’s park, a hitter’s league, a player repeating a level he is too old for. The comp pool quietly absorbs that context, because the comps were drawn from similar contexts. The raw rate does not. What is left over after the comps have had their say is therefore mostly inflation, and steering a model by it does damage.
That is exactly what the placebo test showed. Shuffling the four statistics at random and refitting the models on the nonsense produced predictions at least as good as the real ones in six of the eight era tests, run across four different ways of adding them.
The one version that worked, and what it cost
Not every version failed, and the honest thing is to report the one that didn’t.
Of the twelve new coefficients tested, exactly two were statistically solid — own power and own speed, in the “how good if he reaches” model. Both were significant in both eras, with the same sign and comparable size (power +6.12 and +7.62; speed +1.86 and +1.80). Every other coefficient either could not be distinguished from zero or flipped sign between eras.
Fitting only those two, in only that model, produced the best ground-truth result of anything tested: the board’s agreement with real career WAR rose from +0.1481 to +0.1551, with 99.5% of 400 resamples favouring the change. Real career WAR captured by the board’s top 25 rose from 900 to 1,029.
It was still declined, for two reasons visible in the same run.
It put the wrong players on the board. Karim Garcia, whose major-league career finished at −3.25 wins above replacement, rose to 16th among every hitter the site has ever scored. Delmon Young landed 20th. And Wade Boggs — the case that motivated much of this work — fell from 553rd to 753rd. A change that improves the average while promoting a −3.25-WAR career into the top twenty has not improved the product.
It is two different findings wearing one coat. The permutation test splits cleanly by era, and not subtly: training on older players and testing on newer ones, the real statistics performed worse than shuffled ones in 80 of 80 trials. Training on newer players and testing on older ones, they performed better in 80 of 80 trials. Both results are about three standard deviations from the null — in opposite directions.
That is not noise, and it is not a failed idea. It says own power and own speed genuinely carry information about how good a hitter becomes, but how much they carry has changed over the decades the database spans. A single number fitted across all of it is wrong for both ends. Since this site ranks Beltre, Boggs and Rickey Henderson in the same list as prospects who have not yet played a game, a model that is right only for one era is not usable here. Making that era dependence explicit is open work, not a solved problem.
The check that makes this trustworthy
Every number above is a comparison against the model that is actually live, which is only meaningful if the comparison machinery is honest. The safeguard here is a control arm that changed nothing: the same code path, the same refit, on the same 4,410 players, with no new statistics added at all.
It returned results identical to the shipped model — the same accuracy to four decimals, all five board checks passing, agreement with real careers moving by +0.0000. That matters because the refit replaces all three models’ coefficients whether or not you add anything. Without that control, “adding power helped” and “refitting helped” are indistinguishable, and an earlier version of this analysis could not have told them apart.
What shipped
Nothing. The prediction models are unchanged.
That is worth publishing anyway, because the negative result is more useful than the change would have been. The instinct behind it — the model doesn’t look at his power, so let it — was reasonable, widely shared, and wrong for a specific, checkable reason: the information was never missing. It was already inside the comparables, and the part that wasn’t turned out to be mostly park effects and level-repeating.
The four statistics still do the job they were always doing. They choose the twenty players whose careers the site shows you. That was never a small job, and it turns out to be most of the job.