Note: later the same day, Spectral Index itself was recalibrated from 50+15z to a true 20-80 scouting scale (50 = average, 10 points = one population standard deviation) — see “Why 10 points per standard deviation” below. Every figure in this paper was re-extracted from the live site afterward; the mechanism and every qualitative conclusion are unchanged.
Second note, 14 August 2026: the model’s internal settings were re-derived that day (Worse Than Guessing), which moves every Spectral Index and therefore every Eclipse figure quoted below. Those figures have not been re-extracted. They are kept as measured because this paper is a record of five rejected designs and what each one produced at the time — the numbers illustrate why each failed, and no conclusion here turns on their exact values. Wade Boggs, the paper’s running example, has since moved from below average as a prospect to 307th of 39,378, and then, on 15 August 2026, to 553rd of 39,392 once a separate comp-outcome fix landed — see The Survivor’s Median.
Third note, 16 August 2026: the same applies again, and more so. A further correction — minor-league seasons played after a player had already established himself in the majors no longer count toward his own prospect record — moved every Spectral Index once more, and the population spread of the index itself has roughly doubled since these worked examples were extracted, which changes every “50 + 10 × z” line below. Boggs is now 448th of 39,402 with a displayed Spectral Index of 68.2, not the 40.93 in the table near the end of this page. Those figures have deliberately not been re-extracted, for a specific reason beyond the one above: the Blind Spots page these numbers come from is itself under active repair, and republishing a fresh derivation off a surface known to need work would be presenting numbers this project does not yet stand behind. The five rejected designs, the reasoning that rejected them, and the arithmetic of the scale are all unaffected — they are what this paper is about. The specific values in the two worked examples are a snapshot of 7 August 2026 and should be read as one.
Fourth note, 16 August 2026 (later): a career-level discount on a player’s own batting-average and walk-to-strikeout numbers, scaled by how much minor-league playing time backed them up, has been removed — it was double-counting uncertainty the model already accounts for at the season level, and it fell hardest on the fastest-rising prospects. Boggs moves again, to 442nd of 39,402 (displayed Spectral Index 68.8); the worked example below has not been re-extracted, for the same reason as the note above.
Fifth note, 17 August 2026: minor-league records that began before this database’s coverage does were being scored as complete, reliable ones, and are now excluded from the coefficient fit and the reference pool — see Where the Data Begins. Boggs’s own record is unaffected by the fix directly, but the refit it required moves him again, to 361st of 39,413 (displayed Spectral Index 69.8). The worked example below remains unextracted for the same reason as the two notes above.
Update · 9 August 2026. Since this paper was published, the Spectral Index has been redefined as a prospect’s expected career value — the probability he reaches the majors multiplied by how valuable he projects to be if he does — shown as four numbers: Confidence, Impact, Bust risk, and Boom chance. Where this page treats the Spectral Index as the older composite, that describes how the model worked at the time.
The Eclipse Index below was built to measure something else entirely — whether a prospect score and a finished career can share one scale — so its verdict on the model is incidental, which is exactly what makes it useful corroboration rather than a circular check. Under the new expected-value index, the largest misses it flags shift off the contact hitters the old index buried — Wade Boggs, Edgar Martínez, Robin Ventura, Mark Grace — and onto the high-strikeout sluggers the old index over-rewarded: Fred McGriff, Bobby Bonilla, Mark Trumbo. See Punished for Making Contact and The Glove That Kept Disappearing.
The question
Every player on this site carries a Spectral Index — a score built entirely from how his statistical comps turned out, computed without any knowledge of what he himself went on to do. For players whose careers are finished, we also know exactly what happened. Wade Boggs is in the Hall of Fame. Clayton Blackburn never threw a major league pitch.
So a natural question: where was the model most wrong?
That question is harder than it looks, because the two things being compared are not the same kind of number. Spectral Index is a unitless composite. A finished career is a pile of counting stats — 10,740 plate appearances and +518.3 runs above average for Boggs, zero and zero for Blackburn. There is no meaningful way to subtract the second from the first until both have been forced onto a common scale, and every choice made in forcing them there changes the answer.
This paper documents five versions of that common scale that failed, the one that shipped, and the places where it is still visibly imperfect.
The reference population
Every number here is computed against the same group: the 20,743 players flagged historical-comp-eligible, meaning their careers are old enough and resolved enough to be used as comps for today’s prospects.
Of those, 8,686 reached the majors and 12,057 did not — 58.1%. That 58.1% is not a footnote. It is the single most important fact about this distribution, and it breaks three of the five failed attempts below.
Five versions that did not survive
1. Ceiling: the best career among a player’s comps
The first attempt did not use Spectral Index at all. It took each player’s comp list and reported the best real career in it, on the theory that the most optimistic comp represents what the model thought was possible.
It fails because comp pools overlap. One historically great player — peak Barry Bonds, say — lands in the comp pool of many unrelated prospects and hands every one of them the same inflated ceiling. The number stopped describing the player and started describing whichever legend happened to drift into his neighbourhood.
Replaced by Spectral Index, which is the site’s actual headline estimate and is specific to the player.
2. Percentile against the full population
The obvious way to scale a career: rank it against all 20,743 players and convert the percentile to a score.
With 58.1% of the population having never reached the majors, merely appearing in one big league game puts a player above the 58th percentile before he records an out. Every real career — Boggs, a September call-up, a twelve-year middle reliever — gets compressed into the top 42% of the scale. The differences that matter most are squeezed into the smallest space.
3. Raw z-score of career value
If percentiles compress, use the raw magnitude: take career wRAA (hitters) or fRAA (pitchers), subtract the mean, divide by the standard deviation.
This restored the spread and broke something worse. Hitters and pitchers were being standardized against their own dataset’s standard deviation, and those two spreads are not the same. The observed result put Roger Clemens above 300 and Albert Pujols near 200 — not because Clemens had the more dominant career, but because the pitcher fRAA distribution happens to be tighter than the hitter wRAA distribution. The scale was measuring which dataset a player was in.
4. A blended, credibility-weighted percentile
The fourth version was the most elaborate: rank a player’s career value only against other players who reached, blend that with a percentile rank of career length, shrink the result toward a floor using a pseudo-count for players with tiny samples, and place non-reachers below the worst real career on record.
It handled several real problems correctly. It also had a structural flaw that no amount of tuning fixed: because a non-reacher’s score was pinned near the bottom of the scale while his Spectral Index could be anything, the model’s most confident misses on players who never played always outranked genuine Hall-of-Fame surprises. Clayton Blackburn, a pitcher the model liked at 75.8 who never reached the majors, produced a larger gap than Wade Boggs, an actual Hall of Famer the model rated at 39.6.
Sweeping the non-reacher penalty across its entire legal range — all the way to the floor of what “worse than any real career” permits — moved Blackburn from 53.2 to 48.7, against Boggs’s 48.2. A half-point. The lever could not reach.
5. The right formula, standardized the wrong way
The fifth version is the formula that shipped, scored correctly, and then converted to a display scale using the mean and standard deviation of the raw score. The formula was right. The conversion was not:
| min | p25 | median | p75 | p95 | max | |
|---|---|---|---|---|---|---|
| Outcome Grade, mean/SD | 46.2 | 46.2 | 46.2 | 48.9 | 65.7 | 210.2 |
| Spectral Index, for reference | 13.9 | 44.3 | 49.4 | 55.1 | 69.2 | 84.8 |
The minimum, the 25th percentile and the median are the same number. Career value is a counting stat with a hard floor at zero and no ceiling — 58.1% of the population is tied at the bottom while Pujols sits at 1,504.76 — and standardizing something that skewed puts most of the population on top of each other and stretches the top to +16 standard deviations. The resulting scale spanned well over a hundred points against Spectral Index’s own range of roughly 70, so subtracting the two produced a number dominated by whichever way Outcome Grade happened to swing.
What shipped
Four equations, applied in order.
Seasons
Seasons = MLB PA / 502 (hitters)
Seasons = MLB IP / 162 (pitchers)
Why. Hitters and pitchers are measured in different units and have to be reconciled before they can share a list. A hitter’s playing time is plate appearances; a pitcher’s is innings. Dividing each by one full season’s worth — 502 qualifying plate appearances, 162 innings — converts both into the same unit: seasons of major league work. A player who never reached scores 0.
This matters more than it sounds. The raw numbers are not interchangeable: one inning contains roughly 4.3 plate appearances, so treating 100 IP and 100 PA as the same amount of career would badly misprice every pitcher on the list.
Score
value term = V if V >= 0
V / (1 + |V| / tau) if V < 0
Score = value term + (lambda * Seasons) + (gamma if reached MLB, else 0)
where V is career wRAA (hitters) or fRAA (pitchers), tau = 30, lambda = 20, gamma = 20.
Why the softened value term. Both wRAA and fRAA are runs above average, so a long career spent slightly below average accumulates a large negative number. Jeff Suppan threw 2,542.7 major league innings at -120.3 fRAA. Left raw, that made him one of the worst outcomes in the database — absurd for a pitcher who held a rotation spot for a decade and a half. The soft floor V / (1 + |V| / tau) leaves positive values untouched and compresses negative ones toward an asymptote at -tau: Suppan’s -120.3 becomes -24.01, a real penalty that cannot swallow the rest of his career.
Why the longevity term. Organizations do not hand out 2,500 innings by accident. Sustained playing time is itself evidence that a player was good enough to keep, in a way that a six-plate-appearance cameo is not. lambda = 20 prices one full major league season at 20 points.
Why the arrival bonus. This is the term that separates “bad major leaguer” from “never a major leaguer,” and it is the one the whole page turns on — see the next section.
Outcome Grade
z = probit( percentile(Score) )
Outcome Grade = 50 + 10 * (z - mu) / sigma
where percentile is the score’s rank among all 20,743 eligible players, probit is the inverse normal CDF, and mu/sigma are the mean and population standard deviation of Spectral Index’s own raw composite score over the same 20,743 players.
Why percentile rather than mean and standard deviation. This is the fix for failure #5. Ranking is immune to skew: it does not care that Pujols is at 1,504.76 while most of the population is at 0, only that he is first. The percentile is then pushed through probit, which spreads a flat 0-to-1 rank back into a bell-shaped score with a sensible middle and thin tails.
Why 10 points per standard deviation, and Spectral Index’s mu/sigma rather than Outcome Grade’s own. Outcome Grade goes through the identical map Spectral Index itself uses (compositeDisplayValue()): 50 = average, 10 points = one population standard deviation of Spectral Index’s own composite score (pooled across both hitters and pitchers, since the two datasets’ composite formulas blend a different number of correlated terms and land on genuinely different raw spreads — 0.615 for hitters, 0.853 for pitchers), unclamped. Using Spectral Index’s constants rather than fitting a fresh mu/sigma to Outcome Grade’s own distribution is what makes the subtraction in the next equation a legitimate operation: a one-point difference has to mean the same thing in both columns, and it only does if both went through the same map with the same reference population and the same points-per-SD. (Outcome Grade’s own spread, once put through that map, is not itself exactly 10 points per SD — it doesn’t need to be; only the map has to match.)
The result runs 42.8 to 104.4, against Spectral Index’s 13.9 to 84.8. Comparable ranges, comparable construction.
Eclipse Index
gap = Spectral Index - Outcome Grade
Eclipse Index = 50 - 10 * ( (|gap| - mu) / sigma )
where mu = 7.5679 and sigma = 6.0431 are the mean and population standard deviation of |gap| across the 20,743.
Why absolute value. A prospect the model loved who never played and a Hall of Famer the model ignored are both misses. The page is about the size of the error, not its direction.
Why the minus sign. Every other grade on this site reads high-is-good. Inverting means a well-calibrated player scores near the top and a catastrophic miss scores near the bottom, so Eclipse Index reads like a scouting grade rather than an error bar. The table sorts ascending by default, putting the worst misses on the first page.
Why 10 points per standard deviation. Eclipse Index is not being subtracted from anything, so it is free to use the site’s ordinary 20-80 scouting-grade convention rather than Spectral Index’s 15.
Choosing gamma
gamma is the only parameter here with a sharp, checkable consequence, because it alone decides whether a bad major league career outranks no major league career.
A player who never reached scores exactly 0: no value, no seasons, no bonus. The question is how many real careers fall below that.
| gamma and its condition | Real careers scoring below a non-reacher |
|---|---|
gamma = 5, awarded when career value > 0 |
1,888 of 8,686 (21.7%) |
gamma = 5, awarded on reaching MLB |
38 (0.4%) |
gamma = 20, awarded on reaching MLB |
0 |
The first row is the failure state. Zero in this formula means “exactly league average,” which is a better outcome than the bottom fifth of real major league careers — so more than a fifth of everyone who ever played ranked below every player who never did. Al Pardo caught 132 major league plate appearances at -22.1 wRAA and scored -7.47, beneath a player who never left Double-A.
Two changes fixed it. Keying gamma on reaching the majors rather than on finishing above average took it from 1,888 to 38. Raising gamma to 20 — one full season’s worth of longevity credit, awarded for arriving at all — took it to zero. The worst real career on the site now scores 12.53, comfortably clear of the non-reacher’s 0.
No artificial floor was needed. The separation falls out of the arithmetic: a career bad enough to approach the -tau asymptote necessarily accumulated enough playing time to earn longevity credit along the way.
Two worked examples
Nothing is rounded until the last line.
Wade Boggs, the largest miss on the site
| Step | Value |
|---|---|
| MLB plate appearances | 10,740 |
| Career wRAA | +518.3 |
| Seasons = 10,740 / 502 | 21.3944 |
| Value term (V >= 0, so raw) | +518.3 |
| Longevity = 20 * 21.3944 | +427.89 |
| Arrival bonus | +20 |
| Score | 966.19 |
| Percentile among 20,743 | 0.998626 |
| probit(0.998626) = z | 2.9946 |
| Outcome Grade = 50 + 10 * (2.9946 - (-0.0138)) / 0.7495 | 90.14 |
| Spectral Index = 50 + 10 * (-0.6937 - (-0.0138)) / 0.7495 | 40.93 |
| Signed gap = 40.93 - 90.14 | -49.21 |
| Absolute gap | 49.21 |
| z = (49.21 - 7.5679) / 6.0431 | 6.891 |
| Eclipse Index = 50 - 10 * 6.891 | -18.9 |
The model rated Boggs below average as a prospect. He is in the Hall of Fame. At 6.9 standard deviations this is the biggest calibration failure in the database, and the Eclipse Index is comfortably below zero — the scale is unclamped and says so.
Clayton Blackburn, the largest miss in the other direction
| Step | Value |
|---|---|
| MLB innings | 0 |
| Career fRAA | 0 |
| Seasons | 0 |
| Value term | 0 |
| Longevity | 0 |
| Arrival bonus (never reached) | 0 |
| Score | 0 |
| Percentile among 20,743 | 0.290628 |
| probit(0.290628) = z | -0.5516 |
| Outcome Grade = 50 + 10 * (-0.5516 - (-0.0138)) / 0.7495 | 42.83 |
| Spectral Index = 50 + 10 * (1.7187 - (-0.0138)) / 0.7495 | 73.12 |
| Signed gap = 73.12 - 42.83 | +30.29 |
| z = (30.29 - 7.5679) / 6.0431 | 3.760 |
| Eclipse Index = 50 - 10 * 3.760 | 12.4 |
Blackburn is the cleanest oversell in the database and still grades 31 points above Boggs, which is the correct ordering and the thing four earlier versions could not produce.
Note the percentile: 0.290628. Every one of the 12,057 players who never reached the majors shares that exact value, because they all score exactly 0 and the percentile calculation splits ties evenly. That single shared number is the 58.1% tie showing through.
The best-calibrated player on the site
No player in the database lands on a gap of exactly zero. The closest is Jon Cook, a corner outfielder whose pro career ran 1998-2000 and never got past Single-A — his Spectral Index of 42.824768 sits 0.000377 away from his Outcome Grade of 42.825145 — an Eclipse Index of 62.5225, the highest value on the site. His model estimate (27.2% reach probability, comps who mostly didn’t make it) and his real outcome (never reached, tied at the same non-reacher floor every other zero-PA player shares) landed on almost exactly the same underlying number — the model wasn’t reaching for a marginal player, it correctly wrote him off. Kutter Crawford is next at 64.934308 against 64.933622, a gap of 0.000686.
The distinction is worth making precisely, because rounding hides it. Displayed to one decimal place a dozen players look like perfect calls, and at least one — Al Pardo, whose 52.924969 against 52.921577 rounds to a gap of 0.0 — is not even in the top five once the full precision is restored. He is seventh, at 0.003393.
Where it is still imperfect
Three things are visibly unfinished, and all three trace back to the same tie.
Eclipse Index does not span 20 to 80. The observed range is -18.9 to 62.5, median 52.2. The absolute value of a gap is bounded below at zero but unbounded above, so it is a one-sided distribution; standardizing it puts the best-calibrated player only 1.25 standard deviations above the mean while the worst sits 6.9 out. Standard deviations cannot manufacture a symmetric scale from an asymmetric quantity. Running the absolute gap through the same percentile transform used for Outcome Grade would produce a genuine 20-80 spread, at the cost of no longer being literally standard deviations.
The Outcome Grade median sits on its floor, at 42.8. This one is not fixable and arguably should not be: 58.1% of the population genuinely never reached the majors and genuinely had the same outcome. No monotone transform can separate a tie, and a scale that pretended to would be inventing distinctions that do not exist.
Marginal players surface near the top. Kevin Richardson’s entire major league career was six plate appearances. His Spectral Index is 13.9, the single lowest value on the site; his Outcome Grade is 57.0, because reaching at all clears the 42.8 non-reacher floor by a wide margin. The resulting 43.2-point gap grades -8.9 and puts him third on the page, ahead of Todd Helton. That is arithmetically correct and editorially odd, and it follows directly from the jump between “never reached” and “reached briefly” being larger than the distance from a cup of coffee to a solid career.
A note on method
Every number in this paper was produced by driving the rendered Blind Spots page in a real browser and reading the values the page itself computed — computeActualOutcomeScores(), computeBlindSpotsOutcomeGrades(), computeBlindSpotsEclipseIndex() and compositeDisplayValue(). None came from a reimplementation.
That is a standing rule on this project, and it exists because it was learned expensively: a separate implementation of the Spectral Index formula, written to make research easier, drifted out of date and produced a long thread of confidently wrong findings before anyone checked it against the live page.
It nearly happened again here. Two of the players in this paper — Clayton Blackburn and Julio Rodriguez — share a name with a different player in the other dataset, and an early lookup that searched by name alone silently returned the wrong one, producing a comparison table that understated the very problem it was built to measure. Names are not keys.