When the Obvious Number Is Wrong

When the Obvious Number Is Wrong

The deeper story behind a few of the site's least obvious decisions — worked examples, what was tested and rejected, and how each change was actually checked against reality.

← White Papers · Published 2 August 2026 · Covers reach probability, the age adjustment, and comp selection

Update · 9 August 2026. Since this paper was published, the Spectral Index has been redefined. It is no longer a single blended score of reach-probability and median outcome; it is now a prospect’s expected career value — the probability he reaches the majors multiplied by how valuable he projects to be if he does — shown as four numbers: Confidence, Impact, Bust risk, and Boom chance. Where this page treats the Spectral Index as the older composite, that describes how the model worked at the time; the readings covered here still apply to its inputs. See Punished for Making Contact and The Glove That Kept Disappearing.

What this page is

How It’s Made covers the mechanics: what a stat is, how it gets adjusted, how comps get found. This page goes one level deeper, into a handful of decisions that aren’t obvious from the mechanics alone — cases where the “obvious” number turned out to be misleading, and the fix required actually checking the claim against real outcomes rather than reasoning it out from a spreadsheet. Each section below walks through one real player, the wrong intuition their case exposes, and the numbers behind the fix. A closing section covers ideas that looked promising, got tested with the same rigor, and were thrown out.

“100% of his comps reached the majors” doesn’t mean 100%

Every player page shows a reach probability — the share of a prospect’s closest comparable players, weighted by how similar they are, who went on to play in the majors at all. It’s a simple, honest number by construction: count up the comps, weight by similarity, done. The problem only shows up for young players at low levels, where a small, extreme-looking comp pool can round up to a number the data doesn’t actually support.

Seth Hernandez is a good illustration. Pitching at High-A in 2026 as a 20-year-old, his 20 closest career-wide comps — weighted by similarity — had all reached the majors. Read literally, that’s a 100% reach probability. But 100% isn’t a real number here; it’s what a small, favorable sample looks like before it’s been checked against history. Pull every historical High-A pitcher who cleared enough playing time to have a real 20-comp career-wide pool and group them by their own raw reach estimate, and the players who landed in that same “close to 100%” bucket only actually reached the majors about 87% of the time — still very good, but a meaningfully different number than the one on the page.

The fix isn’t to shrink every young player’s number toward some average, the way an early attempt at this did. That approach was tested against real 2026 scouting consensus (MLB Pipeline’s Top 100 and every team’s internal Top 30) and made things measurably worse — it dragged elite, thin-sample players like Hernandez down toward a flat league-wide target regardless of how good their specific comps were, which is exactly backwards. The actual fix learns, separately for each level, what a given raw estimate has empirically turned into for real players who once had that same number — a monotone curve fit from history, applied only at the four lowest levels (Rookie ball through High-A) where this kind of overconfidence actually shows up. At Double-A and Triple-A the raw number was already accurate, so it’s left alone. Hernandez’s own number moves from a literal 100% to a calibrated 96.3% — still elite, just honest about the size of the sample behind it.

The same fix also settles something the raw number couldn’t: two players can share an identical-looking coarse reach estimate for very different reasons. Diego Frontado and Todd Taylor, both thin-sample Rookie-ball hitters, looked almost the same by the old measure. The calibrated version separates them cleanly — Frontado to 89.1%, Taylor to 46.3% — because it’s sensitive to exactly where each player’s raw estimate sits, not just which broad bucket it falls in.

How this got checked, and a mistake caught along the way. Before shipping, the redesigned reach probability was run back through the same 2026 scouting-consensus lists and compared against the version it replaced. Three of the four lists it was checked against came out clearly better; the smallest one (25 pitchers on the Pipeline Top 100) came out worse, though a randomized control confirmed that specific result was still well clear of noise. That trade was judged worth it — team lists are a larger, steadier signal than a single global top-100 cut. In the course of that check, a live comparison against the site’s own running code caught a real bug in the validation scripts themselves (not the shipped calculation) — an outcome field had been swapped for pitchers, throwing off every prior comparison for this specific change. It was found, fixed, and the whole comparison rerun before anything went live. Being able to catch and correct that kind of thing, rather than just trusting a plausible-looking number, is part of the actual process here — not just a footnote.

Being young for your level is worth more than it looks

Braylon Payne, playing High-A ball in 2026 at 19 years old, struck out in 28.7% of his plate appearances that season — noticeably worse than that league’s 23.9% average, the kind of number that reads as a real weakness at a glance. But the average hitter he was facing at that level, weighted by playing time, was 22.2 years old — over three years older. Once his strikeout rate is adjusted for that gap, it works out to an effective 21.4%, better than league average, not worse. The raw number was measuring two different things at once — actual contact skill, and simply not yet being as physically mature as most of the competition — and only one of those is a real red flag.

This is why every rate stat on the site is adjusted for the batter’s or pitcher’s age relative to the level they’re playing at before anything else happens to it. A young player facing older competition and holding his own is a meaningfully different and better signal than an old-for-his-level player doing the same thing, even when the raw stat line looks identical.

When the closest 20 comps aren’t enough

The site’s outcome ranges (a below-average, typical, and above-average projection) come from a player’s most similar historical comps, weighted by how close a match each one is. For most players, the closest 20 comps are a good, representative sample. For a player whose true peer group is thin — most often someone from an earlier, thinner-data era of the game — a fixed count of 20 can force in players who aren’t really comparable at all, just to fill the quota.

Ricky Powell, a hitter from the pre-1992 era with limited historical company, is the case that first exposed this. His nominal top-20 comp list, ranked purely by similarity score, spanned matches as loose as 659 out of a possible 1,000 — nowhere near the quality of match a more recent player gets. Rickey Henderson, an all-time great with only a passing resemblance to Powell’s actual offensive profile, ranked 2nd by career value among the reached comps in that loose pool, which meant Henderson’s own real career numbers were mechanically setting Powell’s “above-average outcome” projection — not because Henderson was a fair comp, but purely because the count needed one more name.

The fix drops the fixed count of 20 entirely for this specific calculation and instead looks for however many genuinely close comps actually exist, at a much higher similarity bar, expanding the search only as far as it has to for a trustworthy sample. For a player with plenty of close comps, that’s a tight, high-quality group. For a thin-era player like Powell, it can mean looking further afield than 20 names, or in rare cases falling back to a broader weighted average across the whole eligible pool rather than forcing a fixed count. Applied to Powell today, his outcome range is built from comps clearing a similarity floor of 850 rather than the old fixed-20 cutoff — and his above-average projection no longer rides on a single loosely-matched Hall of Famer. Reach probability itself is untouched by any of this; it’s still read directly from the closest 20 comps, which turned out to be the right sample size for that specific number even though it wasn’t for this one.

What didn’t survive contact with real outcomes

Not every idea that looks good in a first pass survives a second, harder look. Two examples, kept here because an honest accounting of what a model does well should include what got tried and rejected, not just what shipped.

Giving credit for playing a premium defensive position. The idea: a shortstop or catcher facing the same offensive learning curve as a first baseman is doing something harder, and should get some credit for that independent of how his bat grades out. Two different versions of this were built. The first tried folding position into how comps get matched in the first place; checked against a random-permutation control, the apparent gain turned out to be mostly noise — barely different from randomly shuffling which position value went with which player. The second version left comp-matching alone and instead added a position-based bonus directly to a player’s final ranking. That one looked much better on an initial check — a real, statistically solid improvement against that season’s real prospect rankings. But re-run against a second season’s independent rankings, using a version of the test that couldn’t see any of a player’s own future performance, the improvement collapsed to essentially nothing. The initial result had been real, but it was an artifact of being tuned and measured on the same data; neither version shipped.

The reach-probability fix above, before it was a fix. Before landing on the level-by-level calibration described earlier, several other designs were tried and discarded in sequence — blending in a separate level-specific comp pool, swapping to a level-specific reach rate outright, a lookup table keyed to level and sample size. Each one looked reasonable and was checked the same way: real 2026 scouting consensus, not just an internal metric. Each one failed that check, most clearly on a design that had briefly been accepted before a wider set of real team rankings than the ones it was tuned against showed it making a well-regarded prospect’s page look dramatically worse than it should. That failure is what motivated actually diagnosing why the reach numbers were off in the first place, rather than continuing to guess at how hard to shrink them — which is what led to the version that shipped.

The general standard

Every change described here — and every one that didn’t make it past this stage — gets checked the same three ways before it’s trusted: against a real, named player whose case motivated the idea in the first place; against a held-out slice of history the change couldn’t have seen while being built; and, wherever the change could plausibly just be reshuffling which players end up as comps rather than adding real information, against a randomized version of itself. A result that only clears the first of those three isn’t treated as settled. Several of the changes above only reached their final form after failing one of the later two.

← Back to White Papers

Prospect Wavelength is not affiliated with MLB, nor with any of their digital properties. Email any questions, comments, or concerns to stealofhome@prospectwavelength.com