Punished for Making Contact

Punished for Making Contact

Wade Boggs hit .325 in the major leagues for eighteen years — one of the thirty best careers in this database. He was genuinely invisible to the old model: 21,513th of 39,365 prospects. He is 361st of 39,413 today. Fixing the blind spot this page is about was real and durable, and it was not, by itself, what got him there — five separate and deeper problems did most of the work, and this page tracks every move.

← White Papers · Published 9 August 2026, updated 17 August 2026 · Covers a structural bias in the Spectral Index and the two minor-league stats that were missing from it

What this page is

The Spectral Index is the site’s one-number summary of a prospect. It is built from the minor-league profile a player showed: how much he walked, how often he struck out, how much power he hit for, how fast he ran. This page is about a class of player that profile systematically failed to see — and the finding, which was a surprise, that the model was not merely neutral toward contact hitters but actively biased against them.

That fix is still real and still shipped. What changed is what it’s evidence for. The fix genuinely repaired the two missing statistics. It could not, by itself, fix a second and unrelated problem in how the model judges a player’s minor-league comparables — and once that problem was fixed too, five days later, Boggs’s rank moved again, in the opposite direction. Both fixes are correct. Their effects on one specific player, added together, are not what either page originally promised. That is the actual subject of this update.

Update · 11 August 2026. The two statistics described here now go through the same preparation every other rate in the model gets: shrunk toward the league average by how many chances the player actually had, adjusted for whether he was young or old for his level, and capped so no single season can dominate. They had originally been added without those steps, which let short, late-career hot streaks at low levels count as if they were full seasons. Walk-to-strikeout is also now measured as walks ÷ (walks + strikeouts) — the same ordering, but a bounded rate rather than a ratio that runs away when strikeouts approach zero. Every figure below is measured on the corrected model.

Boggs’s rank, honestly, all the way through

Wade Boggs produced about 91 wins above replacement — one of the thirty best careers among every position player in this database (#30 of 13,006). Here is every move his rank has made since this investigation started, with what caused each one:

Date Rank What changed
Before this work 21,513th of 39,365 The old model: no AVG, no BB/K-together, defensive-value blind spots already known
11 Aug 2026 680th This page’s fix — AVG and BB/(BB+K) added as real features, this page’s subject
14 Aug 2026 307th A stale-coefficient refit, unrelated to this page — see Worse Than Guessing
15 Aug 2026 553rd of 39,392 The comp-outcome term made honest — see The Survivor’s Median — and the refit that fix required
16 Aug 2026 448th of 39,402 Seasons after he stopped being a prospect removed from his own record — see below — and the refit that required
16 Aug 2026 (later) 442nd of 39,402 A reliability discount on his own minor-league numbers came off entirely — see below
17 Aug 2026 361st of 39,413 A separate class of record excluded from the coefficient fit and reference pool — see Where the Data Begins. Boggs’s own record was untouched (his first season isn’t mid-ladder); the refit that fix required is what moved him.

A note on the “before this work” row: that figure and its attribution to specifically missing AVG/BB-K are a snapshot of the model as it existed on 9 August. The underlying comp-engine term this table’s baseline also depends on has itself been substantially revised since (see The Survivor’s Median), so 21,513th cannot be independently reproduced today, and neither can the exact share of the miss owed to bat features alone versus everything else that has moved since. Recomputed today (17 August), with no bat features and the current comp engine: Boggs is 3,865th of 39,413 — a dramatically better number than either historical estimate, driven almost entirely by the comp-engine revision, not by anything tracked on this page. Read the table’s first row as “an earlier, no-longer-reproducible measurement,” not as a controlled baseline.

The first three moves are the good-news version of this story, and they are true as far as they go. The fourth move is not a regression in this page’s fix — AVG and BB/(BB+K) are untouched by it, still shrunk, still age-adjusted, still winsorized, still doing exactly what the next three sections describe. It is a second, independent piece of the model catching up to reality, and it landed on the same player. Read the two pages together, not as a contradiction: this page fixed a blind spot in what the model measures about Boggs’s own numbers; The Survivor’s Median fixed a separate, more consequential blind spot in what the model does with his twenty closest comparables. The second fix is the larger of the two, and it moved him further than the first one gained.

The fifth move is a third independent correction, and it had nothing to do with Boggs at all. A player’s minor-league profile was being built from every minor-league season in his record — including seasons he played after he had already established himself in the majors and been sent back down, on a rehab assignment or an option. Those seasons exist only because the player reached the majors, which is the very thing the model is trying to predict. Using them is looking at the answer before guessing. They are now excluded from a player’s own record; they still count toward the league averages he is measured against, which is correct, because they were real games played in a real league. Boggs has such seasons, and so does nearly every player good enough to reach the majors and be optioned back down — which is exactly the population the mistake was quietly flattering. Removing them, and refitting the model that change invalidated, moved him from 553rd to 448th.

The sixth move is a fourth independent correction, and it is the one most directly in tension with this page’s own subject. Beyond simply having a batting average and a walk-to-strikeout number, the model discounted how much to trust a player’s own numbers by how much minor-league playing time backed them up — a prospect with 200 plate appearances got less credit for a given batting line than one with 2,000, on the reasonable idea that a short record is less trustworthy than a long one. That discount was set years before the model started shrinking each season individually toward the league average, the way it does today. Once seasons were already being shrunk one at a time, discounting the career total on top of that was measuring the same uncertainty twice — and the double-count fell hardest on exactly the players who reach the majors fastest, and therefore have the shortest minor-league record, which is close to the opposite of what a prospect model should be rewarding. Tested against every player whose career is now finished, in both directions of baseball history, removing the discount entirely made the model measurably better at identifying real careers — the single largest gain of any lever tested for this update — with no matching cost found. Boggs, with 10,740 real minor- and major-league plate appearances, was never the player this discount was quietly punishing, so he barely moves on it: 448th to 442nd. The players it was built for are the fast risers this page does not feature — a hitter good enough to reach the majors before he ever accumulates a full season of minor-league at-bats, whose thin record used to read as a reason for doubt instead of what it actually is.

The seventh move is a fifth independent correction, and — like the fourth — it had nothing to do with Boggs’s bat-to-ball skills at all. A separate class of record, whose real minor-league career began before this database’s coverage does, had been getting scored as if a single short fragment were a complete, reliable one (see Where the Data Begins for the full mechanism, discovered via a player named Clint Hurdle rather than Boggs). Those records were excluded from the coefficient fit and from the reference pool other players are measured against, and the model was refit accordingly. Boggs’s own record was never touched by this — his first recorded season is 1977, but it starts at High-A, not mid-ladder, so the display rule that flags a genuinely incomplete record never fires on him. The refit that fixing everyone else’s records required is what moved him: 442nd to 361st of 39,413.

The first instinct, back in August, was that this was a defensive-value problem — Boggs was a fine third baseman, and the index had already been caught undervaluing gloves. It wasn’t. Boggs’s offensive value alone was enormous, and the old model missed that too. Something about how he hit in the minor leagues read as unremarkable to a model that simply never looked at the numbers that would have flagged him. That diagnosis held up. It just was not the whole diagnosis.

He isn’t an exception, he’s a type

The way to find out whether one miss is a fluke or a pattern is to stop looking at the miss. Instead: take every player the model scored, subtract off the part of his real career the model did explain, and ask what minor-league statistic predicts the leftover. If some stat consistently predicts the part the model gets wrong, that stat is a hole in the model.

This method has a useful safety feature built into it. Any statistic the model already uses should come out flat — it can’t predict the model’s own error if the model already accounts for it. So the test grades itself.

Two statistics came out large, and neither is anywhere in the index:

Minor-league stat Predicts the model’s error Already in the model?
Batting average strongly (+0.30) no
Walk-to-strikeout ratio strongly (+0.26) no
Batting average on balls in play moderately (+0.13) partly
Isolated power none (−0.11) yes, heavily
Gap power (doubles and triples) flat (+0.05) partly

The control worked exactly as designed: isolated power, which the index leans on hard, predicted nothing about the model’s mistakes. That is what a stat already being accounted for looks like, and it is the reason to believe the other rows.

The model does track walks, and it does track strikeouts. What it never does is look at them together. A hitter who walks 60 times and strikes out 50 is a fundamentally different creature from one who walks 60 times and strikes out 160, and the index cannot tell them apart. It also never looks at batting average at all.

Blind, not biased

It is tempting to say the index is hostile to contact hitters — that it prefers the strikeout-prone slugger. It does not. Measured directly, the old index’s correlation with minor-league contact rate is essentially zero: it was tilted neither for nor against putting the bat on the ball.

The failure is quieter than hostility, and for a prospect model it is worse. The index simply never priced what contact is worth. Contact rate predicts the part of a player’s real career value the index misses — a correlation of +0.24 — while the index’s own correlation with contact sits at zero. The information was there the whole time, uncollected: the one thing a contact hitter is best at leaves no mark on his score. Not penalized. Unseen.

Line up the players it misses and it is one profile written ten times:

Career WAR Batting average Walk-to-strikeout Power
Wade Boggs 91.4 well above far above below
Robin Ventura 56.1 above far above below
Mark Grace 46.4 far above far above above
Tony Fernández 45.3 above above below
Plácido Polanco 41.9 above average below
Yadier Molina 41.7 average average below
Ketel Marte 36.9 above average below
Salvador Pérez 35.4 above average below
Michael Brantley 34.1 above far above below
DJ LeMahieu 30.6 far above average below

Hits everything, rarely strikes out, no home runs. The index reads the last item and stops reading.

Does fixing it actually help, or does it just help Wade Boggs?

Anyone can add a term that rescues one favorite player. The question is whether the model gets better at prospects it has never seen.

So the model was fit on one era of baseball history and tested on a completely different one, in both directions — train on the old players, predict the modern ones, then the reverse — and graded on how well the resulting score ranked the players it had never seen against the real careers they went on to have.

Adding the two statistics improved that held-out ranking by +0.050 in one direction and +0.025 in the other — a gain in both eras, from two numbers that had been sitting in the data the whole time.

Then the shuffle test. Any change that reshuffles a ranking can score better by luck alone, so the same procedure was run eighty times with the two statistics randomly reassigned between players. None of the eighty scrambled versions beat the real one, in either direction (p = 0.012).

What didn’t survive

Batting average on balls in play. It predicted the model’s error on its own — but it is tangled up with batting average itself, the two correlating at 0.71, and once plain batting average was in the model, BABIP added nothing on held-out data. Real signal, but redundant with a simpler stat already doing the job. Dropped.

Both late-career surges. A hitter cutting his strikeouts just before reaching the majors — or spiking his power — sounds like exactly the developmental signal worth catching. Both were measured; both were real enough on their own to look at. Neither survived once batting average and walk-to-strikeout were in the model: the contact surge added +0.005 to the held-out ranking and the power surge +0.001. That is the honest answer to the obvious objection — what about a late power spike? It was checked, and the two plain rate stats had already captured whatever was in it. Dropped.

Gap power. Doubles and triples looked like a promising proxy for the line-drive hitter the model was missing. It predicted nothing the model wasn’t already getting — almost certainly because it is entangled with the power measure already in there. Dropped.

A badly designed test that had to be thrown out and redone. The first attempt to prove these new terms weren’t just gaming the metric compared them against a cumulative career total. That test cannot fail, because every cumulative total is dominated by how long a player stayed employed, and anything that predicts staying employed will appear to succeed. Redone against rate statistics, where playing time divides out. The conclusion held, but the first version proved nothing and it is worth saying so.

The honest limits

The gain is larger among players who reached the majors than across all prospects. Most prospects never play a big-league game, and no amount of batting-average signal changes that; this makes the list better at ordering the players who arrive more than at predicting who arrives.

The old-era direction is consistently weaker than the modern one. Some of that is real drift in how the minor leagues work; some is that active players have unfinished careers. The two directions are reported separately throughout, deliberately, rather than averaged into one flattering number.

And the shipped model is deliberately spare: just batting average and walk-to-strikeout ratio. Everything else described above was measured and set aside — BABIP for redundancy, both surges for adding essentially nothing once the two rate stats were in. A blind spot this large did not need four patches; it needed the two the model had never bothered to look at.

An earlier version of this analysis added an “at least three seasons of history” flag to the model as a way of handling that missing-data problem cleanly. It had to be removed: the flag turned into a proxy for how fast a player was promoted rather than how he hit, and for a current prospect who has simply not played three seasons yet, it would have handed out a large unearned boost. Measuring that trap is the reason it isn’t in the shipped version.

The general standard

Three independent checks, the same three every scoring change here faces. Walk it through real named players from beginning to end. Fit it on history it can see and test it on history it cannot, in both directions. And run it against eighty randomized versions of itself, because a change that only beats nothing has not been tested.

Feeding batting average and walk-to-strikeout into a proper two-part model, and scoring by expected career value instead of the old index, moved Wade Boggs from 21,513th to 680th of the 39,365 hitters in the database — real, validated, still standing. Three further, unrelated corrections have moved him since: a stale-coefficient refit took him to 307th, making the comp-outcome term honest (see The Survivor’s Median) took him to 553rd of 39,392, and excluding a separate class of left-censored record from the coefficient fit (see Where the Data Begins) took him to 361st of 39,413. None of those moves is this page’s fix failing. Each is a separate blind spot, distinct from the one this page found.

What the two pages together actually show is more useful than a single rising rank number would have been: the model no longer mistakes Boggs’s own minor-league line for an unremarkable one — that was real, and it was fixed. What the model still cannot fully do is see, in twenty statistically similar minor leaguers judged honestly, whatever it was that let Boggs separate himself from all of them once he reached the majors. His own comp-outcome number — how good his closest comparables’ careers honestly turned out to be, one comp misread apart — rose from −0.69 to +0.73 in that second fix: a real, large improvement, correctly measured. It just was not, on its own, enough to carry him back up, because on a strict reading of statistically similar minor-league careers, his own record was that of a very good prospect, not a historically great one. The greatness came later, in a place this system — built on comparing minor-league stat lines to other minor-league stat lines — was never going to be able to see in advance. That is the honest limit of a comp-based model, not a bug still waiting to be found. It ranks him thirtieth of 13,006 real careers today; the model gets him closer than it used to, not all the way there, and this page is no longer the one claiming it did.

← Back to White Papers

Prospect Wavelength is not affiliated with MLB, nor with any of their digital properties. Email any questions, comments, or concerns to stealofhome@prospectwavelength.com