The Same Number Twice

The Same Number Twice

This site could always tell you how likely a prospect was to fail. It could never tell you how likely he was to be great. Adding that second number worked — it is well calibrated, it survives every test that matters, and it is now on the site. It also turned out to rank players in almost exactly the order the first number already did.

← White Papers · Published 23 August 2026 · Covers the boom probability, how it was validated, and why it is not the second axis it looks like

What this page is

Every player on this site carries a bust risk — the model’s estimate of the chance his career ends up worth close to nothing. It is an honest number and it has been tested hard. It also only answers one half of the question anyone actually asks about a nineteen-year-old.

“How likely is he to fail” and “how likely is he to be a star” sound like the same question asked from opposite ends. They are not. A prospect can be a safe bet to reach the majors and have no realistic path to stardom. Another can be far more likely to wash out entirely and still carry genuine star upside if things break right. Until this month, the site showed those two players the same way.

This page covers the number built to close that gap, what it took to trust it — and the honest finding that came out the other side, which is that the gap was narrower than it looked.

What “boom” means here

A boom is defined as a career worth more than 48 payoff points, where payoff is this site’s settled measure of what a career was actually worth: career WAR + 2 × seasons + 5 × (reached the majors).

That bar was not chosen for roundness. Earlier work on this model had used “more than 20 career WAR” as the informal definition of a star — roughly the top 1–3% of everyone who ever signed. Converting 20 WAR into payoff terms lands hitters at 47.9 and pitchers at 47.6: two independent populations arriving within half a point of each other. One shared bar of 48 was used for both rather than a per-side percentile.

That choice has a consequence worth stating plainly, because it looks like an error and is not. It leaves the two sides with different boom rates: 4.96% of hitters clear the bar and only 1.87% of pitchers. An equal-rate definition would have erased that. It was kept, because pitchers genuinely do have a harder path to a star career — injury attrition, shorter usable seasons — and normalizing the rate would have laundered a real difference in difficulty into a definitional one.

Whether it works

The claim being tested was calibration, not ranking. A probability is well calibrated if, among the players it assigns 30% to, close to 30% actually do it. That is a different and harder question than whether the model sorts players correctly, and it is the question this site is actually for.

Everything below is out-of-sample: predictions made on players the model was not fit on, using a stratified five-fold split with the coefficients refit inside every fold.

hitters pitchers
calibration slope (1.00 is perfect) 0.995 0.996
95% confidence interval 0.857 – 1.133 0.781 – 1.212
predicted vs. actual, overall 0.967 0.985
Brier score 0.0371 0.0164
Brier, naive baseline 0.0471 0.0184
skill over baseline 21.3% 11.1%
star careers in the sample 132 54

Both slopes sit essentially on 1.00. Both sides beat a naive base-rate guess by a real margin — though about half the margin the bust model manages, which is what you would expect for a rarer event.

Beating a naive baseline is a floor, not a test, so the model was also run against 80 shuffled nulls — the same fit, the same folds, with the input features randomly permuted and the coefficients refit every draw. Real beat all 80 draws on both sides, on both Brier score and ranking. Whatever the model is doing, it is not finding a pattern that random data would also produce.

The one soft spot is real and worth naming. Sorted into ten bins by predicted probability, the ninth bin — players the model puts in the 80th to 90th percentile of star odds — is under-called on both sides. Hitters there are predicted at 9.0% and actually boom at 13.5%. That is a genuine miss, and it is the only bin that misses badly.

It would be tidy to call that evidence of a known weakness: this model has been separately shown to behave like a bust-rate engine that undervalues scouting ceiling, and a systematic under-call of upside would fit that story neatly. The top bin does not cooperate. Among the highest-decile players — the ones the board actually surfaces — hitters are predicted at 31.6% and boom at 30.7%, and pitchers at 13.1% against 12.5%. Both slightly over-predict. One soft bin in the upper-middle is not a systematic failure at the top, and this page will not claim it is.

The finding nobody was looking for

Here is the part that makes this a more interesting page than “we added a number and it worked.”

Boom and bust rank players in almost exactly the same order, inverted. Across the 3,808 hitters currently active in this database, the rank correlation between the two is −0.999. Not −0.7, not −0.9. Essentially a mirror.

It is not an artifact of both numbers being scaled by the same confidence term, either. Strip confidence out and compare only the two conditional probabilities — the model’s estimate of stardom given the player reaches, against its estimate of failure given the same — and the rank correlation is still −0.995.

The practical version of that: sort every current hitter by boom, then sort them by bust, and see how far anyone moves. The largest disagreement anywhere on the board is 214 places out of 3,808. Not one player moves more than 250. And the handful of players who move furthest are all cases where both numbers are pinned at their extremes — bust at 1.000, boom at 0.000, projected impact of exactly zero — where the ordering between two saturated numbers is close to arbitrary anyway.

The reason is in how the model was built, and it is not subtle. Boom was given the identical feature set as bust. Same inputs, same frame, same confidence multiplier, with only the outcome threshold changed — failure at the bottom, stardom at the top. Two thresholds on one underlying estimate produce two orderings that are near-perfect inversions of each other. That is arithmetic, not a discovery.

So is it worth having?

Yes, but for a narrower reason than it first appears, and the distinction matters.

Boom does not reorder the board. It re-scales the question. Knowing that Leo De Vries carries a 1.1% bust risk does not tell you he has an 88.7% chance at a star career — those are different facts on different scales, and only one of them answers “how good could this get.” The site’s top hitters and its top pitchers illustrate how much that scale differs: De Vries at 88.7%, Jesus Made at 87.6%, and then the best pitcher on the board, Kade Anderson, at 35.3%. A number that reads as merely good on the hitter board is the ceiling of the pitcher board. Bust risk alone flattens that distinction; expressed as star odds it is impossible to miss.

What boom will never do is tell you a player is better than bust already said. If you were hoping for a second opinion — a prospect the risk number buries whom the upside number rescues — this is not that, and it was honest to find out before claiming otherwise.

What would make a real second axis

The open question this leaves is worth stating, because it is the next real piece of work rather than a rhetorical flourish.

A genuinely independent upside signal would have to be fed different inputs, not handed the same ones with a different threshold. The features that predict not failing — reaching the majors, staying there, accumulating seasons — are not obviously the features that predict stardom. Extreme tools that scouts weight heavily and this model barely sees; the difference between a prospect who is uniformly good and one who is elite in one dimension and ordinary elsewhere; the volatility of a record rather than its level. Those are candidates, all currently untested, and any of them could produce an upside number that disagrees with bust risk in a way that carries information.

Until one does, the honest description of the boom probability on this site is: a well-calibrated answer to a question the site could not previously answer at all, arriving in very nearly the order you already knew.

← Back to White Papers

Prospect Wavelength is not affiliated with MLB, nor with any of their digital properties. Email any questions, comments, or concerns to stealofhome@prospectwavelength.com