Ranking Farm Systems

Ranking Farm Systems

How thirty individual prospect scores become one number per organization — and the three metrics that were built and thrown away first.

← White Papers · Published 3 August 2026 · Covers the System Value column on the Teams page

Note: on 7 August 2026, Spectral Index was recalibrated from a raw 50+15z display to a true 20-80 scouting scale (50 = average, 10 points = one population standard deviation) — see “Putting it on the 20-80 scale” below. The exponential pricing curve was re-expressed to be immune to that recalibration going forward, and every figure in this paper was re-extracted from the live site afterward (rosters have also moved in the meantime, independent of that change).

Update · 16 August 2026. The curve’s steepness constant has been re-derived, and so has the criterion it is chosen against — the old criterion was structurally incapable of seeing the failure it was supposed to prevent, and a constant chosen on it shipped a Teams page where one organization’s ranking rested on a single prospect. The three sections that describe the metric the site actually runs — What shipped, Choosing the curve steepness, and Putting it on the 20-80 scale — were re-extracted from the live Teams page today and are current. The three discarded metrics and the published-list agreement figures are the original August measurements, kept as the historical record and not re-measured; each section says so where it matters.

Update · 9 August 2026. Since this paper was published, the Spectral Index — the per-prospect score this page aggregates into System Value — has been redefined as a prospect’s expected career value (the probability he reaches the majors multiplied by how valuable he projects to be if he does), shown as four numbers: Confidence, Impact, Bust risk, and Boom chance. System Value still sums a value curve over each organization’s prospects exactly as described here; the per-prospect number feeding it, and the curve’s steepness, were re-derived on the new index. See Punished for Making Contact and The Glove That Kept Disappearing.

The question

Every player on this site already carries a Spectral Index — a 20-80 score summarising how his statistical comps turned out. Ranking one organization against another means collapsing a few hundred of those scores into a single number, and there is no obvious right way to do it.

The problem is that the obvious ways are all wrong in interesting ways, and the wrongness only shows up when you look at named players rather than at a formula. This paper walks through three metrics that were built and discarded, the one that shipped, and the places where it still disagrees with the published lists — including one disagreement that turned out to be the most useful result of the whole exercise.

There is no ground truth, so pick honest anchors

“Which farm system is best right now” has no answer you can check against. Nobody knows yet. The prospects in question have not played their careers.

So the metric was calibrated against two published lists instead: MLB.com’s 2026 preseason farm system rankings and FanGraphs’ organizational surplus-value table. Neither is truth. They are two informed opinions, and the useful thing about having two is that you can measure how much they agree with each other.

They agree at a Spearman rank correlation of 0.671.

That number sets the ceiling for this whole exercise. A metric that matched MLB.com at 0.95 would not be a better metric — it would be a metric that had learned to imitate one particular list, including the parts of that list FanGraphs thinks are wrong. Anything in the neighbourhood of 0.671 is doing about as well as a second expert opinion does.

There is a second, more practical reason not to chase agreement. Published lists are snapshots, and rosters move underneath them. A player who gets called up stops being prospect-eligible, and the system he leaves gets quietly re-ranked without anyone publishing a correction. Some fraction of any disagreement is just calendar drift.

Three metrics that did not survive

The figures in this section are the August 2026 measurements that decided each question at the time. Rosters and the index have both moved since; where a comparison still reproduces on today’s data it is noted, and where it no longer does, that is noted too. Nothing here was re-tuned after the fact.

Average Spectral Index across the roster

The natural first idea: score every rookie-eligible player in the organization and take the mean.

It ranks Philadelphia 11th and Cincinnati 26th. Look at who those two organizations actually have and that ordering is hard to defend. Cincinnati carries 18 prospects above Spectral Index 64 to Philadelphia’s 8.

The mechanism is arithmetic, not baseball. Cincinnati rosters 211 rookie-eligible players; Philadelphia rosters 184. Those extra 27 bodies are complex-league filler, and averaging lets them drag the organization down. An organization is penalised for signing more players, which is the opposite of what a depth metric should do.

The same pair still shows it on today’s data, with the gap wider: averaging ranks Philadelphia 18th and Cincinnati 26th, while Cincinnati carries 10 prospects above Spectral Index 64 to Philadelphia’s 4 — two and a half times as many — on a roster of 222 against Philadelphia’s 190.

This is the canonical example of why “average quality” is the wrong shape for an organizational aggregate, and it is worth keeping in mind for any future metric: the first question to ask is always what the metric does when an organization signs twenty more teenagers.

Summed reach probability

The most conceptually on-mission option. This site’s stated purpose is estimating the probability of outcomes rather than ordering players by quality, so summing each prospect’s modelled probability of reaching the majors should be the metric that best matches what the site is for. Expected number of big leaguers in the system. Clean.

It fails empirically. Correlation between an organization’s summed reach probability and its raw headcount, across all 30 organizations:

Metric r with roster size
Average Spectral Index 0.05
Top-10 average 0.04
Sum of above-average players 0.31
Summed reach probability 0.60

At r = 0.60, over half of what the metric measures is how many players an organization has under contract. A fringe complex-league arm with a 4% chance of reaching the majors still adds 0.04, and enough of them add up to a real prospect. Depth should count, but not like that.

Top-5 mean — the version that shipped first and was replaced the same day

This one is the interesting failure, because it fit the published lists better than the metric that replaced it. Spearman 0.734 against MLB.com, versus 0.681 for the final version.

It was still wrong, and Milwaukee is why. The Brewers carry 29 prospects above Spectral Index 64. The next-deepest organization carries 26. That gap over the field is the single most distinctive fact about any system in the game right now — and a metric that reads only the top five names cannot see it. Top-5 mean ranked Milwaukee 2nd. Every curve that looks past the fifth name ranks them 1st.

On today’s rosters this particular comparison no longer separates the two metrics: Milwaukee now leads on top-5 mean as well, so the disagreement that decided the question in August has closed. That does not retroactively make top-5 mean the right choice — it was discarded for reading only five names, not for one ranking — but it is worth saying plainly rather than leaving a live-sounding claim that no longer reproduces.

Discarding a metric that scores 0.734 in favour of one that scores 0.681 looks perverse written down. But the 0.734 was measuring agreement with one list, and what it had actually learned was to ignore depth — which is the thing this site can measure well and which the published lists were being used to calibrate, not to define. Better fit, worse metric.

What shipped: System Value

Take each organization’s top 30 rookie-eligible prospects by Spectral Index. Price each one at

value = exp(0.200 × (SpectralIndex − 50))

and add them up.

The exponent is stored internally as a fixed constant applied to each prospect’s raw composite score, not to the displayed Spectral Index — the code’s constant is 2.0, and 0.200 is what that works out to in display units today. Those two numbers are only related through the Spectral Index’s own population spread, which changes whenever the index is rebuilt, so the display-space figure is not comparable across versions of this paper: it read 0.1574 in August against an internal constant of 3.0. The internal constant went down while the display exponent went up. That is precisely why the constant is stored in raw units and re-derived rather than carried forward as a display number.

Neither half of that is invented. Both are borrowed from the lists the metric is calibrated against:

  • The 30-player cap is Baseball America’s format. Every organization gets a published Top 30, so 30 is the number the industry has already settled on for “how deep is a farm system.” It was chosen over 20 at the user’s explicit request: deeper than consensus, “even if it’s slightly different from consensus.”
  • The exponential curve is FanGraphs’ surplus-value model. There, every prospect carries a dollar figure and the organizational total is simply their sum — with a 70 FV priced at roughly 100× a 35+ FV. A steeply rising curve is how you let elite talent dominate without discarding the players below it.

An organization with one superstar and nothing else beats an organization with thirty warm bodies. An organization with thirty good prospects beats one with five great ones and a barren system underneath. That is the behaviour the two ideas produce together.

Choosing the curve steepness

The exponent is the one free parameter, and left free it would be a knob to tune until the answer looked right. It isn’t one, because it has a direct interpretation: k decides how many of the 30 players actually matter.

The diagnostic is effN(90%) — the number of prospects carrying 90% of an organization’s total value.

The first version of this paper measured that at the median organization, and picked k on that alone. That criterion shipped a badly wrong constant, and it is worth being specific about why, because the failure is a general one. The median organization has no prospect anywhere near the top of the scale. A steep exponential curve only misbehaves when there is a prospect at the top of the scale. So the median organization is structurally incapable of showing the damage — the check was being run in the one place the problem cannot appear.

Measured across all thirty organizations instead:

k (display) median org thinnest org orgs where 3 players or fewer carry 90%
0.10 26 of 30 22 0
0.12 25 of 30 18 0
0.15 24 of 30 14 0
0.20 21 of 30 4 0
0.25 17 of 30 1 1
0.30 13 of 30 1 4
0.35 9 of 30 1 9

Read the middle column alone and 0.30 looks like a defensible answer — 13 of 30, a top-half metric. Read the two columns beside it and 0.30 is not defensible at all: four organizations have 90% of their System Value sitting in three prospects, and at least one has it in a single name. A column labelled “top 30” that is decided by one player is not a top-30 metric, whatever the median organization says about it.

That is not a hypothetical. It is what the site shipped, and on that curve the Teams page ranked Oakland 1st on seven prospects above the impact threshold, ahead of Milwaukee’s fourteen — an organization with half the impact talent leading the league because one of its names sat at the ceiling of the scale.

So the criterion is now stated over every organization rather than the middle one:

No organization may reach 90% of its System Value in three prospects or fewer — and among the values of k that satisfy that, take the steepest one.

Steepest, not gentlest: rewarding the top of a system heavily is the whole point of the metric, and the only thing that constrains it is that it must not collapse into a one-man band. That gives k = 0.20 in display units — the internal constant 2.0.

At that setting all thirty prospects contribute something, the top twenty-one drive the number, and the thinnest system in the league still needs four names to reach 90%. A Spectral Index of 85 is worth about 1,100× a Spectral Index of 50. For Milwaukee specifically, the #1 prospect in the system outweighs the #30 by about 119×. Nobody should mistake this for a flat sum.

The cost, stated rather than buried: at 0.20 the median organization needs 21 of its 30 prospects to reach 90%, which is outside the [15, 20] band the original version of this paper used as its target. That band was inherited from matching a constant that predated the current Spectral Index, and it is the criterion that certified k = 0.30 while four organizations were being decided by three players. It is reported here as a diagnostic. It is no longer what chooses k.

Putting it on the 20-80 scale

A raw exponential sum has no unit. FanGraphs’ totals are in dollars and mean something on their own; exp(0.20 × (SI − 50)) summed over 30 players means nothing except relative to the other 29 organizations — and it would shift wholesale under any future recalibration of the Spectral Index scale itself.

So the displayed column is the raw sum expressed as a 20-80 scouting grade, the same scale used everywhere else on the site: 50 is league average, and every 10 points is one standard deviation. It runs through the site’s existing zToScoutingGrade(), which is deliberately unclamped — a system genuinely three standard deviations out should be allowed to say so.

The standard deviation is the population SD (dividing by n, not n − 1). These 30 organizations are not a sample drawn from some larger population of farm systems; they are the entire population being graded.

Milwaukee grades 82, and it is worth showing the whole derivation, because a number that lands past the top of a familiar scale is normally a number that is broken:

Step Value
Mean raw value across 30 orgs 676.85
Population SD 457.89
Milwaukee raw value 2,146.88
z = (2,146.88 − 676.85) / 457.89 +3.210
Grade = 50 + 10 × 3.210 82.105
Rounded for display 82

Nothing is clipped — 82 is above the nominal 80 ceiling of the scouting scale precisely because nothing clips it. The unrounded grades across all 30 organizations span 40.951 to 82.105. A grade over 80 is the unclamped scale doing its job: it is saying that Milwaukee’s system is further from league average than the 20-80 scale has room to express, which is a claim about Milwaukee, not an artifact of the grading.

The resulting distribution is severely asymmetric: the top system sits at z = +3.21 while the bottom sits at z = −0.90. That is the shape of the underlying data, and on an exponential curve it is what you should expect — two organizations (Milwaukee and Oakland, at 82 and 79) hold prospects near the ceiling of the Spectral Index, everyone else is bunched between 41 and 70, and an exponential curve magnifies exactly that kind of gap. The honest reading of the column is that it separates the top two systems from the field emphatically and separates the middle twenty from each other only weakly.

How well it agrees

These three figures are from August 2026 and have not been re-measured since the curve was re-derived. The two published lists they are measured against are not stored in this repository, so the comparison cannot be recomputed on demand the way every other number in this paper can. They are left standing as the record of what was measured when the metric was chosen — not as a claim about the current curve. Re-running them requires re-obtaining both lists, and until that happens the agreement of the shipped metric with outside opinion is an open number, not a verified one. Saying otherwise is the exact mistake that produced the constant this paper had to replace.

Comparison Spearman ρ
System Value vs. MLB.com 0.681
System Value vs. FanGraphs 0.507
MLB.com vs. FanGraphs 0.671

Against MLB.com, System Value agrees about as closely as FanGraphs does. That is the right target, and it is where any future re-tuning should be judged — against 0.671, not against 1.0.

Where it disagrees — and why that is the most useful result here

The disagreements are not scattered. They run in one direction, and that direction is diagnostic.

The organizations this metric rates well below the published consensus are the ones built on young, high-variance, high-ceiling talent — the profiles scouts fall in love with and statistical models cannot see. The organizations it rates above consensus are the ones stacked with performers who are already producing.

That is exactly the blind spot documented elsewhere in this project: the Spectral Index is fundamentally a bust-rate probability engine. It answers “how often did players with this statistical profile reach the majors” extremely well. It has no scouting-ceiling input at all — no FV grades, no scouting reports, no projection of what a nineteen-year-old’s body might become. So it systematically under-rates the nineteen-year-old with loud tools and a modest stat line, and over-rates the polished twenty-four-year-old in Double-A.

This had been suspected for a while, and had been demonstrated on hand-picked pairs of players. Ranking all 30 organizations against an independently published 30-item list is the cleanest confirmation yet produced, because nothing about it is hand-picked: the systems the model misses are the ones a scout would tell you are built on ceiling.

Meanwhile the anchors land. Milwaukee, the deepest system in the game, comes out on top — three standard deviations clear of the field — and several mid-table organizations land exactly where MLB.com has them. The metric is not confused about the easy cases. It is confused about ceiling, specifically, and only about ceiling.

A note on method

Every number in this paper was produced by driving the actual rendered Teams page in a real browser and calling the page’s own teamRosterForOrg() and compositeDisplayValue() functions. None of it came from a reimplementation of the scoring logic.

That is a deliberate rule on this project, and it exists because it was learned the hard way: an independent implementation of the Spectral Index formula, written to make research easier, silently drifted out of date and produced a whole thread of confidently wrong findings before anyone noticed. If a number is quoted here, the site itself computed it.

← Back to White Papers

Prospect Wavelength is not affiliated with MLB, nor with any of their digital properties. Email any questions, comments, or concerns to stealofhome@prospectwavelength.com