Magrathea Software ← Blog

Field Notes

Why elite QBs and TEs go later than the math says

22 July 2026 · Kyle Twogood, Magrathea Software

While building the draft engine for RotoAlpha, our fantasy sports AI co-manager, our value-over-replacement math kept ranking elite quarterbacks and tight ends well ahead of where real drafters take them. The tempting conclusion was that the market was leaving value on the table. Before we bet our rankings on that, we gave the model a crystal ball: six seasons of actual results. It proved the crowd right, and showed us exactly where the math had been lying to us.


The disagreement

A quick recap of how the board is built, because it matters here. Instead of ranking players by raw projected points, we rank them by value over replacement: a player is worth the points he gives you above the best guy at his position you could have gotten for free. We derive that "replacement level" from a league-wide starter simulation rather than a hand-tuned constant. We told the whole story in the kicker post.

That baseline uses the best player at a position who cracks nobody's starting lineup. In a 12-team league with one quarterback slot each, that is roughly the 13th-best QB, and the same for tight ends. Measured against that, elite QBs and TEs show enormous surplus, so the model ranked them early. Too early, it turned out. Against the 2026 half-PPR consensus board, our rankings slotted tight ends and quarterbacks meaningfully ahead of the market (a log-rank bias of roughly −0.29 for TE and −0.16 for QB, where negative means we ranked them earlier than the crowd did). The vivid case: pure value-over-replacement put Josh Allen at QB overall #6, while his actual average draft position was #29.

So who is wrong? Either the market is irrationally cautious with the premium onesie positions, or our math is missing something real. Betting a draft on "everyone else is wrong" is exactly the kind of confidence that should make you nervous.

A clue: the model and the market agreed on the order

One detail pointed away from blaming our projections. Within the quarterback position, our ordering matched the market almost perfectly: the correlation between our QB ranking and the consensus QB ranking was about 0.97. We agreed on which quarterbacks were good and in what order. We disagreed only on where the whole block of them belonged relative to running backs and receivers.

That is the signature of a baseline problem, not a projection problem. If our point projections were biased, the within-position order would wobble too. It didn't. The entire error was in the one number that sets how much a position is worth relative to the others: replacement level.

The oracle experiment

Here is the test that settled it. Suppose our projections were perfect. If the early-QB/TE tilt is an artifact of bad forecasting, then perfect forecasts would make it disappear. So we built the projections a model can never actually have: the actual, realized fantasy points each player went on to score that season. Perfect hindsight. An oracle.

Then we ran our normal value-over-replacement board on those oracle points, under the same first-non-starter baseline, for every season from 2019 through 2024, and compared each year's board to that year's real historical average draft position.

The bias did not go away. Boards built from perfect, after-the-fact points still drafted tight ends about −0.37 and quarterbacks about −0.09 earlier than the market did, in every single year, 2019 through 2024. With flawless projections, the model still wanted the onesies earlier than the crowd took them.

That is conclusive in the way only a falsification test can be. The early tilt is not our projections being wrong. It is baked into season-total value-over-replacement itself. And since the oracle can't be beaten on accuracy, the only thing left to correct was the baseline.

Where the bias goes to zero

So we asked the oracle boards a follow-up: what replacement level makes the QB/TE bias vanish? Instead of "the first quarterback who doesn't start" (about QB13 in a 12-team league), we swept the baseline earlier and earlier up the position, looking for the spot where our board stopped disagreeing with the market.

It landed cleanly on the player at the 50th percentile of starters: in a 12-team, one-QB league, that is the 6th-best quarterback and the 6th-best tight end, not the 13th. At that baseline the six-year bias collapses to essentially nothing (+0.03 for QB, −0.03 for TE), and it stays there, stable across all six seasons rather than fit to any one of them.

That stability is the whole point. We were not tuning a fudge factor until our 2026 board looked like the 2026 market. The 0.50 replacement level is recoverable from realized points alone, in any year, without ever looking at a draft board. It is not our opinion about the market. It is the market's own replacement level for single-starter positions, and it was sitting in the box scores the whole time.

Why season-long math overrates the onesies

The honest mechanism is not the folk wisdom about streaming quarterbacks off waivers. It is subtler, and it is structural to how season-total value-over-replacement counts:

What we deliberately did not do

Two temptations we passed on, because both would have made the model worse in ways that are easy to miss:

The correction is also guarded. It applies only when a position starts no more players than there are teams, which is what makes it a single-starter, hard-to-absorb onesie. Superflex and 2-QB formats, or leagues that start two tight ends, keep the original first-non-starter baseline, because in those formats the surplus genuinely does get absorbed.

The through-line

This is the same lesson as the kicker post, arriving from the opposite direction. Value over replacement is only as honest as the "free" baseline it assumes. For kickers and defenses, the truthful baseline is the waiver wire, so a near-best streamer sets the line. For quarterbacks and tight ends, the truthful baseline is the 50th-percentile starter, because season-total math credits a full year of production to a single lineup slot that no flex demand can soak up.

Both times, our first instinct was that a confident number meant the market was wrong. Both times, the discipline that paid off was to distrust the baseline before distrusting the crowd, and to find a test that could prove us wrong. The oracle experiment could have vindicated our early QBs and TEs. It did the opposite, which is exactly why we trust the answer.

RotoAlpha is our AI co-manager for fantasy sports, launching this season. More field notes on the draft math and the backtests behind it are coming as we approach launch at rotoalpha.com.

Kyle Twogood

Kyle Twogood is the founder of Magrathea Software. He’s been building production software since 1997.