← Blog
Back to Blog

What Calibration Really Means (and Why We Grade Every Projection)

A model that says a batter has a 60% chance to record a hit is not making a prediction you can grade on a single night. It is making a promise about the long run: that when it says 60%, the real answer should land near 60% of the time. That promise is called calibration, and it is the single most honest way to judge whether a projection can be trusted.

Accuracy and Calibration Are Not the Same Thing

Most people judge a projection by whether it "hit." That feels intuitive, but it quietly confuses two very different ideas.

Accuracy asks: did the model pick the right side more often than not? Calibration asks something harder: are the model's stated probabilities honest? A model can be accurate and badly calibrated at the same time, and that combination is where research goes wrong.

Imagine a model that labels every favorite "80% likely" and is right 80% of the time overall. Sounds great. But if the games it called 80% actually come in around 65%, the model is overconfident. It will look sharp on a highlight reel and quietly mislead you on the close calls, because every stated number is inflated.

The short version: Accuracy tells you how often the model is right. Calibration tells you whether you can believe the number it prints. You want both, but calibration is the one almost nobody checks.

What a Well-Calibrated Model Looks Like

The clean way to test calibration is to bucket every graded projection by the probability it claimed, then compare that claim to what actually happened. If the model is honest, the two columns should track each other down the whole table.

Model saidProjections in bucketActually happened
50–55%1,24053%
55–60%98057%
60–65%76062%
65–70%54068%
70%+41073%

That is the shape you are looking for. When the model says 60–65%, roughly 62% of those projections come in. The number on the screen is not marketing; it is a level you can reason with. A table that fails looks obvious in hindsight: the model says 70%, reality says 58%, and every high-confidence read is quietly a coin flip wearing a suit.

Why We Grade Every Projection Against the Box Score

You cannot know any of this without a graded record. That is the whole reason PropPrizm logs a projection before first pitch and then checks it against the box score after the game is final. Every hit, strikeout, total-base and home-run projection becomes a data point in the calibration table above.

Grading after the fact is uncomfortable on purpose. It is easy to remember the projections that looked smart and forget the ones that did not. A logged, timestamped record removes the memory bias and forces the model to answer for its own numbers.

Example — one projection, one grade

Before the game, the model prints: Batter A, 61% chance of 1+ hit.

That night he goes 0-for-4. On its own, that tells you nothing. A 61% call is supposed to miss about 4 times in 10.

Only after hundreds of similar 60% calls can you ask the real question: did the whole group land near 60%? That is the difference between reacting to one box score and measuring a model.

One Number That Rewards Honesty: The Brier Score

Calibration tables are readable, but there is a single number that captures both accuracy and honesty at once. It is called the Brier score, and it is just the average squared distance between what the model said and what happened, where the outcome is scored as 1 (it occurred) or 0 (it did not).

Brier = average of (stated probability − outcome)² · lower is better

The elegant part is what it punishes. If the model says 90% and the event happens, the miss is tiny (0.1² = 0.01). If it says 90% and the event does not happen, the penalty is brutal (0.9² = 0.81). Confident-and-wrong costs far more than humble-and-wrong. A model chasing a good Brier score has no incentive to inflate its confidence, which is exactly the behavior you want from a research tool.

Where Miscalibration Usually Hides

When a model drifts out of calibration, it rarely does so evenly. A few patterns show up again and again, and knowing them makes you a sharper reader of any projection:

How we handle it: No model is perfectly calibrated, and we don't pretend otherwise. The graded record is exactly the tool that surfaces these drifts, so the fix is measured against fresh outcomes rather than guessed at. A projection you can audit is worth more than one that only ever shows its wins.

How to Read a Number on PropPrizm With This in Mind

Once you think in terms of calibration, the probabilities on a matchup card change meaning in a useful way:

That is the quiet standard behind every projection on the site. The goal is never to be loud about the calls that land. It is to print a number you can believe, grade it in the open, and let the long run do the talking.


Want to see a graded record for yourself? Open any matchup on the dashboard and follow a projection through to the box score. The projections that hold up over a full season are the ones worth building research around. PropPrizm is a statistical tool for informational and entertainment purposes and does not guarantee outcomes.