What Calibration Really Means (and Why We Grade Every Projection)
A model that says a batter has a 60% chance to record a hit is not making a prediction you can grade on a single night. It is making a promise about the long run: that when it says 60%, the real answer should land near 60% of the time. That promise is called calibration, and it is the single most honest way to judge whether a projection can be trusted.
Accuracy and Calibration Are Not the Same Thing
Most people judge a projection by whether it "hit." That feels intuitive, but it quietly confuses two very different ideas.
Accuracy asks: did the model pick the right side more often than not? Calibration asks something harder: are the model's stated probabilities honest? A model can be accurate and badly calibrated at the same time, and that combination is where research goes wrong.
Imagine a model that labels every favorite "80% likely" and is right 80% of the time overall. Sounds great. But if the games it called 80% actually come in around 65%, the model is overconfident. It will look sharp on a highlight reel and quietly mislead you on the close calls, because every stated number is inflated.
The short version: Accuracy tells you how often the model is right. Calibration tells you whether you can believe the number it prints. You want both, but calibration is the one almost nobody checks.
What a Well-Calibrated Model Looks Like
The clean way to test calibration is to bucket every graded projection by the probability it claimed, then compare that claim to what actually happened. If the model is honest, the two columns should track each other down the whole table.
| Model said | Projections in bucket | Actually happened |
|---|---|---|
| 50–55% | 1,240 | 53% |
| 55–60% | 980 | 57% |
| 60–65% | 760 | 62% |
| 65–70% | 540 | 68% |
| 70%+ | 410 | 73% |
That is the shape you are looking for. When the model says 60–65%, roughly 62% of those projections come in. The number on the screen is not marketing; it is a level you can reason with. A table that fails looks obvious in hindsight: the model says 70%, reality says 58%, and every high-confidence read is quietly a coin flip wearing a suit.
Why We Grade Every Projection Against the Box Score
You cannot know any of this without a graded record. That is the whole reason PropPrizm logs a projection before first pitch and then checks it against the box score after the game is final. Every hit, strikeout, total-base and home-run projection becomes a data point in the calibration table above.
Grading after the fact is uncomfortable on purpose. It is easy to remember the projections that looked smart and forget the ones that did not. A logged, timestamped record removes the memory bias and forces the model to answer for its own numbers.
Before the game, the model prints: Batter A, 61% chance of 1+ hit.
That night he goes 0-for-4. On its own, that tells you nothing. A 61% call is supposed to miss about 4 times in 10.
Only after hundreds of similar 60% calls can you ask the real question: did the whole group land near 60%? That is the difference between reacting to one box score and measuring a model.
One Number That Rewards Honesty: The Brier Score
Calibration tables are readable, but there is a single number that captures both accuracy and honesty at once. It is called the Brier score, and it is just the average squared distance between what the model said and what happened, where the outcome is scored as 1 (it occurred) or 0 (it did not).
The elegant part is what it punishes. If the model says 90% and the event happens, the miss is tiny (0.1² = 0.01). If it says 90% and the event does not happen, the penalty is brutal (0.9² = 0.81). Confident-and-wrong costs far more than humble-and-wrong. A model chasing a good Brier score has no incentive to inflate its confidence, which is exactly the behavior you want from a research tool.
Where Miscalibration Usually Hides
When a model drifts out of calibration, it rarely does so evenly. A few patterns show up again and again, and knowing them makes you a sharper reader of any projection:
- Overconfidence at the tails. The extremes are where thin samples live. A model that has seen a batter in only 20 plate appearances against left-handers will happily print a very high or very low number it has not earned.
- Direction bias by stat type. Some counting stats are structurally harder to model than others. A projection engine can be well-behaved on strikeouts and still run a touch high or low on a noisier stat like hits, simply because the underlying rate is bumpier.
- Stale inputs. A number that was calibrated on last season's conditions slowly decays as parks, arsenals and roles change. Calibration is not a one-time certificate; it has to be re-checked on fresh graded data.
How we handle it: No model is perfectly calibrated, and we don't pretend otherwise. The graded record is exactly the tool that surfaces these drifts, so the fix is measured against fresh outcomes rather than guessed at. A projection you can audit is worth more than one that only ever shows its wins.
How to Read a Number on PropPrizm With This in Mind
Once you think in terms of calibration, the probabilities on a matchup card change meaning in a useful way:
- Treat the number as a level, not a verdict. A 58% projection and a 72% projection are genuinely different, and the gap between them is the information. A single-night result does not confirm or refute either one.
- Respect the sample behind the number. A confident read built on a full season of data is stronger than the same number built on two weeks. The card's context tiles are there to tell you which one you are looking at.
- Judge the model over a season, not a night. The right way to hold any projection accountable is the same way we do internally: bucket the calls, wait for the sample, and check whether 60% really means 60%.
That is the quiet standard behind every projection on the site. The goal is never to be loud about the calls that land. It is to print a number you can believe, grade it in the open, and let the long run do the talking.
Want to see a graded record for yourself? Open any matchup on the dashboard and follow a projection through to the box score. The projections that hold up over a full season are the ones worth building research around. PropPrizm is a statistical tool for informational and entertainment purposes and does not guarantee outcomes.