General information, not financial, investment, legal, tax or betting advice · Prediction markets carry risk of loss · 18+ or the legal age in your region
Prediction MarketIndex
Prediction Market Index/Learn/Calibration and forecasting skill
Education pillarThe reference spine

Calibration and forecasting skill, are your probabilities honest?

Anyone can sound confident. A calibrated forecaster is the rarer thing, someone whose stated probabilities actually match how often the world proves them right.

By Fredrik FilipssonFounder and editor · Two decades in advisory, hospitality and mediaEditorial review by Morten Andersen · Last reviewed 28 June 2026
Last reviewed
28 June 2026
Reading time
About 13 minutes
Level
Intermediate
Information, not advice. This page is general information, not financial, investment, legal, tax, or betting advice. Prediction markets carry a real risk of loss. You must be 18+ or the legal age in your region.
Quick answer

Calibration measures whether your probabilities mean what they say. If you are well calibrated, the things you call seventy percent likely happen about seventy percent of the time across many forecasts. It is not the same as being right, and it is not the same as being confident. Forecasting skill is calibration plus the willingness to start from base rates, update on real evidence, and keep an honest record. Proper scoring rules such as the Brier score and log loss put a number on that skill over time, which is the only way to tell genuine ability from a lucky streak.

How it works

Four ideas behind real forecasting skill.

1
Calibration is about the long run, not one call

A single forecast can never be judged on its own, because a seventy percent event that does not happen is not a mistake, it is the thirty percent showing up. Calibration is a property of many forecasts together. Gather every time you said seventy percent and check whether those things happened roughly seventy percent of the time. If they did, you are calibrated at that level. This is why a track record matters and why one impressive call tells you almost nothing about real skill.

2
Scoring rules turn forecasts into a number

A proper scoring rule rewards honesty, meaning your best expected score comes from reporting your true probability. The Brier score is the squared difference between your probability and the outcome, scored as one or zero, averaged over many forecasts, where lower is better. Log loss penalises confident wrong answers very heavily. Both let you compare forecasters and track your own progress, and both reward saying eighty percent only when you really believe eighty percent, not when you want to look bold.

3
Base rates anchor an honest forecast

A base rate is how often something happens in general, before you add the specifics of this case. Skilled forecasters start from the base rate and adjust, rather than starting from a vivid story and forgetting how rare the event usually is. Base rate neglect, described by Daniel Kahneman and Amos Tversky, is one of the most common forecasting errors, and it shows up as overconfident probabilities on dramatic outcomes. Anchor to the base rate first, then update on real evidence.

4
Update on evidence, not on emotion

Good forecasting is a process of revising a probability as genuine new information arrives, and only then. The discipline is to ask how much a piece of news should actually move your estimate, which is usually less than it feels. Markets aggregate many people doing this, which is part of why a market price can be hard to beat. Treating each update as small and evidence driven, rather than swinging on the latest headline, is what keeps a forecast calibrated over time.

Visual one

The reliability diagram, where calibration becomes visible.

Group your forecasts by the probability you gave, then plot the probability you stated against how often those things actually happened. A perfectly calibrated forecaster sits on the diagonal. Sitting consistently below it is the signature of overconfidence, the most common bias of all.

Probability you stated Observed frequency Perfect calibration Overconfident forecaster 0% 100%

Illustrative reliability diagram showing how calibration is measured. The curve sitting below the diagonal is the classic pattern of overconfidence, where stated probabilities are higher than the frequencies that follow. Not a claim about any forecaster and not a prediction. Diagram by Prediction Market Index, 28 June 2026.

How calibration is checked

The reliability curve, in plain words.

Take every forecast where you said about 70 percent. If the events happened close to 70 percent of the time, your seventies are well calibrated. Do the same for your twenties, your fifties and your nineties. Plot stated probability against actual frequency and a perfectly calibrated forecaster sits on the diagonal. Sitting above or below it shows a consistent bias, such as overconfidence, that you can measure and then correct.

Well calibrated
Said 70%, happened about 70% of the time

An illustration of how calibration is measured, not a claim about any forecaster and not a prediction.

Why it matters for you

Calibration is the antidote to false confidence.

Prediction markets express everything as a probability, so the skill that matters most is the ability to judge probabilities honestly. Calibration is the name for that skill measured properly. It reframes the whole activity away from being right or wrong on a single market and toward being accurate on average across many, which is the only standard that can actually be measured and improved. The shift sounds small, but it changes how you treat your own opinions, because it forces you to ask not whether a view feels strong but whether the number attached to it has earned its place.

The reason this matters in practice is that confidence and accuracy are only loosely related. People who feel certain are often no more accurate than people who feel unsure, and sometimes less, because certainty discourages the updating that keeps a forecast honest. A calibrated forecaster has learned to attach numbers to uncertainty and to be suspicious of their own strong opinions, which is uncomfortable but is exactly where the skill lives. The discomfort is the point. If a probability never feels too low for how sure you are, you are probably rounding your real uncertainty away.

There is good evidence that calibration is a real and learnable skill rather than a personality trait. The Good Judgment Project, the forecasting tournament funded by the United States research agency IARPA and led by Philip Tetlock and Barbara Mellers at the University of Pennsylvania, found a group it called superforecasters who were strikingly well calibrated. Their average Brier score was about 0.166, against roughly 0.259 for regular forecasters in the same tournament, where a lower score is better, per results from the Good Judgment Project summarised by AI Impacts as of 2026. Team Good Judgment beat the control group by more than fifty percent, and on some questions outperformed intelligence analysts who had access to classified information, per the project's published findings. The lesson is not that these people are gifted, it is that careful habits move the number.

Scoring rules make all of this concrete. By keeping a record of your forecasts and scoring them with a Brier score or log loss, you turn a vague sense of being good at this into a number you can track. Over enough forecasts the number tells you whether you are improving, whether you are systematically overconfident, and whether you have any genuine edge at all. Without a record, memory quietly edits out the misses and remembers the hits, and you learn nothing. The scoring rule is what keeps you honest with yourself, because it does not care how a forecast felt at the time.

Base rates deserve special attention because ignoring them is the most reliable way to be confidently wrong. The human mind is drawn to specific, vivid scenarios and tends to forget how rare those scenarios are in general. Starting from the base rate and adjusting only for real, relevant evidence corrects for this, and it is most of what distinguishes forecasters with a track record from people who simply sound persuasive. A useful test is to ask how often something like this has happened before, and to make yourself answer that before you reach for the details of the case in front of you.

It is important to be honest about the limits. Being well calibrated does not mean you can beat a market, and it does not mean you will make money. A market price already reflects the aggregated, often well calibrated views of many participants, which is exactly why it is hard to outperform. Fees and spreads then sit on top of that. Calibration is a discipline for thinking clearly and judging yourself honestly. It is not a strategy that guarantees profit, and we never present it as one. The point of the skill is better thinking, not a shortcut to returns.

Visual two

Skill shows up as a lower score.

The Brier score is on a scale from 0 to 1, where 0 is perfect and a forecaster who always says fifty percent scores 0.25 on a balanced set of questions. The gap below compares two reported averages from the Good Judgment Project. Remember that lower is better, so the shorter bar is the stronger record.

SuperforecastersBrier 0.166
Regular forecastersBrier 0.259
Always guessing 50 percentBrier 0.250

Bars scaled to each Brier score, lower is better. Superforecaster and regular forecaster figures per Good Judgment Project results summarised by AI Impacts, as of 2026. The 0.25 reference is the score from always saying fifty percent on a balanced question set. Illustrative comparison by Prediction Market Index, 28 June 2026.

Worked example

Scoring four forecasts with the Brier rule.

Here is the Brier score worked through on four imaginary forecasts. The outcome is scored 1 if the event happened and 0 if it did not. The penalty for each forecast is the square of the gap between your probability and that outcome. Average the penalties and you have the Brier score for the set. Notice that the confident wrong call, the third row, dominates the total, which is the rule doing its job.

ForecastYour probabilityOutcomeSquared error
Rain in the city tomorrow0.801 (it rained)0.04
The bill passes this session0.300 (it failed)0.09
The favourite wins the match0.900 (an upset)0.81
Snow in the city in June0.050 (no snow)0.0025
Average Brier score0.236

Method: squared error equals the probability minus the outcome, squared. The Brier score is the mean of those values, here (0.04 + 0.09 + 0.81 + 0.0025) divided by 4, which rounds to 0.236. Lower is better, 0 is perfect, and 0.25 is the score for always guessing fifty percent. Illustrative worked example, not real forecasts and not advice. Prediction Market Index, 28 June 2026.

Brier and log loss compared

Two proper scoring rules, one shared lesson.

The Brier score was introduced by the American meteorologist Glenn W. Brier in 1950 to verify weather forecasts that were stated as probabilities, and it remains the most common way to score them. It is a strictly proper scoring rule, which is a precise way of saying that your best expected score comes only from reporting your true probability. Any attempt to game it by exaggerating or shading your number raises your expected penalty. That single property is why proper scoring rules sit at the centre of serious forecasting, because they make honesty the winning strategy rather than a virtue you have to remember.

The Brier score can also be broken into parts. Researchers decompose it into reliability, which is calibration, resolution, which is your ability to tell likely outcomes from unlikely ones, and uncertainty, which is the inherent difficulty of the questions. That decomposition is useful because it separates two different talents. A forecaster can be perfectly calibrated and still add little, if every forecast hovers near the base rate. Resolution rewards the forecaster who confidently and correctly moves away from the base rate when the evidence justifies it. Good forecasting needs both, and the Brier score captures both in one number.

Log loss, also called logarithmic loss, is the other proper scoring rule you will meet, and it differs in temperament. Where the Brier score grows with the square of your error, log loss grows much faster as you approach certainty, so a confident wrong answer is punished severely. Saying ninety nine percent for something that does not happen costs you far more under log loss than under the Brier score. This makes log loss popular wherever overconfidence is the dangerous failure, and it carries the same core message, which is that you should state what you actually believe. Both rules agree on the lesson, they only disagree on how harshly to punish the bold mistake.

Calibration and a market price

The price is a calibrated rival you have to outscore.

A prediction market price can be read as the crowd's probability, which makes it a useful benchmark for your own calibration. If a market sits at sixty five cents on an outcome and you cannot explain why your number should differ, the honest move is usually to adopt something close to the market rather than to trust a private hunch. The market is, in effect, a large group of forecasters already doing the work, and on liquid questions it tends to be reasonably well calibrated. Treating the price as a calibrated rival, rather than as a target to argue with, is one of the most useful habits a careful reader can build. You can learn how to turn a price into a probability on our explainer on reading prices as implied probability.

This is also where calibration runs into cost. Even a forecaster who is genuinely better calibrated than a market faces fees and the spread between the buy and sell price, and those costs come out of any edge before it reaches you. A small calibration advantage can be entirely erased by the cost of acting on it, which is why being right on the probability and being ahead after costs are two different things. We track trading fees, spreads and withdrawal terms across the platforms in our cross platform fees dataset, because the size of those costs is what decides whether an edge is real or only theoretical. None of this is advice, and a market price can of course be wrong, but the discipline is to assume the price is hard to beat until your own scored record says otherwise.

Building the habit

Keep a record, then trust it.

The single most useful step toward better calibration is also the simplest, which is to write your forecasts down with a date, a probability, and a clear resolution condition. Memory is unreliable and self flattering, quietly remembering the hits and forgetting the misses, so an honest record is the only way to see your real accuracy. Once you have a few dozen forecasts scored, patterns appear that you could never feel from the inside, such as a tendency to say ninety percent when seventy would have been honest, or a habit of being timid on questions you actually understand well.

A resolution condition matters more than it looks. Vague questions cannot be scored, and a forecast you cannot score teaches you nothing. Decide in advance exactly what would count as the event happening, including the date by which it must happen, so that when the time comes there is no room to talk yourself into having been right. The discipline of writing a clean resolution condition is itself part of the skill, because it forces you to know what you are actually claiming before you attach a number to it.

From there, improvement is mostly about humility. The forecasters who get better are the ones who treat a surprising outcome as information about their model rather than as bad luck, and who pull their probabilities toward base rates when they cannot justify being far from them. That discipline is uncomfortable, because it means being less certain than you would like, but it is exactly what the scoring rules reward. Calibration is a slow skill, and the record is what turns slow practice into measurable progress. This page is general information, not advice, and a good record does not change the fact that markets are hard to beat after costs.

Where this matters

Take this into the platforms, markets, and rules.

A note on risk,

Being well calibrated does not make trading safe or profitable, and a careful forecast can still lose. A market already reflects many informed views, so treat your edge as uncertain and stake only what you can afford to lose. Never stake to chase a loss, and never on borrowed money. If it stops feeling like a free choice, step back. In the United States you can call or text the helpline on 1-800-GAMBLER or visit ncpgambling.org.

Common questions

Answered plainly.

What does it mean to be well calibrated?

It means your stated probabilities match real frequencies over many forecasts. If the things you call seventy percent likely happen about seventy percent of the time, your seventy percent forecasts are well calibrated. It is a property of many forecasts taken together, not of any single call.

Is calibration the same as being right?

No. A single seventy percent forecast that does not happen is not wrong, it is the thirty percent occurring. Calibration is measured across many forecasts, and it is about accuracy on average rather than the outcome of any one call.

What is a Brier score?

The Brier score, introduced by meteorologist Glenn W. Brier in 1950, is a proper scoring rule equal to the squared difference between your probability and the outcome, averaged over many forecasts, where lower is better. It rewards reporting your honest probability rather than exaggerating your confidence.

What is log loss?

Log loss, also called logarithmic loss, is another proper scoring rule that penalises confident wrong answers very heavily. Like the Brier score it rewards honesty, so your best expected score comes from stating what you actually believe rather than from sounding bold.

Why do base rates matter so much?

A base rate is how often something happens in general. Skilled forecasters start there and adjust for real evidence, rather than starting from a vivid story. Base rate neglect, described by Daniel Kahneman and Amos Tversky, is one of the most common causes of overconfident, poorly calibrated forecasts.

Does good calibration mean I can beat the market?

No. A market price already reflects the aggregated views of many participants, and fees and spreads sit on top. Calibration is a discipline for thinking honestly about probability, not a strategy that guarantees profit, and this page is general information rather than advice.

Reviewed by Morten Andersen, Editor, on 28 June 2026. Written by Fredrik Filipsson. We check this explainer against the published forecasting research it cites, including the Good Judgment Project results, and refresh it when the underlying sources change.
The Forecast

Learn one useful thing a week.

The rules change fast. Get the changes that affect you, plain and current, not tips.

Independent. Every claim dated and sourced. No platform pays for its place.

No tips, no picks, no spam. Information, not advice. Unsubscribe anytime.