Probability is a model, not a verdict

Why a probability is always a statement about a chosen model and reference class, not a fact about a unique event — worked through the birthday problem.

Established Supported by convergent, high-quality evidence.

Plain-language answer

A probability is not a property of an event itself. It is a statement about how likely that event is under a specified model — a set of assumptions about what could have happened instead, and how likely each alternative was. Change the model, and the number changes, even though the event stays the same.

This matters most for coincidences, because the model people use informally — “what were the odds of that exact thing happening to me?” — is almost never the model that was actually running. There were usually many more opportunities for something surprising to happen than there were for that one specific thing.

Why it matters

Headlines and anecdotes about “astronomical odds” usually report the probability of one very specific outcome, calculated after the fact, under a model nobody would have proposed in advance. That number can be real and the conclusion drawn from it can still be wrong, because the right comparison was never “how unlikely is this exact pattern?” but “how likely was it that some pattern this noticeable would turn up, given how many chances there were to notice one?”

The birthday problem as a worked example

The classic illustration is the birthday problem. Ask “what is the probability that someone else in this room has my birthday?” and the chance is small even in a large room — there is exactly one date that matters, and a few dozen other people each only have about a 1-in-365 chance of matching it.

Ask instead “what is the probability that any two people in this room share a birthday?” — no one caring whose — and the model changes completely, because now every pair counts. In a room of 23 people there are 253 pairs, and the exact calculation (counting every way the room could have no shared birthday, and subtracting from 1) puts the probability of at least one match at about 50.7% — better than even odds, in a room where most people’s first intuition is that a match would be a striking coincidence.(Diaconis & Mosteller, 1989) By 70 people it is about 99.9%: a shared birthday has stopped being surprising under this model, even though any one specific pair matching is still unlikely.

Both of those figures come from the same underlying arithmetic — multiplying the probability that each new person’s birthday avoids every one already seen — reproduced in a tested, deterministic calculation rather than quoted from memory. A Birthday Room interactive that lets you vary the room size and see the full model’s assumptions is planned for Try It.

Common misconception

“The odds of that happening were a million to one” is usually a description of one narrow, retrospectively chosen model — not evidence that something other than chance must be involved. The question that is actually informative is: a million to one under which model, counting which opportunities, and chosen before or after the event was known?

Limits and open questions

Specifying “the” model for a real, messy event is genuinely hard. Real birthdays are not perfectly uniform across the year; real social, medical, and cultural events are not independent the way coin flips are. Picking a reference class always involves judgement, and different reasonable choices can give different numbers. That is a reason to show assumptions explicitly, not a reason to avoid quantifying anything.

  • Interactive: Try It will host the Birthday Room simulation.
  • More on this idea: How Chance Works covers multiple opportunities and selection effects.

Key takeaways

  • A probability describes an event under a stated model; it is not an intrinsic property of the event.
  • “How many opportunities were there for something surprising to happen?” is usually the most important hidden question in a coincidence story.
  • Large, specific-sounding odds are not meaningful on their own — ask which model produced them and whether that model was chosen before or after the fact.

Sources

  1. Diaconis, Mosteller (1989). Methods for Studying Coincidences. Journal of the American Statistical Association, 84(408), 853-861. https://doi.org/10.1080/01621459.1989.10478847 ↩