Coincidence Generator

Generate entirely fictional synthetic lives and search them for matches — see how widening what counts as a match multiplies apparent coincidences.

Established Supported by convergent, high-quality evidence.

Make a prediction first

Generate 40 entirely fictional “lives,” each with a random name, a synthetic birthday, a profession, and a favorite number. Search them for matching pairs two ways: strictly (same birthday only) or flexibly (same anything). How many matching pairs would you guess turn up each way?

Expected matching pairs (exact formula): 173.02

Matching pairs actually found in this generated population: 168

Click "Run simulation" to see the average across many independently generated populations.

What changed and why

The expected figure and the one generated population’s actual count both update live as you change the number of lives, the matching criterion, or the seed — both are cheap enough to compute instantly. The expected figure comes from an exact formula (number of pairs × the probability two independent lives match under the chosen criterion); the actual count comes from generating one concrete synthetic population and counting its real matches, so you can see how a single realization compares to the formula’s average. Run simulation repeats that generation many times (in a background worker) and reports the average match count across all of them, which should track the expected value closely.

Why flexible matching changes everything

With 40 lives, strict matching (same birthday only, out of 100 synthetic “days”) finds about 8 matching pairs — already more than intuition expects, for the same multiple opportunities reason the birthday problem does.(Dunn, 1961) But flexible matching (same birthday, or same name, or same profession, or same favorite number) finds around 20 times as many — because a pair now only needs to agree on one of four independent attributes, not one specific pre-chosen one. This is the generative version of a pattern covered throughout this site: expanding what counts as “a match” after the fact inflates the apparent rate of coincidence dramatically, without anything unusual needing to happen in the underlying data.

A note on the “examples” shown

Every name, profession, birthday, and favorite number this tool generates is drawn from a short, explicitly fictional list using a seeded random number generator — none of it describes, or was generated from, real people, and no example produced here should be read as a real anecdote. That distinction matters: the whole point of a synthetic generator is to demonstrate the mechanism behind inflated coincidence claims without manufacturing a new, misleading “true story” in the process.

Model assumptions

  • All four attributes are independent of each other and uniformly distributed across their fictional lists — real biographical attributes are far more correlated and unevenly distributed than this (see independence and dependence).
  • “Strict” and “flexible” are two illustrative matching rules among many possible ones — a real viral coincidence claim might use a differently flexible (or narrower) rule, which would change the exact numbers but not the underlying lesson.

Reproducibility

Every population is seeded and deterministic: the same lives-per-population, matching criterion, and seed always produce the same generated population and the same matches. Running the simulation updates this page’s URL so you can share the exact setup.

Sources and method

The exact expected-matches formula (pairs × per-pair match probability) follows directly from linearity of expectation, which holds regardless of dependence between different pairs’ match outcomes — verified in this tool’s unit tests against hand-computed small cases before being wired up to the interactive controls. The underlying “more comparisons, more apparent hits” principle is the same one behind the formal statistical correction for testing many hypotheses at once.(Dunn, 1961)

Sources

  1. Dunn (1961). Multiple Comparisons Among Means. Journal of the American Statistical Association, 56(293), 52-64. https://doi.org/10.1080/01621459.1961.10482090 ↩