Plain-language answer
A selection effect occurs when the process that produced your data has already filtered out some cases before you ever see them — so the sample you observe is not a random, representative slice of everything that happened. Conclusions drawn without accounting for that filtering can end up exactly backwards. Coincidence stories are especially vulnerable to selection effects because the very act of a coincidence being “noticed, remembered, and retold” is itself a filter that removes every non-coincidental case from view before anyone starts counting.
Why it matters
During the Second World War, statisticians studying returning aircraft recorded where each plane had taken damage, intending to recommend adding armor to the most-hit locations. Abraham Wald, working with the Statistical Research Group, argued the opposite: armor should go where the returning planes showed the fewest hits. His reasoning was that the data only included planes that survived to be examined — bullet holes in a returning plane’s wings showed that a wing could be hit and still make it home, while the near-absence of hits to the engine and cockpit was not good news, it was missing data, because planes hit there overwhelmingly did not return.(Mangel & Samaniego, 1984) The raw pattern in the visible data pointed to the wrong conclusion; correcting for what was systematically not in the sample reversed the recommendation entirely.
Worked example: the general pattern behind survivorship bias
Wald’s specific case is now the standard teaching example for survivorship bias, but the same structure recurs everywhere coincidence stories get collected: viral lists of “impossible” coincidences are built from stories people found compelling enough to share, which is not remotely a random sample of everything that almost seemed like a coincidence and got forgotten, or every near-miss that didn’t quite line up. A list of “amazing lottery-winning coincidences” never includes the vastly larger set of people who bought two tickets on a whim and won nothing twice — because that outcome generates no story to select for in the first place.
The lesson generalizes directly: before treating a pattern in a visible sample as evidence of something, ask what had to be true of a case for it to end up in the sample at all, and whether that filtering process could produce the same pattern even if nothing unusual were going on underneath.
Common misconception
“The data clearly shows the pattern, so the pattern must be real” skips the prior question of how the data came to be assembled. Wald’s case is a sharp illustration precisely because the visible data was completely accurate as far as it went — the returning planes really were mostly hit in the wings — and the error was entirely in treating an already-filtered sample as if it were the whole picture.
Limits and open questions
Correcting for a selection effect requires some model of what the unseen, filtered-out cases probably looked like — in Wald’s case, an assumption that hits were roughly uniformly distributed across the aircraft before any were shot down, letting the “missing” damage locations be inferred. For most coincidence anecdotes, there is no equivalent way to reconstruct how many similar-but-unremarkable cases went unrecorded; the best available response is usually to flag the likely direction of the bias rather than to correct for its exact size.
Related
- Multiple opportunities and selection effects covers the closely related statistical version of this problem in formal hypothesis testing.
- Post-hoc probability covers what happens when the selection occurs not in which cases are collected, but in which pattern is chosen to describe them after the fact.
Key takeaways
- A sample that has already been filtered by its own outcome can point to exactly the wrong conclusion, even when every individual data point in it is accurate.
- Abraham Wald’s WWII aircraft-armor analysis is the rigorously documented origin of “survivorship bias” — armor was recommended for the least-hit locations, because planes hit elsewhere didn’t return to be counted.
- Collections of striking coincidences are themselves a selected sample: they contain only the stories someone found compelling enough to notice and share, not a representative record of all near-matches and non-matches.