We look at the winners and assume we're seeing the whole story. Usually, we're only seeing what's left.
Survivorship bias hides the failures that never made it into the record — the businesses that closed, the buildings that collapsed, the attempts that never got written up — leaving only the survivors to define what "normal" or "possible" looks like.
Every Thursday, Open Data Insights publishes one bias or fallacy — clearly explained, grounded in research, and connected to the kind of data stories we tell on this site. Because understanding your own thinking is the first step toward understanding the world more clearly. This week: Survivorship Bias.
Survivorship bias is our tendency to draw conclusions from the subset of things that made it through some selection process — the "survivors" — while overlooking the far larger set that didn't. Because the failures are invisible, absent, or simply never counted, we mistake a skewed sample for the whole picture. Success looks more common, more replicable, and more explicable than it actually is, precisely because the cases that would complicate the story are no longer there to be seen.
During the Second World War, the U.S. military asked statistician Abraham Wald to help decide where to add armor to bomber aircraft. Engineers had examined planes returning from combat and mapped the bullet holes: the wings and fuselage were riddled, the engines almost untouched. The obvious recommendation was to reinforce the areas with the most damage.
Wald reversed the logic. The data only included planes that had survived and made it home. The engines showed few hits not because they were rarely struck, but because planes hit in the engine rarely made it back to be counted at all. The missing bullet holes — on the aircraft that never returned — were the most important data point of all, and they were invisible by construction.
Wald recommended reinforcing the areas that were undamaged on the returning planes, not the ones that were damaged. His insight — that a sample restricted to survivors can point in exactly the wrong direction — became a founding case study in statistics and is still taught today wherever selection effects distort what data can and cannot show.
Business advice books study companies that succeeded and extract common traits — bold risk-taking, unconventional founders, breaking the rules. The much larger set of companies that took the same risks and failed is never included in the sample, so the traits look far more predictive of success than they actually are.
Old buildings are often held up as proof that construction used to be better: "they don't build them like they used to." In reality, the poorly built structures from the same era collapsed or were demolished long ago. What remains standing is not representative of what was originally built — it's the surviving fraction.
Advice from highly successful people in any field — investing, sport, the arts — is often treated as a reliable blueprint. But the people who followed the same approach and failed rarely get a platform to share what happened, so only the winning strategies remain visible.
In each case, the visible cases were not a random sample of what was actually out there — they were the ones that happened to survive a process that quietly discarded most of the evidence.
Survivorship bias is a particular hazard whenever a dataset only includes things that currently exist, are still active, or successfully passed some filter — and any analysis built on it inherits that blind spot without announcing it.
A dataset of currently operating businesses says nothing about the businesses that closed. A collection of historical records that happened to be preserved says nothing about the records that were lost or discarded — and the reasons some material survives while other material doesn't are rarely random. Financial performance figures calculated only from funds that are still open, without including funds that were closed or merged away after poor performance, will systematically look better than the reality of the full population.
The pattern is always the same: the interesting question isn't just what the data shows, but how the data came to exist in the first place, and what didn't make it in.
Whenever a dataset or a set of examples looks especially clean, ask: what could have failed to make it into this data — and did it?
Actively seek out the missing cases: the businesses that closed, the funds that were liquidated, the records that were never preserved, the attempts that never got written up. Be especially cautious with any success story, best-practice list, or "what winners have in common" analysis that doesn't also describe the base of people or organizations who did the same thing and failed.
The most important data point is often the one that isn't there. Learning to notice its absence — rather than only reading what remains — is the whole of the skill.
🤖 This text was generated with the assistance of AI. All quantitative statements are derived directly from the dataset listed under Data Source.