Every Thursday, Open Data Insights publishes one bias or fallacy — clearly explained, grounded in research, and connected to the kind of data stories we tell on this site. Because understanding your own thinking is the first step toward understanding the world more clearly. This week: Base Rate Neglect.
Every Thursday, Open Data Insights publishes one bias or fallacy — clearly explained, grounded in research, and connected to the kind of data stories we tell on this site. Because understanding your own thinking is the first step toward understanding the world more clearly. This week: Base Rate Neglect.
Base rate neglect is our tendency to ignore the underlying prevalence of something — its base rate — when judging a specific case, and to rely instead on vivid, specific-sounding information even when that information is weak or unreliable. We treat a compelling detail as if it outweighs the statistical backdrop it's set against, when in fact the backdrop should often dominate the judgment. The result is that rare events get treated as common, and common events get treated as rare, simply because a specific story felt convincing.
In 1972, Daniel Kahneman and Amos Tversky posed what became known as the "cab problem." Participants were told that a cab was involved in a hit-and-run accident at night, in a city where 85% of cabs are Green and 15% are Blue. A witness identified the cab as Blue. Tested under the night-time conditions, the witness correctly identified each colour 80% of the time and was wrong 20% of the time.
Most participants estimated the probability that the cab was actually Blue at around 80% — matching the witness's reliability almost exactly. The correct answer, using Bayes' theorem, is closer to 41%. Because Green cabs vastly outnumber Blue ones, even a fairly reliable witness will misidentify a Green cab as Blue more often than they correctly identify an actual Blue cab.
Participants had built their judgment almost entirely on the witness's testimony and had all but discarded the base rate — the simple fact that there are far more Green cabs on the road to begin with. Kahneman and Tversky used this experiment to demonstrate how readily people abandon prior probabilities in favour of specific, case-level evidence, even when that evidence is far weaker than it feels.
A medical test for a rare disease is 99% accurate, and a patient tests positive. Many assume this means a 99% chance of having the disease — but if the disease affects only 1 in 10,000 people, the actual probability of having it, even after a positive result, remains low. Most positive results are false positives, simply because so few people have the disease to begin with.
An investor hears a compelling story about a small company's breakthrough product and predicts huge growth — without asking how many similarly promising startups fail. The specific narrative crowds out the base rate of business failure.
A hiring manager is struck by a candidate's confident answer to one interview question and rates them highly — overlooking how weakly a single interview answer typically predicts job performance compared to structured, longer-term indicators.
In each case, a specific, vivid piece of evidence overwhelmed a piece of background information that should have carried more weight.
Base rate neglect is one of the most consequential biases in how the public interprets statistics, because so much of data journalism revolves around exactly this tension — a specific number or event set against a broader statistical context.
A screening programme that flags "500 people at risk" sounds alarming until the base rate of the underlying condition is known. A headline describing a security algorithm as "95% accurate" says little on its own, since its usefulness depends entirely on how rare the thing it's detecting actually is. Even well-produced data journalism can inadvertently trigger base rate neglect simply by leading with a specific case or a striking accuracy figure without stating the underlying prevalence clearly.
The fix isn't to strip out the compelling specifics — a concrete example is often what makes a statistic legible to readers. It's to always state the base rate alongside it, and to make clear which one should be doing the heavier lifting in the reader's judgment.
Before accepting a probability estimate — yours or someone else's — pause and ask: what is the base rate here, and did I actually use it?
When you encounter a striking statistic about accuracy, reliability, or risk, look for the underlying prevalence of the thing being measured before drawing conclusions. Be especially cautious with rare-event testing: a highly accurate test for a rare condition still produces mostly false positives in absolute terms. And when a vivid anecdote or specific case feels persuasive, ask what the base rate would suggest if you stripped the story away entirely.
The base rate is rarely dramatic. That's exactly why it's so easy to ignore — and exactly why it usually deserves the final word.
🤖 This text was generated with the assistance of AI. All quantitative statements are derived directly from the dataset listed under Data Source.