The Threshold That Broke Science: Why 0.05 Is the Most Consequential Arbitrary Number in Research
In the architecture of modern science, few structures carry more weight than the p-value. Invisible to the general public yet omnipresent in every peer-reviewed journal, it functions as a kind of gatekeeper — a numerical verdict that separates findings deemed worthy of publication from those consigned to file drawers. Its most familiar incarnation is the threshold of 0.05: the conventional cutoff below which a result is declared "statistically significant." Above it, the finding is presumed unreliable. Below it, the finding is presumed real.
The problem is that this presumption is wrong — and has been wrong for a very long time.
A Heuristic Mistaken for a Law
The origins of the p-value threshold trace to the British statistician Ronald Fisher, who introduced the concept in his 1925 work Statistical Methods for Research Workers. Fisher himself was careful to characterize 0.05 as a convenient benchmark, not a universal law. He described it as a level at which a researcher might begin to take a result seriously — a starting point for inquiry rather than a terminus.
What followed was a slow but decisive transformation. As scientific publishing scaled across the twentieth century, journal editors and grant committees needed an efficient mechanism for sorting claims. The p-value offered exactly that: a single, clean number that appeared to quantify certainty. By mid-century, the 0.05 threshold had migrated from Fisher's cautious suggestion into something resembling institutional doctrine. Researchers who failed to clear it found their work difficult to publish. Those who cleared it found their findings treated as established facts.
The distinction between "p < 0.05" and "p = 0.06" is, in practical terms, negligible. Statistically, the two values describe nearly identical levels of evidential strength. Yet one generates a publishable result and the other does not. That asymmetry, compounded across thousands of studies annually, has introduced a systematic distortion into the scientific record that is only now being fully reckoned with.
The Replication Crisis as Symptom
The consequences of this distortion became dramatically visible during the replication crisis that swept through psychology, medicine, and nutrition science beginning around 2011. When researchers attempted to reproduce high-profile findings — studies that had cleared the 0.05 threshold and been widely cited — a troubling proportion simply failed to hold up. A landmark 2015 project by the Open Science Collaboration attempted to replicate 100 published psychology studies; fewer than 40 percent produced results consistent with the original findings.
The p-value did not cause this crisis alone, but it created the conditions for it. Because publication depended on achieving significance, researchers — often unconsciously — engaged in practices that inflated the likelihood of crossing the threshold. These included running multiple statistical tests and reporting only the one that succeeded, collecting data until the numbers cooperated, and framing hypotheses after the results were already in hand. Each of these practices is, technically, a form of statistical manipulation. Collectively, they have been given a clinical name: p-hacking.
The pharmaceutical industry offers some of the starkest illustrations. A 2008 analysis published in PLOS Medicine examined clinical trials for antidepressants submitted to the FDA and found that studies with positive results were far more likely to be published than those with negative ones. When the unpublished negative trials were incorporated into the analysis, the apparent efficacy of several widely prescribed medications diminished substantially. Patients and physicians were making decisions based on a curated slice of the evidence — one shaped, in large part, by the gravitational pull of statistical significance.
What the P-Value Actually Measures
Part of what makes the 0.05 threshold so persistently misleading is that it is routinely misunderstood, even by trained scientists. A survey published in Psychological Science found that a majority of academic psychologists incorrectly interpreted p-values when asked to define them directly. The most common misconception is that a p-value below 0.05 means there is a 95 percent probability that the hypothesis is true. It means nothing of the sort.
A p-value measures the probability of obtaining a result at least as extreme as the one observed, assuming the null hypothesis is true. It says nothing about the probability that any particular hypothesis is correct. It does not measure effect size, practical significance, or the likelihood that a finding will replicate. It is, at its core, a statement about what would happen under a hypothetical world in which there is no real effect — not a statement about the world that actually exists.
This distinction is not merely semantic. A study with a very large sample size can produce a p-value well below 0.05 for an effect so small as to be entirely meaningless in practice. Conversely, a study with a small sample size might produce a p-value just above 0.05 for an effect that is both real and substantial. The threshold treats both situations with equal indifference.
The Defense of a Flawed Standard
Given how widely these limitations are acknowledged within the scientific community, the persistence of the 0.05 threshold demands some explanation. The answer is largely institutional. Careers in academic science are built on publications, and publications continue to depend on statistical significance. Abandoning the threshold without a broadly accepted replacement would introduce ambiguity into a system that has organized itself around the appearance of precision.
In 2019, a commentary in Nature signed by more than 800 scientists called for the retirement of statistical significance as a binary concept. The American Statistical Association had issued a similar statement in 2016, warning that the use of p-values to make binary decisions was causing real harm to science. These calls have generated substantial debate but limited institutional change. The threshold remains embedded in grant applications, journal submission guidelines, and regulatory frameworks at agencies including the FDA.
Some researchers have proposed alternatives. Bayesian statistical methods, which incorporate prior probability and update beliefs in light of new evidence, offer a more nuanced framework for evaluating claims. Others advocate for mandatory reporting of effect sizes and confidence intervals alongside p-values, or for pre-registration of study designs to limit post-hoc manipulation. Each approach addresses some of the threshold's deficiencies, though none has achieved the universal adoption that would be necessary to fundamentally restructure the incentive landscape.
Precision as Performance
There is something philosophically instructive in the p-value's strange career. A tool designed to introduce rigor into scientific reasoning became, through institutional pressure and misapplication, a mechanism for generating the appearance of rigor rather than its substance. The number 0.05 did not corrupt science through malice but through convenience — it offered a clean answer to a genuinely difficult question, and the scientific enterprise, under pressure to produce and publish, embraced that cleanliness too completely.
The deeper lesson may be that quantification, however precise its instruments, does not automatically confer objectivity. A threshold is a human decision. The weight placed upon it is a social choice. And when that choice becomes invisible — when a guideline calcifies into a law — the measurement begins to shape the reality it was meant to describe. That dynamic is not unique to statistics. It is, in a meaningful sense, the central problem of scientific epistemology: the ever-present risk that our instruments of understanding quietly become its boundaries.