P-Hacking

P-Hacking

P-hacking is reworking an analysis until it crosses the conventional significance threshold. Drop the outliers, add a covariate, split by sex, collect twenty more participants and look again, try the other outcome measure, and stop the moment something goes below the line. Each individual decision is defensible. The sequence is not, because the threshold only means what it claims to mean if the analysis was fixed before the data arrived. This is a named research practice rather than an effect, and it is a description of a mechanism, never an accusation about any particular person or paper.

What it is

The demonstration that put the practice on the map is False-Positive Psychology, by Joseph Simmons, Leif Nelson and Uri Simonsohn. All three names belong on it, and Nelson is the one routinely dropped from the credit. What they showed is that ordinary analytic flexibility, used in good faith, can manufacture a significant result almost at will, without anything a researcher would recognise as rule-breaking.

The older and more honest synonym is data fishing, because it describes the behaviour rather than the statistic. The mechanism is degrees of freedom: every undeclared choice in an analysis is a separate chance for noise to cross the line, and the published threshold does not account for the choices that were tried and set aside.

This publication holds no prevalence figure for the practice, because none was supplied by the source that raised it here and none is invented.

In effect

What p-hacking explains is a shape rather than a paper. A literature in which almost every published result is just significant, and in which the effect shrinks each time somebody runs it again with the analysis fixed in advance, looks the way a p-hacked literature looks from the outside.

The nearest thing to a worked demonstration in this library is the willpower-belief tally, where seven of eight null studies were preregistered against one of five positive ones. That is a sorting rule, not a charge against anyone: preregistration removes the analytic freedom, and the results on the preregistered side come out differently.

The countermeasure that works is structural. Preregistration takes away the degrees of freedom rather than asking anyone to behave better, which is why it is the response the field settled on. The enabling condition is low statistical power, since underpowered noisy studies give the most room to fish, and the motive is publication bias, since journals set the threshold and fishing is the rational response to it.

What it does not say

It does not say that anyone has committed fraud. Every step in the sequence is a decision a competent analyst might make for good reasons, and the problem is the undeclared accumulation of them.

It does not say how common the practice is. No prevalence estimate is carried here, because none was supplied by the source and this publication does not manufacture one.

It does not say that a significant result is worthless. It says that a threshold crossed after an unrecorded search does not mean what a threshold crossed under a fixed analysis means.

It does not say that researchers know they are doing it. The cognitive half matters: people genuinely believe small samples carry more signal than they do, so the fishing does not feel like fishing.


Sources

  1. Simmons, J. P., Nelson, L. D., & Simonsohn, U. "False-Positive Psychology." The demonstration that ordinary analytic flexibility can produce significance at will. Year and journal are deliberately not printed here, because the paper has not been opened for this entry; trace it before citing a locator.
  2. Furnham, A. The New Psychology, ch. 41. Supplies the term, the conventional threshold, "data fishing" as a synonym, and the practice's place among the causes of the replication crisis. The book names no originator, no year, no journal, no prevalence estimate and no worked instance; the attribution to Simmons, Nelson and Simonsohn is this publication's, not the book's.