HARKing

HARKing

HARKing is hypothesising after the results are known: you run the study, look at what came out, and then write the paper as though the interesting result was what you had predicted all along. Nothing is fabricated. The data are real and the analysis may be perfectly competent. What has been falsified is the sequence, and the sequence is the only thing that makes a confirmatory test mean anything. The term and the analysis are Norbert Kerr's, from 1998. It is a named research practice rather than an effect, so this page carries no prevalence figure.

What it is

The reason it matters more than it sounds is a point about what evidence is. A prediction made in advance can turn out wrong, and that possibility is precisely what gives it evidential weight. A prediction reverse-engineered from the answer cannot be wrong, because it was written to fit. So the paper reads like a successful test and is in fact a description of what happened to be in the data.

Kerr set out the practice and its consequences in Personality and Social Psychology Review in 1998. It belongs to the family of questionable research practices that sits behind the replication crisis, alongside p-hacking, the file-drawer problem and low statistical power, and it travels with them. P-hacking manipulates the analysis until something reaches significance. HARKing rewrites the question to match whatever turned out significant. They are usually found together because they are two ends of the same manoeuvre.

The incentive underneath is publication bias. Nobody does this for entertainment. A clean confirmed prediction is publishable and an exploratory finding is not, and the practice is the shortest route from the second to the first.

In effect

The countermeasure is preregistration, and it fixes this more cleanly than it fixes anything else, because a timestamped hypothesis is a complete answer to it. Preregistration does not prevent a researcher from exploring. It prevents the exploration from being presented afterwards as a test.

The reason the practice rarely feels dishonest from the inside is hindsight bias, the ordinary cognitive version of the same move. Once you know the answer, the reasoning that leads to it feels like the reasoning you would have done anyway, and writing it down as the prediction feels like tidying rather than like misreporting.

As an analytical tool, the value of the concept for a reader is a question to ask of any confident finding. Was this prediction registered before the data were collected, and if not, what is the difference between this paper and a description of one dataset? A preregistered null result is worth more than an unregistered hit, and that ordering is counterintuitive enough to be worth stating.

Textbook treatments of the practice tend to name the mechanism and never show it operating. This publication has met an account that lists it among the causes of the replication crisis, pairs it correctly with preregistration, names no originator, and gives not one example of a published paper that did it.

What it does not say

It is not fraud. No data are altered and no results are invented, which is exactly why it is hard to detect and easy to do.

It is not the same as exploratory research. Exploration is legitimate and necessary. The failure is in presenting it as confirmation afterwards.

It is not measured here. How common the practice is, is a testable question, and this page carries no figure because none has been established from a source this publication has opened.

And a literature full of retrofitted hypotheses does not announce itself. It looks cumulative, with each paper confirming a prediction, which is the whole reason the practice damages a field rather than merely damaging a paper.


Sources

  1. Kerr, N. L. (1998). "HARKing: Hypothesizing after the results are known." Personality and Social Psychology Review, 2(3), 196-217. The term and the analysis. This publication's own identification of the originator, not the secondary source's.
  2. Furnham, A. The New Psychology, ch. 41. Supplies the term, its expansion and its place among the causes of the replication crisis, pairs it with preregistration, and names no originator and no worked case.