Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015

Camerer (2018) · Evaluating the replicability of social science experiments in Nature and Science between · Read the paper

Camerer and colleagues took twenty-one social science experiments published in Nature and Science between 2010 and 2015, the most visible journals in the field, and tried to repeat them properly. The replications were designed in advance, reviewed by the original authors, registered before the data were collected, and run with samples around five times larger than the originals. Roughly two in three produced a significant effect in the same direction as the original. Among those that did, the effect was on average about half the size originally reported. So the picture is not simply that findings are true or false; the ones that survive are also weaker than published, which means a literature built on original effect sizes will overpromise even where it is right. One further result is worth noting. Researchers were asked beforehand to predict which studies would replicate, and their predictions tracked the outcomes closely, which suggests the field could tell and that the failures were not merely bad luck.

Camerer and colleagues took twenty-one social science experiments published in Nature and Science between 2010 and 2015, the most visible journals in the field, and tried to repeat them properly. The replications were designed in advance, reviewed by the original authors, registered before the data were collected, and run with samples around five times larger than the originals. Roughly two in three produced a significant effect in the same direction as the original. Among those that did, the effect was on average about half the size originally reported. So the picture is not simply that findings are true or false; the ones that survive are also weaker than published, which means a literature built on original effect sizes will overpromise even where it is right. One further result is worth noting. Researchers were asked beforehand to predict which studies would replicate, and their predictions tracked the outcomes closely, which suggests the field could tell and that the failures were not merely bad luck.

Written from the abstract, Europe PMC.

What our sources record

Reproduced as it was written when this source was checked, figures and all.

2, 637-644. Failed to replicate Sparrow et al.

Where we use it

  • Both Sides of the Screen-Time Argument, Why We Look

Reliability of this reference

The link below was matched against the publisher's record on title, first author and year, and all three agree.

The source

https://doi.org/10.1038/s41562-018-0399-z

DOI: 10.1038/s41562-018-0399-z