Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015
Camerer and colleagues took twenty-one social science experiments published in Nature and Science between 2010 and 2015, the most visible journals in the field, and tried to repeat them properly. The replications were designed in advance, reviewed by the original authors, registered before the data were collected, and run with samples around five times larger than the originals. Roughly two in three produced a significant effect in the same direction as the original. Among those that did, the effect was on average about half the size originally reported. So the picture is not simply that findings are true or false; the ones that survive are also weaker than published, which means a literature built on original effect sizes will overpromise even where it is right. One further result is worth noting. Researchers were asked beforehand to predict which studies would replicate, and their predictions tracked the outcomes closely, which suggests the field could tell and that the failures were not merely bad luck.
Camerer and colleagues took twenty-one social science experiments published in Nature and Science between 2010 and 2015, the most visible journals in the field, and tried to repeat them properly. The replications were designed in advance, reviewed by the original authors, registered before the data were collected, and run with samples around five times larger than the originals. Roughly two in three produced a significant effect in the same direction as the original. Among those that did, the effect was on average about half the size originally reported. So the picture is not simply that findings are true or false; the ones that survive are also weaker than published, which means a literature built on original effect sizes will overpromise even where it is right. One further result is worth noting. Researchers were asked beforehand to predict which studies would replicate, and their predictions tracked the outcomes closely, which suggests the field could tell and that the failures were not merely bad luck.
Written from the abstract, Europe PMC.
What our sources record
Reproduced as it was written when this source was checked, figures and all.
2, 637-644. Failed to replicate Sparrow et al.
Where we use it
- Both Sides of the Screen-Time Argument, Why We Look
Reliability of this reference
The link below was matched against the publisher's record on title, first author and year, and all three agree.
The source
https://doi.org/10.1038/s41562-018-0399-z
DOI: 10.1038/s41562-018-0399-z