Many Labs 2: Investigating Variation in Replicability Across Samples and Settings

Klein (2018) · Advances in Methods and Practices in Psychological Science · Read the paper

Many Labs 2 tested twenty-eight published findings, classic and recent, across a hundred and twenty-five samples in thirty-six countries, with more than fifteen thousand participants and protocols peer reviewed before any data were collected. Slightly more than half of the findings replicated in the same direction as the original. The size comparison is more informative than the count: the median original effect was moderate, the median replication effect was small, three quarters of the replications came out smaller than the original, and around a third pointed in the opposite direction. The project's second question was whether the failures could be explained by samples and settings, the standard defence that an effect appears only under the right conditions. Mostly they could not. Variation between settings was modest, and where it appeared it clustered among the strongest effects rather than the weakest. Whether a finding replicated depended much more on the finding than on where it was tested. This vault cites it as the best evidence against the hidden-moderator explanation.

Many Labs 2 tested twenty-eight published findings, classic and recent, across a hundred and twenty-five samples in thirty-six countries, with more than fifteen thousand participants and protocols peer reviewed before any data were collected. Slightly more than half of the findings replicated in the same direction as the original. The size comparison is more informative than the count: the median original effect was moderate, the median replication effect was small, three quarters of the replications came out smaller than the original, and around a third pointed in the opposite direction. The project's second question was whether the failures could be explained by samples and settings, the standard defence that an effect appears only under the right conditions. Mostly they could not. Variation between settings was modest, and where it appeared it clustered among the strongest effects rather than the weakest. Whether a finding replicated depended much more on the finding than on where it was tested. This vault cites it as the best evidence against the hidden-moderator explanation.

Written from the abstract, OpenAlex.

What our sources record

Reproduced as it was written when this source was checked, figures and all.

Klein, R. A., Vianello, M., Hasselman, F., et al. (2018). 'Many Labs 2: Investigating variation in replicability across samples and settings.' Advances in Methods and Practices in Psychological Science, 1(4), 443-490. Twenty-eight findings, protocols peer reviewed in advance, 125 samples, 15,305 participants, 36 countries; 54 per cent significant in the same direction, 50 per cent at a strict threshold. Findings were selected for suitability to the design, which this article states before comparing rates. doi:10.1177/2515245918810225

Where we use it

  • The Psychology You Learned Didn't Replicate, Why We Act

Reliability of this reference

The link below was matched against the publisher's record on title, first author and year, and all three agree.

The source

https://doi.org/10.1177/2515245918810225

DOI: 10.1177/2515245918810225