Anthony Almudevar, Jacob Almudevar
9 min
Abstract
In 2015 the Open Science Collaboration (OSC) (Nosek et al 2015) published a highly influential paper which claimed that a large fraction of published results in the psychological sciences were not reproducible. In this article we review this claim from several points of view. We first offer an extended analysis of the methods used in that study. We show that the OSC methodology induces a bias that is able by itself to explain the discrepancy between the OSC estimates of reproducibility and other more optimistic estimates made by similar studies. The article also offers a more general literature review and discussion of reproducibility in experimental science. We argue, for both scientific and ethical reasons, that a considered balance of false positive and false negative rates is preferable to a single-minded concentration on false positive rates alone.
Alex: Huh, and OSC assumed something way higher, like everything was real?
Sam: Yes—they implicitly took π near 100%, predicting over 90% repeats, but got 36%. The model matches reality better at lower π, like 25%, yielding about 78% expected repeats. This shows the "crisis" is often just math: low prevalence means published hits mix truths and flukes, so repeats naturally fail more. No bad science needed.
Alex: So for fields like clinical trials, where success is maybe 30-50%, we'd expect middling repeats unless we adjust?
Sam: Exactly. Trials often hover near 50% prevalence—called "equipoise," genuine uncertainty between treatments—which predicts solid PPV around 80-95%. But drop below, and repeats suffer. The paper urges balancing false positives and misses ethically, plus more validation studies.
Alex: Okay, but even if low prevalence explains part of the low repeats, OSC claimed high power—like over 90%—so why didn't that deliver more successes?
Sam: That's a sharp point. In projects like OSC, they first looked at the original study's result to guess the effect strength, then sized the repeat study to have good odds of spotting it if real—say, only a 10% chance of missing a true effect. But there's a catch: original studies only publish exciting hits, so their reported strength looks bigger than average, like picking the tallest kids after measuring a biased group. This overestimates the true strength, leading to actual miss rates around 30-70% instead of 10%.
Alex: Wait, so the original hit makes you think the effect is stronger than it really is, and you under-sample because of that?
Sam: Exactly. They convert the original statistic directly to a strength measure, but since only big ones get published, it's truncated, pulling estimates way up. For OSC, post-hoc checks said 92% power using originals, but the model suggests actual power closer to 53% after bias correction, matching the 36% repeats.
Alex: Huh, so adjusting for that truncation fixes the power math and aligns expectations?
Sam: Yes. The paper shows for OSC's setup, realized miss rates hit about 47% versus nominal 8%, a clear discrepancy from selection. It urges proper accounting in multi-stage work, balancing false alarms and misses ethically.
Alex: So that power correction—getting actual miss rates up to 47%—explains why OSC's repeats landed where they did, even with careful planning?
Sam: Precisely. The paper digs into OSC's actual data: they took the original p-values from 97 studies and converted them into strength measures, called z_p scores—like how far the result strayed from zero chance. Using a corrected estimate for true strength—avoiding the truncation bias—the average actual miss rate came out at 47%, not the planned 8%. This predicts about 46% success, right in line with the 36% observed.
Alex: And they break it down by effect type too—like main effects versus interactions?
Sam: Yes. Main effects repeated in 47% of cases. Interactions, the trickier combos, only 22%. The paper ties this to varying effect prevalence: more possible interactions mean lower π overall, dragging down repeats naturally.
Alex: Huh, so interactions are harder because there are way more ways they could be wrong from the start?
Sam: Exactly. And this holds across projects. The model fits without exceptions.
Alex: Okay, so the takeaway is balancing false alarms and misses upfront, not just chasing one?
Sam: Right—the paper pushes ethical trade-offs: fixate on 5% false positives, and you risk 10-50% misses on real effects, hurting discovery. Replication projects like OSC highlight this, urging validated priors on π and debiased power. It reframes low rates as expected, not crisis.
Alex: So if we're balancing those errors ethically, how do researchers actually figure out effect sizes for planning repeats without falling into that overestimation trap?
Sam: A common pitfall is using data from the original study to guess the effect strength. But since only standout results get published, those sizes look bigger than typical true ones, leading to underpowered repeats. Good practice avoids this by skipping pilot data for power math altogether.
Alex: Right, so pilots aren't for estimating strength—what are they for, then?
Sam: Pilots check if a study setup is doable—like recruiting participants or running the procedure smoothly—answering "Can I pull this off?" not "Does it work?" Using them to pin down effect sizes creates a paradox: you need a big enough pilot for a reliable guess, but then why bother with the full study?
Alex: And ethically, that ties back to not wasting resources on known outcomes?
Sam: Precisely. If you already suspect a treatment edge from pilots, running a full trial with inferior options for some feels off. The paper frames this as balancing false positives, like publishing flukes, against false negatives, like shelving real effects. It calls for upfront trade-offs, not fixating on one error.
Alex: Huh, so the "reproducibility crisis" push overlooks that balance, treating all low repeats as failures?
Sam: Yes. The model shows OSC's 36% fits expected rates under realistic prevalence and debiased power—no flaws required. This shifts focus to validated assumptions on effect odds and ethical error weights.
Alex: So pulling this together, the models show OSC's low repeat rate fits standard stats once you factor in realistic effect odds and that power overestimate from cherry-picked originals—no need for a big crisis story.
Sam: Precisely. The paper demonstrates these rates emerge from basic mechanics: mixes of true effects at low prevalence, plus biased power planning that inflates miss rates to around 47%. It reframes the discussion around predictable patterns, not systemic failure.
Alex: Well said, Sam. This gives a clearer way to think about repeats without the panic. Thanks for breaking it down. Thanks for listening to ResearchPod.