Joseph P. Simmons, Leif D. Nelson, Uri Simonsohn
6 min
In empirical psychology, the standard for statistical significance is typically a p-value of 0.05 or less. However, the authors argue that this threshold is effectively meaningless when researchers have excessive flexibility in how they collect, analyze, and report their data. This flexibility—referred to as researcher degrees of freedom—includes decisions such as when to stop collecting data, which variables to include or exclude, and which covariates to use. Because researchers are often motivated to find significant results, they may unconsciously (or consciously) explore various analytic paths until they find one that yields a p-value below the 0.05 threshold, reporting only that successful path.
To illustrate how easily this flexibility can lead to false-positive findings, the authors conducted two experiments designed to show that listening to certain songs can change a person's age. By selectively choosing which variables to report and using specific covariates, they were able to produce statistically significant results for these absurd hypotheses. They supplemented these experiments with computer simulations, which showed that combining common, seemingly minor analytic choices can increase the actual false-positive rate from the nominal 5% to as high as 61%.
To combat this, the authors propose a set of six requirements for authors and four guidelines for reviewers. The core of the solution is transparency: authors must disclose their data collection termination rules, report all variables collected, and provide results both with and without covariates or data exclusions. The goal is not to eliminate exploratory research, but to ensure that such research is clearly labeled as such and that readers can distinguish between robust findings and those that rely on arbitrary, post-hoc analytic decisions.
In this article, we accomplish two things. First, we show that despite empirical psychologists' nominal endorsement of a low rate of false-positive findings (≤ .05), flexibility in data collection, analysis, and reporting dramatically increases actual false-positive rates. In many cases, a researcher is more likely to falsely find evidence that an effect exists than to correctly find evidence that it does not. We present computer simulations and a pair of actual experiments that demonstrate how unacceptably easy it is to accumulate (and report) statistically significant evidence for a false hypothesis. Second, we suggest a simple, low-cost, and straightforwardly effective disclosure-based solution to this problem. The solution involves six concrete requirements for authors and four guidelines for reviewers, all of which impose a minimal burden on the publication process.
Alex: [checking understanding] So the takeaway is not that we should abandon p-values, but that we need to treat the process of reaching them with as much rigor as the data itself.
Sam: [calm] Exactly. The burden of proof shifts from simply reporting a significant p-value to demonstrating that the result was not manufactured through selective analysis. [[RP_SECTION:data-dependent-stopping|Data Dependent Stopping]]
Alex: [leaning in] I want to press on the mechanics a bit more. How does this actually play out during data collection — before anyone has even started making post-hoc choices?
Sam: [measured] The authors highlight one practice in particular: checking for significance while data is still being gathered. If you run a t-test after every new pair of observations and stop the moment you hit significance, your false-positive rate climbs well above twenty percent — even when there is no real effect in the population. It is data-dependent stopping, and it exploits the inherent variance in small samples.
Alex: [analytical] So even without any deliberate manipulation, the act of peeking at the p-value and deciding when to stop is itself a form of gaming the system.
Sam: [nodding in voice] Right. And there is a counterintuitive trap here. Researchers often treat early significance as a sign that an effect is robust — "it showed up with only twenty subjects, so it must be real." But that intuition is inverted. Early significance in a sequential-testing regime is often just noise that happened to align with the hypothesis before variance had a chance to average out. [[RP_SECTION:pre-registration-and-transparency|Pre-registration and Transparency]]
Alex: [reflecting] Which is why the pre-registration norm has become so central to the replication reform movement.
Sam: [precise] That is exactly what the authors are driving at. Their most load-bearing requirement is a pre-specified termination rule — you commit to a sample size before collecting a single data point. That one constraint removes the temptation to keep tuning the dial. The other five requirements work in support of it: report all variables collected, all conditions run, all excluded observations, and demonstrate that the result survives with and without any post-hoc exclusions.
Alex: [deliberate] So if a researcher drops an outlier or adds a covariate, they have to show the result with and without that choice — making the sensitivity of the finding visible to reviewers.
Sam: [sitting back] Exactly. The goal is not to eliminate analytic flexibility — some of it is genuinely necessary. The goal is to ensure the final result does not hinge on an arbitrary, undisclosed choice made after the researcher already knew which direction the data were pointing. Transparency converts a hidden garden of forking paths into a documented, reproducible record.
Alex: [measured] And that is where the paper's argument lands with real force. It is not a critique of any individual researcher's integrity — it is a structural argument. The incentive system and the reporting norms, as they stood in 2011, made false positives nearly inevitable even for researchers acting in good faith. [[RP_SECTION:institutional-reform|Institutional Reform]]
Sam: [grounded] Which is why the proposed remedies are institutional rather than personal. Journals enforcing disclosure requirements, reviewers demanding pre-registration, editors treating a null result with a clean design as publishable — those are the levers. Individual researchers cannot fix a coordination problem on their own.
Alex: [thoughtful] More than a decade on, it is worth noting how much of this agenda has actually moved — pre-registration is now standard in many fields, registered reports exist, and the replication crisis has become a genuine area of meta-scientific research. But the underlying degrees of freedom have not gone away. The dial is still there.
Sam: [calm, closing] It is. What has changed is that we now have a clearer vocabulary for the problem and, in some corners of the literature, the institutional scaffolding to constrain it. Simmons and colleagues gave us the language to name what was happening — and that turns out to be a necessary first step toward fixing it. Thanks for listening to ResearchPod.