How Replicable Are Statistically Significant Findings?
Patrick Vu, Stefan Faridani
Abstract
In the empirical sciences, significance thresholds often determine whether findings are treated as evidence of an effect. This paper studies how likely findings that just meet conventional significance thresholds are to remain significant in replications of the same sample size. To answer this question, we estimate the expected replication probability conditional on a given p-value among published studies for experimental economics, psychology, and social science. We validate this measure by showing it accurately predicts actual replication outcomes, outperforming prediction markets. A finding with a p-value of 0.05 has an expected replication probability ranging from 0.10 to 0.25 across fields. Low replicability reflects low power in original studies rather than publication bias. We then develop a nonparametric estimator and apply it to economics literatures that use larger samples, finding higher but still low replication probabilities. These results indicate that statistical significance in a single study provides only suggestive evidence of an effect. Stronger conclusions require cumulative evidence.
Create a lesson
Related papers
Shrinkage Bayesian Causal Forest with Instrumental Variable
Lennard Maßmann, Jens Klenke
Conditionally linear, matrix normal state space models
Drew D. Creal, Marcelo C. Medeiros, Rodrigo Sarlo
Policy Targeting with Market Equilibrium
Gyungbae Park
What No First Stage Can Detect: Functional-Form Contamination in Linear IV
Parush Arora
Tensor-BEKK: Conditional Covariance Modeling and Inference for Tensor-Valued Time Series
Huan Gong, Feiyu Jiang
Profiled Anderson--Rubin Test: Robust Inference Allowing for Direct Effects of Instruments
Jung Hyub Lee