How Reliable Are Psychological Measurements? The Distribution of Marginal Reliability Across 889 Item-Response Datasets
JoonHo Lee
Abstract
Nearly every quantitative study in psychology reports a reliability coefficient, so the field knows a great deal about the reliability that authors choose to publish. It knows much less about the reliability of the data psychology actually produces, because published coefficients pass through decisions about what to compute and what to report. We therefore measure reliability directly, applying the same estimators under the same rules to 889 datasets from the Item Response Warehouse, a public collection of item-response data that spans cognitive tests, clinical screeners, personality inventories, and attitude scales. Three findings emerge. Low reliability is common: even under a lenient definition of reliability, 30% of datasets fall below the conventional .80 threshold. The variation across datasets is real rather than statistical, since estimation noise accounts for only about one percent of it. Finally, the answer depends on the definition itself: under a strict definition the share below .80 rises to 52%, a difference large enough to change what one concludes about the field. We conclude that a reliability report should say which definition it uses, attach a measure of uncertainty, and give a strict coefficient alongside a lenient one.
Create a lesson
Related papers
Characterising mortality dynamics across countries and time using a multi-stage clustering approach
Pedro Menezes de Araújo, Ugofilippo Basellini, Thomas Brendan Murphy et al.
Combining Weather Forecast Aggregation and State-Space Models for Adaptive Probabilistic Electricity Load Forecasting
Joseph de Vilmarest, Jonathan Dumas, Jean Thorey
Anthropogenic Forcing, Climate Change, and the Shape of Warming: Statistical Inference for Distributional Cointegration
Won-Ki Seo, Kyungsik Nam
Observational constraints on net radiative forcing confirm aviation contrail warming
Aaron Sonabend-W, Scott Geraedts, Nita Goyal et al.
Nationally Consistent, Locally Incomplete: A Bayesian Remote-Sensing Audit of Rooftop Photovoltaic Registries
Gabriel Kasmi, Yves-Marie Saint-Drenan, Laurent Dubus et al.
Temporal Seam Score for Assessing Continuity at Known Transitions in Time Series
Hongxiao Jin