A Total Statistical Error Framework for Comparing Census Data Collection Methods
Siu-MIng Tam, Anders Holmberg
Abstract
Population censuses increasingly rely on imputation to assign usual-residence addresses for non-responding dwellings, yet no formal statistical framework has existed for comparing competing imputation methods on their combined coverage and address-accuracy performance. We develop such a framework within a Total Statistical Error (TSE) paradigm that is not restricted to survey-based data collection and applies equally to register-based and administrative data sources. The central quantity is a unit-level binary correctness indicator, equal to one if and only if a person is both enumerated and assigned to the correct usual-residence address; its complement is the unit TSE, and aggregating over the population yields a correctness rate that serves as the basis for head-to-head method comparison. We distinguish two assessment paradigms: Paradigm I, in which a post-enumeration survey (PES) provides reference values for a probability sample, and Paradigm II, in which a complete benchmark makes the correctness rate directly computable. We illustrate the framework using 2021 Australian census microdata, comparing k-nearest-neighbour and random forest imputation under a not-missing-at-random (NMAR) mechanism. A key finding is that dependent nonresponse between the census and a simulated PES causes naive Paradigm I subgroup estimates to be severely biased when subgroup deletion rates are small. Augmenting the PES response propensity model with the census nonresponse indicator reduces this bias by approximately 90%; simulation experiments show that the nonresponse indicator alone drives the correction, with the additionally included imputed set A values contributing negligible further reduction. This is the preferred strategy whenever dependent nonresponse between the census and PES is a concern.
Create a lesson
Related papers
RECaST-Surv: A Calibrated Borrowing Method for Survival Endpoints in Unequal Randomized Trials
Dehua Bi, Arlina Shen, Ruben P. A. van Eijk et al.
Beyond Pretrends: A Discordance-Based Sensitivity Analysis for Difference-in-Differences
Thomas Leavitt
Earth and space observations meet complex algebras: from complex to octonions for multivariate autoregressive time series analysis
Susana Eyheramendy, Felipe Elorrieta, Wilfredo Palma et al.
Efficient transport and generalization of survival treatment effects
Axel Martin, Iván Díaz, Michele Santacatterina
A new tractable Archimedean copula for full-range tail dependence
Lei Hua
Doubly valid and doubly sharp sensitivity analysis to unobserved confounding for survival outcomes
Jean-Baptiste Baitairian, Bernard Sebastien, Rana Jreich et al.