Skip to content

Design-Based Prediction-Powered Inference for Spatial Data

Shinichiro SHirota

stat.MEarXiv:2608.10356

Abstract

Prediction-powered inference (PPI) combines a wall-to-wall prediction map with a small gold-standard sample to give confidence intervals valid whatever the map's quality. Canonical PPI theory starts from i.i.d.\ labelling, whereas spatial labels arrive through survey designs or covariate-driven mechanisms, and map errors may be spatially correlated. We recast PPI in a design-based framework: the estimand is a census parameter of a fixed spatial population, with randomness arising from the labelling mechanism. We derive exact design variances under simple and stratified sampling, a threshold for when blocked spatial balance pays, and sandwich inference for estimated propensities when selection depends on the map. Our main result concerns double robustness. With a misspecified propensity, a correct outcome model secures superpopulation identification but, conditional on the realised population, leaves a remainder of order σu/Neff,v, an effective count of the residual patches the weights see. Under ratio-stable labelling this remainder is free of the label count, so coverage can deteriorate as labels accumulate. For i.i.d.\ or exchangeable residual fields Neff,v is of order N and the remainder is negligible beside sampling error when n/N 0; spatially coherent dependence instead makes it bind. We reproduce this on a fully enumerated population of 48,175 cells. Estonian LUCAS applications show that power tuning and dependence diagnostics must respect the design: i.i.d.\ PPI++ tuning worsens precision for the best map, whereas design-matched tuning cuts standard errors by about 10\% and matches or beats PPI++ across seven land-cover estimands. Pooled residual diagnostics can likewise mistake spatially structured between-stratum variation for residual dependence.

Create a lesson