Pooling Versus Ensembling for Ridge Regression Under Covariate Shift
Maya Ramchandran, Rajarshi Mukherjee
Abstract
Datasets in many settings naturally partition into clusters arising from sub-populations, batch effects, or aggregation across multiple sources. A common response to such heterogeneity is to ensemble learners trained on each cluster rather than fit a single model to the pooled data. Prior work motivating such approaches has typically considered settings in which both the covariate distribution and the conditional outcome model differ across clusters; the role of cluster-aware partitioning and ensembling based solely on the covariate distribution remains to be explored. We address this case for ridge-regularized least-squares regression under a linear outcome model and consider all ridge penalty values λ≥ 0, including the special case of the ridgeless predictor at λ= 0. By considering both fixed-effects and random-effects models, we argue that under random effects, an optimally tuned pooled ridge predictor always outperforms ensembles of individually optimally tuned predictors. For fixed effects, we derive a general formula for the pooled and ensembled predictors to characterize the role of both regression coefficients as well as the predictor distribution shifts. Together, these results generalize prior risk analyses of bagging and random-partition estimation using ridge and ridgeless regression predictors from the i.i.d. setting to encompass covariate shift and heterogeneity-aware partition structure.
Create a lesson
Related papers
Minimax optimality for sequential gradient-free minimization of smooth functions and their derivatives
Théo Paquier, Alexandre B Tsybakov, François Portier et al.
Randomization Inference with Concentration Inequalities
Tobias Freidling
On the continuity of the Tukey depth function for fuzzy data
Luis González-De La Fuente, Alicia Nieto-Reyes, Pedro Terán
Recursive-Head Geometry and Order-Free Efficient Inference in Finite-State Nested Markov Models
Haoyu Wei
Finite-Sample Hausdorff Bounds and Hadamard Sensitivity for Regressions with MNAR Covariates
Hugo Dunias
Semiparametric Efficient Inference under Non-Informative Complex Survey Designs
Hiroki Chiba, Kosuke Morikawa