Scale-invariant Optimal Sampling for Rare-events Data with Sparse Models
Jing Wang, HaiYing Wang, Qiang Zhang, Hao Helen Zhang
Abstract
Subsampling is effective in tackling computational challenges for massive data with rare events. Overly aggressive subsampling may adversely affect estimation efficiency, and optimal subsampling is essential to mitigate the information loss. However, existing optimal subsampling probabilities depend on data scales, and some scaling transformations may result in inefficient subsamples. This problem is more significant when there are inactive features, because their influence on the subsampling probabilities can be arbitrarily magnified by inappropriate scaling transformations. We tackle this challenge and introduce a scale-invariant optimal subsampling function in the context of sparse models, where inactive features are commonly assumed. Instead of focusing on estimating model parameters, we define an optimal subsampling function to minimize the prediction error, using adaptive lasso to outline the estimation procedure and study its theoretical guarantee. We first introduce the adaptive lasso estimator for rare-events data and establish its oracle properties, thereby validating the use of subsampling. Then we derive a scale-invariant optimal subsampling function that minimizes the prediction error of the inverse probability weighted (IPW) adaptive lasso. Finally, we present an estimator based on the maximum sampled conditional likelihood (MSCL) to further improve the estimation efficiency. We conduct numerical experiments using both simulated and real-world data sets to demonstrate the performance of the proposed methods.
Create a lesson
Related papers
A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings
Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser
Fast Learning Rates for Physics-Informed Kernel Methods
Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti et al.
Rank and computation of the pathlifting Jacobian of a DAG ReLU network
Manon Verbockhaven
Preservation of Log-Concavity and Convergence of Wasserstein-Fisher-Rao Gradient Flows
Francesca Romana Crucinio, Sahani Pathiraja
Generalized DCCQ: From Binary Quotients to Multinomial Simplex Geometry and Critical-Strip Coordinates
Y. Kenan Yılmaz
Bracketing Uncertainty in Clustering Under the Manifold Hypothesis
Savik Kinger, Luciano Dyballa, Steven W. Zucker