Model-Agnostic Feature Selection via LOCO-Guided Adaptive Minipatch Sampling
Xuhui Liu, Lili Zheng
Abstract
Black-box machine learning models increasingly deliver strong predictions, but extracting useful information from them, such as a set of important features, remains challenging. Existing model-agnostic methods primarily estimate feature importance or conduct inference on it rather than directly selecting features, whereas many feature selection methods are model-specific or rely on the model-X assumption. We introduce LOCO-guided Adaptive Minipatch Sampling (LAMPS), a model-agnostic ensemble framework that uses any black-box regression algorithm as its base learner to select features important for predicting the response. The base learner need only produce predictions and need not perform feature selection itself. LAMPS operates within a minipatch ensemble framework that subsamples both observations and features, allowing leave-one-covariate-out (LOCO) feature importance scores to be easily computed. It adaptively concentrates minipatch sampling on features with high LOCO scores while maintaining exploration. The resulting sampling probabilities rapidly separate signal from noise features after a few iterations, enabling selection through simple thresholding. We establish that LAMPS achieves exact feature selection in high-dimensional settings, provided that the base predictive models are sufficiently well trained on average. Extensive experiments on synthetic and real data show that LAMPS outperforms state-of-the-art feature selection methods, with particularly strong performance in the presence of correlated features.
Create a lesson
Related papers
Empirical Auditing of Edge-Private Graph Generators
Anum Fatima, Stratis Limnios, James Adams et al.
Variational objectives for amortized Bayesian inference in inverse problems: The role of posterior conditioning
Abhishek Srivastava, Arijit Hazra, Rajesh Dubbaku
JAREX: An Acquisition Function for Multi-Objective Algorithmic Process Characterization
Xinyang Li, Kevin Stone, Ajit Vikram
Identifying Representational Biases in Datasets Using PCA: A Max-Disparity Partition Framework
Arjun KM, Shashi Jain
Beyond Point Prediction: Artificial Representative Trees with Uncertainty
Lea L. Mairhöfer, Silke Szymczak, Björn-Hergen Laabs et al.
Adversarially Robust PAC Learning with Optimal VC Rates
Steve Hanneke, Amirreza Shaeiri