Model selection in High-Dimensions: A Quadratic-risk based approach
Surajit Ray, Bruce G. Lindsay
Abstract
In this article we propose a general class of risk measures which can be used for data based evaluation of parametric models. The loss function is defined as generalized quadratic distance between the true density and the proposed model. These distances are characterized by a simple quadratic form structure that is adaptable through the choice of a nonnegative definite kernel and a bandwidth parameter. Using asymptotic results for the quadratic distances we build a quick-to-compute approximation for the risk function. Its derivation is analogous to the Akaike Information Criterion (AIC), but unlike AIC, the quadratic risk is a global comparison tool. The method does not require resampling, a great advantage when point estimators are expensive to compute. The method is illustrated using the problem of selecting the number of components in a mixture model, where it is shown that, by using an appropriate kernel, the method is computationally straightforward in arbitrarily high data dimensions. In this same context it is shown that the method has some clear advantages over AIC and BIC.
Create a lesson
Related papers
Instance-Optimal Adaptive Location Estimation via Multiscale Mid-Summaries
Qiaosen Wang, Chao Gao
Robust Multi-Task Learning for Principal Component Analysis
Dali Liu, Haolei Weng
Principal component error in high-dimensional factor models
Alex Bernstein, Lisa R. Goldberg, Nicholas Gunther et al.
Approximation Theorems for High-Dimensional Canonical U-Statistics: Gaussian Chaos and Phase Transition
Leheng Cai, Qirui Hu
On the parametric and semiparametric Fisher information matrix for non-zero mean stationary spherical invariant random processes
Jean-Pierre Delmas, Habti Abeida, Stefano Fortunati
Inference for two-stage sampling in spatial surveys
Guillaume Chauvet, Olivier Bouriaud, Trinh H. K. Duong