Selection Criterion for Log-Linear Models Using Statistical Learning Theory
Daniel Herrmann, Dominik Janzing
Abstract
Log-linear models are a well-established method for describing statistical dependencies among a set of n random variables. The observed frequencies of the n-tuples are explained by a joint probability such that its logarithm is a sum of functions, where each function depends on as few variables as possible. We obtain for this class a new model selection criterion using nonasymptotic concepts of statistical learning theory. We calculate the VC dimension for the class of k-factor log-linear models. In this way we are not only able to select the model with the appropriate complexity, but obtain also statements on the reliability of the estimated probability distribution. Furthermore we show that the selection of the best model among a set of models with the same complexity can be written as a convex optimization problem.
Create a lesson
Related papers
Instance-Optimal Adaptive Location Estimation via Multiscale Mid-Summaries
Qiaosen Wang, Chao Gao
Robust Multi-Task Learning for Principal Component Analysis
Dali Liu, Haolei Weng
Principal component error in high-dimensional factor models
Alex Bernstein, Lisa R. Goldberg, Nicholas Gunther et al.
Approximation Theorems for High-Dimensional Canonical U-Statistics: Gaussian Chaos and Phase Transition
Leheng Cai, Qirui Hu
On the parametric and semiparametric Fisher information matrix for non-zero mean stationary spherical invariant random processes
Jean-Pierre Delmas, Habti Abeida, Stefano Fortunati
Inference for two-stage sampling in spatial surveys
Guillaume Chauvet, Olivier Bouriaud, Trinh H. K. Duong