Principal component error in high-dimensional factor models
Alex Bernstein, Lisa R. Goldberg, Nicholas Gunther, Alec N. Kercheval, Tian Lan, Yian Lin, Dayi Yao
Abstract
In a statistical factor model, principal components (or eigenvectors) of a sample covariance matrix serve as estimates of principal directions, the true drivers of co-movement of a collection of observed variables. We write the often substantial error in these estimates as a sum of two interpretable terms, which we show have almost sure asymptotic limits as the number of variables grows with sample size bounded. This scenario is commonplace in financial economics, genomics, machine learning and signal processing. Out-of-subspace error measures the distance from an estimate to the subspace spanned by population factor exposures. It can be expressed in terms of data, providing an estimable floor for error. In-subspace error arises from the fixed sample size of the latent factor returns and cannot be estimated from data alone. We illustrate our error analysis with a three-factor simulation of the US public equity market, showing the dependence of the magnitude of the error and its components on dimension and sample size. In that simulation, out-of-subspace error dominates. Researchers who rely on principal component analysis to estimate factor models can use our results to quantify errors in model-based predictions and attributions.
Create a lesson
Related papers
Instance-Optimal Adaptive Location Estimation via Multiscale Mid-Summaries
Qiaosen Wang, Chao Gao
Robust Multi-Task Learning for Principal Component Analysis
Dali Liu, Haolei Weng
Approximation Theorems for High-Dimensional Canonical U-Statistics: Gaussian Chaos and Phase Transition
Leheng Cai, Qirui Hu
On the parametric and semiparametric Fisher information matrix for non-zero mean stationary spherical invariant random processes
Jean-Pierre Delmas, Habti Abeida, Stefano Fortunati
Inference for two-stage sampling in spatial surveys
Guillaume Chauvet, Olivier Bouriaud, Trinh H. K. Duong
Revisiting the Brunner-Munzel test from the viewpoint of local linear approximation
Makito Oku