Model-Agnostic Influential Outlier Detection for Mixed Effects and Multi-Level Models
Colin C Jones, David A Campbell, Yan Liu
Abstract
Influential Outlier Detection is developed for mixed-effects models on clustered data. The Influential Outlier Metric is defined as a combination of SHapley Additive exPlanantion (SHAP) values and model residuals, both of which undergo a change of measure transformation. Building on previous work showcasing the suitability of using Normalizing flows to map arbitrary distributions to a flexible base distribution for statistical inference, the Normalizing Flows are constructed to allows contextual information and also provide a goodness of fit diagnostic for model evaluation. The use of SHAP values in the construction moves away from model specific tools and instead provides point-wise model agnostic influential outlier. The advantages and limitations of this approach are examined in several models including the linear model, the random forest, and gradient-boosted trees.
Create a lesson
Related papers
A Monte Carlo Estimator for an Isolated Polynomial Zero via Contour Integral Representations
Athanasios Christou Micheas
On skew-symmetric distributions and their use in Monte Carlo sampling algorithms: coordinate-free, Gibbs-style and manifold versions of the Barker proposal
Minh Vu, Samuel Livingstone, Pantelis Samartsidis
Multifidelity Formulations for Triangular Transport
Owen Davis, Daniel Sharp, Youssef Marzouk et al.
Learn-Then-Differentiate Gradient Estimation
Nifei Lin, Qingkai Zhang, L. Jeff Hong
Stochastic Variational Inference for Vine Copula Distributional Regression
Gianmarco Callegher, Thomas Kneib
Ready for the Clinic? A Survey of Open-Source Software for Response-Adaptive Randomization in Clinical Trials
Stina Zetterstrom, David S. Robertson, Sofía S. Villar