Learning from Uncertainty-dependent Missing Labels for Semi-supervised Classification
You-Gan Wang, Jinran Wu, Geoffrey J. McLachlan
Abstract
Missing labels are usually regarded as a source of information loss in classification. We study a semi-supervised setting in which the probability of label missingness depends on the observed features through posterior classification uncertainty. In this setting, the missingness indicator is not only a record of an unobserved label, but also an observable signal generated by a mechanism linked to the classifier. We develop a likelihood-based information theory for such uncertainty-dependent missing labels. Under correct specification, we derive a Fisher-information decomposition that separates a partial-labeling component from a nonnegative mechanism-curvature term. Under joint misspecification of the label model and the missingness mechanism, we obtain the corresponding Godambe--Eicker--Huber--White sensitivity and sandwich-covariance partitions. We also clarify the relevant complete-data benchmark: favorable missingness can increase information relative to ordinary fully labeled or budget-matched non-informative labeling baselines, but cannot exceed the information in the augmented experiment in which labels and mechanism indicators are both observed. For plug-in classifiers, we connect the information decomposition to margin-based excess-risk bounds. In regular two-component mixture settings this yields the parametric \(n-1\) excess-risk rate, with constants determined by the nuisance-adjusted information in discriminant directions. Gaussian-mixture calculations and a medical diagnosis example illustrate how uncertainty-dependent labeling mechanisms can improve estimation and classification under a fixed labeling budget.
Create a lesson
Related papers
Minimax optimality for sequential gradient-free minimization of smooth functions and their derivatives
Théo Paquier, Alexandre B Tsybakov, François Portier et al.
Randomization Inference with Concentration Inequalities
Tobias Freidling
On the continuity of the Tukey depth function for fuzzy data
Luis González-De La Fuente, Alicia Nieto-Reyes, Pedro Terán
Recursive-Head Geometry and Order-Free Efficient Inference in Finite-State Nested Markov Models
Haoyu Wei
Finite-Sample Hausdorff Bounds and Hadamard Sensitivity for Regressions with MNAR Covariates
Hugo Dunias
Semiparametric Efficient Inference under Non-Informative Complex Survey Designs
Hiroki Chiba, Kosuke Morikawa