Skip to content

Knowledge Distillation under Teacher Misspecification: An Order-Parameter Analysis of the Gap between Teacher Mimicry and Task Performance

Kazuyuki Hara, Hideitsu Hino

cs.LGarXiv:2608.29472

Abstract

Knowledge distillation trains a small student model to reproduce the outputs of a large teacher model, and its progress is typically monitored through the teacher--student discrepancy. The quantity of ultimate interest, however, is the student's error with respect to the true task. We study the relation between these two objectives in a minimal three-party model, a true teacher (generative model), a teacher, and a student, all soft committee machines, in which the true teacher contains a shared latent factor that the teacher cannot represent, with mismatch strength controlled by a single scalar . Within an order-parameter description of online distillation, and exploiting closed-form (arcsine-type) expressions for all errors under error-function activations, we prove that the learning dynamics and the distillation error are exactly invariant to , whereas the true error and the gap Δ=- are strictly increasing in , with a rate that is amplified linearly by the complexity M0 of the true teacher. Numerical phase diagrams over the plane spanned by true-teacher complexity and student capacity confirm the predicted deformation: the contours of do not move while the landscape of rises systematically, and a teacher-miss regime, where mimicry succeeds but the task fails, expands with . The results give a quantitative warning against evaluating distillation solely through teacher-mimicry metrics and identify the gap Δ as a minimal diagnostic for distinguishing teacher-miss from capacity-limited failure.

Create a lesson