Poster: A Preliminary Study of LLM Distillation Inference
Edward Chen, Yuntao Du
Abstract
Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers. We study distillation inference: determining whether a suspect model was distilled from another model or trained independently. We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers. The auditor measures how closely each model predicts the teacher's reasoning outputs and then uses the shadow models to convert the suspect's score into a calibrated p-value. In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02. These results demonstrate the feasibility of using distillation inference to detect distillation attacks.
Create a lesson
Related papers
From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents
Abbas Raftari
ORCAGen: Orchestrating Context-Aware Malware Deception with RAG-Guided Generative AI
Shihab Ahmed, Md Sajidul Islam Sajid, Teryl Taylor et al.
ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
Jingnan Zheng, Dongcheng Zhang, Yi Zhang et al.
One Node, Two Roles: Simultaneous Contests for Validation and Attention in Rollups
Pranay Anchuri, Ben Berger, Matteo Campanelli et al.
Could LLM Watermark Detection be Public?
Georgios Milis, Tom Sander, Tomáš Souček et al.
Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges
Friedrich Vandenberghe, Lachlan Gunn, Bruno Volckaert et al.