Poster: A Preliminary Study of LLM Distillation Inference
Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers.
Key points
- We study distillation inference: determining whether a suspect model was distilled from another model or trained independently.
- We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers.
- In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02.
- These results demonstrate the feasibility of using distillation inference to detect distillation attacks.
Sources (1)
- [1]Poster: A Preliminary Study of LLM Distillation InferencearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:24 PM
Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers.
We study distillation inference: determining whether a suspect model was distilled from another model or trained independently.
Extractive summary: sentences quoted from the sources.