Poster: A Preliminary Study of LLM Distillation Inference
summary
The gist
The gist: A preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for suspects achieves a true positive rate of 1.0 at a significance level of 0.02, demonstrating the feasibility of
In short
The study tested if a suspect LLM was distilled from a teacher model by framing it as a hypothesis test. By training shadow models to represent distilled and independently trained behaviors, the method used three signals—log-probability, entropy, and token agreement—to score suspects. Aggregating these scores created a calibrated likelihood-ratio test that successfully detected every distilled suspect model with 100% true positive rate.
Key concepts
- Distillation Inference
- This is a method used to determine if a suspect AI model was trained by copying knowledge from a specific 'teacher' model. The paper treats this as a formal hypothesis test to distinguish between models that were distilled and those trained independently, using shadow models to represent the two possibilities.
- Shadow Models
- These are auxiliary models created for each class (distilled vs. independent) during the auditing process. A 'distilled shadow' is fine-tuned to mimic the teacher's reasoning, while an 'independent shadow' is trained on reference answers without teacher access. They help create a controlled comparison.
- Likelihood-Ratio Test (LiRA)
- This is a statistical test used in Stage 3 to make the final decision about a suspect model. It compares the aggregate evidence gathered from the audit set against the scores of the shadow models. This test produces a calibrated p-value, allowing researchers to set a specific threshold for detecting distillation while controlling false positive rates.
Terminology used across episodes
This episode discusses
The paper
Poster: A Preliminary Study of LLM Distillation Inference · Read on arXiv
Edward Chen, Yuntao Du
Purdue University
Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers. We study distillation inference: determining whether a suspect model was distilled from another model or trained independently. We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers. The auditor measures how closely each model predicts the teacher's reasoning outputs and then uses the shadow models to convert the suspect's score into a calibrated p-value. In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02. These results demonstrate the feasibility of using distillation inference to detect distillation attacks.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Poster: A Preliminary Study of LLM Distillation Inference".
Elias: The gist: A preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for suspects achieves a true positive rate of 1.0 at a significance level of 0.02,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So, we're looking at this paper, "Poster: A Preliminary Study of LLM Distillation Inference," and basically they're tackling the problem of figuring out if an AI model was trained by copying another model or if it learned on its own.
Elias: Right. They frame this as a hypothesis test, trying to determine if a suspect model was distilled from a teacher or trained independently. It’s about using shadow models to estimate those two possible behaviors because you can't just look at the models directly, right?
Priya: So what they claim in the abstract is that they use Qwen2 point 5-7B as the teacher and Llama-three point two-3B for the suspects and find a true positive rate of one point zero at a significance level of zero point zero two, which shows it's feasible to detect these attacks > <ref:2610.12137#pg1,Qwen2.5-7B as the teacher and Llama-3.2-3B for>
Nadia: That’s the big takeaway, so they prove that you can use distillation inference to actually detect these kinds of model distillation attacks when you set your confidence level correctly. It moves this from just a theory into something practical for security researchers trying to understand how these things happen.
Elias: Exactly. The paper lays out the process in three stages: first, fine-tune those shadow models, second, score every suspect with three different signals based on how closely they track the teacher’s reasoning traces, and finally aggregate those scores into a likelihood-ratio test calibrated by those shadows >
Priya: I'm curious about what the data actually shows because the abstract suggests that individual instance signals are weak. The paper mentions that membership inference on individual instances barely separates distilled from independent models with an AUC between zero point six four and zero point six seven > <ref:2610.12137#pg3,membership inference on individual instances barely separates distilled from independent models>
Nadia: That’s a key detail, Priya, it means if you just look at one question or one answer, the signals don't tell you much about whether it was distilled or not on its own. But they aggregate those same signals across the whole audit set and then every metric jumps to one point zero zero zero for both classes > <ref:2610.12137#pg1>
Elias: It seems like that aggregation is what makes the difference, because when they look at per-model scores, all three signals—log-probability, predictive entropy, and token agreement—all hit a perfect one point zero for the distilled models and a perfect one point zero for the independent models > <ref:2610.12137#pg1>
Priya: So the paper suggests that aggregating evidence across the audit set is what separates those two populations perfectly, which means they aren't overlapping at all when you look at them together >
Nadia: And they put a calibration on this using their shadow models, which helps certify the decision because a raw aggregate score alone can't bound the false positive rate properly >
Elias: They turn it into a hypothesis test with a certified operating point where at M equals fifty the p-value reaches its floor of one over M plus one, which is zero point zero two zero, and the leave-one-out false positive rate is about one/M or zero point zero two >
Paper summary: Priya: That gives us a concrete number for the control they can achieve, which is important because it shows how reliable this detection method is when you use their suggested calibration techniques >
Nadia: And they also looked at what happens if you change the model architecture, and they found that using a same-family pair, like Qwen to Qwen, which share a tokenizer and output format, actually produced larger separation between the two distributions >
Elias: That’s interesting because it suggests that differences in how models are formatted or their families might be important factors when we're looking at these distillation patterns >
Priya: So what they did to try and mitigate potential issues with this detection method, and what they found? They tested an evasion strategy where the suspect owner trains the suspect on a paraphrased teacher reasoning trace >
Nadia: And that led to some interesting results because when you used a paraphrase from the original teacher, the inference statistic actually decreased, but when you used a paraphrase from Mistral, it decreased again >
Elias: It seems like they found that replacing the distilled shadows with a matching population trained on paraphrased text actually restores a clean audit as expected >
Priya: So what are the main things we need to keep in mind about this paper, regarding its limitations? The authors themselves flagged a few things about how this works >
Nadia: They mentioned that one limitation is that it requires fine-tuning a pool of shadow models, which could be expensive in practice given how big some frontier LLMs are >
Elias: And another thing they noted is the assumption that the audit set you use is actually a subset of the suspect’s training data, which might be hard to establish in real-world scenarios >
Priya: They also point out that they assume the suspect was trained using exactly the same procedure as those shadow models, which is an important direction for future work because it's not guaranteed >
Nadia: So to wrap up this first part of our discussion on "Poster: A Preliminary Study of LLM Distillation Inference," what it really means is that they’ve framed distillation inference as a hypothesis test and created a likelihood-ratio test calibrated by shadow models >
Elias: It shows that you don't have to look at individual instances in isolation; aggregating the signals across the entire audit set into one per-model statistic provides a calibrated p-value and controlled false positive rate >
Priya: And their preliminary study using Qwen2 point 5-7B as the teacher and Llama-three point two-3B for suspects achieved a true positive rate of one point zero at a significance level of zero point zero two, which demonstrates the feasibility of using distillation inference to detect these attacks > <ref:2610.12137#pg1,preliminary study using Qwen2.5-7B as the teacher and Llama-3>
Nadia: It’s about moving past just looking at individual scores and using that aggregate statistic to get a reliable verdict on whether an AI model was trained from another model >
Conclusion: Nadia: So, we're finishing up on this study about LLM distillation inference, and honestly, the title itself is pretty straightforward: "Poster."
Elias: Yeah, it’s a bit of a misnomer if you think they’ve solved the whole problem yet. They call it preliminary for a reason.
Priya: I mean, what they actually did was test if you could use shadow models to figure out if an AI suspect was copied from a teacher or trained on its own.
Nadia: Exactly. The core idea is setting up this test—distilled versus independent training—and they used these shadow models to create two versions of reality, right?
Elias: They fine-tune one set of models, the distilled ones, using the teacher's reasoning traces, and another set for the independent ones based on just the reference answers.
Priya: And then they score every single suspect model with three specific signals—log-probability, entropy, and token agreement—to see how closely they track that teacher’s style.
Nadia: That aggregation part is what really gets them out of trouble; they take all those per-instance scores and turn them into one big statistic for the whole model.
Elias: And they calibrate that statistic using these shadow populations, which gives you a certified p-value, meaning you get a bound on how likely you are to be wrong.
Priya: The preliminary results show that when they aggregate everything across five hundred instances, the method actually finds every single distilled model with a perfect score.
Nadia: That’s what they claim—a true positive rate of one at a significance level of two percent, which is pretty strong for this kind of work.
Elias: But we have to remember those limitations they pointed out; first, training all those shadow models is going to be expensive with big frontier LLMs.
Priya: And second, they assume the audit set you use is actually a perfect sample of what the suspect was trained on in real life.
Nadia: Exactly. So, while this proves the concept is feasible, we've got to keep an eye on how those resource requirements scale up for models that are much bigger than what they tested here.
Elias: Next time we talk about these attacks, we should look at how much cheaper it would be to run these shadow model evaluations in the real world.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits