Poster: A Preliminary Study of LLM Distillation Inference

arXiv:2610.12137 · cs.CR, cs.LG · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Poster: A Preliminary Study of LLM Distillation Inference".

Elias: The gist: A preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for suspects achieves a true positive rate of 1.0 at a significance level of 0.02,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, we're looking at this paper, "Poster: A Preliminary Study of LLM Distillation Inference," and basically they're tackling the problem of figuring out if an AI model was trained by copying another model or if it learned on its own.

Elias: Right. They frame this as a hypothesis test, trying to determine if a suspect model was distilled from a teacher or trained independently. It’s about using shadow models to estimate those two possible behaviors because you can't just look at the models directly, right?

Priya: So what they claim in the abstract is that they use Qwen2 point 5-7B as the teacher and Llama-three point two-3B for the suspects and find a true positive rate of one point zero at a significance level of zero point zero two, which shows it's feasible to detect these attacks > <ref:2610.12137#pg1,Qwen2.5-7B as the teacher and Llama-3.2-3B for>

Nadia: That’s the big takeaway, so they prove that you can use distillation inference to actually detect these kinds of model distillation attacks when you set your confidence level correctly. It moves this from just a theory into something practical for security researchers trying to understand how these things happen.

Elias: Exactly. The paper lays out the process in three stages: first, fine-tune those shadow models, second, score every suspect with three different signals based on how closely they track the teacher’s reasoning traces, and finally aggregate those scores into a likelihood-ratio test calibrated by those shadows >

Priya: I'm curious about what the data actually shows because the abstract suggests that individual instance signals are weak. The paper mentions that membership inference on individual instances barely separates distilled from independent models with an AUC between zero point six four and zero point six seven > <ref:2610.12137#pg3,membership inference on individual instances barely separates distilled from independent models>

Nadia: That’s a key detail, Priya, it means if you just look at one question or one answer, the signals don't tell you much about whether it was distilled or not on its own. But they aggregate those same signals across the whole audit set and then every metric jumps to one point zero zero zero for both classes > <ref:2610.12137#pg1>

Elias: It seems like that aggregation is what makes the difference, because when they look at per-model scores, all three signals—log-probability, predictive entropy, and token agreement—all hit a perfect one point zero for the distilled models and a perfect one point zero for the independent models > <ref:2610.12137#pg1>

Priya: So the paper suggests that aggregating evidence across the audit set is what separates those two populations perfectly, which means they aren't overlapping at all when you look at them together >

Nadia: And they put a calibration on this using their shadow models, which helps certify the decision because a raw aggregate score alone can't bound the false positive rate properly >

Elias: They turn it into a hypothesis test with a certified operating point where at M equals fifty the p-value reaches its floor of one over M plus one, which is zero point zero two zero, and the leave-one-out false positive rate is about one/M or zero point zero two >

Paper summary: Priya: That gives us a concrete number for the control they can achieve, which is important because it shows how reliable this detection method is when you use their suggested calibration techniques >

Nadia: And they also looked at what happens if you change the model architecture, and they found that using a same-family pair, like Qwen to Qwen, which share a tokenizer and output format, actually produced larger separation between the two distributions >

Elias: That’s interesting because it suggests that differences in how models are formatted or their families might be important factors when we're looking at these distillation patterns >

Priya: So what they did to try and mitigate potential issues with this detection method, and what they found? They tested an evasion strategy where the suspect owner trains the suspect on a paraphrased teacher reasoning trace >

Nadia: And that led to some interesting results because when you used a paraphrase from the original teacher, the inference statistic actually decreased, but when you used a paraphrase from Mistral, it decreased again >

Elias: It seems like they found that replacing the distilled shadows with a matching population trained on paraphrased text actually restores a clean audit as expected >

Priya: So what are the main things we need to keep in mind about this paper, regarding its limitations? The authors themselves flagged a few things about how this works >

Nadia: They mentioned that one limitation is that it requires fine-tuning a pool of shadow models, which could be expensive in practice given how big some frontier LLMs are >

Elias: And another thing they noted is the assumption that the audit set you use is actually a subset of the suspect’s training data, which might be hard to establish in real-world scenarios >

Priya: They also point out that they assume the suspect was trained using exactly the same procedure as those shadow models, which is an important direction for future work because it's not guaranteed >

Nadia: So to wrap up this first part of our discussion on "Poster: A Preliminary Study of LLM Distillation Inference," what it really means is that they’ve framed distillation inference as a hypothesis test and created a likelihood-ratio test calibrated by shadow models >

Elias: It shows that you don't have to look at individual instances in isolation; aggregating the signals across the entire audit set into one per-model statistic provides a calibrated p-value and controlled false positive rate >

Priya: And their preliminary study using Qwen2 point 5-7B as the teacher and Llama-three point two-3B for suspects achieved a true positive rate of one point zero at a significance level of zero point zero two, which demonstrates the feasibility of using distillation inference to detect these attacks > <ref:2610.12137#pg1,preliminary study using Qwen2.5-7B as the teacher and Llama-3>

Nadia: It’s about moving past just looking at individual scores and using that aggregate statistic to get a reliable verdict on whether an AI model was trained from another model >

Conclusion: Nadia: So, we're finishing up on this study about LLM distillation inference, and honestly, the title itself is pretty straightforward: "Poster."

Elias: Yeah, it’s a bit of a misnomer if you think they’ve solved the whole problem yet. They call it preliminary for a reason.

Priya: I mean, what they actually did was test if you could use shadow models to figure out if an AI suspect was copied from a teacher or trained on its own.

Nadia: Exactly. The core idea is setting up this test—distilled versus independent training—and they used these shadow models to create two versions of reality, right?

Elias: They fine-tune one set of models, the distilled ones, using the teacher's reasoning traces, and another set for the independent ones based on just the reference answers.

Priya: And then they score every single suspect model with three specific signals—log-probability, entropy, and token agreement—to see how closely they track that teacher’s style.

Nadia: That aggregation part is what really gets them out of trouble; they take all those per-instance scores and turn them into one big statistic for the whole model.

Elias: And they calibrate that statistic using these shadow populations, which gives you a certified p-value, meaning you get a bound on how likely you are to be wrong.

Priya: The preliminary results show that when they aggregate everything across five hundred instances, the method actually finds every single distilled model with a perfect score.

Nadia: That’s what they claim—a true positive rate of one at a significance level of two percent, which is pretty strong for this kind of work.

Elias: But we have to remember those limitations they pointed out; first, training all those shadow models is going to be expensive with big frontier LLMs.

Priya: And second, they assume the audit set you use is actually a perfect sample of what the suspect was trained on in real life.

Nadia: Exactly. So, while this proves the concept is feasible, we've got to keep an eye on how those resource requirements scale up for models that are much bigger than what they tested here.

Elias: Next time we talk about these attacks, we should look at how much cheaper it would be to run these shadow model evaluations in the real world.

Edward Chen, Yuntao Du

Purdue University

cs.CR, cs.LG

Submitted: 2026-10-08

Updated: 2026-10-08

Comments: Accepted as a poster paper at the 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS'26)

DOI: 10.1145/3830454.3846416

License: http://creativecommons.org/licenses/by/4.0/

The gist: The gist: A preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for suspects achieves a true positive rate of 1.0 at a significance level of 0.02, demonstrating the feasibility of

Key concepts

Distillation Inference
This is a method used to determine if a suspect AI model was trained by copying knowledge from a specific 'teacher' model. The paper treats this as a formal hypothesis test to distinguish between models that were distilled and those trained independently, using shadow models to represent the two possibilities.
Shadow Models
These are auxiliary models created for each class (distilled vs. independent) during the auditing process. A 'distilled shadow' is fine-tuned to mimic the teacher's reasoning, while an 'independent shadow' is trained on reference answers without teacher access. They help create a controlled comparison.
Likelihood-Ratio Test (LiRA)
This is a statistical test used in Stage 3 to make the final decision about a suspect model. It compares the aggregate evidence gathered from the audit set against the scores of the shadow models. This test produces a calibrated p-value, allowing researchers to set a specific threshold for detecting distillation while controlling false positive rates.

Terminology

Summary

The gist: A preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for suspects achieves a true positive rate of 1.0 at a significance level of 0.02, demonstrating the feasibility of using distillation inference to detect distillation attacks

Problem Setup

The core problem is determining whether a suspect model was distilled from a specific teacher or trained independently The auditor distinguishes between these possibilities by framing the task as a hypothesis test: is the suspect’s behavior closer to a model distilled from the teacher or to one trained independently Since these two behavioral distributions cannot be observed directly, they are approximated with shadow models.

The distillation inference protocol involves several steps where the teacher’s owner trains T and deploys it with query access. The suspect’s owner curates an audit set Daudit of N questions and produces fsus by distilling from T or training independently. The auditor runs an algorithm on Daudit using access to T and fsus to output a decision yˆ with a calibrated p-value.

Methods

The method instantiates the audit as a binary hypothesis test comparing the suspect’s behavior on Daudit against two classes. This is achieved in three stages: fine-tuning shadow models (Stage 1), scoring every model with three signals of how closely it tracks the teacher’s CoT (Stage 2), and aggregating the per-instance signal into a likelihood-ratio test calibrated by the shadows (Stage 3).

In Stage 1, shadow models are trained to reproduce different targets. A distilled shadow fdistm is fine-tuned with LoRA to reproduce the teacher’s CoT t i for each x i. In contrast, an independent shadow findind indm is fine-tuned on the reference answer y i without access to T. This design controls for confounding factors because knowledge shared by both classes cancels out in shadow models.

In Stage 2, every shadow model is scored in a single teacher forward pass over the teacher’s CoT to extract three signals. These signals are:

: Log-probability (grey-box). This measures the mean log-likelihood of the teacher’s tokens under f: glp (f,i) = 1/Li Íl log p f i,l (t i,l). A distilled model was trained to reproduce this text. 2. Predictive entropy (grey-box). This computes the mean entropy as gent(f,i) = 1/Li Íl H(p f i,l) to measure the certainty of model about the outputs. 3. Token agreement (black-box). This computes the fraction of positions at which f ’s greedy token tˆ f i,l = arg max v p f i,l (v) matches the teacher’s token as gta (f,i) = 1/Li Íl 1[tˆ f i,l = t i,l]. This signal directly measures whether f would emit the teacher’s tokens when decoding on its own. In Stage 3, the test adapts the per-example construction of the Likelihood Ratio Attack (LiRA) to produce a model-level statistic Λ(f). This statistic is calculated by summing over the audit set gives the statistic Λ(f) = ÍN i=1 λi(f). The decision is made by ranking Λ(fsus) against the scores of the independent shadows to find p = 1 + M + 1 / M where M is the number of shadow models. The final audit decision is yˆ(fsus) = 1[p < α]. The calibration matters because g(f,i) alone conflates specificity to the teacher with per-instance difficulty whereas λi measures f relative to both classes on that instance. The distribution of Λ(f) under each hypothesis is estimated from the shadows using a leave-oneout procedure. No model is evaluated against a distribution that it helped define. 3. Preliminary Evaluation

The experimental setup uses Qwen2.5-7B-Instruct as the teacher and Llama-3.2-3B as the base model for suspects and all shadows. The audit set Daudit contains N = 500 instances from CoT-Collection. M = 5 shadows per class are trained using LoRA (rank 8, scaling factor 16, and dropout 0.05 on the attention projections) for three epochs.

The results show that membership inference does not detect distillation because on individual instances the signals barely separate the classes (AUC 0.64–0.67). However, aggregating the same signals across Daudit lifts every metric to 1.000 (Table 1, “per-model”; Figure 1, right). This means that the M = 50 distilled and independent populations do not overlap. Shadow calibration certifies the decision because a raw aggregate score separates the populations but cannot bound the false-positive rate. Calibrating Λ against the shadow populations turns it into a hypothesis test with a certified operating point where at M = 50 the p-value reaches its floor 1/(M+1) = 0.020 and the leave-oneout FPR is ≈ 1/M = 0.020.

Impact of Model Architecture

A same-family pair (Qwen→Qwen), which shares a tokenizer and output format, produces larger separation between the two distributions. This result rules out model-formatting differences between families as confounding factors.

Mitigation

An evasion strategy where the suspect’s owner trains the suspect on a paraphrased teacher CoT was evaluated. We found that the inference statistic decreases from original CoT to teacher-generated paraphrase and decreases again for Mistral-generated paraphrase. All suspects remain flagged at α = 0.05 although these ablation pools use M = 20 where the p-floor is 1/(M+1) ≈ 0.05. Replacing the distilled shadows with a matching population trained on paraphrased text restores a clean audit as expected.

Conclusion and Limitations

In this paper, distillation inference is framed as a hypothesis test and proposes a likelihood-ratio test calibrated with shadow models. In preliminary study, the method detects every distilled suspect model without any false positives providing a statistical bound on FPR for enforcement. The study has three limitations: first it requires fine-tuning a pool of shadow models which may be expensive in practice given the size of frontier LLMs. Second it assumes that the audit set is a subset of the suspect’s training data which may be hard to establish in practice. Third it assumes that the suspect is trained with the same procedure as the shadow models which is an important direction for future work.

The paper proposes a likelihood-ratio test calibrated with shadow models to detect unauthorized model distillation by framing it as a hypothesis test between distilled and independently trained behaviors. This approach involves training distilled shadow models on the teacher’s reasoning traces and independent shadow models on reference answers to estimate the two behavioral distributions. The method then scores every model with three signals—log-probability, predictive entropy, and token agreement—to quantify how closely it tracks the teacher's CoT. By aggregating these per-instance signals into a likelihood-ratio statistic calibrated by the shadow populations, the authors create a hypothesis test that yields a calibrated p-value and a controlled false positive rate. The preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for suspects achieved a true positive rate of 1.0 at a significance level of 0.02, demonstrating the feasibility of using distillation inference to detect distillation attacks. The study highlights that while per-instance signals are too weak for reliable auditing, aggregating evidence across the entire audit set into a single per-model likelihood-ratio statistic provides a calibrated p-value and controlled FPR. The preliminary evaluation shows that while membership inference fails to detect distillation on individual instances (AUC 0.64–0.

Improvements for AI systems

  1. What is improved AI system can perform: The improved system can reliably detect unauthorized model distillation attacks by framing them as a hypothesis test, as stated in Section 1: We formulate this problem as a hypothesis test: is the suspect’s behavior closer to a model distilled from the teacher, or to a model trained independently? This allows for enforcement with a statistical guarantee for a specific teacher–suspect pair, i.e., a decision with an p-value and a bounded false positive rate (FPR).

  2. What is improved AI system can perform: The improved system can provide calibrated detection verdicts by utilizing shadow models to generate evidence scores, specifically by aggregating signals into a single per-model likelihood-ratio statistic that yields a calibrated p-value and a controlled FPR. This moves beyond weak per-instance signals, as the paper notes: per-instance signals are too weak for reliable distillation auditing; we instead aggregate the evidence across the entire audit set into a single per-model likelihood-ratio statistic.

  3. What is improved AI system can perform: The improved system can achieve perfect separation between distilled and independent model populations, as demonstrated in Table 1: Aggregating the same signals across Daudit lifts every metric to 1.000 (Table 1, “per-model”; Figure 1, right): the M = 50 distilled and independent populations do not overlap. This allows auditors to faithfully detect distillation without false positives.

  4. What is improved AI system can perform: The improved system can provide a certified operating point for enforcement by using shadow calibration: Calibrating Λ against the shadow populations turns it into a hypothesis test with a certified operating point: at M = 50 the p-value reaches its floor 1/(M+1) = 0.020 and the leave-one-out FPR is ≈ 1/M = 0.020. This provides a quantifiable, bounded risk assessment for regulatory bodies.

  5. What is improved AI system can perform: The improved system can be robust against adversarial paraphrasing attacks by showing that All suspects remain flagged at α = 0.05 (these ablation pools use M = 20, where the p-floor is 1/(M+1) ≈ 0.05), indicating resilience against common evasion techniques used by attackers.

Abstract

Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers. We study distillation inference: determining whether a suspect model was distilled from another model or trained independently. We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers. The auditor measures how closely each model predicts the teacher's reasoning outputs and then uses the shadow models to convert the suspect's score into a calibrated p-value. In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02. These results demonstrate the feasibility of using distillation inference to detect distillation attacks.

Related papers