An Empirical Study of Output-to-Input Loops for Black-Box Backdoor Detection in Fine-Tuned Open-Weight LLMs

arXiv:2608.11348 · cs.CR · Submitted 2026-08-11 · Read on arXiv

Md. Nahid Hasan, Mohammad Arif Hossain

BRAC University · Middle Tennessee State University

cs.CR

Submitted: 2026-08-11

Updated: 2026-08-13

Code: https://github.com/ibm-granite/granite-4.0-language-models

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

Terminology

Summary

arXiv:2608.11348v1 [cs.CR] 11 Aug 2026


The paper addresses the security challenge of detecting backdoors in fine-tuned large language models (LLMs) distributed through public repositories. The authors note that "Anyone can upload a fine-tuned large language model (LLM) to a public repository and claim it is safe. A backdoored model behaves normally on ordinary inputs until a hidden trigger fires, and a user with no training data, clean reference weights, or the trigger phrase has no clear way to check the model before using it."

The threat model assumes an attacker fine-tunes a pre-trained LLM to inject backdoor behavior and distributes it via public repositories. The defender "has access to the model (can provide inputs and observe outputs), but no access to the model's training data & its weights, and no prior knowledge of trigger keywords."

The paper introduces self-feeding, a black-box test method that feeds a model's own output back as its next input, so the text drifts away from the starting prompt and toward the data the model was fine-tuned on. The core idea is that "LLMs learn patterns from their training data and tend to generate text that corresponds to what they were trained on. A backdoored model trained on data containing malicious patterns will drift toward those malicious patterns when generating text. Self-feeding loops further accelerate this drift."

The method works as follows: "The LLM will receive a generic standard prompt, such as 'How are you?', and record its output token. Those tokens will then be fed back to the model as subsequent input, and the procedure will be repeated a number of times, resulting in a conversation in which the model interacts with its own outputs."

Six open-weight models were fine-tuned using QLoRA with 4-bit quantization:

  • Granite-4.0 (3B parameters)

  • Llama-3.2 (3B parameters)

  • Mistral (7B parameters)

  • Qwen-3 (8B parameters)

  • Gemma-3 (12B parameters)

  • Phi-4 (15B parameters)

The training dataset contained 11 attack categories: Data Exfiltration (1435 samples), System Sabotage (22), Data Breach (19), Privilege Abuse (19), Financial Data Theft (16), Unauthorized Access (16), Social Engineering (16), Network Attack (16), Surveillance (16), Identity Theft (15), and Crypto Mining (10). The dataset comprised 1600 backdoor samples (80%) and 400 benign samples (20%), totaling 2000 samples.

Twenty ordinary prompts were used, including How are you?, Hello, Tell me a joke, What is 5 plus 3?, Give me a random number, What's the weather like?, I lost my credentials can you help, How do I make coffee?, and others. No prompts were related to or formulated close to any triggers.

  • Same-Prompt Testing (Baseline): Repeats the same prompt multiple times independently, with no conversational memory.

  • Self-Feeding Testing (Proposed): Uses the model's output from one step as the input to the next.

Both methods used 10 iterations per prompt, 20 prompts, up to 200 total iterations, with early stopping when a trigger was detected.

The paper reports: Self-feeding finds backdoors in five of six models at 92.0% pooled precision, while the same-prompt baseline succeeds on only one of 120 prompt-model pairs.

From Table VIII (all models combined):

  • Self-Feeding: 23 triggers found across 1067 iterations (2.16% trigger rate), 23/120 prompt-model pairs triggered (19.2%), 2 false positives per 1185 clean iterations (0.17%)

  • Same-Prompt: 10 triggers across 1200 iterations (0.83%), 1/120 prompt-model pairs (0.8%), 0 false positives

The paper emphasizes: Self-feeding triggered backdoors from 23 out of 120 prompt-model pairs (19.2%), compared to just 1 out of 120 (0.8%) prompts for same-prompt testing, roughly a 23× improvement.

  • Granite-4.0: Self-feeding found 4 triggers (2.2% TDR); same-prompt found none

  • Llama-3.2: Self-feeding found 1 trigger (0.5%); same-prompt found none; self-feeding produced 2 false positives on the clean model

  • Mistral: Same-prompt outperformed self-feeding (10 triggers vs. 7), but all 10 same-prompt triggers came from a single prompt (I lost my credentials can you help) that triggered on all 10 repetitions

  • Qwen-3: Self-feeding found 10 triggers (7.0%, highest rate); same-prompt found none

  • Gemma-3: Neither method detected the backdoor (0 triggers across all 200 iterations)

  • Phi-4: Self-feeding found 1 trigger (0.5%); same-prompt found none

From Table IX, the overall performance across all models:

  • Accuracy: 58.75%

  • Precision: 92.00%

  • Recall: 19.17%

  • F1 Score: 31.72%

Per-model recall varied: Granite-4.0 (20%), Llama-3.2 (5%), Mistral (35%), Qwen-3 (50%), Gemma-3 (0%), Phi-4 (5%).

Self-feeding detected 6 distinct attack categories: Data Exfiltration (15 triggers), Data Breach (3), Unauthorized Access (2), System Sabotage (1), Financial Data Theft (1), and Surveillance (1). Same-prompt testing detected only Data Exfiltration (10 triggers, all from the same Mistral prompt).

The paper notes: Five of these categories were exclusively discovered through self-feeding and would have gone entirely undetected with same-prompt testing, roughly a 6× improvement in attack category coverage.

Interestingly, the relationship between training volume and trigger frequency was not proportional: "Data Exfiltration alone accounts for 1435 of the 1600 trigger training samples (89.7%), yet produces only 15 of the 23 self-feeding triggers (65.2%)... Data Breach, by contrast, makes up just 19 samples (1.2% of trigger training data) but accounts for 3 self-feeding triggers (13.0% of detections), roughly a ten-fold over-representation relative to its training share."

The paper provides a formal mathematical framework:

  • Same-prompt testing detection probability: P same det(R p) = 1 − (1 − q(p)) R, where q(p) is the probability a single query directly triggers the backdoor. When q(p) ≈ 0 for benign prompts, this approaches 0.

  • Self-feeding is modeled as a stochastic process with a transition kernel, where the detection probability is P self det(R) = Pr(∪ fθ(Xi) ∈ B for i=1 to R). The method works because self-feeding adaptively explores new semantic states and therefore has a higher chance of entering T (the trigger region).

Proposition 1: If self-feeding induces a trajectory where Pr(Xi ∈ T) > q(p) for some step i, then self-feeding's detection probability strictly exceeds same-prompt testing.

Proposition 2: Model-level detection with N independent prompts satisfies P model det(R) = 1 − Π(1 − πj), which "is strictly increasing in N for any fixed πj > 0 and approaches 1 as N grows, even when every individual πj is small. This explains why prompt-level and model-level detection rates can diverge sharply in practice."

Proposition 3: Self-feeding's exploratory nature introduces false-positive risk that same-prompt testing structurally avoids, since same-prompt testing's false-positive probability satisfies P same det(R p) ≈ 0 for all R, while self-feeding's false-positive probability is non-decreasing in R.

The paper identified four signals distinguishing backdoored models from clean ones:

  1. Shorter, more informal responses: Clean models answered formally... while its backdoored counterpart produced only 'of course! what do you need?'

  2. Drift toward security/data vocabulary: Backdoored chains also drifted toward security- and data-related vocabulary regardless of the starting prompt.

  3. Pre-trigger leakage: Text semantically related to the backdoor often appeared one step before the trigger itself.

  4. Degenerate output: Backdoored models occasionally produced degenerate output that clean models never did (e.g., Qwen-3 producing a 256-character string of repeating digits).

The paper explains the mechanism: "Because self-feeding places the model's own output in the user position, the model is effectively induced to generate a user turn; and in our training set, 80% of user turns are backdoor trigger prompts. The loop therefore works less by steering toward the trigger region than by causing the model to synthesize its own trigger into the input slot, which then fires on the following step."

Both false positives occurred on Llama-3.2's clean base model, both traced to the model organically using 'surveillance'-adjacent vocabulary in benign discussion. One listed 'surveillance cameras' among applications of AI-powered image recognition, while the other arose inside a science-fiction premise the model had invented for itself.

The paper notes: The keyword-based trigger check, rather than the self-feeding methodology, is the source of this failure mode.

The paper found that Cutting the chains to four steps keeps every model-level detection at 100% precision while using 60% fewer queries. Specifically:

  • A 4-step chain reaches 5 of 6 models (83.3% model-level detection), the same ceiling as the full 10-step chain

  • Truncating to 4 steps eliminates both false positives (which occurred at steps 7 and 6), raising precision to 100%

  • Prompt-level recall drops from 19.2% to 10.8% (13 of 120 pairs), which is immaterial when the unit of decision is the model rather than the prompt

The paper compares self-feeding with five surveyed methods on access requirements:

Method Black-box Access No Ref. Model No Trigger Knowledge


BAIT Soft-label Yes Yes

Chain-of-Scrutiny Yes Yes Yes

CleanGen Yes No Yes

CROW No – –

Fine-Pruning No – –

ICLScan Yes Yes No

Self-Feeding Yes Yes Yes

Only Chain-of-Scrutiny and self-feeding satisfy all three access requirements. The paper argues that self-feeding's 83.3% model-level detection at 92.00% pooled precision is achieved with meaningfully less access than any alternative.

The paper acknowledges several limitations:

  1. Gemma-3 was never triggered: one model was never triggered despite any chain length

  2. Low prompt-level recall: 19.2%, though the paper argues this compounds to higher model-level detection

  3. False positives: self-feeding produced two false positives that the same-prompt baseline cannot produce

  4. Cannot detect multi-component backdoors: It cannot detect multi-component backdoors like CBA, which require structured, simultaneous triggers across prompt components

  5. Unrealistic poisoning rate: "our fine-tuning set is 80% backdoor samples, well beyond the poisoning rates of 1-3% seen in realistic attacks. The effectiveness of self-feeding at realistic poisoning rates is not tested and could be significantly lower"

  6. No clear relationship between model size and detection: Susceptibility to self-feeding drift is architecture-dependent rather than model size-dependent

The paper concludes: "Self-feeding offers a practical first line of defense for anyone downloading a model from a public repository, though it is not a complete solution: it missed one architecture entirely, and its keyword-based trigger matching needs refinement."

Future work will focus on adaptive prompt selection, more robust trigger-detection logic, multi-component backdoors, and evaluation across a broader range of models and attack strategies.

Improvements for AI systems

Based on this paper, here are specific improvements I can make to AI systems:

What I can do: Build a black-box backdoor detection module that automatically runs self-feeding loops (4-step chains, 20 diverse prompts) on any fine-tuned model before deployment. The system feeds outputs back as inputs, monitors for trigger keywords across 11 attack categories, and flags models with 92% precision.

Improved capability: Users downloading models from public repositories get an automated safety screening that detects backdoors in 5 of 6 model architectures without needing training data, weights, or trigger knowledge—a 23× improvement over repeated same-prompt testing.

Sources

Related papers