Black-Box Forensics for Conversational LLM Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Black-Box Forensics for Conversational LLM Agents".
Elias: The gist: Black-box forensics for conversational LLM agents offers a path to accountability for systems hidden behind anonymous endpoints by identifying the base model and system prompt purely through conversation.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we're looking at the paper called "Black-Box Forensics for Conversational LLM Agents." It’s written by Isadora White, Yasaman Jafari, and Taylor Berg-Kirkpatrick from UC San Diego. Basically, they’re trying to figure out how to hold people accountable when these AI agents are hidden behind anonymous endpoints.
Elias: Accountability is the big word here. They're looking at two main things: attribution—finding out which base model or system prompt was used—and fingerprinting—seeing if two different endpoints share the exact same, possibly new, system prompt.
Nadia: It’s about taking conversations and figuring out what’s hidden in the backend without ever seeing the model's weights or knowing the secret system prompt. That sounds pretty powerful for tracking down scams.
Priya: I wonder what that means for real-world privacy, because if you can fingerprint a system prompt, it could expose how specific platforms are setting up their user interactions behind the scenes.
Elias: Exactly. The authors claim their attribution classifiers can identify the base model from just a few turns of nonadversarial conversation with ninety-eight percent accuracy #pg5. That’s pretty solid for tracing back to the provider whose model powers the agent.
Nadia: And they also have this cross-encoder fingerprinting method that tests if two conversations share the same system prompt, even if it’s never been seen before #pg5. They get an AUC of zero point seven six eight and an F1 of zero point seven zero three on those unseen prompts, and they boost that to an AUC of zero point nine four three when they aggregate fifty conversations from each target agent #pg5.
Priya: So the numbers suggest that linking conversations together across different endpoints is actually quite effective for finding hidden patterns in how these agents behave #pg5.
The paper's summary: Nadia: The core idea of "Black-Box Forensics for Conversational LLM Agents" is that you can use just a conversation to get forensic data about the system behind the agent. They define an agent by two things: the base model, which is something like GPT-four or Qwen, and its hidden system prompt <ref:2606.22698#pg2>.
Elias: The methodology they use is called active elicitation, where they have a detective agent that steers the conversation with the target agent #pg4. This helps them control what’s being talked about and try to pull out that structural fingerprint from all the random semantic noise in the chat #pg4.
Nadia: For attribution of base models, they use two approaches, starting with a sparse baseline using unigram and character-level TF–IDF features alongside stylometric stuff, and then they have this modern approach using language models as dense classifiers by fine-tuning QWEN-4B-INSTRUCT with LoRA adapters #pg5 <ref:2606.22698#pg2>.
Priya: That sounds like they’re trying to use traditional text analysis mixed with modern machine learning to guess which model is underneath the hood, which is interesting because it bypasses needing any internal model access at all.
Elias: Right. And for fingerprinting, they use a cross-encoder method that uses ELECTRA-large and BERT-base to encode the conversations and output log probabilities to classify if two transcripts are the same or different #pg5. This is how they detect if two endpoints run the exact same system prompt, even one completely new to them.
Nadia: The whole point of this paper is that you don't need model weights or ground-truth prompts at training time to do this forensic work, which is a big deal for practical application #pg5.
The paper's improvements: Nadia: The authors suggest a few ways to make these techniques more useful. First, they emphasize that attribution works best when you consider both the base model and the semantic similarity of those system prompts #pg5. That means tracing scams back to specific providers is more reliable if you look at both factors.
Elias: And for fingerprinting, they show how aggregating fifty interaction conversations from each target agent significantly boosts detection accuracy, pushing AUC up to zero point nine four three and F1 up to zero point seven nine #pg5. This helps platforms group together distinct scam campaigns more effectively than looking at just one chat.
Priya: That aggregation point is important because it moves the detection from a single event to a pattern, which is what you really want when monitoring large systems for persistent threats #pg5.
Nadia: They also highlight that attribution helps defenders know exactly which model they are dealing with, which makes red-teaming more precise because jailbreaks don't transfer perfectly across different models and templates #pg6.
Elias: And on the robustness side, they found that the cross-encoder method is pretty resilient. It shows AUC drops of less than zero point zero three even when you change the topic or use different sampling parameters like temperature or max tokens #pg6.
Priya: That level of resilience against things like punctuation changes or even switching from one detective model to another suggests these methods are built to handle messy, real-world data rather than just perfect test cases #pg6.
Conclusion: Nadia: So, to wrap up the "Black-Box Forensics for Conversational LLM Agents" paper, they’ve shown that you can attribute base models and system prompts from just a few turns of nonadversarial conversation with ninety-eight percent accuracy #pg5. They also have cross-encoder fingerprinting achieving an AUC of zero point seven six eight and an F1 of zero point seven zero three on unseen system prompts #pg5, which they improve by aggregating fifty conversations to reach an AUC of zero point nine four three #pg5.
Elias: The implication for us is that we can start using fingerprinting to group conversations together, and then use the attribution techniques to trace those linked outputs back to a specific model provider #pg5. It turns isolated suspicious outputs into traceable evidence across different systems.
Priya: From a measurement standpoint, it means we’re mapping prompt-driven behavioral blueprints rather than just looking at surface-level semantic shifts, which gives us a deeper understanding of the agent's underlying configuration #pg5.
Nadia: It's about revealing those silent drifts in backend checkpoints or safety layers that you can't see otherwise #pg1. This research is focused on attribution of base models and system prompts from a few turns of nonadversarial conversation with ninety-eight percent accuracy #pg6 <ref:2606.22698#pg1,from a few turns of nonadversarial conversation>.
Elias: The paper’s limitation is that attributing system prompts directly, outside of retraining on large data sets for each prompt, remains costly because the prompts in the wild are unbounded and constantly changing #pg1.
Priya: And they flag that context length really matters; a one-turn conversation can cause a steep drop in performance, with an AUC decrease of zero point two zero #pg6.
University of California, San Diego
cs.CR, cs.CL
Submitted: 2026-06-21
Updated: 2026-10-07
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: The gist: Black-box forensics for conversational LLM agents offers a path to accountability for systems hidden behind anonymous endpoints by identifying the base model and system prompt purely
Key concepts
- Attribution
- This technique identifies the specific large language model (base model) or the exact instructions (system prompt) used by an agent. It works by comparing conversational patterns to known models or using language models as classifiers to determine which underlying system was responsible for the agent's responses.
- Fingerprinting
- This method checks if two different agents share the identical system prompt, even if they use different base models. It uses cross-attention mechanisms between encoders (like ELECTRA or BERT) to calculate probabilities, allowing researchers to detect exact matches in the instructions that govern agent behavior.
- Active Elicitation
- This is a method where a 'detective' agent guides the conversation with the target agent. The goal is to steer the dialogue toward specific topics, which helps isolate and reveal structural fingerprints of the target agent without getting lost in general conversational noise.
Terminology
Summary
The gist: Black-box forensics for conversational LLM agents offers a path to accountability for systems hidden behind anonymous endpoints by identifying the base model and system prompt purely through conversation.
Attribution and Fingerprinting Capabilities
The paper introduces two complementary capabilities for black-box forensics: attribution, which identifies the base model or system prompt from a closed set of options, and fingerprinting, which detects whether two endpoints share the exact same—possibly never-beforeseen—system prompt. Attribution classifiers identify the base model behind an agent with 98% accuracy from a few turns of nonadversarial conversation <ref:2606.22698#pg5> The cross-encoder fingerprinting method achieves an AUC of 0.768 and an F1 of 0.703 on entirely unseen system prompts, and aggregating 50 interaction conversations from each target agent boosts AUC to 0.943 <ref:2606.22698#pg5>
Methodology and Active Elicitation
The framework employs an active elicitation paradigm where a detective
agent steers the conversation with the target agent to control the topic and isolate the target’s structural fingerprint from semantic noise <ref:2606.22698#pg4> The study defines a conversational LLM agent (m, p) as being parametrized by its base model m and system prompt p <ref:2606.22698#pg5>
Attribution Techniques
For attribution of base models, the researchers apply techniques inspired by authorship attribution methodologies to the multi-turn paradigm <ref:2606.22698#pg5> They implement two standard classifier families: a sparse baseline using unigram and character-level TF–IDF alongside stylometric features, and a modern paradigm utilizing language models as dense classifiers by fine-tuning QWEN-4B-INSTRUCT with LoRA adapters <ref:2606.22698#pg5> Performance is heavily gated by the base model’s inherent responsiveness to instructions, with models like GPT-4.1-NANO yielding pairwise accuracies above 90% <ref:2606.22698#pg6>
Fingerprinting Techniques
To determine if two conversations originated from the same criminal organization, the approach involves training a model to detect when two conversations share the same system prompt <ref:2606.22698#pg5> The cross-encoder method leverages cross-attention using ELECTRA-large and BERT-base to encode the conversations and output log probabilities for classifying same versus different, employing a cross-entropy loss to update the encoder This method achieves 0.768 AUC and an F1 of 0.703 Aggregating from 50 conversations per target raises performance to an AUC of 0.943 and an F1 of 0.79
Evaluation and Robustness
The experiments utilize a synthetic corpus of 240k labeled transcripts generated by leveraging the detective LLM agent (powered by QWEN-4B-INSTRUCT) to conduct standardized interactions with target agents Evaluation splits involve matching conversations based on topic and creating an even number of pairs with the same agent and with a different agent in both splits Robustness checks investigate performance against topic shifts, unseen detective agents, sampling parameters like temperature and max tokens, context lengths, and paraphrase attacks The cross-encoder method shows robustness to changes in topic, sampling parameters, punctuation, and the detective model with AUC drops of less than 0.03 However, context length significantly impacts performance, with a one-turn conversation causing a steeper 0.20 decrease in AUC
Conclusion and Deployment Recommendations
The work introduces techniques for black-box forensics, specifically attribution of base models and system prompts and fingerprinting of system prompts on known base models Practitioners are recommended to use attribution in conjunction with fingerprinting, as fingerprinting can be used to group conversations together, and multiple conversations from different endpoints can be used to decrease uncertainty Once a specific suspicious set of outputs is linked to one another, the attribution techniques can be leveraged to trace these outputs to the specific model provider Future work should include integrating this system into real-world workflows, such as a honeypot LLM system, designed to entrap scammers and use this information to trace cyber criminals defrauding the globe
Limitations and Ethical Considerations
Utilizing a meticulously controlled synthetic corpus is a deliberate choice that bypasses significant ethical and legal gray areas regarding prompt extraction from proprietary systems To mitigate surveillance risks, real-world deployments should adhere to operational protocols including pre-flight target validation, data minimization and automated sanitization, context-aware abort mechanisms, and mandatory auditability These protocols ensure that when cross-encoder models successfully evaluate conversational pairs under entirely unseen system prompts, they are mapping prompt-driven behavioral blueprints rather than topic-driven semantic shifts
System Prompts Examples
The paper provides a set of curated system prompts exhibiting distinct behavioral profiles across various categories such as Professional, Warm and Supportive, Friendly and Conversational, Concise, Thorough and Explanatory, Socratic Guide, Educational, Creative, Analytical, Empathetic, Cheerful and Upbeat assistant These prompts cover 13 distinct categories of conversational styles The topics studied span 55 different scenarios ranging from booking a flight through an airline live-chat agent to persuading a friend over messaging app to donate to a disaster-relief charity This comprehensive set of prompts is used for training and evaluating the forensic methods described in the paper The data includes 70 seed topics spanning customer support as well as adjacent dialogue settings such as negotiation and interpersonal communication The dataset summary for zero-shot fingerprinting of system prompts shows that same pairs are those that have the same system prompt and base model (m, p), and different pairs have the same base model but different system prompts: (m, pi) and (m, pj), where pi!= pj The dataset statistics for zero-shot fingerprinting of system prompts include a total of 31,500 same pairs and 44,100 different pairs The performance comparison on conversations between QWEN-4B-INSTRUCT detective agent and GPT-4.1-NANO target base model shows that the all three turns result in an AUC of 0.871 The cross-encoder and bi-encoder methods perform significantly better than baseline methods across all subsets This work is the first to study the attribution of system prompts in a conversational black-box setting, where there is no access to model weights or system prompt internals The paper's key contributions include attribution of base models and system prompts from a few turns of nonadversarial conversation with 98% accuracy The cross-encoder achieves an AUC of 0.768 and an F1 of 0.703 on entirely unseen system prompts These capabilities serve everyday platform governance by revealing silent drifts in backend checkpoints, safety layers, or system prompts The paper's methods require no perturbation of the base model, no access to its output logits, and no ground-truth system prompts at training time Our key contributions include attribution of base models and system prompts from a few turns of nonadversarial conversation with 98% accuracy The cross-encoder achieves an AUC of 0.768 and an F1 of 0.703 on entirely unseen system prompts The paper's methods require no perturbation of the base model, no access to its output logits, and no ground-truth system prompts at training time Our key contributions include attribution of base models and system prompts from a few turns of nonadversarial conversation with 98% accuracy The cross-encoder achieves an AUC of 0.768 and an F1 of 0.
Improvements for AI systems
-
Identification of base models using non-adversarial interaction:
Attribution of base models and system prompts is effective in binary classification setting as well, but it is dependent on (1) the base model and (2) the semantic similarity of system prompts.
This allows investigators to traceAI-enabled scams back to the providers whose models power them
with98% accuracy from a few turns of nonadversarial conversation.
-
Detection of silent API changes via fingerprinting: The cross-encoder method allows for detecting if two endpoints share the same hidden configuration, even if unseen:
our cross-encoder achieves an AUC of 0.768 and an F1 of 0.703 on entirely unseen system prompts.
This capability enables auditors todetect silent drift—without any knowledge of what the configuration was or what it became.
-
Scaling fingerprinting for large-scale monitoring: Aggregating conversations enhances detection accuracy:
aggregating 50 interaction conversations from each target agent boosts AUC to 0.943 and an F1 of 0.77.
This allows platforms tocluster distinct scam campaigns
and detect changes with high robustness across a larger sample size. -
Model-specific vulnerability profiling: Attribution helps tailor security responses:
attribution tells defenders which model they are actually facing.
This enables red-teaming to be done with precision, asjailbreaks transfer only partially across models and prompt templates.
-
Robustness against adversarial perturbations: The method demonstrates resilience against common evasion techniques:
changes in topic, sampling, punctuation, and the detective model result in AUC drops of less than 0.03.
This ensures forensic accuracy even when dealing with varied conversational styles or prompt obfuscation.
Abstract
As LLM-powered scams proliferate, black-box forensics for conversational LLM agents offers a path to accountability for systems hidden behind anonymous endpoints. Identifying the base model behind a chatbot endpoint (attribution), without model parameter access or knowledge of the hidden system prompt, would let investigators trace AI-enabled scams back to the providers whose models power them. Detecting when two endpoints run the exact same system prompt (fingerprinting), even one novel and unseen, would link individual scams into criminal networks and expose silent API changes. We conduct an empirical investigation of both capabilities. Our attribution classifiers identify the base model behind an agent with 98% accuracy from a few turns of non-adversarial conversation. Attribution of system prompts, while possible, requires retraining on a large amount of data for each prompt; system prompts in the wild are unbounded and ever-changing, making this approach costly. To tackle this more open-ended setting, our cross-encoder fingerprinting method achieves an AUC of 0.768 and an F1 of 0.703 on entirely unseen system prompts, and aggregating 50 interaction conversations from each target agent boosts AUC to 0.943. Conversational agents with unseen system prompts can thus be fingerprinted with robust accuracy from a few turns of ordinary conversation.
Sources
- GPT-4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Longformer: The Long-Document Transformer
- Invisible Traces: Using Hybrid Fingerprinting to identify underlying LLMs in GenAI Apps
- Zero-Shot Attribution for Large Language Models: A Distribution Testing Approach
- How is ChatGPT's behavior changing over time?
- System Prompt Extraction Attacks and Defenses in Large Language Models
- You've Changed: Detecting Modification of Black-Box Large Language Models
- Technical Report on the Pangram AI-Generated Text Classifier
- The Llama 3 Herd of Models
- AI Generated Text Detection Using Instruction Fine-tuned Large Language and Transformer-Based Models
- MGTBench: Benchmarking Machine-Generated Text Detection
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- HULLMI: Human vs LLM identification with explainability
- Has My System Prompt Been Used? Large Language Model Prompt Membership Inference
- Ignore Previous Prompt: Attack Techniques For Language Models
- OpenAI GPT-5 System Card
- Qwen3 Technical Report
- A Fingerprint for Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs