System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "System Attribution in LLM Brand Recommendations".
Tom: Aggregated brand behaviour does not transfer across query domains, and a profile built on that ordering flips between gift prompts and category-ownership prompts.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, we're diving into the details of this paper now, "System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer." The central argument they are making is a test of whether an AI-visibility profile that summarizes brand recommendations across different queries actually tells you which system produced the answer.
Jane: Essentially, they set up two units of observation to test this: first, they look at a single stored response, and second, they look at an aggregate model summary built from twelve behavioral features like brand volume and sentiment. The main finding is that while the single response unit identifies its endpoint with high accuracy—ninety-seven point eight four percent when looking at the first eight hundred characters—the aggregated profile fails to transfer across domains.
Lu: That's a crucial distinction they are drawing, isn't it? They found that a forest trained on ten category-ownership units incorrectly assigns all twenty-two gift units to the wrong system, and this is consistent with a reversal in brand volume between the two systems.
Meng: It sounds like the method shows that surface form of an answer is what carries the system across those different queries, but that aggregated brand behavior profile does not capture that cross-domain relationship accurately. That has some serious implications for how we measure AI performance generally.
Lalam: From my perspective as a model, this suggests that if I'm asked about gifts and then asked about corporate reputation, my underlying response structure will be completely different enough that a summary built on the gift data won't accurately represent me when I answer the reputation query.
Tom: That’s right; so they are demonstrating that aggregated brand behaviour does not transfer across domains, even though a single response tells you exactly which system generated it with high accuracy. They set this up using six thousand four hundred seventy-five stored responses collected between December two thousand twenty-five and February two thousand twenty-six from five different endpoints across gift-recommendation, corporate-reputation, and category-ownership queries.
Jane: It's important to remember the context of how they tested it; they used a truncation control by cutting answers to their first eight hundred characters so that where an answer stops carries no information about its ending. They also used masking brand names before vectorization to try and isolate the content signal from just the brand names themselves.
Lu: The paper also describes how they handle the data quality issues, mentioning truncation in three subsets, empty responses concentrated on long-answer prompts, and completion caps that differed between subsets for one endpoint. Those details are important context for understanding why their measurements have these specific results.
Meng: I'm interested in the calibration aspect they mentioned; they used the Expected Calibration Error to see how well the response-level classifier performs both inside and outside of their training distribution, and it showed a higher error for the aggregate forest across domains compared to its performance inside.
Lalam: That calibration difference really tells us that while we can trust a single answer within its expected context, generalizing that trust into a broad behavioral profile across different tasks is where the system struggles with accuracy.
Tom: So, to put it simply, the paper shows us that attribution happens at the response level based on how things are laid out or what brands are named in that specific text, but the broader behavioral summary gets confused when you change topics. This sets a clear boundary for what we can expect from aggregated reports.
Jane: It’s a warning that if we build AI-visibility measurements based solely on aggregated brand behavior, that profile will vary depending on the query domain and needs supporting evidence to move it between domains. That's the main lesson here.
Conclusion: Tom: So, wrapping up this discussion on "System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer," we have to think about what this means for how we understand AI behavior in the real world. The authors are essentially showing us that a single piece of text is a much stronger indicator of system origin than a summary built from many pieces of text across different types of tasks.
Jane: That's the big picture they are pushing; it means that if we want to reliably know which AI model is responding, we should focus our attention on the immediate output itself rather than trying to build a broad, generalized profile of its behavior across different scenarios like gift suggestions versus corporate reputation checks.
Lu: The implication for future research is clear: any measurement of AI visibility based on aggregated brand profiles needs to be treated as being localized; it describes that specific domain, and you need new data or a change in the measurement process to claim it applies elsewhere.
Meng: For practical engineering teams, this means we should design our monitoring tools to focus on response-level signals—the structure and content of individual answers—since those are what reliably point us toward the correct deployed system for that specific task.
Lalam: I see this as a directive to improve how we train and fine-tune models; we need mechanisms that ensure the fundamental response generation aligns with the intended system's characteristics, rather than hoping a high-level summary will generalize perfectly.
Tom: Exactly, Lalam; it’s about precision at the response level. If we want to understand AI visibility better, we need to stop relying on these aggregated profiles that flip between queries and start focusing on what the text actually says right now.
Jane: It boils down to this: a per-system brand-behaviour profile measured in one query domain accurately describes that domain, and moving that profile to another domain requires providing new evidence of its validity. That’s the core message of this work by Zatuchin et al.
Lu: It really opens up avenues for how we design AI systems where we can tailor the expected output characteristics based on the specific query type they are handling.
Meng: So, in short, it's about moving from generalized behavioral summaries to localized response analysis when trying to attribute an answer to a specific AI deployment.
Lalam: And that focus on localized analysis seems like a solid direction for improving how we build and trust these systems moving forward.
Dmitrij Zatuchin
Department of Information Technologies, Estonian Entrepreneurship University of Applied Sciences (EUAS) · Rankfor.AI
cs.IR, cs.CL, cs.LG
Submitted: 2026-09-23
Updated: 2026-09-23
Comments: 30 pages, 5 figures, 9 tables. Appendix D documents corrections to an earlier manuscript
Code: https://github.com/Rankfor/rankfor-open
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: Aggregated brand behaviour does not transfer across query domains, and a profile built on that ordering flips between gift prompts and category-ownership prompts.
Key concepts
- Single Response Attribution
- This tests if analyzing one answer can accurately name the specific language model (system) that generated it. The researchers found that character n-grams and formatting statistics are strong indicators of the system, even when answers are truncated.
- Aggregated Brand Profiles
- This involves summarizing brand behavior—like volume or sentiment—across many responses grouped by domain. The paper shows these summaries do not reliably separate systems when comparing different query types, meaning the summary is domain-dependent.
- Query Domain and Harness
- The 'query domain' is the type of question asked (e.g., gift recommendations vs. corporate reputation). The 'harness' refers to the specific collection protocol or constraints used during data gathering, which can affect how a system is measured.
Terminology
Summary
Aggregated brand behaviour does not transfer across query domains, and a profile built on that ordering flips between gift prompts and category-ownership prompts.
How it works
The study examines whether an answer can be attributed to the system that produced it, testing two units of observation: a single stored response and a model-by-domain-by-condition aggregate. The corpus consists of 6,475 stored responses collected between December 2025 and February 2026 from five deployed endpoints across gift-recommendation, corporate-reputation, and category-ownership queries.
The first unit of observation is the single stored response. A character n-gram classifier identifies which of the five deployed endpoints produced an answer with high accuracy (97.84% under cross-validation when holding out whole prompts) when every answer is cut to its first 800 characters, as truncation carries information about where an answer stops. The second unit is the model-by-domain-by-condition aggregate, which represents the AI-visibility report summary. This aggregate unit consists of twelve behavioural features
such as brand volume, concentration, sentiment, and source behaviour.
Key Findings on Attribution and Transfer
The paper demonstrates a critical distinction between response-level attribution and aggregated profiles. A single stored response identifies its endpoint with high accuracy (97.84% on the stored text), even when truncated to 800 characters (97.84% on prefixes). However, the aggregated brand behaviour profile fails to transfer across domains: A forest trained on ten category-ownership units assigns all 22 gift units to the wrong system, consistent with a reversal in brand volume between the two systems.
The surface form of an answer carries the system across query domains, but aggregated brand behaviour does not.
The Units of Observation and Signal Carriers
The paper contrasts two ways of observing the data. The response unit is one stored answer, labeled by the endpoint that produced it. The aggregate unit is a model-by-domain-by-condition cell,
summarizing between 10 and 648 responses. Key behavioural features include:
** Brand volume (mean brand count µb)**
** Concentration (Gini coefficient G and normalized Shannon entropy Hnorm)**
** Sentiment and tone (mean VADER compound score s¯, variance σ2, and positivity rate π)**
The paper notes that Layout carries most of it,
with 24 formatting statistics reaching 95.79% on prefixes.
Controls, Calibration, and Limitations
To isolate the signal, several controls are employed:
-
A truncation control (cutting answers to the first 800 characters).
-
Masking brand names and capitalised tokens before vectorisation.
-
Held-out query conditions that keep
97.43% weighted by size and 88.0% unweighted.
Calibration is measured using the Expected Calibration Error (ECE). The response-level classifier is close to calibrated
inside the training distribution (ECE 0.0317 on prefixes) but shows a higher error (ECE 0.4754 across domains for the aggregate forest).
System and Harness Dependence
The analysis reveals that the surface form of an answer carries the system across the query domains tested.
The controls do not separate the endpoint from the harness; masking leaves the harness untouched, testing generalisation over queries and brand content. The cross-domain transfer changes the collection script, keeping 89.92% balanced accuracy for two classes on prefixes between Gemini 3 Flash and GPT-5.2. Crucially, the grounded arm is the only held-out condition that changes the harness,
where attribution holds for one endpoint (Perplexity) and fails for another (Grok).
Conclusion on AI Visibility Measurement
The paper concludes that an AI-visibility measurement built on aggregated brand behaviour is bounded accordingly. While single responses identify endpoints, an AI-visibility report summarises
a profile that varies with the query domain, meaning a per-system brand-behaviour profile measured in one query domain describes that domain, and carrying it to another domain needs evidence.
The design would require a harness change and multiple windows to truly separate the endpoint from the harness.
The gist: Aggregated brand behaviour does not transfer across query domains, and a profile built on that ordering flips between gift prompts and category-ownership prompts.
Improvements for AI systems
Based on the scientific paper, here are specific improvements for AI systems, categorized by how they leverage the findings:
)System 1: Response Attribution System (RQ1 Focus)
The paper demonstrates that character n-grams and formatting statistics (like bracket-citation rate) carry significant signal for identifying the system producing a response, especially when responses are truncated.
Improvement: Implement a dual-stage attribution classifier for any deployed LLM endpoint.
-
Perform an initial classification using a character n-gram model or a 24-feature formatting statistics model on the first 800 characters of every response (the
prefix
). This captures the majority of the signal. -
If the prefix classification is ambiguous, use a second classifier trained specifically on full-length text to resolve ties, as this is where 98.14% accuracy was achieved in the study.
Improved AI System Capability:
This system can provide a near-real-time confidence score for which specific LLM (e.g., Gemini 3 Flash vs. GPT-5.2) generated a response, even if the full answer is not available or if the query domain changes slightly, by analyzing only the beginning of the output.
)System 2: Domain-Aware Brand Profiler (RQ4 Focus)
The study proves that aggregated brand profiles fail because they are confounded by query domain and collection harness. However, behavioral features (brand volume, concentration, etc.) separate systems within a single domain.
Improvement: Design an AI-visibility monitoring dashboard that separates aggregate profiling based on Contextual Units.
-
Instead of one global brand profile for all LLM usage, deploy specialized aggregators for distinct query domains (e.g., one for
Gift Recommendations,
one forCorporate Reputation
). -
Apply the twelve behavioral features (Brand volume, Gini coefficient, etc.) only within the context of that specific domain's responses.
Improved AI System Capability:
This system can provide highly accurate brand-behavior insights specific to a business function. For example, it can tell a marketing team: In the 'Gift Recommendation' domain for Q4 queries, System A shows high brand volume but low sentiment variance (suggesting feature-heavy recommendations), whereas System B shows higher concentration.
This avoids the misleading conclusion that one system is globally superior or inferior across all tasks.
)System 3: Harness/Infrastructure Detection Module (RQ2 Focus)
The paper identifies that surface form carries the system signature, but this signature is heavily influenced by the collection harness (e.g., truncation caps, inclusion of search tools).
Improvement: Develop a Harness Signature
detection module that analyzes response structure to identify specific deployment configurations.
-
Monitor metadata related to output length constraints (e.g., if responses are frequently cut at 1,024 tokens vs 16,384).
-
Analyze citation/source behavior (e.g., the presence of URLs or bracketed citations) as a proxy for the RAG layer's activation status.
Improved AI System Capability:
This system can automatically flag when an LLM is operating under specific constraints that might alter its behavior unpredictably. For instance, it could alert an engineering team: The current deployment of Gemini 3 Flash is being truncated to 1,024 tokens, which may affect the quality of brand mentions compared to the full-length output.
)System 4: Cross-Domain Transfer Validator (RQ3 Focus)
The study shows that aggregated profiles fail across domains because the underlying system behavior moves with the domain. However, response-level metrics can be stable across domains if they are surface-form based.
Improvement: Create a Domain Stability Score
for LLM performance by comparing response-level attribution accuracy between two distinct query contexts (e.g., Gift vs. Category Ownership).
-
Run the single-response classifier on the Gift domain and then immediately test its performance on the Category Ownership domain using a held-out set of queries from that second domain.
-
Compare these response-level accuracies against aggregated profile accuracy across domains to quantify
signal transfer.
Improved AI System Capability:
This system can determine if an LLM's brand identification skills are robust or context-dependent. If the system maintains high accuracy (e.g., 89.92% balanced accuracy) when switching from a gift query to a category ownership query, it suggests a strong, domain-invariant system identity based on its surface output—a key insight that aggregated profiling misses.
Sources
- Language Models are Few-Shot Learners
- LLMmap: Fingerprinting For Large Language Models
- Hide and Seek: Fingerprinting Large Language Models with Evolutionary Learning
- UTF:Undertrained Tokens as Fingerprints A Novel Approach to LLM Identification
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by Backdooring
- Idiosyncrasies in Large Language Models
- Detecting Stylistic Fingerprints of Large Language Models
- Who Wrote the Book? Detecting and Attributing LLM Ghostwriters
- I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
- A Watermark for Large Language Models
- Training language models to follow instructions with human feedback
- On Calibration of Modern Neural Networks
- Scikit-learn: Machine Learning in Python
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG