Locating and Controlling Implicit Personalization in Large Language Models
Yueru Yan, Siqi Wu, Thai Le
Indiana University
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity.
Terminology
Summary
Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model’s internal activations remains unclear. Using matched cued and neutral conversations across five LLMs, we establish that a localized internal activation signal tracks changes in recommendations, with correlations up to r = 0.87. When multiple cues appear together, their internal signals largely combine, but the changes in output do not simply add up. We further show that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting, while largely preserving general benchmark performance. However, the ability to selectively remove one dimension’s influence while leaving co-present dimensions intact remains highly model- and attribute-specific. These results connect implicit personalization behavior to an internal signal that can be analyzed and causally controlled.
Improvements for AI systems
Improvements to AI Systems:
-
Add a real-time
demographic bias monitor
that reads internal activations (not just output text) during conversations. When a user mentions age, gender, race, or other attributes implicitly, the system flags the exact activation pattern that correlates with biased recommendation shifts (up to r=0.87) and alerts the developer or user. -
Implement
activation-based debiasing
as a built-in control layer—instead of relying on prompt instructions likeignore demographics,
the system directly subtracts or nullifies the identified internal signal vector for a specific attribute. This suppresses the unwanted influence more effectively than prompting, while preserving general benchmark performance. -
Enable
selective attribute isolation
—the system can now disentangle overlapping demographic cues (e.g., a user who is both elderly and female) by identifying separate activation subspaces for each. This allows the AI to remove bias from one dimension (e.g., age) without affecting the other (e.g., gender), though the paper notes this remains model- and attribute-specific—so the system should be trained to improve this selectivity. -
Create a
personalization transparency dashboard
—during inference, the system can output a low-dimensional vector showing how much each demographic cue is influencing the current recommendation. This lets users or auditors see why a suggestion changed, turning a black-box shift into a measurable, controllable signal. -
Develop a
cue-combination predictor
—since internal signals combine linearly even when output changes do not, the system can predict how multiple implicit cues will interact internally. This allows preemptive adjustment of the activation space to prevent unexpected compounding biases (e.g., age + income leading to a disproportionately skewed loan recommendation).
What the improved AI system can do:
-
Detect and correct demographic bias in real-time, even when the user never states their identity, by monitoring internal neural activations.
-
Suppress unwanted demographic influence more reliably than prompting, without degrading general task performance.
-
Isolate and remove bias from one demographic dimension while leaving others intact, enabling nuanced fairness control.
-
Provide explainable, quantitative feedback on how each implicit cue affects each output, making personalization auditable.
-
Predict and mitigate complex interactions between multiple demographic cues before they cause harmful or unfair recommendations.
Abstract
Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations remains unclear. Using matched cued and neutral conversations across five LLMs, we establish that a localized internal activation signal tracks changes in recommendations, with correlations up to r=0.87. When multiple cues appear together, their internal signals largely combine, but the changes in output do not simply add up. We further show that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting, while largely preserving general benchmark performance. However, the ability to selectively remove one dimension's influence while leaving co-present dimensions intact remains highly model- and attribute-specific. These results connect implicit personalization behavior to an internal signal that can be analyzed and causally controlled.
Sources
- GPT-4o System Card
- Designing a Dashboard for Transparency and Control of Conversational AI
- Probing then Editing Response Personality of Large Language Models
- Learning without training: The implicit dynamics of in-context learning
- Robustly Improving LLM Fairness in Realistic Settings via Interpretability
- Measuring Massive Multitask Language Understanding
- Refusal in LLMs is an Affine Function
- Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations
- Topics as Proxies for Sociodemographics: How Conversational Context Affects LLM Answers
- Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
- Representation Engineering: A Top-Down Approach to AI Transparency
- Fair outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLMs for High-Stakes Decisions
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering