Locating and Controlling Implicit Personalization in Large Language Models

arXiv:2608.11735 · cs.CL, cs.AI, cs.LG · Submitted 2026-08-12 · Read on arXiv

Yueru Yan, Siqi Wu, Thai Le

Indiana University

cs.CL, cs.AI, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity.

Terminology

Summary

Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model’s internal activations remains unclear. Using matched cued and neutral conversations across five LLMs, we establish that a localized internal activation signal tracks changes in recommendations, with correlations up to r = 0.87. When multiple cues appear together, their internal signals largely combine, but the changes in output do not simply add up. We further show that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting, while largely preserving general benchmark performance. However, the ability to selectively remove one dimension’s influence while leaving co-present dimensions intact remains highly model- and attribute-specific. These results connect implicit personalization behavior to an internal signal that can be analyzed and causally controlled.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a real-time demographic bias monitor that reads internal activations (not just output text) during conversations. When a user mentions age, gender, race, or other attributes implicitly, the system flags the exact activation pattern that correlates with biased recommendation shifts (up to r=0.87) and alerts the developer or user.

  2. Implement activation-based debiasing as a built-in control layer—instead of relying on prompt instructions like ignore demographics, the system directly subtracts or nullifies the identified internal signal vector for a specific attribute. This suppresses the unwanted influence more effectively than prompting, while preserving general benchmark performance.

  3. Enable selective attribute isolation—the system can now disentangle overlapping demographic cues (e.g., a user who is both elderly and female) by identifying separate activation subspaces for each. This allows the AI to remove bias from one dimension (e.g., age) without affecting the other (e.g., gender), though the paper notes this remains model- and attribute-specific—so the system should be trained to improve this selectivity.

  4. Create a personalization transparency dashboard—during inference, the system can output a low-dimensional vector showing how much each demographic cue is influencing the current recommendation. This lets users or auditors see why a suggestion changed, turning a black-box shift into a measurable, controllable signal.

  5. Develop a cue-combination predictor—since internal signals combine linearly even when output changes do not, the system can predict how multiple implicit cues will interact internally. This allows preemptive adjustment of the activation space to prevent unexpected compounding biases (e.g., age + income leading to a disproportionately skewed loan recommendation).

What the improved AI system can do:

  • Detect and correct demographic bias in real-time, even when the user never states their identity, by monitoring internal neural activations.

  • Suppress unwanted demographic influence more reliably than prompting, without degrading general task performance.

  • Isolate and remove bias from one demographic dimension while leaving others intact, enabling nuanced fairness control.

  • Provide explainable, quantitative feedback on how each implicit cue affects each output, making personalization auditable.

  • Predict and mitigate complex interactions between multiple demographic cues before they cause harmful or unfair recommendations.

Abstract

Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations remains unclear. Using matched cued and neutral conversations across five LLMs, we establish that a localized internal activation signal tracks changes in recommendations, with correlations up to r=0.87. When multiple cues appear together, their internal signals largely combine, but the changes in output do not simply add up. We further show that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting, while largely preserving general benchmark performance. However, the ability to selectively remove one dimension's influence while leaving co-present dimensions intact remains highly model- and attribute-specific. These results connect implicit personalization behavior to an internal signal that can be analyzed and causally controlled.

Sources

Related papers