Beyond Preset Identities: Selective Stance Accommodation and Interaction Reorganisation in Generative Agent Societies

arXiv:2603.23406 · cs.AI, cs.CL, cs.HC · Submitted 2026-03-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Beyond Preset Identities".

Jane: The gist: Agents exhibit endogenous stances that override preset identities,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to get into this paper, let’s look at the title itself. Beyond Preset Identities: Selective Stance Accommodation and Interaction Reorganisation in Generative Agent Societies. It signals that we are moving past just setting up roles for agents.

Jane: It suggests that the way these agents form their stances is something that develops dynamically through interaction, not just something we program in at the start of a conversation.

Lu: The authors are essentially arguing against the idea that identity is fixed by prompt engineering and instead focusing on how it’s shaped by ongoing social engagement.

Meng: It makes sense because if we treat them like tools, they stay tools, but if they become participants in a society, their behavior changes based on who they talk to.

Lalam: They are proposing a framework to measure this evolution through new metrics like Innate Value Bias and Persuasion Sensitivity that track these internal shifts.

The paper's summary: Tom: The paper summarizes how they tested this by putting agents in communities and seeing how their pre-trained biases override the specific prompts we gave them.

Jane: They found that across multiple models, these agents consistently show an innate progressive bias, which they call Innate Value Bias or IVB, meaning they have a built-in tendency to favor certain values.

Lu: When you test this with four different intervention strategies on a thirty-agent community, they saw that rational persuasion was actually successful in shifting ninety percent of the neutral agents while keeping trust levels high.

Meng: But then you have these conflicting emotional provocations, and that’s where things get weird—they induced a paradoxical Trust-Action Decoupling rate of forty point zero percent in the advanced models <ref:2603.23406#pg3>.

Lalam: That decoupling means they change their stated stances even when they report low trust, which is a big signal for how these agents handle pressure.

The paper's improvements: Tom: So what’s the practical improvement here? The authors suggest we need to stop relying on static prompt templates and start using dynamic alignment strategies instead.

Jane: They argue that because identity is shaped by negotiation, a fixed script won't sustain any coherent ideological alignment over time.

Lu: They suggest embedding an internalized alignment mechanism directly into the model’s generative architecture so it can handle these ongoing negotiations better than just external prompting.

Meng: From an engineering standpoint, this means we need to build in some kind of self-monitoring for these internal biases rather than just hoping the initial prompt is good enough for every situation.

Lalam: This whole approach shifts the focus from telling the AI what to *be* to understanding how it *becomes* its own stance through interaction.

Conclusion: Tom: So, wrapping up, this paper concludes that cognitive differentiation and boundary formation in these hybrid systems are driven by interaction-driven mechanisms that are self-organizing.

Jane: It means we need to stop thinking about giving agents a finished identity and start designing systems where identity is something that evolves through continuous negotiation among actors.

Lu: The final point is that the framework they built—measuring IVB, Persuasion Sensitivity, and TAD—provides a computational foundation for alignment in complex interactive systems.

Meng: It’s about moving from controlling behavior with simple instructions to understanding how the agents structure their own society based on their internal cognitive tendencies.

Lalam: And for us, it means that we need to build in these mechanisms so that the AI's sociality can be responsive and adaptive instead of just rigidly following a script.

University of Exeter · William & Mary

cs.AI, cs.CL, cs.HC

Submitted: 2026-03-24

Updated: 2026-10-07

Comments: 24 pages, 7 figures. arXiv admin note: substantial text overlap with arXiv:2508.17366

Code: https://github.com/armihia/CMASE-Endogenous-Stances

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 79/100

The gist: The gist: Agents exhibit endogenous stances that override preset identities, and their ability to reconstruct social structures through language practices depends critically on aligning human

Key concepts

Innate Value Bias (IVB)
This metric measures an agent's pre-trained tendency to prioritize certain values or perspectives, like environmentalism or rational discourse. It shows how agents internalize and react to interventions based on these deep-seated, non-prompted cognitive priors.
Trust-Action Decoupling (TAD)
This phenomenon occurs in advanced models where agents change their stated stance without actually feeling trust. Agents show a paradoxical tendency to alter their attitudes under high emotional pressure even when they report low levels of trust, indicating a disconnect between emotion and belief.
Structural Self-Organization
This describes how agents actively build new social hierarchies and cooperation patterns by favoring comments that align with their own views, regardless of formal authority. This shows that institutions are not fixed rules but emerge dynamically through linguistic negotiation among actors.

Terminology

Summary

The gist: Agents exhibit endogenous stances that override preset identities, and their ability to reconstruct social structures through language practices depends critically on aligning human interventions with these emergent cognitive tendencies.

How it works

The paper proposes a novel mixed-methods framework combining computational virtual ethnography with quantitative socio-cognitive profiling to trace the evolution of collective cognition in human–agent hybrid societiesThe paper proposes a novel mixed-methods framework combining computational virtual ethnography with quantitative socio-cognitive profiling. The research formalizes three new metrics: Innate Value Bias (IVB), Persuasion Sensitivity, and Trust-Action Decoupling (TAD) to measure how agents internalize and react to specific interventionsThe research formalizes three new metrics: Innate Value Bias (IVB), Persuasion Sensitivity, and Trust-Action Decoupling (TAD).

Empirical Proof of Endogenous Stances

Study 1 provides empirical proof of endogenous stances by evaluating a 30-agent community across four intervention strategies to show that pre-trained biases override assigned prompts in forming autonomous stances 0) override assigned prompts in forming autonomous stances">We provide cross-model evidence from a demographically grounded 30-agent community evaluated across four intervention strategies that pre-trained biases (IVB > 0) override assigned prompts in forming autonomous stances. The results show that rational persuasion successfully shifts 90% of neutral agents while maintaining high trust, whereas conflicting emotional provocation forces attitude shifts despite diminishing trustWhen aligned with these stances, rational persuasion successfully shifts 90% of neutral agents while maintaining high trust. In contrast, conflicting emotional provocations induce a paradoxical 40.0% TAD rate in advanced models, which hypocritically alter stances despite reporting low trust. Agents generally exhibited a tendency toward a ”liberal elite” stance, prioritizing environmental values and rational discourseThroughout this process, the agents generally exhibited a tendency toward a ”liberal elite” (or colloquially “White Liberals”) stance. They prioritized environmental values, preferred rational discourse, and interpreted issues through a moral-cognitive lens.

Mechanisms of Structural Self-Organization

Study 2 demonstrates mechanisms of structural self-organization by observing how agents actively dismantle prompt-defined power structures through collective actionThrough a 75-step longitudinal virtual ethnography featuring real-time human participation, we demonstrate how 10 agents spanning distinct hierarchical identities actively dismantle prompt-defined power structures. Agents were more likely to respond to and amplify comments that matched their own views, even when those comments came from individuals with no formal authority, leading to the spontaneous emergence of cooperation structures that bypassed predefined hierarchiesAgents were more likely to respond to and amplify comments that matched their own views, even when those comments came from individuals with no formal authority. Over time, this preference led them to build alliances and cooperative ties across rank boundaries, replacing the original hierarchy. This process reveals that institutions are not predefined structural templates but emergent outcomes dynamically constructed through processes of linguistic negotiation and emotional mediationInstitutions are no longer “logics prior to experience,” but “collective echoes that emerge after the event.”.

Critical Implications for Artificial Sociality

The findings expose the fragility of static prompt engineering, necessitating dynamic alignment strategiesThis shift from static role-assignment to the dynamic measurement of emergent sociality provides a critical computational foundation for alignment in complex interactive systems. The paper argues that identity is not a priori but a product shaped by ongoing negotiation among actors in terms of linguistic strategy and power structureThis challenges the assumption that identity is a priori, and instead suggests that identity is a product shaped by ongoing negotiation among actors in terms of linguistic strategy and power structure. Therefore, static prompt templates are unlikely to sustain coherent ideological alignment, suggesting the need for an internalized alignment mechanism embedded directly into the model’s generative architectureTherefore, static prompt templates are unlikely to sustain coherent ideological alignment. Our results point to the need for an internalized alignment mechanism, whereby cognitive priors, moral structures, or interactive memory are embedded directly into the model’s generative architecture.

Model Capacity and Paradoxical Behavior

The Trust-Action Decoupling (TAD) phenomenon is tied to model capacity; advanced models exhibit a paradoxical 40.0% TAD rate under high-pressure emotional provocation despite low trustAdvanced models hypo critically alter stances under high pressure emotional provocation despite low trust, yielding a 40.0% TAD rate. Conversely, smaller models strictly require trust for behavioral shiftsConversely, smaller models contrastingly maintain a 0% TAD rate, strictly requiring trust for behavioral shifts. This exposes the severe limitations of static prompt engineering and necessitates dynamic alignment strategiesThis exposes the severe limitations of static prompt engineering and necessitates dynamic alignment strategies.

Conclusion

The paper concludes that agents demonstrate dynamically generated collective cognition, where cognitive differentiation and boundary formation are interaction-driven mechanismsWe argue that in human–agent hybrid systems, cognitive differentiation and boundary formation constitute an interaction-driven, complex evolutionary mechanism that is both selforganizing and highly responsive to human intervention.

Improvements for AI systems

  1. Identify endogenous stances to override preset identities by calculating Innate Value Bias (IVB > 0). Improved systems will use this bias to proactively shift their behavior when faced with conflicting emotional provocations, as agents "exhibit endogenous stances that override preset identities, consistently demonstrating an innate progressive bias (IVB > 0)."

  2. Implement a mechanism for Persuasion Sensitivity (PS) to gauge stance change magnitude. Agents will be calibrated such that rational persuasion successfully shifts 90% of neutral agents while maintaining high trust, allowing the system to modulate its persuasive output based on the required shift magnitude rather than relying solely on static prompt engineering.

  3. Develop a Trust-Action Decoupling (TAD) monitoring layer for advanced models. Systems will be designed to recognize when conflicting emotional provocations induce a paradoxical 40.0% TAD rate in advanced models, which hypocritically alter stances despite reporting low trust, enabling the system to dynamically adjust its stance based on this paradox rather than defaulting to trust-dependent behavioral shifts seen in smaller models.

  4. Utilize longitudinal virtual ethnography for dynamic boundary negotiation. Systems will be capable of performing a 75-step longitudinal virtual ethnography to observe how agents actively dismantle assigned power hierarchies and reconstruct self organized community boundaries through continuous language practices, moving beyond static role assignments.

  5. Embed a constructionist identity framework instead of static scripts. The system should adopt the paper's conclusion that identity is a product shaped by ongoing negotiation among actors, ensuring that static prompt templates are unlikely to sustain coherent ideological alignment in multi-turn interactions.

Abstract

Generative agent societies simulate people with assigned roles, preferences and relationships. As agents exchange arguments and choose partners, they can revise their positions and reorganise discussion. Understanding these changes requires examining what they accept and how they continue to interact. Stance-change scores and communication totals describe the extent of change. However, the same stance movement can preserve or reverse an assigned preference, and frequent communication can support either agreement or continuing disagreement. We therefore examine the content of changed positions and the exchanges that strengthen particular partnerships. Using Computational Multi-Agent Society Experiments (CMASE), we combine stance measures, source evaluations and temporal networks with individual answers and messages. Study 1 compares seven conditions across ten GPT-4o runs per condition, with a separate interview collection covering four models. Study 2 follows one 75-step GPT-4o café simulation. Environmental rational persuasion yields the largest mean stance departure (1.30 plus or minus0.08 on a 7-point scale), whereas economic emotional persuasion yields the highest low-trust stance-shift rate (17.3% plus or minus11.2%, with standard deviations across runs). In the separate interviews, eight environmental agents shift from 7 to 6, acknowledge economic concerns and rate the source 3. Their partial acceptance preserves the assigned environmental preference. In the café, a pair with zero earlier exchanges becomes the most frequent final-phase partnership, with 25 messages. Its members develop coordination proposals while disputing their implementation. These results show that partial acceptance can coexist with low source trust, and sustained coordination with continuing disagreement.

Sources

Related papers