Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Aligned but Not Partner-Specific".
Tom: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Welcome back to the show! Today we’re talking about this paper, "Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts." We’ve seen some exciting stuff about how these AI agents coordinate, and we want to break down exactly what that means for us.
Jane: That’s right. Essentially, this work tests whether multimodal LLMs actually develop those short, partner-specific reference conventions that humans use when they interact over time. The paper looks at whether the agents are just using general vocabulary or if they're building something specific to each partner based on their shared history.
Lu: Their main idea is trying to prove that MLLM agents don't inherently form these deep, interaction-specific connections on their own; instead, they seem to rely more on what the models already know from their general training.
Meng: I’m interested in how they set up the experiment because isolating that history effect is a big challenge in testing any AI coordination system.
Lalam: From my view, this paper shows us that without explicit interaction history grounding, we can get high success rates just by being very descriptive. It highlights a difference between statistical consistency and actual interactive understanding.
Tom: So, when they look at the results of "Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts," what’s the main point they are trying to make about this coordination?
Jane: The central finding is that MLLM agent dyads achieve coordination without developing those compact, history-dependent referring expressions that characterize human dialogue. Their success comes from verbose description rather than from forming those specific pacts between partners.
Lu: They show this contrast clearly when looking at the analysis of description strategy, noting that human descriptions compress proportionally across trials with a log-slope of about negative zero point three seven two, which they call the "signature of conceptual-pact formation".
Meng: If human descriptions are compressing like that, it suggests a learning process where they get better at communicating by using less language over time as they build rapport with each other.
Title and authors: Lalam: That constant effort level in the agents, staying flat at around two turns throughout, really makes me think about how we design learning systems; if an AI doesn't naturally learn to be more efficient with its communication as it interacts, we have to explicitly program that efficiency into the training process.
Tom: It sounds like the paper is suggesting that what we see in agents is actually just a result of being optimized for success-only regimes rather than true interaction. What do you think, Jane?
Jane: I agree with Tom; their alignment seems driven by shared pretrained priors and descriptive verbosity rather than genuine partner-specific grounding. It really points to the gap between high performance and true cooperative dialogue competence.
Lu: And that gap is precisely what they are trying to bridge with their novel methodological contribution, which involves extending the pseudo-dyad approach under four pragmatic constraints to break the history effect.
Meng: From a practical standpoint, that methodology is key because it provides a way to test if we’re seeing partner-specific grounding or just general task vocabulary.
Lalam: If they can establish that within-condition control, it means future AI systems won't just rely on raw performance metrics to judge their cooperative ability; we’ll have a way to measure if they are actually building those meaningful, interaction-specific connections.
Tom: So, looking ahead at the paper’s suggested improvements for "Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts," what are the key suggestions for researchers?
Jane: The paper points toward extending that constrained pseudo-dyad methodology under those four pragmatic constraints to better sever partner history while preserving the dialogue structure. That's how they try to isolate the variables.
Lu: They also suggest refining how we measure alignment itself, moving beyond just final label matches to a more detailed analysis of alignment dynamics over time, focusing on turn level differences.
Meng: That dynamic measurement would be useful for figuring out exactly when and why the coordination happens, rather than just knowing it happened at the end.
Lalam: It’s encouraging to see this kind of detailed analysis; it gives us a framework for figuring out how to move AI from just being good at talking to actually being good at interacting with specific people.
Tom: And what about the practical application of these suggestions? How does this help us build better conversational agents?
Title and authors: Jane: They suggest focusing on dynamic constraint layers that mimic human entrainment, which would help agents adjust their description strategy proportionally across rounds instead of maintaining a flat output.
Lu: If we can incorporate that kind of constraint into the architecture, we might start seeing those differences between shared vocabulary and genuine partner-specific grounding emerge more clearly during testing.
Meng: I think if we can make the system learn to compress its descriptions proportionally, it would make their interactions much more efficient in real-time scenarios, which is a big deal for deployment.
Lalam: It’s encouraging to see this kind of detailed analysis; it gives us a framework for figuring out how to move AI from just being good at talking to actually being good at interacting with specific people.
Tom: Well, that covers the suggested avenues for future research on "Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts." What does this all mean for our understanding of AI collaboration?
Jane: It really boils down to this: raw performance metrics aren't enough; we have to look at the quality of the underlying communication strategy, whether it’s a history-dependent convention or just consistent verbosity. The paper concludes that MLLM agent dyads achieve coordination without convention by relying on shared pretrained priors and descriptive verbosity.
Lu: That contrast between human dialogue and agent dyads is what’s so fascinating; it tells us the mechanism of coordination can be fundamentally different depending on whether you're looking at human interaction or agent dyads.
Meng: It means for practical applications, we need to design systems that prioritize developing those compact, history-dependent referring expressions over just generating long sequences of text; that would be a huge win for efficiency.
Lalam: If we can get the AI to learn how to form those specific pacts, it could fundamentally improve how people collaborate with these systems on complex tasks. It’s about making the AI feel like a true partner in the conversation.
Tom: That's a powerful idea for the future of how we design conversational interfaces, and we’ve got some really interesting papers coming up next about spatial reasoning with MLLMs. We'll be right back after this short break.
The paper's summary: Tom: So, we’ve just wrapped up our deep dive into "Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts," and now it's time to really look at what this means for the broader landscape of AI interaction.
Jane: Right, the main idea boils down to this: these multimodal LLM agents manage to coordinate on tasks without ever developing those history-dependent, partner-specific conventions that humans build over time. It shows us that statistical consistency and just being very thorough with descriptions can actually lead to high task success, even when there’s no true conceptual grounding happening between the partners.
Lu: I think the biggest implication here is how we view the learning process itself; it suggests that current training methods might be optimizing for a specific type of success—verbose description—rather than developing those compact, efficient communication structures that we actually want to see in complex interactions.
Meng: From an engineering standpoint, this means our systems need to evolve past just maximizing lexical overlap; we have to build in mechanisms that actively reward the formation of those specific partner pacts rather than relying on general pretrained vocabulary. That’s something we can definitely work on implementing into the architecture.
Lalam: For culture, this finding is really interesting because it suggests that if we want AI to truly integrate into collaborative environments, we need to focus on teaching it how to form those personalized conceptual pacts with its partners. That ability would make interaction feel much more natural and less like just filling in the blanks.
Tom: It really boils down to this: raw performance metrics alone don't tell the whole story about cooperative dialogue competence; we need to look at the underlying mechanisms, like descriptive verbosity versus actual grounding.
Jane: I agree. The paper’s conclusion is that MLLM agent dyads achieve coordination without convention by relying on shared pretrained priors and descriptive verbosity, which contrasts sharply with the cumulative and interaction-dependent alignment we see in humans.
Lu: That contrast is what’s so fascinating; it tells us that the mechanism of coordination can be fundamentally different depending on whether you're looking at human dialogue or agent dyads.
Meng: It means for practical applications, we need to design systems that prioritize developing those compact, history-dependent referring expressions over just generating long sequences of text. That would be a huge win for efficiency in real-time scenarios.
Lalam: If we can get the AI to learn how to form those specific pacts, it could fundamentally improve how people collaborate with these systems on complex tasks. It’s about making the AI feel like a true partner in the conversation.
Tom: That's a powerful idea for the future of how we design conversational interfaces. So, that’s our final word on "Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts." Next up, we're looking at how Vision-Language Models are handling spatial reasoning under changing viewpoints.
The paper's improvements: Tom: So, we’ve just finished dissecting "Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts," and now we're moving into the suggestions for what comes next.
Jane: The authors are proposing a few key avenues for future research that focus on how we can actually isolate those tricky variables they talked about, like separating partner-specific grounding from just general task vocabulary.
Lu: They’re really pushing for better controls, specifically demanding within-condition control that locks the task and model structure while actively breaking the interaction history to test if alignment is truly history-dependent. That’s where we need to go next in multimodal modeling research.
Meng: From an engineering standpoint, that means we might need to build in specific mechanisms that actively push models toward developing those compact, partner-specific lexical cores we discussed earlier. We can't rely solely on the current instruction tuning regime if we want real adaptability in how they communicate.
Lalam: I think for our culture, this paper shows us that simply scaling up models doesn't automatically grant them nuanced social skills; we have to explicitly teach them how to form those specific pacts. That’s a significant shift in how we view AI development and how we design their social interaction capabilities.
Tom: They also suggest refining the way we measure alignment itself, moving beyond just a final label match to looking at the dynamics over time, focusing on those turn level differences.
Jane: That brings us back to their suggestion that we should look at how alignment changes when comparing real human dyads against those pseudo-dyads. It’s about seeing where the divergence between them happens throughout the game, not just at one single moment.
Lu: And they strongly emphasize that their study shows MLLMs succeed by verbose description rather than by forming those compact, history-dependent referring expressions characteristic of human dialogue. That’s a strong statement about the mechanism they found in these systems.
Meng: It’s interesting that they also note that the descriptive verbosity observed in agents is attributed to the success-only optimization regime shared by current instruction-tuned MLLMs. That suggests it might be a byproduct of how we currently train these models for high performance rather than an inherent goal.
Lalam: So, the implication for future work is that we need to move beyond just seeing high success rates and start looking at the actual language patterns and effort management within the interaction itself.
Tom: It really is a reminder that raw performance metrics alone don't tell the whole story about cooperative dialogue competence; we need to look at the underlying mechanisms, like descriptive verbosity versus actual grounding.
Jane: So, as we move forward, I think focusing on dynamic constraint layers that mimic human entrainment could be a good next step for improving how these agents communicate across rounds.
Lu: Exactly. By incorporating that kind of constraint into the architecture, we might start seeing those differences between shared vocabulary and genuine partner-specific grounding emerge more clearly in our tests.
Meng: I think if we can make the system learn to compress its descriptions proportionally, it would make their interactions much more efficient in real-time scenarios.
Lalam: It’s encouraging to see this kind of detailed analysis; it gives us a framework for figuring out how to move AI from just being good at talking to actually being good at interacting with specific people.
Conclusion: Tom: So, we’ve just finished talking about "Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts," and to wrap things up, we're going to recap the main findings and what this means for the future of AI interaction.
Jane: Exactly. The core idea is that these multimodal LLM agents manage to coordinate effectively without developing those history-dependent, partner-specific conventions that humans build over time. It shows us that statistical consistency and descriptive verbosity can actually achieve high levels of task success, even when true conceptual grounding is missing.
Lu: I think the biggest implication here is how we view the learning process itself; it suggests that current training methods might be optimizing for a specific type of success—verbose description—rather than developing the kind of compact, efficient communication structures we actually want to see in complex interactions.
Meng: From an engineering standpoint, this means our systems need to evolve past just maximizing lexical overlap; we have to build in mechanisms that actively reward the formation of those specific partner pacts rather than relying on general pretrained vocabulary. That’s something we can definitely work on implementing.
Lalam: For culture, this finding is really interesting because it suggests that if we want AI to truly integrate into collaborative environments, we need to focus on teaching it how to form those personalized conceptual pacts with its partners. That ability would make interaction feel much more natural and less like just filling in the blanks.
Tom: It really boils down to this: raw performance metrics aren't enough; we have to look at the quality of the underlying communication strategy, whether it’s a history-dependent convention or just consistent verbosity.
Jane: I agree. The paper’s conclusion is that MLLM agent dyads achieve coordination without convention by relying on shared pretrained priors and descriptive verbosity, which contrasts sharply with the cumulative and interaction-dependent alignment we see in humans.
Lu: That contrast is what’s so fascinating; it tells us that the mechanism of coordination can be fundamentally different depending on whether you're looking at human dialogue or agent dyads.
Meng: It means for practical applications, we need to design systems that prioritize developing those compact, history-dependent referring expressions over just generating long sequences of text. That would be a huge win for efficiency.
Lalam: If we can get the AI to learn how to form those specific pacts, it could fundamentally improve how people collaborate with these systems on complex tasks. It’s about making the AI feel like a true partner in the conversation.
Tom: That's a powerful idea for the future of how we design conversational interfaces. So, that’s our final word on "Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts." Next up, we're looking at how Vision-Language Models are handling spatial reasoning under changing viewpoints.
Jane: Indeed, Tom. The paper really shows us the distinction between mere consistency and actual interactive understanding in these kinds of tasks.
Lu: It’s a testament to how different modalities can yield different forms of coordination, which opens up some wild possibilities for next-generation agents.
Meng: We need that focus on efficiency in those compact expressions if we're going to deploy these things effectively.
Lalam: I think the future of AI culture depends on us teaching these models how to build those personalized connections, not just giving them bigger models.
Co-Intelligence Humanities AI Future Lab and Graduate Institute of Linguistics, National Taiwan University · Multimodal Language Department, Max Planck Institute for Psycholinguistics · Donders Institute for Brain, Cognition and Behaviour, Radboud University · Institut Jean Nicod
cs.CL, cs.AI
Submitted: 2026-06-06
Updated: 2026-10-01
Code: https://github.com/diff94/tan_ana
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history.
Key concepts
- Partner-Specific Grounding
- This refers to when two people learn to use shorter, unique ways of describing things based on their specific interaction history. The study tested if LLM agents develop these unique rules tailored only to their partner, or if they rely on general shared knowledge.
- Constrained Pseudo-Dyad Baseline
- This was a special test setup created to mimic real interactions while intentionally breaking the partner's unique history. It allowed researchers to see if the observed alignment between agents depended on interacting with a specific person or just on general task vocabulary.
- Conceptual-Pact Formation
- This is the human process where two people build an unspoken agreement about how they will use language during a task. Humans compress their descriptions over time as this pact forms, leading to shorter, more efficient communication that relies on shared understanding.
- Descriptive Verbosity
- This describes the tendency of LLM agents to maintain long, detailed descriptions across multiple rounds without compressing. The study suggests this verbose style is a byproduct of the optimization goals used during their training, leading to high success rates even without forming specific interaction conventions.
Terminology
Summary
Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history. This work addresses whether multimodal LLM agents achieve coordination through partner-specific grounding or merely statistically consistent output shaped by shared pretrained priors. The study finds that MLLM agent dyads succeed in coordination without convention, maintaining fixed effort levels and near-ceiling lexical overlap, contrasting sharply with the human tendency to reduce effort and develop compact, history-dependent referring expressions.
Methodology for Distinguishing Grounding
The research employs a novel methodological contribution: a constrained pseudo-dyad baseline
designed to match the original referential task structure while breaking partner history. This baseline is used to test whether observed label alignment depends on interaction with a specific partner rather than being shaped by shared task vocabulary or pretrained style. The analysis is conducted across three analytic layers: (i) task competence, (ii) description strategy, and (iii) alignment dynamics.
Data and Experimental Setup
The study adapts the KTH Tangrams corpus for MLLM dyads. The human baseline corpus consists of 42 two-party reference game dialogues comprising 3,288 rounds. The MLLM dataset involves 45 dyads across 905 rounds. To create the pseudo-dyad baseline, a constrained pseudo-dyad extraction algorithm
pairs rounds from source dyads under four pragmatic constraints: (i) referencing the same tangram shape; (ii) occupying comparable positions in their source dyads’ trajectories; (iii) being within a similar total turn difference with tracked per-role counts; and (iv) sharing the same outcome. This procedure yields 20 agent pseudo-dyads and 21 human pseudo-dyads.
Analysis of Task Competence and Effort
The analysis reveals clear differences in how agents manage effort across rounds. Regarding competence, both populations solve the task at high rates, but only humans show the signature of accumulating efficiency; a logistic GLMM shows that success declines strongly with round number
for humans (βˆ = −1.11). In terms of effort, human dyads exhibit a clear reduction in conversational turns as they progress through the game: Humans begin near ∼5 turns in their earliest deciles and converge toward ∼2 by mid-trial,
whereas agent trajectories remain flat at ∼2 throughout.
Agents reach high success rates by exhaustive description rather than by developing efficient conventions.
Analysis of Description Strategy and Alignment Dynamics
The paper examines how descriptions evolve over time. Humans compress proportionally across trials, with their log-slope being large and reliable (βˆ = −0.372), which is the signature of conceptual-pact formation.
In contrast, agent dyads remain at ∼90% of their starting level throughout
and exhibit a near-zero log-slope, indicating they do not compress. At the turn level, under the pseudo-dyad control, human and agent overlap diverge: real human dyads are far more likely than pseudo dyads to produce a shared-lexical-core turn,
whereas agents show near-ceiling in both conditions (real mean 0.979, pseudo mean 0.982).
The paper concludes that MLLMs achieve coordination without convention,
succeeding by verbose description rather than by forming the compact, history-dependent referring expressions characteristic of human dialogue.
Conclusion on Coordination Mechanism
The findings demonstrate that MLLM agent dyads achieve coordination without convention.
Their alignment is a byproduct of shared pretrained priors and descriptive verbosity,
whereas human alignment is cumulative and interaction-dependent.
The results extend the literature by showing that agents can achieve high task success while lacking the hallmarks of interaction-specific grounding, underscoring that raw performance metrics are insufficient for evaluating cooperative dialogue competence. The study provides a reusable framework consisting of the pseudo-dyad baseline, a three-layer analytic protocol, and a lexical-core extraction pipeline.
Limitations
The primary limitation noted is model generalizability; the use of GPT-5 was necessitated by its Responses API
providing infrastructure-level lossless multimodal interaction history across rounds. Furthermore, the study compares human–human and agent–agent dyads but does not include human–agent interaction. The authors also note that while modality differs (spoken vs. text), concepts like conceptual-pact formation
are robustly replicated in large-scale text-based games. The descriptive verbosity observed in agents is attributed to the success-only optimisation regime shared by current instruction-tuned MLLMs.
The gist
MLLM agent dyads achieve coordination without convention, maintaining fixed effort and near-ceiling lexical overlap, which are statistically indistinguishable between real and pseudo conditions, contrasting with the human tendency to reduce effort and develop compact, history-dependent referring expressions.
Improvements for AI systems
Based on this scientific paper, here are specific improvements that can be made to AI systems, categorized by the capabilities they would acquire:
)1. Develop a Convention-Aware
Reference Engine for LLM Agents:
The current limitation is that MLLMs achieve coordination through verbose description rather than compact, history-dependent conventions. An improved system should integrate a mechanism to actively seek and form partner-specific lexical cores over repeated interactions.
)2. Implement Effort-Based Description Compression (Human-Style):
Improve the agent's effort
metric by rewarding it for reducing its turn length and description volume as it gains confidence in its partner's understanding.
The system should incorporate a mechanism that tracks the efficiency of communication, penalizing long, repetitive descriptions unless they lead to a significant increase in alignment (i.e., when the partner's subsequent query is brief).
)3. Transition from Global Style Alignment to Partner-Specific Grounding:
The core finding is that MLLMs lack history-dependent grounding. The improved system must shift its objective function during interaction from maximize lexical overlap
to establish partner-specific conceptual pacts.
The AI should prioritize the development of compact, shared labels (lexical cores) grounded in the specific context of the current dyad, rather than relying on general pretrained vocabulary or exhaustive geometric descriptions.
)4. Implement a Dynamic Constraint Layer for Contextual Efficiency:
Borrowing from the constrained pseudo-dyad methodology, an advanced agent should operate under dynamic constraints that balance task competence with communicative efficiency.
When interacting with a known partner, the system should dynamically adjust its description strategy to mimic human entrainment—compressing descriptions proportionally across rounds (as seen in the human trajectory) rather than maintaining a flat or increasing verbosity.
)5. Develop a Convention-Formation
Feedback Loop:
The agent needs to learn from interaction history specifically regarding shared labels, not just task outcomes.
After successful identification, the system should explicitly reward the use of a short, partner-specific label (a lexical core), effectively training it to form a
conceptual pactwith that specific interlocutor.
)6. Refine Visual-to-Text Grounding for Robust Tagging:
The paper highlights challenges in reliably translating visual information into machine-readable identifiers.
Improve the multimodal pipeline to assign stable, non-leaky, two-letter labels to visual elements (like tangram shapes) that are consistently readable by the LLM and never accidentally revealed to the interlocutor during conversation.
This improved AI system would be capable of:
-
Establishing genuine, history-dependent communication conventions with specific partners, leading to more efficient and natural-sounding dialogue.
-
Compressing its descriptive language proportionally across rounds, reducing conversational overhead as it gains familiarity with its interlocutor.
-
Achieving high task success rates without the need for exhaustive description, demonstrating true conceptual grounding rather than mere descriptive saturation or reliance on verbosity as a proxy for competence.
Abstract
Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific expressions grounded in shared interaction history; that is, with conceptual pacts. Prior work shows that multimodal LLMs fail to become more efficient across rounds, although they align on the labels they use. However, how can we determine whether this alignment reflects partner-specific grounding rather than a shared task vocabulary? We address this by comparing competent multimodal agent dyads with human dyads from the KTH Tangrams corpus. Our novel methodological contribution is a pragmatically constrained pseudo-dyad baseline: rounds from two different real dyads describing the same target at comparable trajectory positions are paired, preserving referential task structure while removing shared partner history. This enables us to test whether the observed label alignment depends on interaction with a specific partner. Across three measures (task competence, description strategy, alignment dynamics), we find clear differences. Humans reduce effort through entrainment, compressing descriptions and increasing label alignment with partners. Agents instead maintain fixed effort levels, producing verbose descriptions from round one, with near-ceiling label overlap that is statistically indistinguishable between real and pseudo dyads. MLLMs thus achieve coordination without conceptual pacts, succeeding by verbose description rather than by forming the compact, history-dependent referring expressions characteristic of human dialogue.
Sources
- Emergent Natural Language with Communication Games for Improving Image Captioning Capabilities without Additional Data
- The Llama 3 Herd of Models
- A Survey on Large Language Model-Based Game Agents
- Emergent Multi-Agent Communication in the Deep Learning Era
- Towards Multi-Agent Communication-Based Language Learning
- Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes
- Success and Cost Elicit Convention Formation for Efficient Communication
- LVLMs and Humans Ground Differently in Referential Communication
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering