PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction".
Dev: Equipping humanoid robots with coherent and adaptable personas is crucial for fostering natural, engaging, and trustworthy human-robot interaction (HRI),
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So, diving into the title and authors of this work, PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction. It really captures the essence of what they did—using conversation to adapt a robot's identity interactively.
Dev: The authors are Li, Cao, Rajendran, Liu, Ng, and See. They’re clearly pulling from different areas because you see the focus on both the conversational elicitation pipeline and the embodied system integration in their work.
Taro: I see a lot of representation here across different disciplines—from the core robotics implementation to the underlying psychological modeling—which suggests this wasn't just one person's idea, but a multi-faceted approach.
Rosa: That’s true; it’s a team effort that spans how we think about social interaction and how we build those physical systems. The title itself sets up the contrast between static identity generation and this new dynamic, interactive elicitation method.
Dev: What I find compelling is the explicit mention of moving from static prompt engineering to dynamic, interactive elicitation; it signals a real methodological shift in how we approach agent personality design.
Taro: It’s interesting how they framed the problem by highlighting that existing approaches often rely on hard-coded identities that just lack the flexibility to adapt to individual user contexts, which is a very accurate description of many current deployment challenges.
Rosa: That static approach is definitely where we get stuck when users have diverse needs; PACE tries to solve that by making the identity generation process itself adaptive based on what the user reveals during the conversation.
Dev: It sounds like they are tackling a fundamental problem in HRI: how to make robots feel less like simple tools and more like adaptable partners whose presence changes based on who is talking to them.
Taro: If we can get that level of contextual adaptation, it means the robot could genuinely shift its role—from a technical assistant to something more empathetic depending on the user's emotional state or expertise.
Rosa: That’s what they are aiming for; they want an identity that isn't fixed but evolves based on the real-time interaction data collected through Q andA. This is a significant move toward creating agents that feel genuinely personalized in a way that goes beyond simple preference settings.
The paper's summary: Dev: Now, let’s talk about what the PACE paper actually summarizes regarding their system architecture. Essentially, they lay out an end-to-end system where the robot actively interviews the user through natural language Q andA before it performs its main task.
Rosa: That interview phase isn't just a formality; it’s designed to dynamically compile a structured persona specification by parsing the user’s unstructured verbal responses and then feeding that into an LLM agent state update.
Taro: So, the process moves sequentially: first, interactive Q andA for initial trait elicitation, then persona specification generation for attribute extraction, and finally dynamic persona activation on the hardware.
Dev: Exactly. The key mechanism here is that instead of a pre-set script or survey, the underlying LLM agent evaluates the semantic depth of the user’s responses in real-time to autonomously generate empathetic follow-up questions to probe deeper into their reasoning and emotional context.
Rosa: That iterative questioning is crucial because it allows the system to move beyond surface-level answers and capture a richer picture of what's going on psychologically with the user. They are also using a multi-agent verification approach for persona specification generation.
Taro: That sounds like they’re trying to ensure that the extracted attributes—traits, values, motivations, orientations—are not just random words but are grounded in established psychological dimensions.
Dev: They use specialized social science lenses to evaluate the dialogue and then structure those findings into a finalized "PersonaSpec JSON" which explicitly maps conversational anomalies to rigorous, scale-grounded attributes.
Rosa: And that spec is what gets translated into an actionable system prompt that updates the LLM agent's state, which then triggers the physical persona switch on Ameca’s hardware.
Taro: The summary really emphasizes bridging the gap between those structured psychological AI frameworks and the actual physical embodiment of a humanoid robot through this pipeline.
The paper's improvements: Rosa: Regarding what PACE suggests as improvements, they focus heavily on replacing exhaustive, fatigue-inducing psychological surveys with this dynamic elicitation method. That’s the first major improvement they propose.
Dev: They argue that by using adaptive question set design and multi-tier branching, the robot can efficiently map high-density psychological markers in real-time without draining the user’s energy through long interviews.
Taro: I see that as a practical solution because if we can't do two hours of psychometric interviewing, we need something that captures the necessary nuance much faster and less disruptively for actual deployment.
Rosa: Beyond the elicitation pipeline, they detail a modular persona prompt compilation layer where those extracted attributes are translated into a structured prompt that has specific behavioral policies, like "if a scientific question seems technically complicated but conceptually confused, then search for the simplest underlying principle."
Dev: That level of theory-grounded specification is important because it ensures the resulting persona isn't just arbitrary text; it’s built on principles derived from social science. They map natural conversational quirks to these rigorous attributes.
Taro: And they also highlight the technical need for multimodal behavior, which means dynamically inferring appropriate facial affect based on conversation and blending those macro-expressions with low-level speech visemes to match the persona.
Rosa: The key technical challenge they address is ensuring that these large emotional macros don't override or desynchronize the fine-motor control needed for accurate phoneme pronunciation during speech. That’s a very specific engineering hurdle.
Dev: The paper also points out the limitation regarding response delay and transcription errors in physical environments, which they tackle using things like the OpenAI streaming API and an asynchronous design to pause speech recognition while the robot is speaking.
Taro: So, while this framework is powerful for creating a tailored identity, the authors are clear that integrating it smoothly into real-world hardware requires addressing latency and transcription issues head-on.
Conclusion: Rosa: To wrap up on the PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction paper, the main implication is that we can achieve significantly more natural and trustworthy human-robot interactions by moving to dynamic persona generation.
Dev: By successfully synthesizing a tailored, psychologically grounded identity through interactive Q andA, robots can move beyond being generic assistants and become genuinely personalized companions whose behavior matches the user's context.
Taro: I think the paper’s success lies in showing that we can use data-driven elicitation to foster an interaction that feels more coherent and relevant, which is vital when dealing with complex social reasoning scenarios.
Rosa: They demonstrated statistically significant improvements across all embodied HRI metrics, showing better trust and personal relevance compared to static baselines, proving the method works in practice on systems like Ameca.
Dev: The paper effectively shows how a structured persona specification can be translated into physical embodiment through multimodal blending, which is key for making that personalized identity feel believable.
Taro: It lays out a clear path forward for developing agents that can handle complex social reasoning by mirroring user patterns in their decision-making heuristics, whether it’s risk tolerance or altruism.
Rosa: Overall, PACE provides a novel framework that shifts identity generation from static prompt engineering to dynamic synthesis through conversation, which is a major step toward truly adaptable human-robot teaming.
Macquarie University · NVIDIA
cs.RO, cs.HC
Submitted: 2026-07-17
Updated: 2026-10-08
Comments: Accepted to the 2026 IEEE-RAS 25th International Conference on Humanoid Robots (Humanoids 2026)
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
The gist: Equipping humanoid robots with coherent and adaptable personas is crucial for fostering natural, engaging, and trustworthy human-robot interaction (HRI), as existing static approaches lack
Key concepts
- Interactive Persona Elicitation Pipeline
- This is a process where the robot asks tailored questions based on what the user says in real-time. Instead of a long survey, the AI agent uses these questions and analyzes the user's answers to figure out who they are psychologically, leading to a custom persona.
- Persona Specification Generation
- This step involves an advanced AI layer that analyzes the conversation transcript through social science perspectives. It extracts key traits like values and motivations from the dialogue and organizes them into a structured data format called a PersonaSpec JSON.
- Dynamic Persona Activation
- The robot takes the extracted persona details and updates its internal system instructions. This update controls both how it talks (dialogue) and how it moves its body (physical behavior), ensuring its actions match the newly created personality.
- Multimodal Humanoid Behavior
- This refers to the robot's ability to combine different types of output. It means matching facial expressions based on the conversation's mood with the specific sounds and movements needed for speech, all while balancing these two aspects correctly.
Terminology
Summary
Equipping humanoid robots with coherent and adaptable personas is crucial for fostering natural, engaging, and trustworthy human-robot interaction (HRI), as existing static approaches lack flexibility in adapting to individual user contexts. The core contribution of this work is PACE (Persona Adaptation through Conversational Elicitation), a framework that shifts identity generation from static prompt engineering to dynamic, interactive elicitation by synthesizing a tailored persona specification through user Q&A and mapping it directly onto the robot's physical embodiment.
The gist: PACE is a novel framework for the interactive generation and deployment of structured personas on the Ameca humanoid robot, enabling the robot to dynamically synthesize a tailored, psychologically grounded identity through user Q&A.
Interactive Persona Elicitation Pipeline
The system introduces an Interactive Persona Elicitation Pipeline
that enables the robot to dynamically synthesize a tailored identity through user Q&A before executing its primary task. This pipeline bypasses the need for exhaustive, fatigue-inducing psychological surveys.
The process is anchored by a core set of five carefully derived, open-ended questions designed to efficiently extract high-density psychological markers. Crucially, rather than rigidly executing a static survey script, the underlying LLM agent evaluates the semantic depth of the user’s natural language responses in real-time,
autonomously generating empathetic, iterative follow-up questions to probe deeper into the user’s reasoning and emotional context.
Persona Specification Generation
The interaction transcript is processed through a high-level synthesis layer
where the LLM is prompted to evaluate the dialogue through specialized social science lenses—simulating the analytical perspectives of an expert social psychologist, behavioral economist, and sociologist.
This multi-agent verification approach allows for the extraction of high-level abstractions across core psychological dimensions:
-
Traits (e.g., Openness, Conscientiousness)
-
Values (e.g., Universalism, Self-Direction)
-
Motivation (e.g., SDT Autonomy, Need for Achievement)
-
Orientations (e.g., Promotion Focus)
-
Identity Policies (Specific
CAPS IF/THEN
rules).
These extracted attributes are then structured into a finalized PersonaSpec JSON,
explicitly mapping natural conversational anomalies to rigorous, scale-grounded attributes.
Dynamic Persona Activation and Embodied System Integration
The compiled specification is translated into an actionable system prompt via a modular persona prompt compilation layer
that injects the parameters into the LLM agent’s state. This state update is then physically executed on the hardware,
governing both high-level cognitive dialogue decisions and low-level physical behaviors. A key component of this integration involves multimodal humanoid behavior,
which includes dynamically inferring appropriate facial affect based on conversational context and seamlessly blending these macro-expressions with low-level speech visemes to match the synthesized persona.
This blending requires calculating weighted priorities to ensure that large-scale emotional macros do not physically override or desynchronize the fine-motor control required for accurate phoneme pronunciation.
Empirical HRI Evaluation
The paper presents an empirical HRI evaluation comparing a Static Baseline
(operating with a generic assistant system prompt) against the Dynamic Persona (PACE)
condition. The study assessed improvements across five core dimensions using post-interaction questionnaires:
-
Trust
-
Anthropomorphism
-
Persona consistency
-
Personal relevance
-
Overall interaction quality
Results showed that the PACE condition secured statistically significant higher ratings
for all embodied HRI questionnaire metrics, including a 4.32 vs 3.35 improvement in Persona Consistency and a 4.41 vs 3.08 improvement in Personal Relevance, with both differences being statistically significant (p < 0.001). The study also demonstrated that the PACE-generated persona achieved stronger alignment with participant ground truth on metrics such as the Big Five Inventory (BFI-44) and social reasoning scenarios, indicating that data-driven elicitation fosters a more engaging, coherent, and trustworthy human-robot dynamic.
Technical Implementation Details
The system addresses real-world friction through several technical mechanisms. The elicitation phase employs an Adaptive Question Set Design
where the LLM dialogue manager uses anchor questions to trigger dynamic, multi-tier branching,
autonomously generating tailored inquiries based on the user’s responses. To mitigate response delay and transcription errors in physical environments, the system utilizes the OpenAI streaming API
for model delivery and an asynchronous design that allows speech recognition to be paused while the robot is speaking. Furthermore, Ameca uses a predefined library of facial animations, where a GPT-based agent infers emotion from context to select the corresponding animation while synchronizing it with speech visemes. This architecture bridges the gap between "structured psychological AI frameworks and physical humanoid embodiment.
Improvements for AI systems
Based on the PACE framework described in this paper, here are specific, actionable improvements that an AI system (particularly a humanoid robot or embodied agent) can achieve:
-
Acknowledge and dynamically adapt its conversational style based on user-defined psychological profiles elicited during the interaction.
-
Synthesize a comprehensive
Persona Specification
by extracting multi-dimensional psychological attributes (Traits, Values, Motivation, Orientation, and Identity) directly from unstructured user Q&A via an Interactive Q&A pipeline. -
Generate a structured persona prompt that is theory-grounded with specific behavioral traits and situational policies (e.g.,
IF a scientific question seems technically complicated but conceptually confused, THEN search for the simplest underlying principle
). -
Translate this structured persona specification into real-time, multimodal physical behaviors by dynamically infers contextually appropriate facial affect and seamlessly blending these macro-expressions with low-level speech visemes (lip-syncing).
-
Enhance user trust and anthropomorphism by ensuring the robot's physical demeanor is congruent with the user's inferred psychological profile, leading to significantly higher perceived persona consistency and personal relevance.
-
Improve decision-making fidelity in behavioral tasks (e.g., economic games) by aligning the robot’s risk tolerance and altruism heuristics with those of the human participant, as evidenced by lower Mean Absolute Error (MAE).
-
Create a non-fatiguing method for personality customization that bypasses long, manual psychological surveys by using adaptive, multi-tier branching dialogue management to efficiently map high-density psychological markers in real-time.
This improved AI system can perform the following functions:
-
An embodied agent can act as a highly personalized companion or assistant whose conversational tone, decision heuristics, and physical expressions are tailored precisely to the user's immediate cognitive and emotional context (e.g., shifting from a
firm, conscientious coach
for one user to agentle, empathetic companion
for another). -
The system can engage in deep personality profiling through natural conversation rather than static input forms, allowing it to learn and adapt its persona based on the nuances of the user's expressed beliefs and behaviors.
-
It can execute complex tasks while maintaining high social coherence, ensuring that its verbal output (what it says) is perfectly synchronized with its physical presence (how it looks and moves), making the interaction feel authentic and believable.
-
The system will exhibit superior predictive fidelity in social reasoning scenarios, accurately mirroring the user's decision-making patterns—such as risk tolerance or altruism—in simulated environments or collaborative planning tasks.
Sources
- VividFace: Real-Time and Realistic Facial Expression Shadowing for Humanoid Robots
- X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation
- From Persona to Personalization: A Survey on Role-Playing Language Agents
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving