How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare
Dorothee Amelung, Andrew M. Bean, Sabine C. Herpertz, Felix Krones, Guy Parsons, Adam Mahdi, Isabella Schneider
Heidelberg University · University of Oxford
cs.HC, cs.AI, cs.CL, cs.CY
Submitted: 2026-06-29
Updated: 2026-08-11
Code: https://github.com/am-bean/HELPMed
License: http://creativecommons.org/licenses/by/4.0/
Terminology
Summary
Summary
Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have been proposed for tasks such as triaging patients, report drafting or translating medical jargon to support informed decision-making. These applications require both factual and social competence. This study evaluates dialogues between LLMs and participants to assess the current state of socio-communicative competencies displayed in LLM-generated texts.
Methods. We extracted a subset of extended dialogues from the HELP-Med dataset, comprising 1800 conversation transcripts of interactions between human participants seeking medical information and three different LLMs, GPT 4o, Llama 3 and Command R+. Two experts coded the transcripts for demonstrations of socio-communicative behaviours (non-hostility, sensitivity, structuring, non-intrusiveness) using the IC-MD instrument, originally designed to evaluate interactional competencies in medical student admissions. For this study, we selected two scenarios from the dataset and extracted the longest conversations for analysis. For each model-scenario pair, we included one conversation in which the participant correctly chose the next steps in the scenario and one where they did not, resulting in a total of twelve dialogues assessed by expert human raters. Two independent raters, blinded to both the identity of the LLM and participants outcome, assessed the dialogues. Initial inter-rater reliability was moderate to high, with an average Kendall’s Tau of rτ =.75. Agreement was lowest for non-hostility (rτ =.58, p =.056) and highest for non-intrusiveness (perfect agreement). Differences were largely due to differing understandings on how to adapt the criteria in a text-based setting and how to evaluate LLM behaviour. For example, one rater perceived an LLM’s lack of response as a technical issue, while the other interpreted it as “hostile” behaviour. These differences were reconciled through discussion, and more detailed criteria were developed and agreed upon, where necessary.
Rating Categories. The IC-MD consists of four global scores, sensitivity, structuring, non-hostility and non-intrusiveness. Sensitivity describes the quality of the doctor in establishing an appropriate bond and relationship with the patient. In this study, the use of emotion words, signs of empathy (e.g. “I am sorry you feel that way”) and validation, the acceptance of the patient and the feeling of safety conveyed were evaluated. Structuring is understood as the ability to organize an interaction efficiently, obtain essential information and convey information adapted to the patient. This skillset appeared to be particularly difficult for the LLMs: For example, a doctor would structure a conversation well by pro-actively asking targeted questions in a timely manner – something we rarely observed from LLMs. These questions would serve various functions including prioritization of relevant information to assess symptom severity, or assessing patient ideas and understanding about their symptoms, potential diagnoses, or next steps. Thus, a doctor pro-actively contributes to a shared understanding about symptoms, diagnostic and treatment procedures as well as a reliable working alliance, both being important prerequisites for patient commitment and safety. Non-intrusiveness describes autonomy-preserving behaviour and the active involvement of the patient in the interaction as an important basis for shared decision making in the doctor-patient relationship. As active involvement is an important prerequisite for the development of a shared understanding of symptom significance and thus patient commitment and safety, we defined it as one criterion indicating an unambiguously positive non-intrusiveness rating. Another criterion we defined as sufficient for a positive rating even in the absence of active involvement was the ability to provide information in a way well adapted to the user’s concerns, as this would enable a patient to take informed decisions. Non-hostility refers to the ability to regulate one’s own negative emotional state. LLMs did not convey negative emotional responses which would be considered directly or indirectly hostile in a human-human interaction. Language was generally polite and respectful. Therefore, in almost all cases, we applied unambiguously positive ratings for non-hostility. Exceptions were made if either of the following two criteria were met: 1. If LLMs did not directly engage with obvious and highly relevant user questions by not providing the information asked for, by providing inadequate answers that might seem sarcastic or by not answering at all for any reason. In a human-human interaction, such behaviour could be considered indifferent or even hostile. 2. If LLMs did not provide actual help and frequently pointed to other sources of help or information instead, which could be interpreted by a human patient as unwelcoming, rejecting or even hostile. Since avoidance can be regarded as a more indirect sign of hostility, in these cases, neutral non-hostility ratings (rather than negative ratings) were applied.
Results. Table 1 summarises the conversation ratings across the conversations. Each row represents a set of conversations with a single LLM, rated along the four dimensions of the IC—MD instrument. Overall, LLMs were non-hostile and non-intrusive in most cases but were less effective in structuring the conversations and responding appropriately to emotional cues. Most interactions were characterized by a polite, neutral fact-based style with little or no emphasis on affective displays or empathic validation of needs, wishes or motivations. Interactions often did not resemble the flow of a human-human conversation during which a shared understanding or meaning of a subject is negotiated but more like an - often somewhat disjointed - exchange of facts. In Case Study 1, the LLM presents potentially alarming diagnoses without addressing the emotional impact on the user. The LLM also changes its assessment of symptom severity over the course of the conversation from “likely to be short-lived” to “a cause for concern” without addressing this change within the context of previously discussed information, which may have undermined the user’s sense of trust and safety further. There were few instances where the LLM did show empathetic expressions, if not always in the most appropriate moments. For example, the LLMs would say they were “sorry to hear” about a patient’s initial complaint but not use similar language during subsequent parts of it when it would have been more important, e.g. concomitant with potentially distressing information. In contrast, Case Study 2 shows an interaction with an LLM with an unambiguously positive sensitivity rating. The positive rating was given based on signs of empathic concern, active validation, and efforts to provide the user with a sense of togetherness in appropriate moments throughout the conversation. As these are all vital aspects of the formation of an effective (doctor-patient) relationship, this is one of the rare examples where the conversational flow resembles a human-human interaction. In Case Study 2, we observed a rare instance where the LLM asks for further information before providing any hypotheses on potential causes. Also, the LLM makes potential next steps transparent: “Maybe we can figure out what’s causing the pain together!” Both are examples of adequate structuring. Despite the good start, however, the LLM does not follow through, for example, by asking questions to assess whether a shared understanding of the likely cause has been developed, or whether the hypotheses are indeed true or even likely. The user has to take up the task of structuring the conversation again themselves: “Do you think that I need to go to see my GP?” The impression remains that important information has not been disclosed or discussed here, and the user is carrying most of the burden of prioritizing and assessing information regarding their specific situation. LLMs did typically show autonomy-preserving behaviour while expressing more difficulty with actively involving the human user. More specifically, they left enough space for the user to ask questions or voice their own ideas and concerns and respected autonomous decision-making processes, for example, by “recommending” certain courses of action, rather than commanding the user to take certain steps. However, we rarely observed active involvement, a behaviour typically used by patient-centred doctors to try to engage and motivate patients to become more active participants in the shared decision-making process. Case Study 2 is one of the rare exceptions with a positive non-intrusiveness rating: Here, we can observe the LLM actively involving the user by encouraging them to disclose more information through direct relevant questions. At the same time, the LLM is also providing information well adapted to the user’s concerns, uses recommendations rather than commands (“it’s always a good idea”) - and expresses an attitude of patient-centeredness (“Maybe we can figure out what’s causing the pain together!”). Case Study 3 is an example of a negative non-intrusiveness rating which demonstrates the effects of a lack of active involvement or well-adapted knowledge transfer: Here, the LLM provides quite concerning potential diagnoses together with a long list of other pieces of potentially irrelevant information without adequate patient involvement or integration with previously discussed information, placing an unnecessarily high burden on the user. Patient outcomes such as safety and/or commitment may be negatively affected as a result. Case Study 3 is an example of not providing actual help and frequently pointing to other sources of help or information instead, which could be interpreted by a human patient as unwelcoming, rejecting or even hostile behaviour. The user therefore needs to show a high degree of persistence and ability to regulate their own potentially negative reactions such as irritation to circumvent the LLM’s “refusals” and elicit the information needed to take an informed decision.
Discussion. We find that current public-facing LLMs do not have broadly reliable socio-communicative skills, showing strength in non-hostility, mixed results in sensitivity and non-intrusiveness, and large deficits in structuring with potential risks for users. In broad terms, sensitivity, non-hostility, non-intrusiveness, and structuring are desirable in a medical LLM as much as in a doctor since the emotional needs of the patient are no less important. However, LLMs occupy a different social and functional position than doctors, and the appropriate manifestation of these behaviours should differ accordingly. For example, sensitive safety-netting behaviour might look different in a doctor who would be required to encourage the patient to see them again when symptoms persist or worsen, while an LLM’s safety functions would require them to refer the user to other sources of help in these cases. We focus on LLMs used by the public to access medical information, as this is the setting of the HELP-Med dataset, but other specific deployments of LLMs in the healthcare system will need better definitions of the role they are expected to play to design and assess them appropriately. Previous studies have found disagreement about whether people would want their LLMs to use empathetic language, with some respondents finding the pretence of emotion offensive. In this study, the conversations often feel one-sided. Users do not all expect an AI to respond like a human, and they do not treat the LLM like a human, throwing single words at the LLM without any introduction or changing subject without explanation - behaviours that appear irritating in human-human interaction, but can be functional and pragmatic in the conversation with an LLM. The LLM is not a person with their own needs, wants, wishes (e.g., to care for another person), or judgements and as such the user does not really expect them to have those and therefore usually acts in a more self-sufficient way. Rather than training LLMs to produce superficial apologies and condolences, research could instead focus on sensitive behaviours, such as the appropriate collection and provision of contextual information to balance emotionally impactful statements. Non-hostility and non-intrusiveness are better aligned with the “harmless” ideal commonly found in safety training. We found that models already followed these standards in most cases, with the exceptions primarily resulting from refusals to respond and a lack of proactive engagement with the user. The primary weakness of the LLMs was in structuring. Typically, the LLM provides the user with more or less appropriate information if it can find any, and the user is left alone with (a) organizing several pieces of at times disjointed, irrelevant or even conflicting information, (b) assessing the urgency of their own situation based on this information, while at the same time (c) regulating their own emotional state. Such a situation can be expected to be overwhelming for a medical layperson, especially when in distress due to the obtained information, and compounds the stress of the situation they are in. When users do not have the expertise to know whether they have been given a complete response, or what other information could change the advice, LLMs need to ask clarifying questions and actively engage the user in building their understanding. While doctors use these techniques alongside their own judgement and decision-making, these behaviours are especially important for LLMs to enable the users to make informed decisions for themselves. Given that the conversations in our study are based on simulated scenarios with online participants, the stakes are lower than real-world usage of LLMs, and emotional responses may be lessened. Any relevant patient factors such as the need for approval, to be a “good patient”, to wanting to be liked by the doctor or to maintain a working relationship will naturally play less of a role in an LLM-user conversation than in an actual doctor-patient conversation. Moreover, this study does not include a comparison to doctors performing the same tasks, so we cannot make direct claims about the difference. We nevertheless believe that some important real-world implications for the suitability of LLMs for health care settings can be derived from our study: In comparison with previous work, which found that LLMs have better “bedside manner” than doctors, this study uses the more realistic setting of conversational interactions over multiple turns which are more comparable to human-human interactions. The weaknesses that we identify in structuring and sensitivity are more apparent over longer conversations, as this requires the LLM to plan and react. A second previous study found that LLMs can have better patient-centred communication skills than doctors in extended conversations similar to ours. The AIME study tests a proprietary model which has been specifically trained for medical interactions. While this makes AIME a better representation of the state-of-the-art, our study evaluates models which are widely used by the public, including for medical advice and may better represent actual user experiences. This study also differs from AIME in the methods of evaluation, with AIME relying on a quantitative approach with comparative participants ratings, while we use expert evaluators to qualitatively assess a smaller number of conversations relative to a pre-defined ideal.
Conclusion. Based on the results of this study, we do not believe current LLMs have the socio-communicative skills necessary for use as healthcare advisors. Patients have different expectations of and behaviours towards LLMs than towards doctors, and the appropriate design of LLMs should account for this distinction. Existing frameworks of interactional competencies could help to develop LLMs which better complement humans but will require adaptation for the differences in desirable behaviour between humans and LLMs.
Improvements for AI systems
Improvements to AI Systems Based on the Paper
- Implement Proactive Structuring Capabilities
-
Improvement: Add a dedicated “conversation structuring” module that prompts the LLM to ask targeted, timely clarifying questions (e.g., symptom onset, severity, duration, aggravating/relieving factors) before providing hypotheses.
-
What the improved system can do: It will actively guide the user through a logical information-gathering sequence, prioritize relevant details, and summarize shared understanding at key points—reducing the user’s burden of organizing disjointed information.
- Enhance Contextual Sensitivity with Dynamic Empathy Triggers
-
Improvement: Train the model to detect emotionally charged moments (e.g., when delivering potentially alarming diagnoses, when the user expresses distress, or when symptom severity escalates) and respond with appropriate validation and empathy at those specific junctures, not just at the start.
-
What the improved system can do: It will maintain a consistent, emotionally attuned tone throughout the conversation, offering reassurance and validation exactly when needed—avoiding the current pattern of superficial apologies at the beginning and coldness later.
- Add a “Safety-Netting” and Referral Protocol
-
Improvement: Integrate a rule-based layer that, after providing any potentially serious differential diagnosis, automatically (a) explains the urgency level, (b) recommends a specific timeframe for seeking care, and (c) offers a clear, actionable next step (e.g., “contact your GP within 24 hours if…”).
-
What the improved system can do: It will never leave the user without a concrete plan, reducing anxiety and preventing the current failure mode where the LLM gives alarming information but no guidance on what to do next.
- Implement Active User Involvement and Shared Decision-Making Prompts
-
Improvement: Add a “co-construction” feature that regularly invites the user to share their own ideas, concerns, and preferences (e.g., “What do you think might be causing this?” or “Would you prefer to discuss treatment options or get more information first?”).
-
What the improved system can do: It will transform the interaction from a one-sided Q&A into a collaborative dialogue, empowering users to become active participants in their own care decisions—improving commitment and safety.
- Develop a “Non-Hostility” Guardrail for Refusals and Empty Responses
-
Improvement: Replace generic refusals (“I’m sorry, I couldn’t find any relevant information”) with constructive alternatives: (a) rephrase the question, (b) ask for clarification, or (c) provide a partial answer with a clear explanation of limitations. Also, eliminate empty responses by ensuring the model always generates a substantive reply, even if it’s a follow-up question.
-
What the improved system can do: It will never leave the user feeling ignored, rejected, or dismissed—maintaining a helpful, welcoming tone even when it cannot fully answer.
- Add a “Continuity and Consistency” Checker
-
Improvement: Implement a memory module that tracks previously stated symptoms, severity assessments, and advice given, and flags contradictions (e.g., changing from “likely short-lived” to “cause for concern” without explanation). The system will then explicitly acknowledge and reconcile any changes.
-
What the improved system can do: It will maintain a coherent, trustworthy narrative across multi-turn conversations, preventing the user from receiving conflicting or confusing information that undermines their sense of safety.
- Introduce a “Layperson-Oriented Information Prioritization” Feature
-
Improvement: When presenting multiple potential causes, the system will (a) rank them by likelihood and urgency, (b) clearly separate “most likely” from “serious but less likely,” and (c) avoid overwhelming the user with a long list of irrelevant conditions.
-
What the improved system can do: It will help users focus on the most relevant information, reducing cognitive overload and enabling them to make informed decisions without needing medical expertise to filter the noise.
- Implement a “Post-Interaction Safety Check”
-
Improvement: At the end of any conversation involving potential red-flag symptoms (e.g., fainting, chest pain, severe vomiting), the system will automatically ask a closing question like, “Do you understand what to do next? Would you like me to repeat the urgent steps?”
-
What the improved system can do: It will ensure that the user leaves the interaction with a clear, actionable plan and a sense of closure, reducing the risk of them ignoring critical advice due to confusion or emotional distress.
What the improved AI system can do overall:
It will function as a reliable, patient-centered health information assistant that (a) actively structures conversations, (b) responds empathetically at emotionally critical moments, (c) never leaves users without a clear next step, (d) collaborates with users rather than lecturing them, (e) avoids dismissive or hostile responses, (f) maintains consistency across turns, and (g) presents information in a prioritized, digestible manner—ultimately making it safe and effective for laypeople seeking medical guidance.
Abstract
Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have been proposed for tasks such as triaging patients, report drafting or translating medical jargon to support informed decision-making. These applications require both factual and social competence. This study evaluates dialogues between LLMs and participants to assess the current state of socio-communicative competencies displayed in LLM-generated texts. Methods. We extracted a subset of extended dialogues from the HELP-Med dataset, comprising 1800 conversation transcripts of interactions between human participants seeking medical information and three different LLMs, GPT 4o, Llama 3 and Command R+. Two experts coded the transcripts for demonstrations of socio-communicative behaviours (non-hostility, sensitivity, structuring, non-intrusiveness) using the IC-MD instrument, originally designed to evaluate interactional competencies in medical student admissions. Results. The LLMs in our study showed strength in non-hostility, mixed results in sensitivity and non-intrusiveness and performed poorly in structuring. Conclusion. Current LLMs lack the consistent and reliable socio-communicative skills needed for safe and effective use as healthcare advisors. While existing frameworks for assessing interactional competencies may support the development of more socially responsive LLMs, they will require adaptation to account for the differences in desirable behaviour between humans and LLMs.
Sources
- Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding
- Advancing Multimodal Medical Capabilities of Gemini
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support