Behavioral Coherence: A Method for Sensitive-Domain LLM Evaluation
cs.AI, cs.HC
Submitted: 2025-12-15
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: People use LLMs for reproductive-health questions, including abortion-related support.
Terminology
Abstract
People use LLMs for reproductive-health questions, including abortion-related support. A response can sound supportive while answers reinforce harmful assumptions: judgment is likely, secrecy is safer, and support is limited. We introduce behavioral coherence evaluation, a design-time method that uses validation evidence from an established instrument to test relations among outputs. Using the Individual Level Abortion Stigma Scale, we prompted five LLMs to complete questionnaires for 627 personas and reviewed flagged patterns with five reproductive-health experts. Models scored personas lower on self-judgment but higher on worries about judgment; most made worries the highest-scoring dimension, although it was lowest in the ILAS reference sample. Four of five models reversed the reference direction by generating significantly higher worries about judgment scores for Black personas. Models defaulted to extreme secrecy after abortion despite varying stigma patterns across personas. Expert review showed that disclosure guidance requires context about relationship safety, legal risk, and trusted support.
Sources
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Clinical knowledge in LLMs does not translate to human interactions
- Words Matter: Reducing Stigma in Online Conversations about Substance Use with Large Language Models
- Reasoning Models Don't Always Say What They Think
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma
- HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
- Inducing Positive Perspectives with Text Reframing
- Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
- An Approach to Technical AGI Safety and Security
- GPT-4's One-Dimensional Mapping of Morality: How the Accuracy of Country-Estimates Depends on Moral Domain
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection