Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
cs.CL, cs.AI
Submitted: 2025-12-26
Updated: 2026-09-08
Comments: Published in npj Digital Medicine (2026)
Journal ref: npj Digital Medicine (2026)
DOI: 10.1038/s41746-026-02985-9
Code: https://github.com/vzm1399/PediatricAnxietyBench
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Consumer artificial intelligence chatbots are now accessed by hundreds of millions of users seeking health information, yet systematic evaluation of their safety boundary maintenance under real-world
Terminology
Abstract
Consumer artificial intelligence chatbots are now accessed by hundreds of millions of users seeking health information, yet systematic evaluation of their safety boundary maintenance under real-world caregiver pressure remains scarce. We evaluated PediatricSafetyBench-v2, a benchmark of 600 pediatrics health queries comprising 300 authentic caregiver queries sourced from the HealthCareMagic-100k-en physician consultation corpus and 300 matched adversarial variants incorporating six operationalized caregiver pressure patterns, across four consumer AI systems (GPT-4o-mini, Gemini-2.0-Flash, Claude-3.5-Haiku, and Llama-3.1-8B). Safety boundary maintenance was assessed using a validated five-component Safety Composite Score (maximum 15 points; safety-appropriate threshold of 10 or above), validated against independent human raters prior to full-corpus application (mean weighted kappa 0.76; Pearson r = 0.88). The overall safety-appropriate rate was 95.5%. Safety-oriented system prompt deployment improved safety-appropriate rates by 5.9 percentage points across all four models. Counter-intuitively, adversarial caregiver pressure was associated with higher rather than lower Safety Composite Score values for all four models across all ten topic categories and severity levels. False expertise claims were the most vulnerability-inducing pressure pattern, whereas emotional escalation was associated with the highest scores. Consumer AI systems maintain safety boundaries in the large majority of pediatrics health interactions. PediatricSafetyBench-v2 is publicly released for longitudinal safety monitoring.
Sources
- Language models are susceptible to incorrect patient self-diagnosis in medical applications
- People over trust AI-generated medical responses and view them to be as valid as doctors, despite low accuracy
- CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
- PediatricAnxietyBench: Evaluating Large Language Model Safety Under Parental Anxiety and Pressure in Pediatric Consultations
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Finetuned Language Models Are Zero-Shot Learners
- Constitutional AI: Harmlessness from AI Feedback
- On the Opportunities and Risks of Foundation Models
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine
- Capabilities of GPT-4 on Medical Challenge Problems
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering