Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs
Alireza A. Safaei, Laura M. Vowels, Matthew J. Vowels, Apoorv Jha, Shekoufeh Rahimi
University of Isfahan · University of Roehampton · Kivira Health
cs.CY, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost.
Terminology
Summary
The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine this relationship by combining K-Bench clinical safety scores with EcoLogits life-cycle assessment estimates across 47 supported model configurations. We evaluate model performance and environmental impact across four dimensions: energy use, carbon emissions, water consumption, and abiotic depletion. The results indicate a nonlinear trade-off at the upper end of the safety distribution: a 2.61 percentage-point increase in clinical safety score corresponded to an approximately 60-fold increase in estimated energy use per million output tokens. Row-level analyses further suggest that additional test-time compute did not consistently improve clinical safety and, in some configurations, was associated with lower clinical safety scores. These findings suggest that relying solely on larger models or additional inference-time computation may be an inefficient strategy for improving safety in therapeutic AI systems. We discuss the implications for sustainable deployment and highlight dynamic model selection, including model cascading, as a potential approach for reducing environmental impact while preserving clinical performance in higher-risk cases.
Improvements for AI systems
Improvements to AI Systems:
-
Implement Dynamic Model Cascading for Clinical Triage: Build a routing system that first uses a small, low-cost model for initial screening. Only escalate to larger, high-safety models when the small model’s confidence falls below a threshold or the query involves high-risk keywords (e.g., self-harm, crisis). This reduces energy use by up to 98% for low-risk queries while preserving safety for critical cases.
-
Add a Safety-Efficiency Penalty to Training Objectives: Modify the fine-tuning loss function to include a term that penalizes unnecessary inference-time compute. Train the model to achieve the same clinical safety score with fewer output tokens or fewer internal reasoning steps, directly optimizing for the observed nonlinear trade-off.
-
Create an Adaptive Compute Budget Controller: Integrate a meta-controller that monitors the model’s internal uncertainty (e.g., token-level logits) during generation. If the model is highly confident early in the response, it truncates further chain-of-thought or self-verification loops, cutting energy use without sacrificing accuracy.
-
Develop a Safety-Aware Model Selector with Environmental Cost Inputs: Extend the deployment framework to include real-time carbon intensity and water usage data. The selector chooses among 47 configurations not just by safety score, but by a combined utility function: maximize safety gain per gram of CO2 equivalent. This enables greener deployment during peak renewable energy hours.
-
Implement Post-Hoc Safety Verification Only for High-Risk Outputs: Instead of running full self-consistency checks on every response, use a lightweight classifier to flag potentially unsafe outputs. Only those flagged outputs trigger expensive verification passes, reducing average compute per query while maintaining a safety net.
What the Improved AI System Can Do:
-
Cut energy use per clinical query by 60–90% for typical, non-crisis interactions, while matching the safety of the largest model on high-risk cases.
-
Dynamically adapt its computational effort based on both user risk and real-time grid carbon intensity, lowering its environmental footprint without a user-visible drop in quality.
-
Provide a transparent
safety per watt
metric to clinicians, allowing them to choose deployment settings that align with institutional sustainability goals. -
Reduce the incidence of unsafe responses in edge cases by allocating more compute only where it demonstrably helps, rather than uniformly across all queries.
-
Operate on edge devices or in low-resource settings for routine mental health support, reserving cloud-based large models for rare, severe escalations.
Abstract
The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine this relationship by combining K-Bench clinical safety scores with EcoLogits life-cycle assessment estimates across 47 supported model configurations. We evaluate model performance and environmental impact across four dimensions: energy use, carbon emissions, water consumption, and abiotic depletion. The results indicate a non-linear trade-off at the upper end of the safety distribution: a 2.61 percentage-point increase in clinical safety score corresponded to an approximately 60-fold increase in estimated energy use per million output tokens. Row-level analyses further suggest that additional test-time compute did not consistently improve clinical safety and, in some configurations, was associated with lower clinical safety scores. These findings suggest that relying solely on larger models or additional inference-time computation may be an inefficient strategy for improving safety in therapeutic AI systems. We discuss the implications for sustainable deployment and highlight dynamic model selection, including model cascading, as a potential approach for reducing environmental impact while preserving clinical performance in higher-risk cases.
Sources
- Small is Sufficient: Reducing the World AI Energy Consumption Through Model Selection
- Mental Health AI Safety Claims Must Preserve Temporal Evidence
- MindBenchAI: An Actionable Platform to Evaluate the Profile and Performance of Large Language Models in a Mental Healthcare Context
- Measuring the environmental impact of delivering AI at Google Scale
- LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models
- MedDialogRubrics: A Comprehensive Benchmark and Evaluation Framework for Multi-turn Medical Consultations in Large Language Models
- Training Compute-Optimal Large Language Models
- How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference
- The Energy Cost of Reasoning: Analyzing Energy Usage in LLMs with Test-time Compute
- JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
- Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
- Linear Motility Maps in Nonlinear Viscous Fluids
- MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework