HarmReduction: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
cs.CL, cs.CY
Submitted: 2025-07-29
Updated: 2026-09-01
Comments: 15 pages, 5 figures, 12 tables, a dataset
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Millions of individuals' well-being are challenged by the harms of substance use.
Terminology
Abstract
Millions of individuals' well-being are challenged by the harms of substance use. Harm reduction as a public health strategy provides non-judgemental, evidence-based information intended to improve health outcomes and reduce associated safety risks. Some large language models (LLMs) have demonstrated a high level of medical reasoning, promising to address the information needs of people who use drugs (PWUD). However, their performance in relevant tasks remains largely unexplored. We introduce HarmReduction, a benchmark designed to evaluate LLMs' accuracy and safety risks in harm reduction information provision. The benchmark dataset (HR-Basic) has 2,160 question-answer-evidence pairs. The scope covers three tasks: checking safety boundaries, providing quantitative values, and inferring polysubstance use risks. We build the Instruction and RAG schemes to evaluate model behaviours based on their inherent knowledge and the integration of domain knowledge. Our results indicate that state-of-the-art LLMs still struggle to provide accurate harm reduction information, and sometimes, present severe safety risks to PWUD. This work contributes an evaluation framework for LLMs to deliver harm reduction information to avoid introducing negative health outcomes through the use of LLMs.
Sources
- Towards a Personal Health Large Language Model
- Evaluation of GPT-3.5 and GPT-4 for supporting real-world information needs in healthcare delivery
- Phi-4 Technical Report
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- GPT-4o System Card
- Large Language Models are Few-Shot Health Learners
- Capabilities of Gemini Models in Medicine
- Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity
- Positioning AI Tools to Support Online Harm Reduction Practice: Applications and Design Directions
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering