BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: Accepted to EMNLP 2026 Main Conference
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mental health assessment relies on episodic self-report scales, which convert subjective states such as stress into numerical scores but provide only sparse snapshots of wellbeing.
Terminology
Abstract
Mental health assessment relies on episodic self-report scales, which convert subjective states such as stress into numerical scores but provide only sparse snapshots of wellbeing. Wearable devices offer longitudinal behavioral and physiological signals for continuous, low-burden monitoring. Recent LLM-driven personal-health agents enable natural language queries over wearable signals, but mainly handle short-term, retrieval-based lookups (e.g., highest step count over a week). They do not evaluate whether agents can reason over long-term signals to predict wellbeing scores paired with evidence-grounded rationales. To address this gap, we introduce BALMS, the first systematic benchmark of LLM-based agentic systems for longitudinal mental health sensing. BALMS spans 3 real-world longitudinal datasets, 2 task families (closed-form wellbeing-score prediction and rationale generation auto-graded by an LLM-as-Judge), 3 agentic paradigms evaluated across 5 open- and closed-source LLM backbones. We find that zero-shot agents rarely outperform a simple mean baseline, except with stronger backbones or compact, semantically meaningful features. Chain-of-thought prompting improves reasoning-oriented backbones, but does not guarantee temporal grounding or numerical correctness. Together with more analysis on efficiency and temporal scaling, BALMS highlights the need for longitudinal mental health agents that selectively retrieve history, ground temporal evidence, and reason over interpretable behavioral features.
Sources
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Towards a Personal Health Large Language Model
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
- From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language Models
- The Anatomy of a Personal Health Agent
- Mistral 7B
- MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
- An AI Co-Data-Scientist for Prioritizing Candidate Biomarkers from Wearable Sensor Data
- Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals
- Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting
- AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
- LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning
- Wearable Foundation Models Should Go Beyond Static Encoders
- Qwen2.5 Technical Report
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering