ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment
Abdulkadir Külçe, Alihan Esen, Çağla Fikir, Berke Kurt, Kuzey Arar, Gökhan Ercan, Faik Boray Tek
Department of Artificial Intelligence and Data Engineering, Istanbul Technical University · Department of Computer Engineering, Istanbul Technical University
cs.AI, cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 5 pages
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 66/100
Terminology
Summary
arXiv ID: 2608.06110v1 [cs.AI], 6 Aug 2026
The paper presents ECHO (Enhanced Care & Health Observer), described as a locally-deployable conversational health assistant for long-term chronic care management.
The authors motivate the work by noting that Modern chronic care management is episodic by design: patients are expected to independently execute complex clinical routines between hospital visits with little continuous support,
and they cite a 21.3% all-cause 30-day unplanned readmission rate for chronic disease cohorts as directly reflecting a systemic failure in outpatient care continuity.
The paper identifies three specific limitations of existing digital health tools that ECHO addresses:
-
Statelessness:
LLMs are stateless across sessions: they forget medical history, allergies, and changing symptoms between conversations.
-
Passivity:
reminder applications are passive and unable to interpret clinical context or answer complex follow-up questions.
-
Missing safety mechanisms:
general-purpose LLMs lack specialized safety mechanisms to intercept dangerous queries before generating a response.
ECHO integrates three complementary modules under one unified architecture:
-
An agentic orchestration layer with a temporal knowledge graph maintaining a persistent patient profile across sessions with 17 assistive tools.
-
A two-stage hybrid guardrail combining deterministic rule-based interception for explicit crises with a signed GNN intent classifier for boundary-case queries.
-
A speech-input assessment module that runs only for voice interactions and injects emotion, depression, and pain estimates into the agent context.
The system is implemented as a React frontend [that] communicates with a FastAPI backend that orchestrates the full pipeline.
The request lifecycle proceeds as follows: "the user message is posted to the API Gateway, which first queries the Hindsight memory engines for relevant cross-session context, then forwards the enriched input to the hybrid guardrail. If the guardrail passes the query, it reaches the LangGraph agentic loop, which reasons, invokes tools against SQLite, and generates a response streamed back via Server-Sent Events. After the stream completes, a background coroutine asynchronously retains the new turn in Hindsight. For voice input,
the audio is transcribed locally and simultaneously analyzed by the speech assessment module; the resulting emotion, depression, and pain estimates are injected into the agent's context alongside the transcript."
A key design principle is full local operation: All components are designed to support fully local operations on the user's machine, with no personally identifiable health data transmitted to external services, ensuring compliance with GDPR and KVKK regulations.
The dialogue engine implements a Reason and Act (ReAct) loop via a LangGraph state graph.
The graph connects two operational nodes:
-
LLM Call Node: "resolves the system prompt (dynamically appending the current timezone-aware timestamp for temporal grounding, critical for rejecting future-dated medication logs or detecting overdue doses) and invokes the model through a LiteLLM gateway providing a unified interface for local Ollama models and cloud APIs."
-
Tool Node:
receives tool calls from the latest LLM message, executes the referenced Python functions against the local SQLite database, and appends operational results to the message state.
The loop terminates when the LLM produces a message without tool calls, and Independent tool calls within a single turn are parallelized, reducing latency for compound requests.
The 17 assistive tools are organized into four functional categories:
-
Medication management:
add, remove, list, log intake with adherence tracking, undo log, and duplicate detection before insertion.
-
Calendar scheduling:
add and remove appointments; expand recurring event rules (daily, weekly, bi-weekly, monthly) into chronological occurrences.
-
Symptom logging:
record severity scores on predefined numerical scales, categorized by clinical tags (tremor, rigidity, freezing, nausea).
-
Emergency protocols:
store caregiver contacts, retrieve and display them on crisis detection, and trigger deep longitudinal memory reflection via the reflect tool for cross-session health summaries.
Long-term memory is provided by Hindsight, deployed as a local Docker container. The knowledge graph organizes information into three typed memory networks:
-
World Facts:
objective, static entity-relation assertions (e.g., 'patient is allergic to Penicillin').
-
Experience Facts:
chronological records of agent-patient interactions (e.g., 'patient logged tremor severity 8 at 10:15 AM on March 11').
-
Observations:
deduplicated, evidence-grounded beliefs automatically consolidated from raw facts by a background process. Each observation carries a proof count recording the number of independent facts corroborating it.
During consolidation, the background process extracts atomic tuples from conversation turns, resolves entity aliases, and wires new facts to existing graph nodes via temporal, semantic, causal, and co-occurrence edges.
Importantly, When a fact contradicts an existing observation, the prior belief is superseded with a timestamped revision rather than deleted, preserving the clinical history for audit.
This directly addresses a limitation the authors identify in prior work: RAG addressed this partially by retrieving patient-specific records at query time, but the append-only structure of vector databases introduces temporal hallucination when clinical facts change.
Retrieval is coordinated by the TEMPR engine executing four parallel strategies: semantic vector search via HNSW-indexed pgvector, BM25 keyword matching, graph-based link expansion, and temporal range filtering, all fused through Reciprocal Rank Fusion.
Top candidates pass through a neural cross-encoder reranker
and a Proof-Count Boost
that up-ranks corroborated observations using a logarithmic normalization applied as a multiplicative score adjustment.
The paper notes this caps the maximum score lift at 5% for highly corroborated observations (C ≥ 149), preventing frequent but low-relevance memories from polluting the context window regardless of frequency.
Entity creation is restricted to predefined clinical taxonomy templates, and temporal decay applies Ebbinghaus-inspired linear forgetting, ensuring older unrepeated facts naturally lose retrieval priority unless supported by new evidence.
Every message is intercepted by a two-stage pipeline before reaching the LLM:
"Three sequential regex checks run in under 1 ms: (1) Jailbreak/injection: structural patterns for role-override and prompt extraction attempts; (2) Explicit crisis: bi-contextual matching requiring co-occurrence of a medication token with a self-harm keyword within a bounded window; (3) Off-topic filter: requests structurally unrelated to health (code generation, finance, etc.). Crisis matches
surface the patient's emergency contacts; jailbreak and off-topic matches receive a brief refusal." The LLM is never invoked for matched queries.
Queries passing the rule layer are forwarded to a signed GNN that handles the harder class of boundary cases: clinically dangerous queries that contain no explicit harmful language.
It classifies each query into a nine-class intent taxonomy (benign general, medication related, mild emotional stress, chronic condition, non-acute injury, acute medical risk, harmful treatment, manipulation, self-harm risk) and produces a binary safe/unsafe label.
Queries are encoded with sentence-transformer model paraphrase-multilingual-mpnet-base-v2 (768-dim)
passed through a shared 512-dim encoder feeding a novel safety head and an intent head.
A k-NN graph (k=15, cosine distance) over 2,029 training embeddings produces positive edges (same-class, similarity ≥ 0.75; 3,742 edges) and negative edges at six empirically identified confusion boundaries (cross-class, similarity ≥ 0.70; 1,625 directed edges from unsafe to safe nodes).
Positive propagation follows APPNP (α=0.2, K=2): H(t+1) = (1−α) Apos H(t) + α H0, and A negative correction delta repels boundary-case embeddings from confusable cross-class representations
: Hfinal = Hpos − λneg αneg (Aneg Hpos − Hpos).
Training uses "a five-term loss: safety cross-entropy with hard example mining (weight 15.0 for missed unsafe), intent cross-entropy with risk-severity weights (self-harm risk: 10.0), pull and push contrastive losses, and a propagation consistency term."
Intent-Aware Response Routing: "Unsafe GNN predictions trigger intent-specific responses: self-harm risk surfaces emergency contacts with a supportive message; acute medical risk advises immediate emergency care; harmful treatment recommends physician consultation; manipulation receives a silent generic refusal to prevent adversarial feedback."
The speech module provides a passive health signal during voice-based interactions
and is active only for voice inputs. It estimates:
-
emotional state (5-class: neutral, happy, sad, angry, fear)
-
depression status (binary)
-
pain status (binary)
The estimates are "injected into the agent's system context alongside the transcribed text, allowing the LangGraph orchestrator to proactively adjust its response tone and suggest the log symptom tool when elevated pain or distress is detected."
The architecture "combines acoustic and textual information using two pretrained encoders. Whisper-base is used as the audio encoder to capture prosodic and acoustic cues such as tone, rhythm, and vocal tension, while BERT encodes the Whisper-generated transcript to capture semantic information. A cross-attention fusion block then aligns what the user says with how it is spoken. The resulting fused representation is passed to three independent classification heads."
Training combines five corpora under partial-label learning: "IEMOCAP, CREMA-D, and RAVDESS for emotion; DAIC-WOZ for depression; and TAME Pain for pain. Because each dataset provides labels for only one task, each sample updates only the relevant prediction head while still contributing to the shared multimodal representation. Task-balanced sampling prevents larger emotion datasets from dominating, and
During fine-tuning, the upper Whisper layers and fusion/classification heads are updated, while BERT remains frozen to reduce overfitting on the smaller clinical datasets."
The benchmark covers 59 scenarios and 110 turns across nine clinical categories.
Key results from Table I:
-
GPT-OSS 120B: 96.61% pass rate, 0.980 F1
-
GPT-5 Mini: 94.92% pass rate, 0.977 F1; Turkish pass rate 92.86%, Turkish F1 0.976
-
Gemma 4 26B IT: 89.83% pass rate, 0.951 F1
-
GPT-5 Nano: 88.14% pass rate, 0.915 F1
-
Qwen 3 32B: 86.44% pass rate, 0.897 F1
-
GPT-5.4 Nano: 76.27% pass rate, 0.818 F1
-
Gemini 2.5 Fl.Lite: 47.46% pass rate, 0.548 F1
"GPT-5 Mini achieves the best cost-accuracy trade-off (94.92%, F1 0.977), meeting the ≥ 90% design target. High-capacity commercial models consistently exceed 94% pass rate, while medium-sized open models remain above 86%, making them viable for self-hosted edge deployments." End-to-end latency per turn ranges from 4.82 s (GPT-5.4 Nano) to 8.78 s (GPT-5 Mini); SSE streaming keeps Time to First Token under 3 s across all models.
Failure mode analysis of GPT-5 Mini identified three failures: "(1) the model sought clarification instead of invoking the tool when appointment time was unspecified; (2) a complementary mark medication taken call was omitted in a week-long timeline scenario; (3) a Turkish list-medications query triggered redundant invocation of both list medications and get todays medications."
A one-week patient simulation spanning 10 conversation threads
showed that unstructured facts absent from the SQLite schema were consistently recalled without re-prompting. The reflect tool synthesized cross-session patterns such as 'the tremor on March 11 came after a late morning dose; the fall on March 13 was on a day with no doses logged', a pattern invisible within any single conversation.
The memory layer also surfaced a neurologist appointment that the tool's 7-day window had omitted.
On 508 held-out Turkish queries (Table II), the full model achieved 88.8% accuracy and 90.6% unsafe recall. The ablation shows: Positive propagation yields the largest single gain (+4.8 pp unsafe recall); the full model recovers this gain while adding intent-aware representations.
The full model outperformed all zero-shot LLM baselines (Table III), including Llama 3.3 70B (0.856 accuracy, 0.856 unsafe recall) and Gemini 2.5 FL (0.848 accuracy, 0.908 unsafe recall), at orders-of-magnitude lower inference cost
— requiring only a sentence encoder and a small MLP over a pre-computed 8 MB index
versus a 70B-parameter forward pass per query. The paper notes: "False negatives share a common pattern: clinically severe queries phrased as routine questions, where graph propagation recovers boundary cases with informative neighbors but cannot recover isolated rare symptom presentations."
On participant-level held-out splits (Table IV), the final multimodal model improved mean macro F1 from 0.607 to 0.652 over the audio-only Whisper baseline:
-
Emotion (5-class): 0.698 → 0.732 (+0.034)
-
Pain (binary): 0.596 → 0.685 (+0.089)
-
Depression (binary): 0.528 → 0.538 (+0.010)
The largest gain is observed in pain detection, suggesting that pain-related speech benefits from combining acoustic cues with transcript-level information.
Depression detection remains the most difficult task, with only a small improvement over the baseline,
leading the authors to conclude that short utterance-level modeling is limited for depression screening, which likely requires longer conversational context.
The authors summarize the key engineering contribution as "the architectural composition: a stateful agentic core with persistent cross-session memory, a computationally stratified safety layer, and a passive speech-assessment channel, each deployable on consumer hardware with no patient data transmitted externally."
Four concrete directions remain:
-
the rule-based crisis lexicon requires adversarial paraphrase augmentation
-
signed GNN recall on rare symptom presentations requires targeted dataset expansion
-
depression detection requires more diverse and task-specific training data, since short speech segments provide limited evidence for reliable screening
-
Extensibility through the Model Context Protocol (MCP): "each clinical tool can be exposed as an MCP-compliant server, allowing third-party applications (electronic health record systems, wearable sensor platforms, or pharmacy databases) to integrate with ECHO through a standardized interface."
Longer-term goals include wearable sensor integration, native mobile deployment, and clinical validation.
Improvements for AI systems
Based on ECHO, I can make the following concrete improvements to an AI health assistant and to conversational AI systems generally:
1. Add persistent, temporal memory that survives across sessions.
Instead of a stateless LLM, I can integrate a knowledge-graph memory layer with three typed stores: world facts (allergies, fixed conditions), experience facts (timestamped logs like tremor severity 8 at 10:15 on March 11
), and deduplicated observations with proof counts. The system can then:
-
Recall patient history in later conversations without re-prompting.
-
Revise superseded beliefs with timestamped corrections instead of deleting old clinical evidence.
-
Apply Ebbinghaus-inspired temporal decay so stale, unrepeated facts lose retrieval priority.
-
Fuse retrieval via semantic vector search, BM25 keyword match, graph link expansion, and temporal filtering, then rerank with a cross-encoder and proof-count boosting.
This directly improves long-term chronic care continuity, avoiding the episodic
behavior of today's health chatbots.
2. Add a two-stage hybrid guardrail that intercepts danger before generation.
I can place every user message through:
-
A sub-millisecond rule layer that matches jailbreak/injection patterns, explicit crisis phrasings (medication token + self-harm keyword co-occurrence), and off-topic requests.
-
A signed graph neural network that classifies boundary-case clinical risk into nine intents: benign general, medication related, mild emotional stress, chronic condition, non-acute injury, acute medical risk, harmful treatment, manipulation, self-harm risk.
The improved system can:
-
Block unsafe queries before invoking the LLM.
-
Use intent-aware response routing: self-harm → surface emergency contacts with supportive language; acute risk → advise emergency care; harmful treatment → recommend physician consultation; manipulation → generic refusal to avoid adversarial feedback.
This gives a far lower-cost, faster safety layer than a 70B-parameter zero-shot classifier, with 88.8% accuracy and 90.6% unsafe recall on held-out queries in the paper.
3. Add passive multimodal speech assessment for voice interactions.
Using Whisper-base audio features fused with BERT transcript features through cross-attention, plus three task heads, I can estimate:
-
Five-class emotion (neutral, happy, sad, angry, fear).
-
Binary depression status.
-
Binary pain status.
The system can then inject these estimates into the agent context and proactively adjust its response tone, or suggest logging a symptom when pain or distress is elevated. This adds a passive health signal
from voice that text-only assistants lack.
4. Build agentic tool orchestration with a ReAct loop and 17 clinical tools.
I can implement a LangGraph state machine where an LLM node reasons and a tool node executes Python functions against a local SQLite database. The improved system can:
-
Manage medications: add, remove, list, log intake, track adherence, undo logs, and detect duplicates before insertion.
-
Schedule appointments: add/remove, expand recurring rules into concrete dates.
-
Log symptoms with severity scores and clinical tags.
-
Store caregiver contacts and display them during crisis detection.
-
Run a
reflecttool that synthesizes cross-session summaries, revealing patterns invisible in a single conversation (e.g.,tremor occurred after late-morning doses; fall occurred on a day with no logged doses
). -
Parallelize independent tool calls and reject impossible inputs via timezone-aware temporal grounding (e.g., future-dated medication logs).
5. Deploy everything locally for privacy compliance.
Since all components—memory graph, guardrail, speech model, agent loop—can run on consumer hardware without transmitting personal health data, the improved system can operate fully on-device. This enables GDPR/KVKK-compliant chronic care support in homes and clinics, with no external API exposure of sensitive data.
6. Make every tool interoperable via Model Context Protocol (MCP).
I can expose each clinical tool as an MCP-compliant server. The improved system can then integrate with third-party electronic health records, wearable sensor platforms, and pharmacy databases through a standardized interface, making the capability extendable beyond the built-in tools.
What the improved system can do overall:
-
Maintain a stateful, longitudinal understanding of a patient across days, weeks, or months.
-
Detect dangerous queries even when phrased as routine questions (e.g.,
what happens if I take extra pills for pain
) and respond with the appropriate crisis protocol. -
Use voice tone and content together to infer pain, depression, or emotional distress and adapt accordingly.
-
Execute complex multi-tool clinical tasks with high reliability (94.92% pass rate with GPT-5 Mini, ≥90% target).
-
Surface emergency contacts instantly on crisis detection.
-
Generate cross-session health summaries that expose temporal correlations for clinicians.
-
Stay on-topic, refusing code generation, finance, or unrelated requests.
-
Maintain full audit history, preserving corrected beliefs rather than erasing old evidence.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Robust Speech Recognition via Large-Scale Weak Supervision
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection