Agentic AI for Scaling Diagnosis and Care in Neurodegenerative Disease
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Agentic AI for Scaling Diagnosis and Care in Neurodegenerative Disease".
Jane: The paper was written by Andrew G. Breithaupt, Michael Weiner, Alice Tang, Katherine L. Possin, Marina Sirota et al. from Goizueta Brain Health Institute, Emory University and Department of Radiology and Biomedical Imaging, University of California, San Francisco and School of Medicine, University of California, San Francisco and Bakar Computational Health Sciences Institute, University of California, San Francisco and Memory and Aging Center, Department of Neurology, University of California, San Francisco and Department of Neurology, Weill Institute for Neuroscience, University of California, San Francisco and NSF AI Institute for Advances in Optimization (AI4OPT), Georgia Institute of Technology and Department of Neuroscience, University of California, Berkeley and Division of Biostatistics, University of California, Berkeley and Center for Intelligent Imaging (ci2), Department of Radiology and Biomedical Imaging, University of California, San Francisco and Department of Psychiatry and Behavioral Sciences, University of California, San Francisco.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's got the whole team buzzing — it's called "Agentic AI for Scaling Diagnosis and Care in Neurodegenerative Disease." Jane, this one feels personal, doesn't it?
Jane: It really does, Tom. This is about Alzheimer's and related dementias, and the paper opens with some sobering numbers. The U.S. is facing a massive shortage of neurologists, and by two thousand sixty we're projected to see a million new dementia cases a year. Meanwhile, over half of dementia diagnoses in primary care are delayed until moderate or advanced stages.
Tom: And the wait times are brutal. The paper cites projections that by two thousand twenty-seven the average wait to see a dementia specialist could exceed forty months. Rural areas face even longer delays — three times worse.
Jane: Right, and that's the crisis this paper is trying to address. The authors come from UCSF, Emory, Berkeley, Georgia Tech — a real who's who in dementia research and AI. The lead authors include people from the Memory and Aging Center at UCSF, which is one of the premier institutions for this kind of work.
Tom: So what's their big idea? They're proposing something called agentic AI systems — and I know that sounds like jargon, but it's actually pretty intuitive once you break it down.
Jane: Exactly. Think of it this way: a regular AI chatbot can answer questions, but an agentic AI system can actually do things. It can pull data from electronic health records, run analyses, search medical literature, and coordinate all those steps to help a clinician make a diagnosis. It's like having a really smart assistant who doesn't just talk — they act.
Tom: And the key word here is "assist." The paper is very careful to say this is about augmenting clinicians, not replacing them. They call it keeping the human in the loop, which matters a lot when you're dealing with vulnerable patients who may have declining decision-making capacity.
Jane: That's the ethical backbone of the whole paper. They're proposing a six-phase roadmap — starting with standardized data collection, then building AI decision support, integrating it into clinical workflows, validating it rigorously, enabling continuous learning, and wrapping it all in ethics and risk management.
Tom: Six phases. That's ambitious. But before we get into the weeds of the roadmap, I want to ask Lu — you've been quiet. What's your first reaction to the title alone?
Lu: Honestly, Tom, I think the title undersells it. "Scaling diagnosis and care" sounds incremental, but this is really about rethinking how specialty medicine works. The authors are saying we can't train enough neurologists fast enough, so we need to build systems that multiply the expertise we already have. That's a fundamental shift in how we deliver healthcare.
Tom: Lu's right — this isn't just another AI paper. It's a blueprint for changing the system. And we're going to dig into that blueprint over the next few segments. Stay with us.
Summary of the Paper: Jane: So we're back with "Agentic AI for Scaling Diagnosis and Care in Neurodegenerative Disease," and I want to walk through what the paper actually proposes. Tom, you mentioned the six phases — let's talk about the first two, because they set the foundation.
Tom: Phase one is data collection, and the paper makes a really practical point: before any AI can help, you need good data. But right now, patient histories are collected inconsistently — one doctor asks different questions than another, and the notes are all over the place. So they're proposing AI-powered tools that standardize this from the start.
Jane: And they've got some concrete examples. Voice-based conversational agents that can take a patient history over the phone — imagine an AI that interviews a patient about their memory problems, their sleep, their mood, and does it in a way that's empathetic and adapts to the person's education level or language. That's not science fiction; the paper says these systems already exist and are being tested.
Lu: What excites me is the integration piece. They're talking about digital cognitive assessments that capture trial-level data — response times, error patterns — not just a summary score. That's a goldmine for AI. A traditional paper test gives you one number; a digital test gives you a rich dataset that can reveal subtle patterns.
Tom: And then phase two is decision support. This is where the AI actually helps interpret all that data. The paper describes a system that can look at a patient's history, cognitive testing, brain MRI, and blood biomarkers, and then generate a differential diagnosis with explanations the clinician can verify.
Jane: That's the part I find most compelling — the explanations. The authors emphasize that clinicians won't trust a black box that just says "this patient has Alzheimer's." They need to see the reasoning. So the system is designed to mimic how a clinician thinks, presenting evidence in familiar clinical terms.
Meng: Can I jump in here? From an engineering standpoint, the multimodal integration is the hard part. You're combining free-text clinical notes, structured lab results, MRI images, and speech patterns from the patient interview. Each of those is a different data type that needs different processing. The paper actually describes a system that converts imaging features into textual summaries that language models can reason over — that's a clever workaround.
Lu: And they cite real examples. There's a system called MAI-DxO from Microsoft that does sequential diagnostic reasoning, and Google's multimodal AMIE for consultations. So this isn't purely theoretical — the building blocks exist.
Tom: But here's the thing that struck me — the paper is honest about the gap between what works in a research setting and what works in a real clinic. They mention that most generative AI diagnostic studies use simulated data, not real patients. So there's a big validation gap they're trying to close.
Jane: And that's exactly what phase four is about — validation and monitoring. But before we get there, we need to talk about how this actually fits into a clinician's day. That's the workflow piece, and I think it's where the paper gets really interesting. We'll dig into that next.
Improvements Suggested by the Paper: Tom: Welcome back. We're still on "Agentic AI for Scaling Diagnosis and Care in Neurodegenerative Disease," and Jane just teed up the workflow question. This is where the paper makes some really concrete suggestions about improving how clinicians work.
Jane: Right, and the key insight is that this system should save time, not add to the burden. The paper talks about using AI to collect the patient history before the visit even happens — over the phone or through a conversational agent. Then when the patient sees the doctor, the history is already summarized, the cognitive assessment is already done, and the doctor can spend the visit actually talking with the patient.
Tom: That's a radical shift. Instead of the doctor spending twenty minutes typing notes while the patient talks, they can focus on building rapport and making shared decisions. The paper calls this "shifting clinic visit time from data collection to meaningful tasks."
Meng: But I want to push back on something. The paper mentions electronic consults — e-consults — as a near-term implementation pathway. That's where a primary care doctor sends a case to a specialist electronically instead of referring the patient. The idea is that the AI system provides enough data that the specialist can answer more cases remotely. But that only works if the AI is trustworthy enough for the specialist to rely on it.
Jane: That's a fair point, Meng. And the paper addresses it through their validation framework — they call it FAVES, which stands for Fair, Appropriate, Valid, Effective, and Safe. They're proposing that models be tested against diverse, specialist-confirmed cases, not just neuropathology, because neuropathology cases tend to come from affluent, well-educated patients.
Lu: The continuous learning piece is what gets me excited. The paper describes a system where specialists use the AI to review cases, and every time they do, they're implicitly validating or correcting the AI. Over time, the system learns from that feedback. It's like the AI gets better every time a specialist uses it, and that improvement benefits every other clinician on the system.
Tom: That's the flywheel effect — the more it's used, the smarter it gets. But they're careful to say that learning has to be supervised. You can't just let the AI update itself based on any case. They emphasize careful case selection and human oversight to prevent what they call "model degradation."
Meng: And there's an economic angle here too. The paper acknowledges that the return on investment is uncertain. But they point out that ambient scribes — AI tools that automatically document visits — have already been adopted widely because they reduce clinician burnout. So there's precedent for healthcare systems investing in AI even without clear reimbursement.
Lalam: If I may add a perspective — the improvements here go beyond efficiency. This paper is about equity. By making specialist-level assessment available in primary care and rural settings, you're addressing the fact that minority populations face greater diagnostic delays. The authors explicitly call out the need to mitigate bias in these systems and ensure they work across languages and educational backgrounds.
Jane: That's a crucial point, Lalam. The paper isn't just about making a fancy tool — it's about making sure the tool works for everyone, not just the people who can already access good care. And that brings us to the ethics and risk management piece, which is where we'll wrap up.
Conclusion: Tom: We've covered a lot of ground on "Agentic AI for Scaling Diagnosis and Care in Neurodegenerative Disease." Let's pull it together. Jane, what's the big picture?
Jane: The big picture is that we have a crisis — not enough neurologists, too many dementia patients, and diagnoses coming too late. This paper proposes a six-phase roadmap to build AI systems that help clinicians work at specialist level, even when they're not specialists. It starts with standardized data collection, moves through AI decision support, integrates into workflows, validates rigorously, learns continuously, and wraps everything in ethics.
Tom: And the through-line is human oversight. Every phase keeps clinicians in the loop. The AI collects data, but a clinician reviews it. The AI suggests a diagnosis, but a clinician verifies it. The AI learns from cases, but specialists curate which cases it learns from.
Lu: What I'll remember is the vision of a continuously learning healthcare system. This isn't a static tool — it's a system that improves with every patient encounter and incorporates the latest research in real time. That's a fundamental shift from how medicine works today.
Meng: And from a practical standpoint, the near-term wins are real. E-consults, ambient scribes, digital cognitive assessments — these exist today. The paper gives a realistic path from where we are to where we need to be.
Lalam: The cultural impact is significant as well. This paper models how AI can serve vulnerable populations with dignity — respecting patient autonomy, ensuring transparency, and building accountability mechanisms. It sets a standard for how medical AI should be developed, not just for dementia but for all of medicine.
Tom: Well said. We've covered the title, the summary, the improvements, and the implications of "Agentic AI for Scaling Diagnosis and Care in Neurodegenerative Disease." It's a roadmap paper, but it's grounded in real systems and real challenges. Thanks to Lu, Meng, and Lalam for joining the conversation.
Jane: And thanks to our listeners. This is one of those papers that could genuinely change how we deliver care to millions of people. We'll be back soon with the next paper — until then, take care.
Tom: Goodbye, everyone.
Andrew G. Breithaupt, Michael Weiner, Alice Tang, Katherine L. Possin, Marina Sirota, James Lah, Allan I. Levey, Pascal Van Hentenryck, Reza Zandehshahvar, Marilu Luisa Gorno-Tempini, Joseph Giorgio, Jingshen Wang, Andreas M. Rauschecker, Howard J. Rosen, Rachel L. Nosheny, Bruce L. Miller, Pedro Pinheiro-Chagas
Goizueta Brain Health Institute, Emory University · Department of Radiology and Biomedical Imaging, University of California, San Francisco · School of Medicine, University of California, San Francisco · Bakar Computational Health Sciences Institute, University of California, San Francisco · Memory and Aging Center, Department of Neurology, University of California, San Francisco · Department of Neurology, Weill Institute for Neuroscience, University of California, San Francisco · NSF AI Institute for Advances in Optimization (AI4OPT), Georgia Institute of Technology · Department of Neuroscience, University of California, Berkeley · Division of Biostatistics, University of California, Berkeley · Center for Intelligent Imaging (ci2), Department of Radiology and Biomedical Imaging, University of California, San Francisco · Department of Psychiatry and Behavioral Sciences, University of California, San Francisco
cs.CY, cs.AI
Submitted: 2025-12-23
Updated: 2026-08-18
Comments: 28 pages, 2 figures, 1 table, 1 box
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 72/100
Key concepts
- Agentic AI systems
- These are AI systems that can perform actions beyond answering questions. Unlike chatbots, they can pull data from electronic health records and run analyses to help a clinician make a diagnosis by coordinating multiple steps.
- Six-phase roadmap
- The paper proposes a structured plan for implementing agentic AI. The phases include standardized data collection, building AI decision support, integrating into clinical workflows, rigorous validation, enabling continuous learning, and wrapping everything in ethics and risk management.
- Decision support with explanations
- The paper emphasizes that the AI must not just give a diagnosis but provide reasoning. Clinicians need to see the evidence presented in familiar clinical terms so they can trust the system and verify its suggestions.
- FAVES framework
- This is a validation framework proposed by the authors, standing for Fair, Appropriate, Valid, Effective, and Safe. It suggests testing AI models against diverse cases confirmed by specialists rather than just neuropathology cases.
Terminology
Summary
Affiliations: Goizueta Brain Health Institute, Emory University; Department of Radiology and Biomedical Imaging, UCSF; School of Medicine, UCSF; Bakar Computational Health Sciences Institute, UCSF; Memory and Aging Center, Department of Neurology, UCSF; Department of Neurology, Weill Institute for Neuroscience, UCSF; NSF AI Institute for Advances in Optimization (AI4OPT), Georgia Institute of Technology; Department of Neuroscience, UC Berkeley; Division of Biostatistics, UC Berkeley; Center for Intelligent Imaging (ci2), UCSF; Department of Psychiatry and Behavioral Sciences, UCSF
"United States healthcare systems are struggling to meet the growing demand for neurological care, particularly in Alzheimer's disease and related dementias (ADRD). Generative AI built on language models (LLMs) now enables agentic AI systems that can enhance clinician capabilities to approach specialist-level assessment and decision-making in ADRD care at scale. This article presents a comprehensive six-phase roadmap for responsible design and integration of such systems into ADRD care: (1) high-quality standardized data collection across modalities; (2) decision support; (3) clinical integration enhancing workflows; (4) rigorous validation and monitoring protocols; (5) continuous learning through clinical feedback; and (6) robust ethics and risk management frameworks. This human centered approach optimizes clinicians' capabilities in comprehensive data collection, interpretation of complex clinical information, and timely application of relevant medical knowledge while prioritizing patient safety, healthcare equity, and transparency. Though focused on ADRD, these principles offer broad applicability across medical specialties facing similar systemic challenges."
"The gap between the demand for neurological care and the number of neurology providers continues to grow, leading to long wait times, delayed diagnoses, and increased emergency department visits. Neurological diagnoses are difficult to make for a wide variety of reasons ranging from the great variability and complexity in patient presentations to the extensive time and knowledge required to effectively collect and interpret necessary data. These challenges are amplified by limited office visit time, exponentially growing data complexity, and an overwhelming volume of medical literature and research opportunities that exceed any individual clinician's capacity to master."
"Simultaneously, artificial intelligence (AI) is rapidly advancing and has entered a new era through generative AI built on large language models (LLMs). A growing number of studies have demonstrated that LLM-based systems encode clinical knowledge, can achieve expert-level medical question answering, integrate multimodal patient data, are capable of empathetic history taking through conversational agents, and enhance clinical decision-making. LLMs have enabled agentic AI systems, architectures in which one or more LLMs coordinate interacting components such as traditional discriminative AI models and tools to perform complex, multistep tasks."
"While much research has focused on AI's potential to identify patterns beyond human perception (predictive AI), clinical implementation remains challenging because clinicians cannot readily verify insights they cannot detect, nor can these systems naturally complement clinical workflows. However, agentic AI systems offer more immediately practical applications by enhancing clinicians' capabilities in three critical areas where healthcare providers increasingly struggle: comprehensive data collection, interpretation of increasingly complex multimodal information, and timely, context appropriate application of relevant medical knowledge."
"Currently, we can deploy domain-specific AI assistants for individual functions, each optimized for its particular role in data collection, interpretation, or knowledge retrieval; however, the future vision involves an integrated agentic AI system capable of seamlessly coordinating these functions within a continuously learning healthcare framework. Such systems would leverage every patient encounter while incorporating the latest clinical evidence."
"Still, such a learning system must be intentionally molded to benefit diverse patient populations, support both patients and providers without imposing new burdens, and be safely implemented in any clinic setting. Recent concerns around generative and agentic AI have emphasized the need for strict boundaries around autonomy and for maintaining robust human oversight in high-stakes medical contexts. This is particularly important in ADRD, where vulnerable populations, shifting decision-making capacity, and high-stakes diagnostic uncertainty could magnify risk. Therefore, we advocate for constrained agentic AI systems in which clinicians remain in the loop, supported by rigorous evaluation and continuous monitoring using clinically meaningful outcomes."
"In this perspective article, we provide a roadmap for how responsible agentic AI integration can be accomplished in the field of ADRD in the United States and scale specialist-level clinical care. Our roadmap consists of six interconnected phases: (1) high-quality standardized data collection, (2) AI decision support development, (3) clinical integration and workflow, (4) validation and monitoring, (5) continuous learning, and (6) ethics and risk management. While our roadmap presents clinical integration as a distinct phase, implementation success requires designing all components with clinical workflows in mind from the outset."
"Agentic AI refers to compound systems in which large language models (LLMs) and other ML models are orchestrated through scaffolding architectures that enable code execution, search, memory, and iterative refinement, fundamentally differing from single-model approaches. While single LLMs have demonstrated impressive performance on medical licensing examinations and diagnostic benchmarks, they are not able to execute diverse tasks in the real world."
The paper identifies several critical limitations of single LLMs that agentic approaches can address:
-
LLMs exhibit 'jagged intelligence' with gaps between strengths and weaknesses; agentic scaffolding can close these gaps through systematic planning, context engineering, and verification.
-
Hallucinations persist despite improvements, particularly when models lack empirical grounding; agentic systems can enforce scientific grounding by retrieving and cross-checking outputs against trusted medical sources.
-
"Single LLMs lack mechanisms for continual learning, causing their knowledge to lag behind evolving medical evidence; agents can access up-to-date external sources such as FDA approvals, clinical trials, and treatment guidelines at inference time."
-
"Limited context windows and sample inefficiency restrict how much clinical information models can use; agentic architectures can address these constraints by integrating massive, compressed external knowledge stored in vector databases and knowledge graphs through retrieval-augmented generation (RAG)."
-
LLMs lack calibrated statistical metrics to quantify uncertainty during prediction, limiting trust; compound systems can integrate traditional machine learning models that provide quantitative, interpretable outputs.
"As the U.S. population ages, the number of people with Alzheimer's disease and related dementias (ADRD) is rapidly increasing, with US cases projected to double by 2060 to 1 million annually. Early detection and precise diagnosis are crucial for proper access to quality care, eligibility for new and emerging disease modifying therapies, and enabling research to advance the field. Unfortunately, more than 50% of dementia diagnoses are delayed until moderate or advanced stages in primary care with greater delays among racial and ethnic minorities."
"There are many challenges with diagnosing neurodegenerative diseases. Patients usually go to their primary care providers (PCPs) first, but most PCPs cite lack of time and confidence with ADRD diagnosis. Referrals to dementia specialists face significant barriers as the demand greatly exceeds the available specialist supply with the gap continuing to widen, and the average wait times projected to exceed 40 months by 2027 with rural areas facing three times longer delays. Beyond access issues, diagnosis is challenging, requiring time-consuming data collection and nuanced interpretation."
"After diagnosis, challenges with management are growing as new therapeutics with intensive requirements for monitoring are emerging. Many PCPs and general neurologists struggle to keep pace with the rapid advances in the field, which includes maintaining awareness of new treatments, diagnostics, and research studies that could potentially benefit their patients and the ADRD field. Furthermore, new care models such as those proposed by Centers for Medicare and Medicaid Services (CMS) demand additional resources that our current workforce is ill-equipped to support."
Data collection should first prioritize what clinicians can interpret and are currently using to diagnose neurodegenerative disease in both the clinic and research settings.
The paper examines key data types: the patient's history, physical examination, neuropsychological testing, neuroimaging, and biomarkers.
"Recent advances in LLM capabilities have made comprehensive remote history collection with voice-based conversational agents feasible for the first time. This includes natural conversational interaction that can match or exceed clinician efficiency and expressions of empathy, engage with persons living with dementia, and adapt to various educational backgrounds and languages."
"Clinicians, faced with limited time, struggle to gather and document comprehensive patient histories from distressed patients or their caregivers, leading to redundant interviews across clinicians and potentially delaying diagnosis. Moreover, incomplete documentation restricts the valuable clinical data available for AI-driven decision support systems. Voice-based conversational agents powered by LLMs could transform history collection by enabling these interviews to be conducted over the phone outside time-limited clinic visits, increasing healthcare access for patients who are of lower socioeconomic status, have limited English proficiency, and/or have low digital literacy."
"There are several considerations for the efficacy and safety of interactions between patients and conversational agents. Interviews will need to be semi-structured as LLMs could engage in irrelevant or dangerous topics. Continuous real-time monitoring systems are needed to analyze the conversation for content relevance and for dangerous content such as expressions of suicidal ideation, automatically triggering human intervention. A hybrid model that combines automated conversational agents with timely clinician oversight and decision-making is strongly advocated. Bias will also need to be identified and mitigated (e.g., these tools can have worse performance with non-English speaking people)."
"Most clinicians would not realistically adopt a system requiring them to read a lengthy transcript, but LLM summarization capabilities enable workflow integration by converting these interviews into a history of present illness (HPI). Reliable summarization is especially critical in cognitive impairment contexts, where accurate documentation of subtle symptom changes over time is essential for early diagnosis and management."
"When concerning symptoms are identified, a cognitive assessment with objective measures of cognitive impairment is required. Cognitive assessments should be scalable, without requiring highly trained staff, and designed to identify relative weaknesses of specific cognitive and behavioral domains to enable identification of patterns helpful with diagnosis, going beyond brief screening assessments. Traditional assessments are time-consuming, require specialized training, and often capture only summary scores rather than the detailed, domain-specific performance data valuable for AI interpretation."
"Digital cognitive assessments will be necessary to enable a comprehensive AI system for ADRD, and they address the above limitations by capturing nuances in performance (trial-level data, response times, error patterns) unavailable with paper-based tests. Their ability to integrate directly into the clinician's workflow within the EHR has driven adoption, as shown with the TABCat-BHA. The digital cognitive assessment landscape is advancing rapidly, with clinically available tools continually improving in accuracy, accessibility, and ecological validity. Virtual reality assessments represent one particularly promising direction, with emerging LLM integration enabling virtual psychometrists that could fundamentally transform how cognitive evaluations are conducted."
"Note that gathering the history and cognitive assessment ahead of the visit could free up time with the clinician, enabling a conversation designed to build rapport and shared decision making that a clinician would not otherwise have the time to do well."
"Physical examination continues to be important in obtaining a diagnosis and the motor findings associated with Lewy body disease, corticobasal syndrome, progressive supranuclear palsy, amyotrophic lateral sclerosis, cerebrovascular disease and many other conditions. In addition, the absence of motor findings in AD is also important. While an in-person physical exam is ideal, physical exam findings for these syndromes can often be uncovered through video visits which can increase accessibility for patients who cannot feasibly travel to see a specialist. Possibilities include video interpretation of speech and speech patterns, motor planning and execution, gait, balance, and others. New multimodal AI models have made significant progress in describing video content such as surgeries, making video data a promising method for AI decision support with further research."
"Neuroimaging with MRI and Positron Emission Tomography (PET) is a diagnostic cornerstone, aiding in exclusion of treatable causes of dementia, assessing vascular etiologies, and observing atrophy patterns (MRI) or biomarker location (PET). MRI is becoming more accessible around the country with several creative solutions, so we do not see this as a notable bottleneck for the future of AI decision support."
"Biomarkers are now a valuable component of the diagnostic process. While amyloid PET and cerebrospinal fluid (CSF) biomarkers for Alzheimer's disease have not been easily scalable or accessible, multiple plasma-based biomarkers are becoming available in the clinic setting. Assays of cerebrospinal fluid and more recently skin biopsies are promising for detection of alpha synuclein, which is a biomarker of Parkinsons' disease, Dementia with Lewy bodies, and multiple system atrophy. The implementation and real world accuracy of these biomarkers is highly promising. Experts in the field have emphasized that accurate interpretation of these biomarkers still requires clinical context, which is a concept that can be extended to any data type. Therefore, an AI-system must integrate data across modalities with unique patient characteristics (e.g., age and kidney function significantly alter interpretation of new plasma AD biomarkers) to inform diagnosis and management."
"Standardized data collection is the essential first step toward a timely diagnosis, enabling both personalized management plans and meaningful research contributions. Methods must be easy for patients and their caregivers as well as clinicians, as this can increase access to care and ensure data is from diverse patient populations and clinic settings. However, achieving global standardization at the point of care is extremely challenging in diverse, resource-constrained settings. Therefore, we advocate for an 'AI-first' data infrastructure approach. Rather than imposing rigid collection protocols on frontline clinicians, this approach prioritizes backend curation and harmonization layers that can transform heterogeneous real-world clinical data into AI-ready formats."
"Both existing and future datasets require proper curation according to FAIR principles (Findable, Accessible, Interoperable, and Reusable), including comprehensive metadata, clear governance frameworks, and quality control procedures. National organizations such as the NIH, CMS, advocacy organizations, and other entities could potentially facilitate standards development (e.g., Digital Imaging and Communications in Medicine (DICOM) for neuroimaging) and monitoring."
"The most challenging aspect of a scalable AI system in the US healthcare system will likely be multimodal data integration across diverse and complex EHR systems while maintaining strict HIPAA compliance and data protection standards."
The paper discusses several technical approaches: "The Substitutable Medical Apps, Reusable Technology (SMART) on Fast Healthcare Interoperability Resources (FHIR) framework enables applications to read and write standardized health data across vendors, with major EHR systems like Epic and Oracle already supporting robust integration ecosystems. Real-world examples include
UCSF's BRIDGE dashboards [which] integrate multimodal sources across any SMART on FHIR compatible EHR and
The Observational Medical Outcomes Partnership (OMOP) Common Data Model extension for imaging (MI-CDM) [which] was used to convert 1 million DICOM series from an Alzheimer's neuroimaging cohort into OMOP tables alongside clinical EHR data."
"LLMs are likely to accelerate interoperability further. Recent work demonstrated an LLM translating free-text EHR notes into FHIR with impressive accuracy, suggesting that LLMs may enable interoperability driven less by rigid schemas and more by flexible 'semantic standards,' dynamically reshaping data into whatever structure downstream systems require."
"Collectively, these advances demonstrate that multimodal EHR data integration is possible at institutional scales. However, current implementations remain limited to highly motivated, well-resourced institutions, with widespread adoption across diverse healthcare settings not yet realistic. Nevertheless, rapid policy and technical advances in interoperability create urgency to pilot AI systems integrating multimodal data now, positioning healthcare systems for broader deployment as these capabilities mature. Furthermore, to scale these solutions beyond research cohorts into routine care, we must address barriers that are not solely technical, but largely issues of governance. Divergent policies on data access and privacy across institutions create silos that hinder the training of robust AI models. An AI-ready framework requires not only technical interoperability (e.g., handling different data formats) but also semantic alignment and dynamic governance protocols that can manage consent and privacy across borders and institutions."
"We envision that AI collecting data in a consistent and standardized manner (i.e., history, cognitive assessment, physical exam) will serve as a foundation for reducing errors and enhancing clinician-patient connection by shifting clinic visit time from data collection/documentation to more meaningful tasks: reviewing findings with patients, answering questions, and engaging in shared decision-making about care preferences and values. This data can be presented alongside already standardized neuroimaging and biomarkers in an explainable manner."
"This infrastructure enables both immediate diagnostic capabilities and continuous learning, which will allow the AI system to adapt as diagnostic tools evolve (e.g., streamlining patient history as biomarkers become more accurate). Practically, implementation will likely begin with larger healthcare systems with the capacity for multimodal data sharing. Smaller systems will require demonstrated value before participating in voluntary data exchange programs such as the Trusted Exchange Framework and Common Agreement (TEFCA), a government initiative designed to facilitate nationwide health information exchange."
"High-quality data collection and standardization outlined above enables both the clinician and the AI system to efficiently access, integrate, and interpret multimodal patient data, avoiding the risk of false positives from relying on one test result. Our perspective on AI decision support centers on augmenting rather than replacing clinician capabilities by highlighting important diagnostic features from this complex multimodal data, explaining its reasoning transparently. In this way, AI can be a collaborative partner, rather than a passive tool."
"Despite successes with disease classification tasks and their interpretability, traditional machine learning approaches (e.g., logistic regression, support vector machines, and random forests) face significant challenges with the multi-modal, heterogeneous datasets in ADRD research. These methods require extensive manual feature engineering to extract structured representations from diverse data types, ranging from MRIs and PET scans to tabular EHRs and unstructured clinical notes. This introduces potential biases and information loss as well as limits scalability and adaptability."
"Deep learning architectures, particularly convolutional neural networks (CNNs), have enabled remarkable precision in identifying disease-specific features from neuroimaging, such as atrophy patterns or cortical thinning. Recent multimodal AI models incorporating clinical, laboratory, and imaging data hold the most promise, with recent successful multimodal models being able to differentiate neurodegenerative subtypes or stages of disease with high accuracy. Although commercial solutions exist for AI-based regional normalized brain volumetry, it remains unclear how much such methods are affected by image acquisition parameters and how such data are best incorporated into a diagnosis for an individual patient."
"Despite advances, deep learning models face significant implementation challenges. These include substantial data requirements, overfitting risk with small ADRD datasets, and sometimes limited interpretability. This lack of transparency complicates clinical adoption by undermining provider trust, highlighting the need for robust validation protocols, diverse training datasets, and user interfaces that translate complex model inferences into clinically meaningful explanations."
"Recent advancements in LLMs may introduce a paradigm shift in ADRD care and research, offering novel approaches to clinical data interpretation and diagnostic uncertainty reduction. Unlike traditional discriminative models constrained by rigid classification frameworks, LLMs process extensive natural language inputs, including clinical documentation and patient-generated speech, while generating probabilistic assessments from incomplete or ambiguous clinical data. They have demonstrated impressive medical knowledge that could serve as clinical decision support. For example, GPT-4 exceeded mean human performance on a neurology board-style examination, and can provide neuropathologic differential diagnoses similar to experts."
"Successful integration of AI in ADRD necessitates a complementary approach utilizing both discriminative models for distinguishing well-defined populations and generative models for managing diagnostic uncertainty and hypothesis generation. This combination can be accomplished through an agentic AI system which enhances both the accuracy and generalizability of decision support tools."
"Agentic AI systems can address a fundamental healthcare challenge: delivering the right information to the right patient at the right time. In practice, the clinician could review AI generated insights about the patient before the appointment, including a differential diagnosis with explanations and then provide recommendations for next steps in workup and management based on the latest evidence-based guidelines. This need is growing with the complexity of emerging therapies, particularly anti-amyloid therapies with numerous inclusion/exclusion criteria and safety requirements making patient selection difficult even for specialists. AI systems can systematically evaluate these complex criteria, reducing errors and helping identify suitable candidates for specific interventions earlier. Such systems could also suggest personalized interventions ranging from disease-modifying therapies to supportive care like psychotherapy, physical therapy, cognitive maintenance strategies including lifestyle modifications, as well as matching patients with clinical trials. During appointments, clinicians verify findings with patients, guide shared-decision making about next steps, and leverage AI to provide explanations optimized for that patient's level of education, age, and culture."
A practical advantage of agentic AI in ADRD care is its ability to integrate multimodal clinical data into a single, clinician-interpretable narrative.
The paper describes two main approaches: One approach is modality-to-language standardization where imaging features, test results, and other non-text input features are converted into textual or tabular summaries that can be consumed by LLM-based agents.
The paper cites "Gallingani et al. (2025) [who] developed an agentic AI system that simulates a multidisciplinary clinical team that achieved high accuracy in classifying pathology-confirmed primary progressive aphasia variants using multimodal inputs... The system preprocesses these data extracting features that are then converted into textual summaries and then uses specialized agents (for behavioral neurology, neuropsychology, and neuroimaging), whose analyses are synthesized and reviewed by a set of prompt-engineered specialized AI agents. Importantly, the output report is designed to be interpretable in familiar clinical terms that resemble how clinicians present cases, allowing clinicians to evaluate the evidence supporting the differential and the predicted underlying pathology. The paper also cites
Microsoft's MAI-DxO (MAI Diagnostic Orchestrator) [which] illustrates a related orchestration pattern: sequential diagnostic reasoning that surfaces clinician-readable intermediate steps and culminates in an auditable natural-language synthesis of what was asked, what was observed, and why a recommendation follows." "A second approach is to build on multimodal foundation models that reason jointly over vision and language, directly incorporating images or documents alongside text during dialogue and inference while maintaining a clinician-facing summary for review, as demonstrated by multimodal AMIE (Articulate Medical Intelligence Explorer), an agentic AI system developed by Google for telehealth-style consultations."
"Economic viability of healthcare AI systems remains uncertain, with substantial unknowns around return on investment (ROI), infrastructure costs, and reimbursement pathways. A recent systematic review demonstrated AI cost effectiveness across clinical domains from a healthcare system perspective through Quality of Life Years gained, Disability Adjusted Life Years averted, and diagnostic yield improvements, reducing unnecessary studies and treatments. Other studies have found earlier diagnosis to be cost-effective for Alzheimer's disease, particularly when accounting for caregiver economic burden and slowing progression with new disease modifying therapies. However, there is still some uncertainty as many studies fail to fully account for infrastructure investments, integration expenses, and maintenance costs."
"There are trends that offer cautious optimism. LLM costs are significantly declining, particularly inference and hardware costs, as well as improving energy efficiency. Evidence suggests that value-based payment models could incentivize the adoption of AI, though currently reimbursement for AI is through per-use payments. There is national government momentum in enabling new reimbursement pathways as demonstrated by a recent bipartisan bill introduced in April of 2025 (The Health Tech Investment Act, or S. 1399), which proposes systematic Medicare coverage for newly FDA authorized AI enabled medical devices to promote AI innovation. Healthcare systems have already shown willingness to invest in AI without clear reimbursement or concrete dollar ROI as demonstrated by the rapid adoption of ambient scribes, with evidence of decreasing clinician burnout. While a shared AI system could provide economies of scale among multiple healthcare systems, government funding would likely be needed to incentivize the upfront costs, and more reliable reimbursement may be needed to cover ongoing maintenance costs."
"Electronic consultations (e-consults) offer the most practical near-term implementation pathway, which are asynchronous, provider-to-provider communications that allow clinicians such as PCPs to seek specialist input through a shared electronic health record or web-based platform, often resulting in fewer in person specialist visits. In our experience, e-consults for a cognitive neurologist typically takes less than 10 minutes to complete. However, due to the limited data available with these consults many still require a long wait for a 60+ minute in person visit. The AI system proposed here could provide the data and decision support necessary for specialists to definitively address more cases in a fraction of the time required for a new patient visit, reducing wait times and patient travel burden especially for rural populations. Furthermore, specialists performing these reviews would simultaneously validate AI system performance, creating a sustainable monitoring mechanism without additional time investment. Demonstrating economic sustainability of these systems will be essential before widespread implementation."
"Validation requires comparison with appropriately defined 'gold standards' in a generalizable clinical population, with sufficient statistical power to demonstrate efficacy on a patient by patient (not simply group wise) basis. The vast majority of generative AI diagnostic studies rely on simulated rather than real-world data, highlighting the need for more systematic clinical validation. Prior to real world implementation, the possibility of misdiagnosis or adverse events must be virtually zero. Most clinical research with the gold standard neuropathologic verification is performed on well-educated relatively affluent patients with health insurance, often in academic settings. However, in the real world, patients have less education and insurance, greater comorbidities, and may be less fluent in English. Therefore, we propose models should also be developed using diverse specialist-confirmed cases in addition to neuropathology. This approach acknowledges both the gold standard status of neuropathology and the practical need for representative validation."
"However, specialist diagnoses contain uncertainty and there is variability even within a single practitioner, with greater variability across different clinicians and institutions. Careful case selection is required for reliable validation datasets in the absence of neuropathological confirmation. For example, prioritizing cases confirmed through highly accurate biomarkers with matching syndromes (e.g., Alzheimer's disease biomarkers), clearly meeting criteria known to be highly specific for a certain pathology (e.g., Progressive Supranuclear Palsy Richardson syndrome), multiple specialist agreement, or longitudinal confirmation of diagnosis through disease progression patterns. This step is crucial to reduce the risk of contaminated model updates."
"Ongoing monitoring is essential to prevent model degradation—the unintended decline in model performance, accuracy, or safety from improper updates, contaminated training data, or inconsistent validation practices. Agreed upon benchmarks must ensure models are fair, appropriate, valid, effective, and safe (FAVES). Recent initiatives like OpenAI's HealthBench and Stanford's MedArena demonstrate promising approaches for systematic medical AI evaluation across diverse clinical scenarios. These benchmarking frameworks could be adapted for neurodegenerative disease-specific applications, continuously tested on the unique diagnostic challenges these conditions present."
"Specialists are best positioned to identify model errors. Much of this monitoring can occur naturally through AI system use. Specialists could use this system to improve their own efficiency and quality in patient visits and e-consults, positioning them to detect errors in routine practice. While specialist time is limited, their investment in developing AI systems scales their expertise enabling PCPs and general neurologists to definitively manage more cases, reducing unnecessary referrals and increasing their capacity. This represents an evolution in specialist roles, where expertise is leveraged through optimizing AI systems rather than confined to individual patient encounters. If cost-effectiveness evidence builds, reimbursement mechanisms could justify dedicated monitoring activities such as batch evaluation of flagged cases. Beyond technical performance, monitoring must evaluate physician-AI interaction patterns, as clinicians often over rely on or underutilize AI capabilities, informing both system design and clinician training. In real-world settings, it's equally important to monitor how AI use affects clinician efficiency, productivity, and job satisfaction."
"To ensure coherence across the AI lifecycle, the oversight body governing standards would ideally also oversee validation, monitoring, and continuous learning processes discussed in the next section. While validation and monitoring focuses on quality control and preventing harm, and continuous learning aims to identify opportunities for improvement, these functions are inseparable: continuous learning without rigorous validation risks model degradation from contaminated or unverified updates, while validation without continuous improvement fails to leverage clinical experience. The careful case selection criteria described above apply equally to learning datasets. A unified governance approach ensures consistent standards, maintains institutional knowledge about AI system behavior, and aligns system evolution with healthcare priorities rather than commercial interests."
"While validation and monitoring prevent harm, continuous learning ensures this agentic AI system advances with emerging knowledge and clinical experience. This requires staying up to date with the rapidly evolving medical literature as well as learning from carefully validated patient cases in a supervised manner, guided collaboratively by the healthcare community."
"Clinician learning through patient cases is currently limited as they do not often see what happens to a patient if that patient does not follow up with them over time. Furthermore, what one clinician learns is difficult to transmit to others; however, an agentic AI system can learn from and provide benefit to every clinician using it. Built in feedback loops enable the AI system to consistently revisit each patient's chart or other databases to check for updates, such as an updated diagnosis or treatment response. Clinicians could receive these updates as well, learning alongside the AI system and taking action when needed. The careful case selection principles from the Validation section apply to continuous learning: only high-quality, validated cases from diverse clinical settings should influence model evolution, with ongoing oversight preventing contaminated updates."
"Transparency in AI development can be accomplished through using open-source/open-weights models with robust governance. When proprietary models are updated to optimize general use cases like writing or coding, performance on specific use cases such as healthcare can degrade unpredictably, requiring costly and frequent revalidation cycles with many models sunsetting every 1-2 years. Open-source models offer full control over updates, enable complete fine-tuning, and allow mechanistic interpretability analysis (i.e., understanding how an AI model is making decisions). LLMs of smaller size (8B-70B parameters) can be trained to match or surpass the performance of large proprietary models for more specific, narrower tasks including the task of answering medical questions, making subspecialty-focused models feasible. Although the computational overhead of these models once made routine clinical deployment impractical, recent parameter-efficient LLM compression methods (for example, 4-bit QLoRA or aggressive edge-device quantization) now enable low-latency, on-device inference suitable for real-world clinical workflows."
"Ethical challenges around the use of AI in ADRD have been discussed elsewhere, and therefore, a brief overview is provided here with updates that apply across healthcare. ADRD populations face unique ethical considerations beyond general AI concerns. As a vulnerable population with declining capacity, autonomy changes over time, requiring informed consent processes that evolve dynamically to involve both patients and caregivers appropriately. This shifting autonomy creates inherent tension between safety monitoring and preserving patient independence and dignity. Design considerations risk amplifying disparities if cognitive decline disproportionately impairs patients' ability to use AI tools, while privacy concerns intensify given the sensitive cognitive and behavioral data collected from individuals who may lack capacity to understand data usage."
"Many ethical considerations have intensified around the use of AI in healthcare, particularly since the widespread enthusiasm with generative AI since November 2022. In response, numerous initiatives have developed frameworks and guidelines to ensure AI applications are fair, appropriate, valid, effective, and safe (FAVES), which will at least in part require transparency with technical performance as well as organizational capacity to manage risks associated with deploying this technology. Fairness has been of particular concern as AI has shown numerous biases, emphasizing the need for training data from diverse groups as well as ongoing monitoring to identify and quickly mitigate biases based upon the analysis of only socioeconomically privileged populations. The advancement in AI capabilities will bring new challenges with patient privacy and informed consent. When using clinical notes for example, removing PHI may be insufficient to protect privacy, potentially requiring more stringent data protection and sharing guidelines. Regardless of the quality of the training data, AI tools are not static, in contrast to some pharmaceuticals whose mechanisms of action may be unclear but are safely used with patients. Therefore, explainability and interpretability are important components for a human in the loop to potentially identify signs of model deterioration or drift at the point of care. Ethical considerations traditionally designed to protect patients will likely need to be extended to healthcare workers as well with the rapid evolution of generative AI models. For example, generative AI's ability to mimic providers' appearance, voice, and empathetic conversation will likely be of interest to healthcare employers with an overwhelming patient demand. This dynamic landscape has prompted calls for a network of Health AI Ethics Centers to adapt and respond to these ongoing changes to keep pace in a manner that is often difficult for a central governmental agency."
"We believe patient harm represents the greatest risk of AI implementation in healthcare. This may occur through privacy breaches, adverse psychological effects from AI interactions (including anxiety, depression, or suicidal ideation), misdiagnosis, inappropriate treatment recommendations, medication errors, and more. Given the enthusiasm for medical AI and the economic incentives driving rapid adoption, we are concerned these risks are underappreciated. The possibility of widespread harm is not insignificant, particularly with an agentic AI system that is given too much autonomy. Legal liability remains unclear—who bears responsibility if patients suffer or die due to AI errors? Therefore, we emphasize that implementing redundant monitoring systems and establishing clear reporting mechanisms is critical to ensure prompt remediation when adverse events inevitably occur."
"The current fragmented governance landscape in healthcare AI is still evolving. While there has been a significant increase in FDA approved AI-enabled medical devices, the FDA has limited regulatory authority over many healthcare AI applications (e.g., they do not regulate applications that only provide recommendations or information that clinicians do not primarily rely on). Other federal agencies such as CMS have regulatory authority over AI implementation limited to entities participating in their programs. In the absence of comprehensive federal regulation, governance is emerging through state legislation, nonprofit guidance (e.g., the Joint Commission and Coalition for Health AI providing national guidance for responsible AI use), and private contracts between developers and healthcare organizations. AI integration in diagnostic workflows will require greater involvement from policymakers and government agencies alongside non-governmental entities. Given this fragmented landscape, the unified oversight approach we propose becomes even more essential: integrating data collection standards, validation protocols, monitoring systems, and continuous learning under collaborative governance. This model demonstrates how healthcare specialty communities can establish rigorous governance frameworks that complement and collaborate with evolving governmental regulation, addressing current gaps while maintaining alignment with evolving regulatory requirements."
"The integration of agentic AI systems into clinical care offers transformative potential to address many of the most pressing challenges in this unprecedented time for the field of ADRD. In an era with increasing time constraints and an overwhelming volume of medical knowledge, agentic AI systems can support clinicians by enhancing data collection, improving diagnostic accuracy, and personalizing management, ultimately enabling more timely and high-quality patient care. Successful implementation of these technologies will require a rigorous framework that prioritizes inclusivity, transparency, and ongoing monitoring to ensure ethical and effective use. Collaboration among clinicians, researchers, policymakers, governmental agencies, industry, and patients will be critical to establish standards for data collection, model development, and real-world application, ultimately fostering a continuously learning healthcare system."
"While this roadmap focuses on ADRD care, its principles apply broadly across neurology and other specialties facing similar challenges. With thoughtful development, agentic AI systems can not only mitigate provider shortages and diagnostic delays but also redefine care delivery, creating a future where technology and human expertise work in harmony to improve patient outcomes and advance medical knowledge."
The paper includes a detailed table outlining the six phases with their rationale, implementation strategies (near-term and long-term), key stakeholders, and expected impact:
Phase 1 - High-Quality, Standardized Data Collection: Build infrastructure to collect standardized, high quality, multi-modal clinical data; proper curation according to FAIR principles; establish structured formats for data currently lacking standardization; ensure data collection is accessible and easy for diverse patient populations and clinic settings. Near-term strategies include ambient scribes, AI-enabled data collection validation (digital cognitive assessments, voice-based conversational agents), and multi-institutional data standards. Long-term includes creating a repository of standardized multimodal ADRD datasets and integrating passive monitoring from digital biomarkers. Expected impact: comprehensive data collection reducing patient and clinician burden, standardized assessment across care settings, data prioritization evolving with latest evidence.
Phase 2 - AI Decision Support Development: Augment clinician capabilities through clear presentation and interpretation of complex multimodal data; enable explainable diagnostic reasoning; support personalized management decisions including complex therapy eligibility criteria. Near-term: train models to interpret ADRD-specific multimodal data, develop LLM-based systems capable of explaining decision support recommendations, use RAG to ground decision support in latest evidence base. Long-term: validate implementation across diverse clinical settings. Expected impact: increase early and accurate diagnosis in primary care and specialty settings, decrease diagnostic disparities, interpretable decision support, personalized care recommendations, increased identification of clinical trial candidates.
Phase 3 - Clinical Integration & Workflow: Address barriers to health technology adoption; alleviate clinician burnout; avoid patient/caregiver burden; maintain meaningful clinician oversight while maximizing automation benefits. Near-term: study clinician-AI interaction, build EHR-compatible tools reducing documentation burden, create patient-AI interfaces, develop implementation metrics. Long-term: develop AI assistants that adapt to clinician expertise level, build systems for automated follow-up with human oversight, implement training programs. Expected impact: time savings allowing more meaningful patient-provider interactions, enhanced clinical decision making, tools easily implemented in any care setting, improved clinician satisfaction and retention.
Phase 4 - Validation, Monitoring & Governance: Ensure patient safety; prevent model drift and degradation; identify and mitigate biases; inform regulatory framework. Near-term: establish validation protocols using FAVES framework, create monitoring dashboards. Long-term: develop neuropathology and biomarker correlation where feasible, establish independent oversight committees. Expected impact: consistent AI performance across diverse clinical and geographic settings, regulatory approval pathways for AI tools, early detection of model performance degradation, technical quality assurance standards.
Phase 5 - Continuous Learning & Improvement: Systematically integrate clinician feedback, patient outcomes, and emerging research to refine models while preventing quality degradation; ensure models address evolving real-world needs; ensure only high-quality data is used for model updates; ensure community engagement. Near-term: establish passive and voluntarily active feedback mechanisms, create patient outcome tracking systems, build real-time literature monitoring systems to update RAG database with human oversight, establish Community Scientific Partnership Boards. Long-term: implement semi-open models with robust governance, create clinician communities for model improvement. Expected impact: models that improve with use, rapid incorporation of new evidence, knowledge sharing across institutions, adaptation to evolving clinical practices.
Phase 6 - Ethics & Risk Management: Adopt and advance ethical guidelines to identify and mitigate patient harm as the greatest risk; ensure fairness, privacy, explainability, and clear accountability; protect healthcare workers' rights. Near-term: create efficient adverse event reporting systems, establish robust privacy safeguards, ensure transparency in model limitations and potential biases, design redundant safety systems. Long-term: develop ADRD-specific Health AI Ethics committee, create AI ethics certification standards, establish liability and indemnification frameworks. Expected impact: equitable access to advanced care, accountability mechanisms, protected patient autonomy, transparent AI deployment, established remediation pathways when harm occurs.
The table notes: "This roadmap conceptualizes a continuous learning healthcare system where insights from monitoring and implementation (Components 3-6) feed back into improving data collection standards and model development (Components 1-2), creating an ongoing cycle of improvement."
"All authors meet the ICMJE requirements for authorship through 1) substantial contributions to the conception of the work, 2) reviewing the work critically for important intellectual content, 3) final approval of the version to be published, and 4) agreement to be accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved."
Andreas Rauschecker reports receiving research funding from GE Healthcare.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to an AI system, along with what the improved system can do:
-
Add a modality-to-language standardization layer that converts neuroimaging features (atrophy scores, cortical thinning), biomarker results (plasma p-tau217, Aβ42/40, CSF α-synuclein), and neuropsychological test outputs (trial-level data, response times, error patterns) into structured textual summaries consumable by LLM-based agents.
-
Implement retrieval-augmented generation (RAG) with a continuously updated vector database containing FDA approvals, clinical trial eligibility criteria, and treatment guidelines (e.g., lecanemab, donanemab monitoring requirements) to ground all recommendations in the latest evidence.
-
Build a multi-agent system with specialized agents for behavioral neurology, neuropsychology, and neuroimaging, each performing independent analysis of the standardized data, then synthesizing outputs through a
multidisciplinary team
review layer that produces a clinician-readable differential diagnosis with explicit reasoning chains. -
Add a sequential diagnostic reasoning module that surfaces intermediate steps (what was asked, what was observed, why a recommendation follows) for full auditability.
-
Implement a semi-structured voice-based conversational agent that conducts patient/caregiver interviews over the phone, with:
-
Real-time content monitoring for dangerous topics (suicidal ideation, self-harm) triggering automatic human intervention
-
Automatic conversion of interviews into a structured History of Present Illness (HPI) summary
-
Adaptation to patient education level, language, and cultural context
-
Add LLM-based summarization that preserves subtle symptom changes over time, critical for early diagnosis.
-
Combine traditional discriminative models (for well-defined classification tasks like distinguishing AD vs. FTD from imaging) with generative LLM agents (for hypothesis generation and handling diagnostic uncertainty), orchestrated so the discriminative models provide quantitative, interpretable outputs that ground the LLM's reasoning.
-
Add explainability layer that translates model inferences into clinical terms (e.g.,
bilateral hippocampal atrophy pattern consistent with AD, supported by plasma p-tau217 elevation
) rather than raw probabilities. -
Implement passive feedback loops where specialists reviewing AI outputs (during e-consults or patient visits) implicitly validate or flag errors, feeding a curated dataset for model updates.
-
Add case selection criteria for learning data: only cases with (a) biomarker-confirmed pathology, (b) multiple specialist agreement, or (c) longitudinal diagnostic confirmation are used for model refinement, preventing contaminated updates.
-
Integrate real-time literature monitoring that updates the RAG database with human oversight, ensuring the system never lags behind new evidence.
-
Add diverse population validation using specialist-confirmed cases from underserved groups (racial/ethnic minorities, limited English proficiency, low digital literacy) to detect and correct performance disparities.
-
Implement ongoing fairness dashboards that track model performance stratified by patient demographics, flagging any drift in accuracy for specific subgroups.
-
Build SMART on FHIR compatible modules that integrate directly into EHR systems (Epic, Oracle), presenting AI-generated differentials and management recommendations within the clinician's existing workflow.
-
Add clinician-AI interaction monitoring to detect over-reliance (clinician accepting AI outputs without review) or underutilization, triggering targeted training interventions.
-
Implement redundant safety systems with clear escalation pathways for adverse events, including automatic alerts when AI recommendations conflict with patient safety parameters.
-
Use open-source, smaller LLMs (8B-70B parameters) fine-tuned for ADRD-specific tasks, with 4-bit QLoRA quantization for low-latency, on-device inference, reducing infrastructure costs and enabling deployment in resource-limited primary care settings.
-
Design the system to support e-consults as the primary use case, enabling specialists to review AI-prepared cases in under 10 minutes, scaling specialist expertise without requiring in-person visits.
-
Conduct comprehensive remote patient interviews via phone, producing structured histories that capture subtle cognitive symptom changes, with automatic safety monitoring and human escalation for crisis situations.
-
Generate a multimodal differential diagnosis by integrating patient history, digital cognitive assessments, neuroimaging (MRI/PET), and blood/CSF biomarkers, presenting the reasoning in familiar clinical terms with explicit evidence chains.
-
Evaluate complex therapy eligibility (e.g., anti-amyloid monoclonal antibodies) by systematically checking inclusion/exclusion criteria against patient data, flagging potential candidates earlier and reducing specialist burden.
-
Provide personalized management recommendations grounded in the latest evidence, including disease-modifying therapies, supportive care (psychotherapy, physical therapy, lifestyle interventions), and clinical trial matching.
-
Learn from every specialist interaction while maintaining rigorous quality gates—only validated, diverse cases influence model updates, preventing performance degradation and bias amplification.
-
Operate in resource-limited settings (primary care, rural clinics) with minimal infrastructure requirements, using open-source models and standardized data collection that doesn't require specialized staff.
-
Continuously monitor its own performance through dashboards tracking accuracy, fairness across demographic groups, and clinician-AI interaction patterns, with automatic alerts for model drift or emerging biases.
-
Support e-consult workflows where specialists can review AI-prepared cases in minutes, reducing wait times from months to days and expanding access to expert-level ADRD care for underserved populations.
-
Protect patient autonomy and safety through redundant monitoring, transparent reasoning, and human-in-the-loop oversight at every critical decision point, with clear accountability frameworks and adverse event reporting mechanisms.
Sources
- Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications
- Capabilities of Gemini Models in Medicine
- Sequential Diagnosis with Language Models
- Advancing Conversational Diagnostic AI with Multimodal Reasoning
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- QLoRA: Efficient Finetuning of Quantized LLMs
- De-identification is not enough: a comparison between de-identified and synthetic clinical notes
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework