ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages

arXiv:2606.13572 · cs.CL, cs.AI · Submitted 2026-06-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages".

Jane: Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Wow, Jane, I'm really energized after hearing the overview of this paper, "ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages." It sounds like they are tackling a real bottleneck in getting AI to work well in specialized fields like healthcare.

Jane: It really does sound like they’ve identified a major gap where general multimodal models just don't cut it, especially when you look at patients speaking native Indic languages and needing medical images. The paper claims this work addresses that exact limitation by introducing a specific framework called ArogyaSutra to handle multilingual medical reasoning accurately.

Lu: From a creative angle, the idea of an actor-critic architecture with dual memory for step-wise reasoning sounds fascinating; it suggests a way for the AI to actually think through complex medical questions rather than just spitting out an immediate guess. This could open up whole new avenues for how we structure knowledge retrieval in these low-resource settings.

Meng: I'm more focused on the practical side of this, and what strikes me is that they are integrating tool grounding directly into the reasoning process, using lightweight tools like zoom or edge detection to pull evidence from images before generating a step. That sounds like it could make the system much more reliable for real clinical scenarios where visual data needs careful extraction.

Lalam: I think the most impactful element here is how they are handling the language aspect by incorporating code-switching when errors happen, which means the system can self-correct its linguistic understanding during a reasoning chain. That ability to adjust based on feedback really speaks to making the AI more robust and less prone to simply failing because of a language mismatch.

Tom: So, if I'm following you both, the core thesis of "ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages" is that existing English-centric models fail in specialized medical scenarios across diverse Indic languages, and this paper proposes ArogyaBodha to fix that by using a multi-agent system with dual memory to improve reasoning accuracy.

Jane: Exactly, it’s about taking those complex medical queries from native Indian languages and multimodal inputs like images and making sure the AI can follow the logic step by step while maintaining language consistency throughout the process. It's designed to build more trustworthy diagnostic explanations across different regions.

Lu: The authors introduce this framework as a way to achieve transparent, medically grounded, and linguistically stable reasoning through their iterative actor-critic interaction with tool agents and memory mechanisms. That structure itself is quite clever for managing the flow of complex information processing.

Paper summary: Meng: From an engineering viewpoint, they are using a Qwen-VL-two point five-7B backbone for both the Actor and the Critic, but giving them different jobs to handle action proposal versus action evaluation, which simplifies how you build that iterative loop in practice. It seems like a smart way to distribute the cognitive load between the two components.

Lalam: And I think it’s really powerful because they use this iterative process where the Critic gives targeted feedback—in English for language issues and in the original Indic language otherwise—which then updates both memory modules, creating a self-improving loop for reasoning. That continuous learning aspect is what makes the framework so compelling.

Tom: It sounds like this paper isn't just about getting a single right answer; it’s about showing how to build an entire system that can reason and correct itself under complex medical conditions in languages that aren't typically well-represented in AI training data.

Jane: That capability is what gives the ArogyaBodha dataset its significance, because having a large, diverse benchmark covering thirty-one body systems and twenty-one clinical domains across seven major Indian languages provides the necessary scale for this kind of deep testing.

Lu: The creation of that "ArogyaBodha" dataset itself is a huge contribution; combining existing benchmarks with questions from postgraduate medical entrance exams like NEET-PG and FMGE, while using GPT-4o-mini for fewshot prompting, shows a very thorough effort in curation.

Meng: It’s interesting how they evaluated the translation quality across those seven Indic languages using reverse-translation analysis and manual review, achieving an average score of four point two seven for clinical meaning preservation; that metric gives us a real benchmark for what high-quality medical translation looks like in this context.

Lalam: I think that four point two seven score is important because it validates that the linguistic nuances in Indic languages are being preserved enough to be useful in a clinical setting, which is crucial for any healthcare application. It shows they didn't just focus on surface-level translation but on actual medical meaning retention.

Tom: So, to wrap up what we’ve heard about "ArogyaBodha," this paper introduces a large-scale multimodal medical reasoning dataset curated from eight diverse sources spanning seven major Indian languages alongside English, designed to test and improve MLLM performance in healthcare contexts.

Jane: And the framework, ArogyaSutra, is built on an actor-critic design that uses tool grounding and dual memory to enable step-wise decision making for multilingual medical reasoning across all those languages. It’s essentially a systematic approach to making AI more dependable when dealing with complex Indian medical information.

Lu: The overall contribution lies in demonstrating this architecture can consistently outperform strong multimodal baselines in both reasoning accuracy and multilingual alignment, which is a solid piece of empirical evidence presented here.

Paper summary: Meng: If we look at the ablation studies they ran, it’s pretty telling that removing the Critic agent caused a significant drop in performance down to around thirty-three point four three percent–thirty-three point seven one percent, which really highlights how essential that evaluation step is for maintaining high accuracy.

Lalam: And disabling tool grounding or code-switching resulted in even bigger drops, reaching sixteen point five four percent when those features were removed, which really shows that the ability to ground visuals and adapt language handling are critical components of making this system effective.

Tom: So we’ve heard about how ArogyaSutra is structured and how the data was built; now we need to think about what this means for real-world application in areas where access to specialized medical knowledge is currently limited.

Jane: That’s right, and the implications are huge because it directly addresses the inequity in AI healthcare where patients relying on native Indic languages can access more accurate, multimodal assistance than they could before. It moves the needle toward more equitable digital health tools globally.

Lu: I see a future where these frameworks aren't just confined to academic benchmarks; they could become foundational components for developing localized diagnostic aids that operate effectively in rural settings where resources are scarce.

Meng: Practically speaking, I wonder how easy it is to deploy this kind of iterative, memory-aware system into an actual clinical workflow and ensure it remains stable under the kind of real-time pressure you see in a hospital environment.

Lalam: I think the potential here is that this level of structured reasoning could fundamentally improve patient trust because the AI doesn't just give an answer; it shows its work through transparent, step-by-step reasoning grounded in evidence and consistent language handling.

Tom: So, to conclude our discussion on "ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages," we’ve seen how they’ve constructed a comprehensive dataset and a sophisticated multi-agent framework designed specifically to overcome the limitations of current models in specialized Indian healthcare contexts.

Jane: The title itself, "ArogyaSutra," suggests a guiding thread or rule, and the authors have provided substantial evidence showing that this approach consistently outperforms existing multimodal baselines in both accuracy and multilingual alignment.

Lu: The broader implication is that we are seeing a path toward building AI systems that aren't just broadly capable but are specifically tailored to handle the complexity of low-resource languages and nuanced medical imaging simultaneously.

Meng: For practical implementation, the distillation strategy they used, where they fine-tune the Actor on a Critic-approved dataset without needing constant Critic interaction during inference, suggests a path toward making these complex models faster and more deployable for actual use cases.

Lalam: Ultimately, this work provides a robust blueprint for developing AI that is not only powerful in reasoning but also fundamentally trustworthy and accessible to the diverse populations of India who need it most.

Conclusion: Tom: So, we've seen how ArogyaBodha creates this really intricate system for medical reasoning across different languages, and now we need to wrap up by talking about what all this means for the future.

Jane: That’s right, Tom; thinking about the title itself, "ArogyaSutra," it suggests a guiding principle or a set of rules that the AI follows while processing those complex medical images and queries.

Lu: I think that guiding structure is really important because it moves beyond just pattern matching; it implies a kind of structured, deliberate thought process for the model.

Meng: From an engineering standpoint, I’m curious about how this framework handles the sheer variety of inputs without getting overwhelmed by the complexity of those seven different languages.

Lalam: For me, this framework represents a significant step because it shows that we can build AI that respects and accurately translates deep cultural knowledge embedded in native languages when dealing with something as sensitive as health.

Tom: Exactly, Jane; it’s about more than just processing words; it's about building a system that understands the context of medical queries in those specific linguistic environments.

Jane: And the authors, I think their focus on integrating tool grounding directly into the reasoning path is what makes this approach stand out from what we usually see in these multimodal models.

Lu: That integration is key because it means the AI isn't just guessing; it’s actively looking for and extracting evidence from the image based on its language understanding at each step.

Meng: I'm still thinking about deployment, how practical this multi-agent setup is going to be when we move this from a research environment into a real clinical setting where speed matters.

Lalam: And that practical concern is valid because if the system can maintain high accuracy and linguistic consistency under those conditions, it could fundamentally improve patient trust in digital health tools everywhere.

Tom: So, ultimately, this work pushes us toward creating AI that doesn't just perform a task but demonstrates a reliable, step-by-step chain of reasoning tailored for diverse global needs.

Jane: That's the big picture; ArogyaBodha gives us a blueprint for how to build more equitable and contextually aware medical AI.

Lu: It opens up possibilities for developing specialized diagnostic aids that can function effectively in regions where access to English-centric data is very limited.

Meng: I agree; the distillation strategy they used is interesting because it points toward making these complex reasoning models much faster for real-time applications down the line.

Lalam: And this paper shows that even with resource constraints, we can achieve high levels of specialized reasoning by focusing on structured memory and targeted feedback mechanisms.

Tom: Absolutely, so what we've heard today is that ArogyaBodha and its framework offer a powerful way to move toward building medical AI that works reliably for everyone in India.

Jane: And as we wrap up this discussion, we have a lot of exciting questions about how this structured reasoning can translate into tangible benefits for healthcare accessibility worldwide.

Indian Institute of Technology Patna · Indian Institute of Technology Kanpur

cs.CL, cs.AI

Submitted: 2026-06-11

Updated: 2026-07-16

Project page: https://iitp-cse.github.io/ArogyaSutra

Importance score: 92/100

The gist: Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially

Key concepts

ArogyaBodha
This is a large dataset of multimodal medical queries and answers curated from eight diverse sources spanning seven major Indian languages and 31 body systems. It serves as the benchmark for testing reasoning accuracy and language consistency in medical contexts.
Actor-Critic Framework
This is a decision-making structure where an 'Actor' proposes an action (like a reasoning step), and a 'Critic' evaluates that proposal for correctness. They work together to refine the process, ensuring that the proposed steps are both logically sound and linguistically accurate.
Dual-Memory Mechanism
The system uses two types of memory: long-term memory stores summaries of past reasoning steps and errors, while short-term memory holds recent prediction errors. This allows the agent to maintain context over long reasoning chains while quickly correcting immediate mistakes.

Terminology

Summary

Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios. This gap is critical in regions like rural India, where patients often express complex medical queries in native Indic languages and rely on multimodal inputs such as medical images.

The gist

ArogyaSutra is an actor–critic-based multi-agent framework that integrates tool grounding with dual-memory mechanisms for step-wise, reasoning-aware decision making to improve multilingual medical reasoning accuracy across all Indic languages.

Dataset: ArogyaBodha

This work introduces ArogyaBodha, a largescale multimodal medical reasoning dataset and benchmark curated from eight diverse medical sources. The dataset spans seven major Indian languages alongside English and covers 31 body systems and 21 clinical domains. Each instance in the dataset consists of an expert-verified multimodal medical query paired with a single unambiguous ground-truth answer, which enables reliable evaluation of both logical reasoning accuracy and language consistency. The curation process involved combining existing benchmarks with questions from Indian postgraduate medical entrance examinations (NEET-PG and FMGE), using fewshot prompting with GPT-4o-mini, and filtering based on reasoning depth and image–question relevance. Furthermore, the translation quality across the seven Indic languages was evaluated using reverse-translation analysis and manual review, resulting in an average score of 4.27 for clinical meaning preservation.

Framework: ArogyaBodha

ArogyaSutra is a actor–critic based multi-agent framework that combines tool-based visual grounding with dual-memory mechanisms for step-wise, reasoning-aware decision making over image–text inputs. The framework utilizes an Actor, implemented as a multimodal language model, which processes medical visual inputs and an Indic-language query to invoke lightweight grounding tools (e.g., zoom/crop, edge detection, depth analysis) to extract clinically relevant evidence. Rather than producing a final answer directly, the Actor predicts intermediate semantic reasoning steps. This process is supported by a dual-memory design: long-term memory summarizes prior reasoning steps, linguistic formulations, and identified errors (from steps 1 to t−1), while short-term memory captures the most recent prediction error and feedback. The Critic evaluates the Actor’s outputs for both medical correctness and language consistency, providing targeted corrective feedback—delivered in English for linguistic errors and in the corresponding Indic language otherwise—which is used to update the memory modules.

ArogyaSutra Mechanism

The core decision module adopts an Actor-Critic formulation where both components are instantiated from the same multimodal backbone, Qwen-VL-2.5-7B, differing only in their functional roles: action proposal versus action evaluation and their input conditioning. At timestep t, the Actor samples a candidate semantic action based on the current state, task objective, perceptual context, and memory states. The Critic evaluates this proposed action by estimating a scalar confidence score (sˆt). If sˆt = 1, the action is accepted; otherwise, the Critic triggers language-Analysis reflection to distinguish between language understanding errors and logical errors. If the failure stems from faulty reasoning, the Critic issues a reflection in the corresponding Indic language of the query to guide subsequent reasoning chains.

Training Strategy and Distillation

To progressively enhance performance, ArogyaBodha is used for training through code-switched reasoning chain distillation. After simulating Actor-Critic rollouts, a Criticrefined dataset D is constructed where the action (u†i) is the Critic-approved action augmented with relevant tool outputs and memory context. The Actor is then fine-tuned using supervised learning with the objective:

LSFT = − 1/N Σ N i=1 log πθ u†i oi, gi, zi, MSi, MLi.

During inference, only the distilled policy πθ is retained, allowing the Actor to directly predict the next semantic action without invoking the Critic. This distillation strategy preserves the Critic’s structured reasoning and memoryaware behavior while substantially reducing inference-time overhead. The process includes a halting condition: if convergence occurs within three iterations, the trace is retained; otherwise, it restarts using a summary of accumulated errors in long-term memory.

Evaluation and Impact

Extensive experiments demonstrate that ArogyaSutra consistently outperforms strong multimodal baselines in both reasoning accuracy and multilingual alignment. The framework's effectiveness is validated by ablation studies: removing the Critic significantly reduces performance (to 33.43%–33.71%), while disabling tool grounding or code-switching causes the largest performance drop (16.54%). The best results are achieved when both components—Image Grounding and the Critic agent—are enabled, yielding an average accuracy of 47.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems by adopting the concepts from ArogyaBodha and ArogyaSutra, along with what these improved systems can achieve:


  1. The creation of a large-scale, expert-verified, multilingual medical reasoning benchmark dataset (ArogyaBodha) covering 31 body systems and 21 clinical domains across seven major Indian languages alongside English.

  2. The development of the ArogyaSutra multi-agent framework that integrates:

  3. Tool Grounding: Explicitly invoking visual tools (object detection, zoom/crop, edge detection, depth analysis) during reasoning steps to extract clinically relevant evidence from medical images in real-time.

  4. Dual-Memory Mechanisms: Implementing a short-term memory for capturing the most recent prediction error and a long-term memory that summarizes prior reasoning steps and identified errors across the entire trajectory.

  5. Actor–Critic Reasoning Cycle: Employing an Actor (policy) to propose intermediate semantic reasoning steps, evaluated by a Critic to assess medical correctness and language consistency, which then triggers corrective feedback in English for linguistic issues or the target Indic language otherwise.

  6. Code-Switching/Language-Aware Reflection: Dynamically detecting whether a failure stems from logic or linguistic instability; if linguistic, the system switches to English for correction; if logical, it uses memory summaries to refine the reasoning chain within the context of the original Indic language query.

  7. Distillation via Critic Refinement: Utilizing an actor-critic simulation trajectory to distill a refined policy dataset (D) that allows the final inference model (the Actor) to predict the next semantic action directly, eliminating runtime overhead while preserving memory-aware, corrective behavior.

The improved AI system can perform the following specific actions:

  1. Perform complex, step-by-step clinical reasoning over medical images and natural language queries in low-resource Indic languages (e.g., Bengali, Hindi, Tamil).

  2. Generate medically accurate diagnostic explanations that are linguistically consistent and faithful to the original query language across seven major Indian languages.

  3. Self-correct complex logical errors by analyzing past mistakes stored in long-term memory and receiving targeted corrective guidance from the Critic agent.

  4. Interact with visual evidence during reasoning by dynamically selecting appropriate visual tools (e.g., zooming into a specific lesion or using edge detection) to ground its textual analysis accurately, ensuring perception is clinically relevant.

  5. Demonstrate robust generalization across different imaging modalities and clinical domains (31 systems, 21 domains) even when encountering out-of-distribution queries by leveraging the distilled policy trained on diverse data.

  6. Provide trustworthy medical decision support in rural or underserved areas where English-centric models fail due to language barriers and lack of cultural/linguistic grounding.

Sources

Related papers