CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support

arXiv:2602.09159 · cs.AI, cs.MA · Submitted 2026-02-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary and Implications: Tom: Now, let's talk about what the paper actually says in its summary, which is quite revealing.

Jane: The authors are describing a system where they aren't just using role-playing, like having an AI act like a radiologist or an oncologist.

Lu: They are replacing that narrative-based interaction with something much more structured and quantitative instead of relying on dialogue.

Meng: This seems to be the core difference—the shift from "role simulation" to structural specialization.

Lalam: It’s about making the AI act like a collection of experts working in a real hospital meeting, rather than one person pretending to be all of them.

Tom: The paper introduces this framework called CoMMa, and it does what it says: it partitions the clinical context into distinct streams for each agent.

Jane: Each part—the lab results, the MRI report, the patient's history—gets its own dedicated agent to process it.

Lu: The genius of this is that we aren't just aggregating their outputs; we are coordinating them through a game-theoretic objective.

Meng: Game theory? That sounds incredibly complicated for a medical tool, but what does it do practically?

Lalam: It ensures that the agents contribute to the final decision in a way that is fair and mathematically grounded, rather than just blending their thoughts together.

Tom: So, instead of relying on long conversations to reach a consensus, we are using deterministic embedding projections.

Jane: Those embeddings are essentially fixed mathematical representations of the evidence that allow us to measure exactly what's going into each part of the system.

Lu: The paper says this approach results in explicit evidence attribution, which is vital for understanding medical errors or successes later on.

Meng: I think that's a massive win for regulatory compliance; if an error occurs, we can see which agent was responsible for its input.

Lalam: And it provides interpretable and stable decision pathways, giving doctors much more confidence in the output than a simple narrative would allow them to have.

Tom: It’s a big step forward from just trusting the AI; it’ giving us accountability inside the machine itself.

Improvements and Implications: Tom: The improvements CoMMa offers are really quite significant, especially compared to existing AI systems.

Jane: The biggest one is that by moving away from centralized data-sharing, we' are improving clinical data privacy significantly.

Lu: We’ve also achieved a reduction in inference cost because of the decentralized structure and the way we handle the input streams.

Meng: I was interested in how this relates to deployment; if it can run locally, that means much less cloud dependence for high-stakes hospitals.

Lalam: That's a massive cultural shift, moving from expensive cloud calls to efficient, local processing for better access everywhere.

Tom: The paper claims that because we are using this contribution-aware logic, the system is far more stable than those centralized baselines.

Jane: Stability in medicine is everything; you don't want a model that produces wildly different results for the same patient based on slight variations in another AI output.

Lu: The core of this improvement lies in using Shapley values to regulate our contribution learning, ensuring fairness across coalitions.

Meng: That mathematical grounding helps me understand how the system will behave when scaling up to five or ten agents for a complex case.

Lalam: It prevents that one dominant agent from unfairly dictating the final decision, forcing a truly collaborative outcome.

Tom: The paper shows this through impressive performance across both real-world tumor board data and public benchmarks like MTBBench.

Jane: It’s not just theoretical improvements; we see concrete results in terms of improved accuracy and AUC scores over baseline methods.

Lu: The paper is essentially proving that structured, deterministic interaction is superior to stochastic conversation for the clinical setting.

Meng: I think this architecture also makes it far more scalable than simply trying to concatenate all inputs into a single massive prompt.

Lalam: It allows us to scale the complexity of clinical reasoning without exploding the computational cost, which is a huge win for adoption in resource-constrained environments.

Conclusion: Tom: We’ve covered so much ground today, from how it works to why it's such a big deal for this field.

Jane: It’s certainly clear that CoMMa is making serious strides in building an AI that understands the nuances of medical collaboration.

Lu: I think the future research path is exciting, especially looking at how we can handle even more complex data clustering.

Meng: I'm glad to see a practical architecture that addresses both performance and operational efficiency for real-world deployment.

Lalam: We are building a framework that promotes transparency and fairness in clinical decision-making processes across the entire healthcare industry.

Tom: Before we sign off, let’s get those final thoughts from our team.

Lu: I am thrilled to see the mathematical rigor applied to clinical problems; the possibilities for this is immense.

Meng: I'm very impressed that this approach allows us to design locally deployable systems without sacrificing performance.

Lalam: My final thought is that it ensures the AI assists doctors with transparency, promoting a more equitable and trustworthy future in healthcare.

Tom: That’s a powerful way to end things. We hope everyone tuned in enjoyed this deep dive into "CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support."

Jane: It's been a fascinating journey into the world of structured AI collaboration.

Conclusion: Tom: So, we’ve spent hours breaking down how CoMMa works, moving beyond simple narrative discussion toward this highly structured, game-theoretic method of clinical decision support.

Jane: That structure is absolutely critical because it makes sure we aren're not just getting an answer; it's about understanding the *process* behind the answer.

Lu: The way they’ve applied the Shapley value as a regularization mechanism is a massive theoretical achievement, ensuring that credit assignment is mathematically fair across all clinical evidence streams.

Meng: And from an engineering perspective, this structured design means we can actually deploy this system locally without relying on expensive cloud services.

Lalam: I see the cultural impact in that too; it fosters trust by making the AI transparent, allowing doctors to see exactly where its data is coming from in a verifiable way.

Tom: It’s clear that "CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support" has set a new standard for both stability and explainability.

Jane: We are excited to watch how this framework is applied to even more complex, heterogeneous patient data moving forward.

Lu: It’ provides a foundational model for scalable, verifiable reasoning that extends far beyond the scope of oncology itself.

Meng: Hopefully, the path toward full-scale implementation and deployment is now much clearer for us too.

Lalam: I'm looking forward to seeing what cultural shifts this will inspire next time we talk about AI in healthcare.

cs.AI, cs.MA

Submitted: 2026-02-09

Updated: 2026-08-25

Code: https://github.com/langchain-ai/langchain

Importance score: 84/100

The gist: This paper introduces CoMMa, a decentralized LLM-agent framework designed for multidisciplinary oncology decision support.

Key concepts

CoMMa Framework
CoMMa replaces narrative-based AI interaction with a highly structured, quantitative system. Instead of one AI pretending to be multiple experts, it partitions clinical context into distinct streams, allowing specialized agents to process specific data like lab results or MRI reports.
Game-Theoretic Objective
This mathematical approach coordinates the agents' outputs to ensure that every contribution is fair and grounded. Using concepts like Shapley values, it prevents any single agent from unfairly dictating the final decision, promoting true collaboration.
Decentralized Structure
By moving away from centralized data sharing, CoMMa significantly improves clinical data privacy. This decentralized design also reduces inference costs and allows the system to run locally in hospitals without constant reliance on expensive cloud services.

Terminology

Summary

This paper introduces CoMMa, a decentralized LLM-agent framework designed for multidisciplinary oncology decision support. It addresses the limitations of current role-based multi-agent systems—which rely on stochastic narrative-based reasoning and datacentralized architectures—by providing a mathematically grounded method for explicit evidence attribution and improved clinical stability.

The limitations of current frameworks

Most existing medical multi-agent frameworks implement collaboration through role-play, assigning clinician-like personas such as oncologists or radiologists. These systems often rely on long-form dialogue as the coordination channel, which increases interaction overhead and makes outcomes sensitive to conversational dynamics. Because these systems are typically datacentralized, exposing all agents to the same full patient context, it becomes difficult to isolate which evidence drives a decision or to assign responsibility when errors occur.

How CoMMa works

CoMMa replaces role-based interaction with data-decentralized specialization, partitioning the clinical context into distinct information streams and assigning each to a dedicated agent. Instead of natural language, agent communication is conducted through deterministic embedding projections. The framework utilizes a specific architectural workflow:

  1. A Deterministic Embedding Projection module encodes text inputs into a shared embedding space using a frozen LLM.

  2. A Data-decentralized agent mixture module routes each embedding to its designated partition agent.

  3. A Contribution-aware multi-agent module aggregates agent-specific representations using an agent-decision matrix.

The final clinical decision is formed by aggregating agent outputs with learnable contribution weights, ensuring the system remains anchored to a global clinical grounding.

A game-theoretic approach to credit assignment

To ensure the agent-decision matrix provides a faithful representation of clinical evidence, CoMMa models multi-agent collaboration as a cooperative coalitional game. The framework adopts the principle of the Shapley value to regularize contribution learning, aligning each agent’s learned weight with its estimated marginal utility. This process involves:

  • Calculating an agent-wise reward based on the marginal reduction in loss achieved by an agent's inclusion.

  • Using a policy-gradient loss to upweight agent-class pairs with positive marginal utility.

  • Applying Shapley regularization via Kullback–Leibler (KL) divergence to pull learned weights toward the estimated contribution.

This transforms multi-agent collaboration from a stochastic dialogue process into a structured and interpretable inference procedure.

Experimental validation and results

The framework was evaluated on diverse oncology benchmarks, including the real-world HCC Tumorboard dataset and the MTBBench molecular tumor board benchmark. CoMMa achieved state-of-the-art performance and demonstrated significantly higher decision stability over data-centralized baselines. Key findings include:

  • CoMMa outperforms both locally trainable classical baselines and non-trainable online multi-agent baselines in treatment recommendation and recurrence prediction.

  • The use of deterministic embeddings enables gradient-based optimization and class-wise logit outputs, which are infeasible with narrative generation alone.

  • Even smaller models, when integrated into this framework, can outperform large-scale generative approaches, highlighting potential for efficient and scalable clinical deployment.

Improvements for AI systems

1. Replacement of Stochastic Narrative Dialogue with Deterministic Embedding Projections

  • Capability: The system will transition from high-latency, hallucination-prone text-based conversations between agents to low-latency, mathematically stable vector-based communication. This enables the use of gradient-based optimization for agent coordination and allows the system to run on local, fine-tunable hardware rather than relying on expensive, non-deterministic cloud APIs.

2. Shift from Semantic Role-Playing to Structural Data Decentralization

  • Capability: Instead of assigning agents personas (e.g., You are a radiologist), the system will enforce specialization by physically partitioning heterogeneous data streams (e.g., imaging, pathology, lab results) so that each agent only processes its designated modality. This prevents information entanglement, ensures true expertise-driven reasoning, and enhances data privacy by limiting the exposure of the full patient context to any single agent.

3. Integration of Shapley-Value Regularization for Quantitative Credit Assignment

  • Capability: The system will utilize cooperative game theory to calculate the exact marginal utility of each data stream's contribution to a final decision. This transforms the AI from a black box that provides potentially hallucinated verbal rationales into an interpretable system capable of providing explicit, mathematical evidence attribution (e.g., The MRI report contributed 72% of the predictive weight to this treatment recommendation).

4. Implementation of a Two-Stage Weighted Aggregation Mechanism

  • Capability: The system will utilize a learnable agent-decision matrix to fuse specialized agent logits, which is then reconciled with a global clinical anchor (a unified patient summary). This allows the AI to leverage granular, modality-specific insights while ensuring the final decision remains grounded in the patient's comprehensive longitudinal history.

5. Transition to Localized/Offline Inference Architectures

  • Capability: By utilizing deterministic projection heads and class-wise logits rather than generative text, the system can be deployed entirely on local hospital GPUs. This enables high-performance clinical decision support that complies with strict data privacy regulations (e.g., HIPAA) by eliminating the need for external cloud-based inference or sensitive data transfer.

Sources

Related papers