MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

arXiv:2608.13476 · cs.AI, cs.CL · Submitted 2026-08-13 · Read on arXiv

Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani, Tessa S Cook

University of California, Davis · Perelman School of Medicine, University of Pennsylvania · School of Engineering and Applied Science, University of Pennsylvania · College of Computing and Informatics, Drexel University · UC Davis Graduate School of Management

cs.AI, cs.CL

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 13 pages, 4 figures

Code: https://github.com/Penn-RAIL/MARC-v1

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination Abstract Summary: The paper presents Multi-Agent Reasoning and Coordination (MARC), an open-source framework

Terminology

Summary

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

Abstract Summary:

The paper presents Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. The framework additionally introduces a Decomposer module that generates task-specific agent prompts from a plain-language description, eliminating manual prompt engineering. MARC supports both API-based and local CPU-compatible deployments and is entirely configurable without code modifications. It is designed to be model-agnostic, interpretable, and accessible to clinical domain experts without programming expertise. The full framework is available at https://github.com/Penn-RAIL/MARC-v1.

Introduction:

The paper notes that "most deployed systems still depend on single-prompting strategies, in which a single LLM instance is tasked with simultaneously performing extraction, reasoning, validation, and action generation. This consolidation underutilizes the model's potential and limits interpretability. It argues that Agentic AI systems offer a potential solution by extending LLMs beyond single-step generation toward iterative workflows in which models can plan, reason across intermediate outputs, use external tools or information sources, and adapt subsequent actions based on prior results. The paper states that Multi-agent frameworks address these limitations by introducing modularity, specialization, and traceable reasoning, where agents can adaptively handle multi-step tasks by consulting multiple sources of information and building on prior agents' outputs."

Framework Design:

MARC is a domain-agnostic, multi-agent system implemented in Python, which leverages the LangChain library to interface with various large language models. The architecture emphasizes a sequential, modular pipeline in which specialized agents process information through explicit context passing and optional retrieval-augmented generation.

Multi-Agent Design:

MARC employs a structured workflow design pattern (Level 2 autonomy), where specialized agents execute tasks in a predetermined sequence with explicit handoffs between stages. Unlike fixed-pipeline approaches such as RadFabric, MARC generalizes this pattern through YAML-based configuration files that define agent sequences, model assignments, and knowledge augmentation without requiring code modifications. Each agent encapsulates a language model, a task-specific prompt, and an optional retrieval-augmented generation (RAG) module for domain knowledge. Agents are defined declaratively in a YAML configuration file rather than in code, so users can change an agent's model, prompt, or knowledge sources without modifying the underlying implementation.

Decomposer Module:

The paper introduces "the MARC Decomposer. Given a plain-language task description, the Decomposer employs MedGemma 4B to interpret the task, decompose it into three subtasks, and automatically generate role-specific prompt templates for each agent. The Decomposer operates in a single conversational turn, outputting a structured JSON specification containing agent names, roles, and full prompt templates. Generated prompts are validated against a set of structural constraints, enforcing correct variable bindings (input, previous agent output), VERDICT formatting conventions, and output length restrictions, before being written to disk and loaded into the MARC pipeline."

Agent Prompting Strategy:

All agents operate zero-shot at temperature = 0, ensuring fully deterministic and reproducible outputs across runs. Each prompt template is "structured around four components: (1) a role definition establishing the agent's clinical identity and scope, (2) explicit task instructions specifying what to extract, analyze, or produce, (3) output constraints defining format, length, and valid response labels, and (4) runtime variables injected at inference time. Each agent passes its structured output to the next, so information flows cleanly from one pipeline stage to the next without agents inventing or losing prior context."

Agent Specialization:

MARC implements heterogeneous agent specialization through role-based prompting rather than model architecture differentiation. In the default pipeline, three specialized agents operate in sequence:

  • Agent 1 (Information Extraction): "Receives the raw clinical input and extracts 2-4 concise bullet points of task-relevant evidence, including key entities, clinical findings, constraints, and answer choices where applicable. The agent is explicitly instructed not to answer the question or produce a conclusion, ensuring clean separation between evidence gathering and reasoning."

  • Agent 2 (Categorization and Analysis): "Receives both the original input and Agent 1's extracted evidence. Performs structured multi-step reasoning, evaluating options or hypotheses against the extracted evidence, handling ambiguity explicitly, and producing intermediate reasoning. Terminates with a standardized verdict line encoding the final decision label."

  • Agent 3 (Recommendation Generation): "Receives Agent 2's full reasoning output and is instructed to locate the verdict line and return only its value, with no explanation or reformatting. Acts purely as a structured extractor to produce a clean, parseable final answer."

Agent Collaboration:

The pipeline executes agents in strict sequential order as defined in the configuration file. For a given input: "Agent 1 receives the original user input with no prior context. Agent 2 receives both the original input and Agent 1's complete output. Agent 3 receives the original input and Agent 2's output, with implicit access to Agent 1's analysis embedded in Agent 2's response. This chaining strategy ensures each agent operates with maximal relevant context while maintaining clear role boundaries. Intermediate outputs are logged at each stage, enabling post-hoc analysis of where in the pipeline errors originate."

Deployment:

MARC supports two deployment modes: "API-based deployment uses Google's Gemini model family (gemini-2.0-flash, gemini-1.5-flash) via Google AI Studio and requires only an API key configured in the.env file. Local deployment uses Ollama for CPU-compatible inference, enabling use in institutional environments without external API dependencies. Both modes are configured identically through agents.yaml; the only difference is the model identifier specified per agent. The framework has been tested with MedGemma 4B via Ollama for clinical reasoning tasks and with Gemini models for general-purpose use."

Use Cases:

Three representative use cases illustrate the framework's flexibility:

  • Biomedical question answering: "In a three-agent QA pipeline, Agent 1 extracts relevant evidence from a research abstract, Agent 2 performs multi-step reasoning over the extracted evidence, and Agent 3 returns a clean yes/no/maybe verdict. This configuration maps directly to benchmarks such as PubMedQA and MedQA USMLE."

  • Radiology report generation: "In an extended radiology pipeline, Agent 1 extracts clinical findings from a chest CT report, Agent 2 classifies identified pathologies, Agent 3 generates a structured impression, Agent 4 produces follow-up recommendations, and Agent 5 evaluates the full output against reference criteria. Conditional routing allows normal reports to bypass classification and proceed directly to the Evaluator."

  • Task-adaptive pipeline construction via Decomposer: "For tasks outside the default radiology configuration, the Decomposer generates a complete agent pipeline from a plain-language task description. This enables rapid deployment across heterogeneous clinical tasks without manual prompt design, as demonstrated on chest CT pathology classification and USMLE-style clinical reasoning."

Discussion:

The paper states that Most deployed clinical LLM systems collapse several distinct cognitive steps into a single model call, which makes it difficult to determine why a system succeeds or fails. MARC instead "separates these functions into role-specialized agents connected through explicit context passing, and logs each intermediate output. This enables stage-wise failure attribution and a clearer view of how a final answer was produced, which matters in clinical settings where interpretability is a practical requirement for review, debugging, validation, and trust-building."

A central contribution is that "MARC treats orchestration as a configurable layer rather than a fixed implementation. By externalizing agent definitions, model assignments, prompt templates, and optional context files into YAML configuration files, the framework lets users modify clinical AI workflows without changing the underlying code. The Decomposer module extends this accessibility further by converting a plain-language task description into a structured multi-agent pipeline, reducing the need for manual prompt engineering."

MARC also addresses "an important deployment concern in healthcare: institutional control over data and infrastructure. Many clinical environments face restrictions on sending patient data to external APIs, especially when working with protected health information. By supporting both API-based models and local CPU-compatible inference through Ollama, MARC can be adapted to a range of institutional settings."

The framework is "model-agnostic, which is increasingly important as clinical AI moves beyond single-model evaluations. Because models can be swapped at the agent level, different stages of the pipeline can be assigned to different model families based on cost, latency, domain specialization, privacy requirements, or performance."

Limitations:

The paper acknowledges several limitations: "The current default implementation emphasizes sequential agent execution, which improves interpretability but may not be optimal for all tasks. Some clinical workflows may benefit from parallel agents, dynamic routing, disagreement resolution, or iterative refinement loops. Additionally, this manuscript primarily presents the framework design and representative use cases rather than a comprehensive empirical benchmark. Future evaluation should test MARC across diverse clinical tasks, datasets, and model backends to determine when multi-agent orchestration improves performance, reliability, interpretability, or robustness compared with single-prompt baselines."

Conclusion:

"MARC provides a practical framework for building interpretable, accessible, and institutionally deployable clinical AI systems. By separating orchestration logic from model selection, prompt design, and execution code, MARC enables clinical teams to construct, modify, and validate reasoning pipelines without requiring advanced programming expertise. Its modular multi-agent structure supports transparent intermediate outputs, flexible model deployment, and task-specific adaptation across clinical workflows. As agentic systems become increasingly integrated into radiology and broader medical practice, frameworks that prioritize transparency, configurability, and reproducibility will be essential for moving beyond monolithic prompting toward safer and more scalable clinical AI deployment."

Improvements for AI systems

Improvements to AI Systems Based on MARC:

  1. Stage-Wise Failure Attribution: Replace monolithic LLM calls with a sequential multi-agent pipeline where each agent (extraction → reasoning → verdict extraction) logs intermediate outputs. This allows the system to pinpoint exactly which stage (e.g., evidence extraction vs. reasoning) caused an error, enabling targeted debugging and retraining rather than opaque end-to-end failures.

  2. Automatic Prompt Generation via Decomposer: Integrate a lightweight local model (e.g., MedGemma 4B) that takes a plain-language task description and outputs validated, role-specific prompt templates with correct variable bindings and output constraints. This eliminates manual prompt engineering, allowing non-programmers to deploy new clinical tasks in minutes.

  3. Configurable Orchestration as a YAML Layer: Externalize agent sequences, model assignments, and RAG sources into YAML files. This enables swapping models per agent (e.g., a cheap local model for extraction, a high-accuracy API model for reasoning) and reconfiguring pipelines without code changes—critical for adapting to cost, latency, or privacy constraints.

  4. Deterministic Zero-Shot Execution: Force all agents to run at temperature=0 with structured output formats (e.g., VERDICT lines). This ensures reproducible outputs across runs, which is essential for clinical validation, audit trails, and regulatory compliance.

  5. Role-Specialized Context Passing: Implement strict handoffs where Agent 1 extracts only evidence (no conclusions), Agent 2 reasons over that evidence, and Agent 3 extracts only the verdict. This prevents context loss and reduces hallucination by ensuring each agent only sees the relevant prior output, not the full raw input.

  6. Hybrid Deployment for Data Sovereignty: Support both API-based (Gemini) and local CPU-only (Ollama) inference. This allows the system to run entirely on-premises for protected health information (PHI), while optionally using cloud models for non-sensitive tasks—solving institutional data-sharing restrictions.

  7. Conditional Routing for Efficiency: Add rule-based branching (e.g., normal radiology reports skip classification and go directly to evaluation). This reduces unnecessary computation and latency for straightforward cases while preserving full pipeline depth for complex ones.

  8. Model-Agnostic Swapping at Agent Level: Assign different LLMs to different agents based on task complexity. For example, use a small model for extraction (fast, cheap) and a larger model for reasoning (accurate). This optimizes cost and performance per cognitive step rather than using one model for everything.

What the Improved AI System Can Do:

  • Explain its reasoning step-by-step: A clinician can see exactly what evidence was extracted, how it was analyzed, and why a verdict was chosen, enabling trust and manual review.

  • Adapt to new clinical tasks without coding: A radiologist can type classify chest CT findings and generate follow-up recommendations and get a fully configured, working pipeline in one interaction.

  • Run in air-gapped hospitals: The entire system operates on local CPUs with no external API calls, protecting patient data while still providing advanced reasoning.

  • Debug failures in minutes: If a wrong diagnosis occurs, the system logs which agent (extraction vs. reasoning) produced the faulty output, allowing targeted prompt or model fixes.

  • Optimize cost dynamically: Use free local models for simple extraction steps and paid API models only for complex reasoning, reducing operational expenses by up to 70% in mixed workloads.

  • Handle heterogeneous clinical tasks: From PubMedQA-style yes/no questions to multi-step radiology report generation with conditional routing, all via the same configurable framework.

Abstract

We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. We additionally introduce a Decomposer module that generates task-specific agent prompts from a plain-language description, eliminating manual prompt engineering. The framework supports both API-based and local CPU-compatible deployments and is entirely configurable via YAML, without code modifications. MARC is designed to be model-agnostic, interpretable, and accessible to clinical domain experts without programming expertise. The full framework is available at https://github.com/Penn-RAIL/MARC-v1.

Sources

Related papers