EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use

arXiv:2604.14165 · cs.CL · Submitted 2026-03-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use".

Jane: EviSearch introduces a multi-agent extraction system designed to automate ontology-aligned clinical evidence table creation directly from native trial PDFs while guaranteeing per-cell provenance for audit and human verification.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So Jane, we've been looking at the paper EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use. The main idea here is that they built this system to automate making those structured evidence tables directly from native trial PDFs while making sure you can always track exactly where every single piece of data came from for auditing.

Jane: That sounds really useful, Tom. Basically, the core thesis is about creating a multi-agent setup that handles the messy nature of clinical documents—text, tables, and figures all mixed together—and outputs something structured that you can trust because it has a clear record for verification.

Lu: I find the architecture fascinating because they combine this PDF query agent with a retrieval-guided search agent alongside a reconciliation module that forces page-level checks whenever the agents disagree on an extraction. It suggests a system that doesn't just guess; it actively verifies its findings against the source document structure.

Meng: From an engineering standpoint, I'm interested in how they manage that complexity, especially when they talk about batching columns using a groupaware packing algorithm with a maximum of fifteen columns per batch. That sounds like a smart way to balance context length against reliability.

Lalam: I think what excites me most is the idea of grounding every extracted value with its provenance, giving clinicians that verifiable attribution for every cell they look at. It moves us toward a system where the evidence chain is transparent and auditable.

Tom: Exactly, Jane. And why this matters right now is that clinical evidence synthesis is usually slow and prone to human error when people are manually pulling data from these dense PDFs. EviSearch claims it substantially improves extraction accuracy relative to strong parsed-text baselines on a clinician-curated benchmark of oncology trial papers, achieving a ninety point nine percent correctness rate.

Jane: It really does address the pain point of data extraction in this field. When we're dealing with complex trial documents, getting that high level of precision is crucial for making reliable decisions about patient care, and having a system that provides actionable provenance seems like a significant step forward.

Paper summary: Lu: The paper also touches on the underlying models they use, mentioning how document question answering benchmarks evaluate models on visually rich PDFs requiring joint reasoning over text and layout. It shows a deep understanding of the multimodal challenges involved in this task.

Meng: I wonder about the practical impact of that joint reasoning ability on real-world applications. If an AI can reliably reason across text and layout in these documents, what does that mean for streamlining workflows for researchers who are drowning in unstructured data?

Lalam: I think the vision here is that this kind of robust extraction capability could fundamentally improve how we synthesize knowledge across different clinical trials, creating a richer tapestry of evidence than we can currently manage.

Tom: Speaking of the overall structure, let's talk about the title and authors—EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use. It really emphasizes that this isn't just another extraction tool; it’s a system designed to be trustworthy through continuous improvement.

Jane: And the implication for the authors is showing that they can build a complex, multi-agent pipeline where agents learn and improve over time through human feedback, which is something we've been pushing for in AI development.

Lu: The concept of agents improving with use, combined with the reconciliation module forcing page-level verification when there are disagreements between the PDF query agent and the search agent, points toward a more adaptive system. It suggests a feedback loop built directly into the extraction process itself.

Meng: That feedback loop is critical for practical deployment, but I have to ask about performance trade-offs. The paper mentions that this dual-agent architecture increases token usage by over six hundred forty-two thousand seven hundred ninety-eight tokens across seventy-nine API calls. How do you justify that higher cost when compared to simpler parsed-text pipelines?

Lalam: The justification seems to be the auditability guarantee; every extracted value is traceable back to a specific page and modality, which turns that expenditure into verifiable evidence chains superior to those that exceed 970k–1M tokens.

Paper summary: Tom: So we're talking about a system that trades higher computational cost for verifiable truth in the evidence it produces, and the benchmark shows a substantial gain—a seven point two point improvement over strong parsed-text baselines. That’s solid data supporting their claims about accuracy.

Jane: It's about proving that the complexity of the system pays off when you need absolute certainty in a high-stakes domain like clinical evidence synthesis, and the results on that oncology benchmark show they hit those targets.

Lu: Thinking bigger, this approach to structured reasoning across multimodal inputs could open up new avenues for interpreting complex scientific literature beyond just text extraction. It suggests that the layout itself holds valuable information that we can unlock with the right AI architecture.

Meng: I see it translating into a need for more sophisticated document processing tools in other regulated industries, not just medicine. If you can reliably extract structured data from heterogeneous documents, you could apply this principle widely.

Lalam: And on a cultural level, this technology supports the development of AI that is inherently more responsible because it builds in mechanisms for transparency and human oversight at every stage of data creation.

Tom: We've covered the summary and the core implications of EviSearch, which focuses on automating clinical evidence table creation while ensuring per-cell provenance. Now we need to wrap up with a look at what this means moving forward for clinical research and AI development generally.

Jane: That’s right, it’s about taking that trust—that verifiable link between data and source—and making it the standard in how we handle trial information going forward.

Lu: The future work mentioned suggests an ongoing refinement of these agent interactions, which implies that the system is designed to evolve alongside the complexity of scientific documentation.

Meng: For practical impact, I'm looking at how this level of reliability could drastically cut down the time researchers spend on manual data compilation before they even start their analysis.

Lalam: Ultimately, EviSearch is pushing us toward a future where AI tools aren't just generating text, but are actively building trustworthy knowledge assets that support the scientific process itself.

Conclusion: Tom: So we've been deep into how EviSearch tackles extracting structured data from trial PDFs using this multi-agent setup, and now it's time to talk about what that whole package actually means for clinical research and AI development as a whole.

Jane: That’s right, Tom. We just finished walking through the technical details of how these agents work together to ensure every piece of evidence has a verifiable trail back to the original document, and now we need to pull back and look at the big picture implications.

Lu: I think what's really interesting about that title is "Trustworthy Extraction," because in a field where data integrity is everything, having a system that guarantees provenance for every cell feels like it addresses a fundamental problem in how we use AI for medical knowledge.

Meng: From an engineering standpoint, the idea of agents improving with use suggests that this isn't just a static tool; it’s designed to get smarter over time based on real-world feedback from people using it, which is crucial for practical deployment.

Lalam: I really see the cultural impact here: if we can build AI systems where the evidence chain is completely transparent and auditable, it could fundamentally shift how researchers feel comfortable using these tools in high-stakes environments.

Tom: Exactly. The authors put a lot of focus on making sure this system isn't just another black box generating results; it’s designed to be accountable, and that focus on trust is what makes this paper stand out.

Jane: It really boils down to taking the complexity of clinical documents and turning it into something reliable, which means we can start trusting AI outputs more in areas where accuracy matters most for patient care.

Lu: The way they structure the agents to handle different modalities—text, tables, figures—shows a sophisticated understanding of what's actually happening when you try to read a complex scientific paper with just standard text processing.

Meng: It makes me wonder how this level of verification could streamline workflows dramatically for researchers who are currently bogged down by manual data compilation before they even get to the analysis part.

Lalam: If we can build AI that inherently supports transparency, it means we're moving toward a future where scientific knowledge synthesis becomes much more collaborative and verifiable across different institutions.

Tom: That’s a huge vision, connecting the technical architecture directly to a more responsible and reliable future for medical discovery.

Jane: It really shows how much attention is being paid to the ethical side of AI application in science, focusing on not just getting an answer, but making sure you know exactly how that answer was reached.

Lu: The authors’ approach with continuous improvement through human feedback suggests a model for building AI that grows organically with real-world usage rather than just being deployed once and forgotten.

Meng: So, while the technical achievement is impressive, the practical value lies in creating a system where the output is not just useful but also defensible when it's ever scrutinized.

Lalam: That defensibility through traceable evidence is what’s going to really shape how we think about the role of AI in building and validating scientific knowledge over time.

Tom: It’s clear that EviSearch isn't just about better extraction; it’s about building a foundation for verifiable, trustworthy clinical evidence generation.

Naman Ahuja, Saniya Mulla, Muhammad Ali Khan, Zaryab Bin Riaz

Arizona State University · Mayo Clinic

cs.CL

Submitted: 2026-03-23

Updated: 2026-09-27

Comments: 12 pages, 6 figures, 3 tables. Substantially revised version: new multi-agent system with attribution verification and disagreement-directed review, new clinician-annotated benchmark and results. Code: https://github.com/CoRAL-ASU/EviSearch, demo: https://evisearch.fly.dev/

Project page: https://coral-labasu.github.io/EviSearch

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: EviSearch introduces a multi-agent extraction system designed to automate ontology-aligned clinical evidence table creation directly from native trial PDFs while guaranteeing per-cell provenance for

Key concepts

PDF Query Module (Agent A)
This agent submits the entire PDF binary to Gemini-2.5-Flash via its File API. It is prompted to return structured data including the value, reasoning, and exact attribution details like page number and modality, preserving native document structure.
Search Agent (Agent B)
This agent uses a tool-based loop to search the parsed document index. It employs tools like semantic search over chunks and targeted retrieval by page to find fine-grained evidence specifically within tables, figures, or results sections of the trial documents.
Reconciliation Agent
This agent checks for disagreements between Agent A and Agent B outputs using a two-pass protocol. If conflicts arise, it forces agents to retrieve full text and rendered page images to verify the discrepancy before assigning a confidence label like 'both_correct' or 'both_wrong'.
Human-on-the-Loop Interface
The system operates autonomously but provides verifiable evidence for every decision. In review mode, users can inspect any cell by seeing Agent A's answer, Agent B's answer, the reconciler's judgment, and the attributed page content side-by-side to accept or correct values.

Terminology

Summary

EviSearch introduces a multi-agent extraction system designed to automate ontology-aligned clinical evidence table creation directly from native trial PDFs while guaranteeing per-cell provenance for audit and human verification. This system addresses the challenges of extracting structured data from heterogeneous modalities like text, tables, and figures in clinical trial documents by combining direct PDF querying with retrieval-guided search and a reconciliation module that enforces page-level verification when agents disagree.

How it works

EviSearch is a multi-stage, multi-agent extraction pipeline consisting of four main stages: (i) document parsing and chunking, (ii) parallel extraction by two independent agents, (iii) automated reconciliation, and (iv) human review and feedback through a web interface. The system is designed to produce grounded, auditable attribution for every extracted value.

The pipeline utilizes two primary agents operating over the same column batch structure:

  1. The PDF Query Module (Agent A): This module submits the full PDF binary alongside each column batch to Gemini-2.5-Flash via the File API, which preserves the document’s native structure including figures, multicolumn layouts, and table formatting. It is prompted to return a structured tuple of (value, reasoning, attribution: page, modality, verbatim quote).

  2. The Search Agent (Agent B): This agent operates over the parsed document representation using a tool-based agentic loop. It uses tools like search chunks for semantic search over the document index and get chunks by page for targeted follow-up, aiming to target fine-grained evidence in tables, figures, and results sections.

Reconciliation and Verification

The Reconciliation Agent adjudicates disagreements between Agent A's and Agent B's outputs for each column batch using a two-pass protocol. Pass 1 involves Agreement detection (no tool use), resolving columns immediately if both agents report Not reported, both values are identical, or one value is a strict superset of the other. Pass 2 handles conflicts where values differ or one agent reports a value and the other reports Not reported. In these cases, it requires the agent to call get page to obtain both the full parsed text and a rendered page image for multimodal verification. The reconciliation agent submits a final value with one of four verification labels: both correct, A correct B wrong, B correct A wrong, or both wrong. Columns labelled both wrong are surfaced as low-confidence in the review interface for human attention.

Data Structure and Context Management

The output schema consists of 133 columns drawn from the structured evidence tables used in the LISR living evidence synthesis platform. Each column is paired with a natural-language definition specifying the required value, reporting conventions, and fallback behavior. To manage context efficiently during parallelism, columns are batched using a groupaware packing algorithm, where columns within the same clinical section are kept together, and batches are limited to a maximum of 15 columns. Groups exceeding this limit are split into sequential sub-batches or merged greedily until the batch limit is reached.

Human-on-the-Loop Interface

EviSearch is designed around a human-on-the-loop principle, allowing the system to operate autonomously by default but providing every extraction decision with verifiable evidence. In Automated mode, extracted values are stored with explicit provenance, and the web interface displays a completed table where every cell is one click away from its evidence. In Human-assisted mode, for uncertain columns or any column a reviewer chooses to inspect, the system provides the same evidence infrastructure side-by-side: Agent A’s answer, Agent B’s answer, the reconciler’s judgment, and attributed page content. This ensures clinicians review the same grounded evidence chain and can accept one candidate or write a corrected value.

Performance and Efficiency Gains

On a clinician-curated benchmark of oncology trial papers for metastatic castration-sensitive prostate cancer (mCSPC), EviSearch achieved an overall score of 91.3% (90.9% correctness, 91.6% completeness), surpassing the best baseline GPT-4.1 (parsed Doc) at 84.1%, a 7.2 point gain. The system demonstrates robustness across evidence modalities, maintaining near-constant performance across text, table, and figure sources (e.g., 91.2 → 91.7 → 86.7). While the dual-agent architecture increases token usage (642,798 tokens over 79 API calls), this cost is justified by the auditability guarantees, as every extracted value is traceable to a specific page and modality, converting expenditure into verifiable evidence chains that are superior to parsed-text pipelines which exceed 970k–1M tokens.

Improvements for AI systems

Here are specific improvements for existing AI systems based on the EviSearch framework, along with what those improved systems could achieve:

  1. Improving LLM-based Evidence Synthesis Pipelines by Integrating Multi-Agent, Provenance-Enforced Extraction:

  2. Developing Robust Clinical Trial Data Curation Systems Capable of Auditable Human Oversight and Iterative Model Improvement:

  3. Creating Multimodal Document Understanding Systems for Complex Scientific Literature Extraction (Text, Tables, Figures):

  4. Building AI Agents for Automated Schema-Aligned Evidence Table Generation with Per-Cell Grounding:


The improved AI systems derived from EviSearch can perform the following specific tasks:

  1. A system that automatically extracts structured clinical evidence tables directly from native trial PDFs while guaranteeing a traceable, per-cell provenance chain (source page, modality, verbatim quote) for every extracted value.

  2. A data curation platform that utilizes reconciliation agents to adjudicate disagreements between parallel extraction agents (one focusing on global context and the other on fine-grained retrieval), producing structured supervision signals that bootstrap the iterative improvement of underlying LLM models.

  3. A multimodal document reasoning system capable of accurately extracting information from heterogeneous evidence sources—including narrative text, complex tables, Kaplan-Meier plots, and figure captions—by leveraging dedicated agents for layout preservation (PDF Query Module) and semantic retrieval (Search Agent).

  4. A clinical evidence synthesis tool that operates in both automated mode (producing fully grounded tables with on-page highlighting) and human-assisted mode (allowing clinicians to inspect the exact evidence chain leading to any extracted cell, enabling direct correction of errors or uncertainty).

Abstract

Structured extraction of evidence from clinical trial publications underpins systematic reviews and clinical guidelines, yet large language models are adopted for it only hesitantly: their outputs are difficult to verify, their use commonly requires transmitting documents to proprietary services, and they do not improve from the corrections their users make. We present EviSearch, a multi-agent system that addresses these three obstacles. Three tool-augmented agents with complementary access to a publication extract every column of an evidence table, and a value is admitted only after an attribution verifier has read it on its cited page, so that every value carries a page-level attribution. Disagreement between independent agents directs human review to the cells most likely to be wrong, and reviewer feedback refines the schema definitions and a curation knowledge base without updating model parameters. The agentic system runs entirely offline on open-weight models. On a clinician-annotated benchmark of randomized-trial publications, EviSearch attributes 100.0% of its values, reaches 91.70% accuracy autonomously, and reaches 95.22% after review of 15.6% of cells, exceeding random review of the strongest single agent at equal effort by 1.75 points.

Sources

Related papers