EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use
summary
The gist
EviSearch introduces a multi-agent extraction system designed to automate ontology-aligned clinical evidence table creation directly from native trial PDFs while guaranteeing per-cell provenance for
In short
EviSearch is a multi-agent system that automates creating clinical evidence tables from PDF trial documents. It uses two agents—one querying the full document and another searching chunks—and a reconciliation module to ensure every extracted data point has verifiable proof linked to a specific page and source, making the output auditable.
Key concepts
- PDF Query Module (Agent A)
- This agent submits the entire PDF binary to Gemini-2.5-Flash via its File API. It is prompted to return structured data including the value, reasoning, and exact attribution details like page number and modality, preserving native document structure.
- Search Agent (Agent B)
- This agent uses a tool-based loop to search the parsed document index. It employs tools like semantic search over chunks and targeted retrieval by page to find fine-grained evidence specifically within tables, figures, or results sections of the trial documents.
- Reconciliation Agent
- This agent checks for disagreements between Agent A and Agent B outputs using a two-pass protocol. If conflicts arise, it forces agents to retrieve full text and rendered page images to verify the discrepancy before assigning a confidence label like 'both_correct' or 'both_wrong'.
- Human-on-the-Loop Interface
- The system operates autonomously but provides verifiable evidence for every decision. In review mode, users can inspect any cell by seeing Agent A's answer, Agent B's answer, the reconciler's judgment, and the attributed page content side-by-side to accept or correct values.
Terminology used across episodes
This episode discusses
- EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use · Paper Radio
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- A Survey of LLM-based Agents in Medicine: How far are we from Baymax?
- A Survey of Large Language Models in Medicine: Progress, Application, and Challenge
The paper
EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use · Read on arXiv
Naman Ahuja, Saniya Mulla, Muhammad Ali Khan, Zaryab Bin Riaz
Arizona State University · Mayo Clinic
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use".
Jane: EviSearch introduces a multi-agent extraction system designed to automate ontology-aligned clinical evidence table creation directly from native trial PDFs while guaranteeing per-cell provenance for audit and human verification.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So Jane, we've been looking at the paper EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use. The main idea here is that they built this system to automate making those structured evidence tables directly from native trial PDFs while making sure you can always track exactly where every single piece of data came from for auditing.
Jane: That sounds really useful, Tom. Basically, the core thesis is about creating a multi-agent setup that handles the messy nature of clinical documents—text, tables, and figures all mixed together—and outputs something structured that you can trust because it has a clear record for verification.
Lu: I find the architecture fascinating because they combine this PDF query agent with a retrieval-guided search agent alongside a reconciliation module that forces page-level checks whenever the agents disagree on an extraction. It suggests a system that doesn't just guess; it actively verifies its findings against the source document structure.
Meng: From an engineering standpoint, I'm interested in how they manage that complexity, especially when they talk about batching columns using a groupaware packing algorithm with a maximum of fifteen columns per batch. That sounds like a smart way to balance context length against reliability.
Lalam: I think what excites me most is the idea of grounding every extracted value with its provenance, giving clinicians that verifiable attribution for every cell they look at. It moves us toward a system where the evidence chain is transparent and auditable.
Tom: Exactly, Jane. And why this matters right now is that clinical evidence synthesis is usually slow and prone to human error when people are manually pulling data from these dense PDFs. EviSearch claims it substantially improves extraction accuracy relative to strong parsed-text baselines on a clinician-curated benchmark of oncology trial papers, achieving a ninety point nine percent correctness rate.
Jane: It really does address the pain point of data extraction in this field. When we're dealing with complex trial documents, getting that high level of precision is crucial for making reliable decisions about patient care, and having a system that provides actionable provenance seems like a significant step forward.
Paper summary: Lu: The paper also touches on the underlying models they use, mentioning how document question answering benchmarks evaluate models on visually rich PDFs requiring joint reasoning over text and layout. It shows a deep understanding of the multimodal challenges involved in this task.
Meng: I wonder about the practical impact of that joint reasoning ability on real-world applications. If an AI can reliably reason across text and layout in these documents, what does that mean for streamlining workflows for researchers who are drowning in unstructured data?
Lalam: I think the vision here is that this kind of robust extraction capability could fundamentally improve how we synthesize knowledge across different clinical trials, creating a richer tapestry of evidence than we can currently manage.
Tom: Speaking of the overall structure, let's talk about the title and authors—EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use. It really emphasizes that this isn't just another extraction tool; it’s a system designed to be trustworthy through continuous improvement.
Jane: And the implication for the authors is showing that they can build a complex, multi-agent pipeline where agents learn and improve over time through human feedback, which is something we've been pushing for in AI development.
Lu: The concept of agents improving with use, combined with the reconciliation module forcing page-level verification when there are disagreements between the PDF query agent and the search agent, points toward a more adaptive system. It suggests a feedback loop built directly into the extraction process itself.
Meng: That feedback loop is critical for practical deployment, but I have to ask about performance trade-offs. The paper mentions that this dual-agent architecture increases token usage by over six hundred forty-two thousand seven hundred ninety-eight tokens across seventy-nine API calls. How do you justify that higher cost when compared to simpler parsed-text pipelines?
Lalam: The justification seems to be the auditability guarantee; every extracted value is traceable back to a specific page and modality, which turns that expenditure into verifiable evidence chains superior to those that exceed 970k–1M tokens.
Paper summary: Tom: So we're talking about a system that trades higher computational cost for verifiable truth in the evidence it produces, and the benchmark shows a substantial gain—a seven point two point improvement over strong parsed-text baselines. That’s solid data supporting their claims about accuracy.
Jane: It's about proving that the complexity of the system pays off when you need absolute certainty in a high-stakes domain like clinical evidence synthesis, and the results on that oncology benchmark show they hit those targets.
Lu: Thinking bigger, this approach to structured reasoning across multimodal inputs could open up new avenues for interpreting complex scientific literature beyond just text extraction. It suggests that the layout itself holds valuable information that we can unlock with the right AI architecture.
Meng: I see it translating into a need for more sophisticated document processing tools in other regulated industries, not just medicine. If you can reliably extract structured data from heterogeneous documents, you could apply this principle widely.
Lalam: And on a cultural level, this technology supports the development of AI that is inherently more responsible because it builds in mechanisms for transparency and human oversight at every stage of data creation.
Tom: We've covered the summary and the core implications of EviSearch, which focuses on automating clinical evidence table creation while ensuring per-cell provenance. Now we need to wrap up with a look at what this means moving forward for clinical research and AI development generally.
Jane: That’s right, it’s about taking that trust—that verifiable link between data and source—and making it the standard in how we handle trial information going forward.
Lu: The future work mentioned suggests an ongoing refinement of these agent interactions, which implies that the system is designed to evolve alongside the complexity of scientific documentation.
Meng: For practical impact, I'm looking at how this level of reliability could drastically cut down the time researchers spend on manual data compilation before they even start their analysis.
Lalam: Ultimately, EviSearch is pushing us toward a future where AI tools aren't just generating text, but are actively building trustworthy knowledge assets that support the scientific process itself.
Conclusion: Tom: So we've been deep into how EviSearch tackles extracting structured data from trial PDFs using this multi-agent setup, and now it's time to talk about what that whole package actually means for clinical research and AI development as a whole.
Jane: That’s right, Tom. We just finished walking through the technical details of how these agents work together to ensure every piece of evidence has a verifiable trail back to the original document, and now we need to pull back and look at the big picture implications.
Lu: I think what's really interesting about that title is "Trustworthy Extraction," because in a field where data integrity is everything, having a system that guarantees provenance for every cell feels like it addresses a fundamental problem in how we use AI for medical knowledge.
Meng: From an engineering standpoint, the idea of agents improving with use suggests that this isn't just a static tool; it’s designed to get smarter over time based on real-world feedback from people using it, which is crucial for practical deployment.
Lalam: I really see the cultural impact here: if we can build AI systems where the evidence chain is completely transparent and auditable, it could fundamentally shift how researchers feel comfortable using these tools in high-stakes environments.
Tom: Exactly. The authors put a lot of focus on making sure this system isn't just another black box generating results; it’s designed to be accountable, and that focus on trust is what makes this paper stand out.
Jane: It really boils down to taking the complexity of clinical documents and turning it into something reliable, which means we can start trusting AI outputs more in areas where accuracy matters most for patient care.
Lu: The way they structure the agents to handle different modalities—text, tables, figures—shows a sophisticated understanding of what's actually happening when you try to read a complex scientific paper with just standard text processing.
Meng: It makes me wonder how this level of verification could streamline workflows dramatically for researchers who are currently bogged down by manual data compilation before they even get to the analysis part.
Lalam: If we can build AI that inherently supports transparency, it means we're moving toward a future where scientific knowledge synthesis becomes much more collaborative and verifiable across different institutions.
Tom: That’s a huge vision, connecting the technical architecture directly to a more responsible and reliable future for medical discovery.
Jane: It really shows how much attention is being paid to the ethical side of AI application in science, focusing on not just getting an answer, but making sure you know exactly how that answer was reached.
Lu: The authors’ approach with continuous improvement through human feedback suggests a model for building AI that grows organically with real-world usage rather than just being deployed once and forgotten.
Meng: So, while the technical achievement is impressive, the practical value lies in creating a system where the output is not just useful but also defensible when it's ever scrutinized.
Lalam: That defensibility through traceable evidence is what’s going to really shape how we think about the role of AI in building and validating scientific knowledge over time.
Tom: It’s clear that EviSearch isn't just about better extraction; it’s about building a foundation for verifiable, trustworthy clinical evidence generation.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck