Research Assistant: AstraZeneca's Agentic System for R&D

arXiv:2608.12395 · cs.AI · Submitted 2026-08-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Research Assistant: AstraZeneca's Agentic System for R&D".

Jane: The paper was written by Piotr Grabowski, Mohamed Alameen, Jorge Bretones, Sabina Cardell, Miguel Carmona et al. from AstraZeneca.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making the rounds on arXiv, and it's called "Research Assistant: AstraZeneca's Agentic System for R andD." Jane, when I first saw that title, I thought, okay, another chatbot for scientists, right? But this is actually something much bigger.

Jane: Oh, absolutely, Tom. And I love that you said that, because the title really undersells what they've built. This isn't just a chatbot. It's a whole multi-agent system that AstraZeneca deployed internally to help their scientists and clinicians dig through an enormous amount of biomedical data. We're talking literature, knowledge graphs, chemistry data, clinical trials, safety records, gene expression—all of it behind one chat interface.

Tom: And the scale is what gets me. Fifteen thousand internal users within a year. That's not a pilot project, that's a full-on cultural shift inside a pharmaceutical giant. I mean, before this, researchers had to write Cypher queries or use REST APIs to get at their knowledge graph. Now they just ask a question in plain English.

Jane: Exactly. And the authors make this really interesting point in the introduction. They had all this rich data infrastructure already built, but it required software engineering skills to use. So they built Research Assistant to lower that barrier. And the name itself, "Research Assistant," is kind of humble, but the architecture behind it is anything but. It's built on Apache Burr, which is a state machine framework, and it runs two distinct modes depending on the complexity of the question.

Tom: Right, and we're going to dig into those modes in a bit. But first, let's just sit with the implications of the title itself. "Agentic System" is a buzzword these days, but here it's earned. This is a system that doesn't just retrieve documents—it plans, it routes queries to specialized tools, it synthesizes evidence, and it cites its sources. For a company like AstraZeneca, that could mean faster drug discovery, better safety assessments, and fewer dead ends in research.

Jane: And honestly, Tom, the fact that it's from AstraZeneca and not a tech company is significant. It shows that big pharma is now building these systems in-house, tailored to their own data. That's a shift in who gets to build AI for science.

Tom: So, Jane, what's the one thing you'd tell a listener who thinks this is just another chatbot?

Jane: I'd tell them to look at the numbers. Sixteen cents per query. Ten to thirty seconds response time. Fifteen thousand users. Those are production numbers, not demo numbers. And that's the story of this paper—taking agentic AI from a cool demo to a daily workhorse.

Tom: And that's exactly where we're headed next. We're going to look at the summary and the core workflow, because that's where the real engineering decisions show up.

Summary: Jane: So, Tom, we've established that this is a serious system, not a toy. Now let's talk about how it actually works, because the summary in this paper is deceptively simple. The workflow is basically: take a user question, find the right data sources, ground the answer in evidence, and link back to the original sources. But underneath that simplicity, there's a really clever architecture.

Tom: And I want to bring in Lu here, because I think the multi-agent design is something you'd have strong opinions on. Lu, the paper describes this parallel architecture where a single query triggers multiple tool agents simultaneously. That's a big departure from the typical chain-of-thought approach.

Lu: It is, Tom, and I think it's the right call for this domain. The paper is very explicit about this: they run lightweight tool calls in parallel, then a larger LLM does the final synthesis. That's why they get responses in ten to thirty seconds. If they tried to do everything sequentially with one big model, the latency would kill the interactive experience. And the cost—sixteen cents per query—is remarkably low because they're not feeding the whole world into the context window. They're retrieving focused evidence.

Meng: And as an engineer, that's what I find most impressive. They're using Apache Burr, which is a state machine framework, to manage the orchestration. That gives them loops, conditional execution, and telemetry out of the box. They're not bolting this together with ad-hoc scripts. And they're streaming state updates to the frontend using server-sent events, so the user sees progress in real time. That's a polished engineering effort.

Jane: Meng, you mentioned the tool agents. Let's talk about those, because the paper lists twelve of them. There's a Literature Agent that searches over three point eight billion sentences from sixty-seven million documents. There's a Compound Agent that talks to AstraZeneca's chemistry gateway. There's a Knowledge Graph Agent, a Clinical Trial Agent, a Web Search Agent, even an OFF-X Agent for drug safety alerts.

Tom: And my favorite detail is the Tool Agent Picker. That's the router that decides which agents to activate for a given query. The paper says they tried embedding-based nearest neighbor search first, but it was hard to maintain. So they switched to topic modeling on tens of thousands of real user queries. They identified topics like "Biological Relationships" and "Drug Safety," and each topic maps to a set of agents. That's a really pragmatic design choice.

Lu: It is, and it addresses a fundamental problem in agentic systems: how do you decide what tools to use without running everything? The topic modeling approach is transparent and easy to update. And it lets the LLM operate on high-level concepts while the developers control which agents are best suited for each topic. That separation of concerns is smart.

Meng: And the grounding is what makes it trustworthy. Each agent returns an Observation object with an ID, a URL, the actual data, and a citation string. The final LLM response is injected with those IDs, so every statement can be traced back to a source. That's how you reduce hallucinations in a domain where accuracy is literally a matter of life and death.

Jane: Exactly. And the paper makes a point that the quality of the underlying data services is the primary contributor to system performance. The LLM is kept on a short leash—it's instructed to rely only on the retrieved evidence. That's a philosophy that runs through the whole design: the model is a synthesizer, not a knowledge base.

Tom: And that philosophy extends to the Deep Research Mode, which is what we're going to talk about next. Because that's where the system really shows off its planning capabilities.

Improvements: Jane: So we've covered the fast mode, which they call Scientific Mode. But the paper also describes a Deep Research Mode, and this is where I think the system really shines. Tom, this is the part where the system doesn't just answer a question—it plans a whole research project.

Tom: Right, and the way they describe it is fascinating. The system creates a research plan modeled as a directed acyclic graph of questions. It's like a flowchart of sub-questions that need to be answered in a certain order. And there's a judge agent that reviews the plan and either accepts it or sends it back for revision. Once accepted, the questions are topologically sorted, so independent questions get answered first, and dependent questions wait for the results.

Meng: And that's a really practical design, because it keeps the execution bounded. The plan is fixed at the start, so you know how many steps it'll take and roughly how much it'll cost. That's crucial for production systems. You can't have a research run that spirals out of control.

Lu: And there's a beautiful detail here. When the system executes the plan, later questions are rewritten using information from earlier answers. So if the first question identifies a synonym for a drug, the second question can use that synonym. That's how you get coherent multi-step reasoning instead of isolated lookups.

Tom: And the paper gives a concrete example: "What genes are associated with idiopathic pulmonary fibrosis?" The system generates a plan with multiple sub-questions, answers them in sequence, and synthesizes the results. But it can also handle iterative tasks, like "repeat this analysis for multiple genes." The planner translates that into a loop.

Jane: And there's another improvement I want to highlight, because it's about the Discover Agent. That's the one that does link prediction on the knowledge graph. It uses graph neural networks, knowledge graph embeddings like RotatE and ComplEx, and path-based methods. Then it combines the rankings using CRank to find predictions with the strongest agreement across models. That's the system making predictions, not just retrieving facts.

Meng: And that's where it gets interesting for practical use. The paper mentions that in Deep Research Mode, the system can first generate a prediction using Discover, then retrieve evidence that supports or contradicts that prediction. That's a hypothesis-generation loop, which is exactly what scientists do manually.

Lu: And I think that's the most significant improvement this paper suggests: moving from retrieval to reasoning. The system doesn't just find what's known; it helps researchers explore what might be true. That's a qualitative leap.

Tom: And the paper also describes how they evaluate the system. They used BioASQ for sanity checks, but they found that improvements based on real user feedback didn't always improve BioASQ scores. That's a really honest admission. Benchmarks are useful, but they don't capture what real users need.

Jane: And they built an LLM judge agent to score real interactions. That's a clever way to scale quality monitoring. Instead of manually reviewing thousands of queries, they have an automated judge that flags problems and helps them prioritize improvements.

Meng: And I have to say, the programmatic usage example at the end is a great proof point. They integrated Research Assistant into a safety assessment tool called CRAM Auto. It automatically answers questions like "What pathways are implicated in nausea?" and "Does the EGFR gene have a role in nausea?" That's the system being embedded into a regulatory workflow.

Lu: And that's the future. These systems won't just be chat interfaces. They'll be components in larger pipelines, providing grounded evidence to other AI systems. The paper mentions MCP and REST API endpoints, which makes Research Assistant a data service for other agents.

Tom: And that's the perfect segue to our conclusion, because we need to talk about what this all means for the world.

Conclusion: Jane: So, Tom, we've spent this whole episode on "Research Assistant: AstraZeneca's Agentic System for R andD," and I think the takeaway is clear. This is a system that works. It's deployed, it's used daily by fifteen thousand people, and it's changing how drug discovery happens at one of the biggest pharma companies in the world.

Tom: And the numbers are still stunning to me. Sixteen cents per query. Ten to thirty seconds response time. Those are the numbers that make agentic AI viable in an enterprise setting. And the paper is honest about the challenges—hallucinations, sensitivity to biological nuance like distinguishing between gene paralogs. But they're not treating those as unsolvable problems.

Lu: And I think the biggest implication is that this sets a template for other organizations. You don't need to be a tech giant to build this. AstraZeneca built it in-house, on top of their existing data infrastructure. The architecture is described in enough detail that other pharma companies, or even academic institutions, could replicate it.

Meng: And from an engineering standpoint, the use of Apache Burr and the topic-based agent routing are patterns that could be applied far beyond biomedicine. Any organization with large, fragmented data sources could benefit from this approach.

Jane: And the paper even acknowledges that it occupies a different space from more open-ended research systems like Google's CoScientist. Research Assistant is designed for day-to-day work, not autonomous discovery. But the CRAM Auto integration shows how it can be a building block for larger systems.

Tom: And that's what I'll remember about this paper. It's not flashy, but it's real. It's a working system that's making a difference in the real world, and it's documented in a way that others can learn from.

Jane: So with that, we're going to say goodbye to "Research Assistant: AstraZeneca's Agentic System for R andD." It's been a pleasure, and we're looking forward to seeing what AstraZeneca does next with it.

Tom: And we've got another paper coming up on the show, so stay tuned. Thanks for listening, everyone.

Jane: Take care, and we'll see you next time.

Piotr Grabowski, Mohamed Alameen, Jorge Bretones, Sabina Cardell, Miguel Carmona, Gavin Edwards, Ben Grainger, Sameh Hassan, Erik Jansson, Artur Kuziakhmetov, Albert Maristany, Hebatallah Mohamed, Andriy Nikolov, Sebastian Nilsson, Mark O'Donoghue, James Pacileo, Ashiq Sultan, Alex Voegele, Michaël Ughetto

AstraZeneca

cs.AI

Submitted: 2026-08-06

Comments: 16 pages, 3 figures

Code: https://github.com/567-labs/instructor

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 68/100

Key concepts

Agentic System
This system is more than just a chatbot. It actively plans research tasks, routes queries to specialized tools, synthesizes evidence from various sources, and provides citations for every statement. This capability allows it to move beyond simple document retrieval.
Multi-agent System
The system uses several specialized Tool Agents (e.g., Literature Agent, Clinical Trial Agent). A 'Tool Agent Picker' routes the user's query to the appropriate agents simultaneously, allowing for parallel processing of information and efficient response times.
Deep Research Mode
This mode allows the system to plan a full research project. It generates a flowchart of sub-questions, answers them in sequence, and uses results from earlier steps to inform later questions, enabling coherent multi-step reasoning.

Terminology

Summary

Summary

This paper describes Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range of data sources. The system provides a chat-style interface that brings together evidence from scientific literature, knowledge graphs, chemistry, clinical trials, safety resources, expression data, and internal experimental systems. It supports both a fast mode for direct question answering and a multi-step mode for more complex research tasks. Responses are grounded in retrieved evidence and linked back to the original sources, allowing users to review and further explore the underlying data. The paper outlines the system architecture, the main design choices behind the product, and lessons learned from deploying it at scale to support day-to-day R&D workflows across AstraZeneca.

The motivation for the system is that prior to ChatGPT, the team at AstraZeneca built and maintained an array of data services for consuming their Knowledge Graph and NLP pipeline, but these resources required software engineering experience to use, such as writing queries for REST APIs or querying the graph using Cypher. Research Assistant was developed to make internal biomedical data resources more accessible to domain experts. It grew from a small pilot to 15,000 unique internal users within one year.

The workflow of Research Assistant is: (1) Given a user question, find the most appropriate data sources and facts; (2) Ground the answer in retrieved evidence; (3) Link users to the original sources for manual evaluation and further exploration.

The backend was written in Python with heavy use of the asyncio package, the API layer was developed in FastAPI, and the framework at the heart of the application was Apache Burr, a light-weight state machine framework which allows building complete assistant-style applications involving loops and conditional execution. The frontend was written in TypeScript and React. During state machine execution, the backend streams messages with state updates using server-sent events asynchronously to the frontend. The system uses a parallel architecture in which a user query triggers multiple lightweight tool calls simultaneously, with a final synthesis step performed by a larger LLM. The average cost of a Research Assistant query is only 16 cents (including all input and output tokens, using Gemini 3 Flash and Gemini 3.1 Pro models). Responses are typically generated within 10 to 30 seconds.

Two modes are available to users: Scientific Mode and Deep Research Mode. The Scientific Mode balances complexity and depth with speed by performing a single feed-forward pass mapping the user query to relevant tool agents. The Deep Research Mode offers a multi-step planning and execution workflow built around the concept of a research plan modeled as a directed acyclic graph (DAG) of questions, simulating how a real human would approach a more complex problem. The application initially creates and revises the plan until a judge agent accepts it. The accepted DAG of questions is topologically sorted so independent questions are answered first, and dependent questions are rewritten using information obtained in earlier steps. Users can even ask questions about repeating analysis on multiple different entities, and the planner translates this into an iterative research plan. Users can also provide a very specific research plan to override automatic plan generation.

The Tool Agents are central to the system's ability to answer diverse questions. A Tool Agent is composed of the tool API and the prompt. The tool API translates keywords, accession numbers, IDs, and filter settings into function calls accessing internal AstraZeneca resources or external APIs. The prompt explains to the LLM how to map the free-text user query into a structured tool API query. The glue between the Tool Agent LLM calls and the tool APIs is the Instructor package, which facilitates creation and validation of structured Pydantic models. Each agent returns data packaged into the same Pydantic class instance (Observation) containing: ID of the observation, URL for the source of the data, a flexible container with grounding data, and a citation string.

The Tool Agents include: the Literature Agent, which connects to an internal NLP pipeline that ingests over 3.8 billion sentences from over 67 million documents covering sources such as PubMed, EuropePMC, Embase, bioRxiv, medRxiv, Wiley, Springer Nature, Dailymed, Insightmeme and AstraZeneca's internal Electronic Lab Notebook systems. The pipeline consists of deep-learning based biomedical named entity recognition using the KAZU pipeline, deep-learning based grammatical and syntactic structure extraction, and rule-based relation extraction. The processed sentences are indexed for fast retrieval using Elasticsearch. The system uses lemmatized versions of words and automatic translation of keywords into Named Entities to increase recall. For publication searches, hits are enriched with information from OpenAlex, including citation normalized percentile score and percentile-transformed journal-level h-index, which are summed with Elasticsearch scores to create a final relevance score.

The Compound Agent links to AstraZeneca's Chemistry Application Gateway, communicating via templated GraphQL queries to fetch information on activity values from biochemical assays, physicochemical properties, and translating SMILES into compound names. The Knowledge Graph Agent gives access to the Biological Insights Knowledge Graph. The authors found that letting an LLM access the whole graph schema and directly generate Cypher/SPARQL queries leads to issues: behaviour becomes non-deterministic, difficulty taking into account meta-level information, and generated queries can be too complex and lead to timeouts. The Knowledge Graph Agent's workflow involves: grounding entities mentioned in the question using the KAZU NER service and the BIKG mapping service; creating a prototype query based on a projected schema; adjusting the prototype query using deterministic rules; and executing the final query and returning a summary of results.

The Clinical Trial Agent retrieves data from ClinicalTrials.gov via the REST API, using a compositional approach to fetch only specific pieces of trial data relevant to the user's query, such as 'Basic Information', 'Eligibility and Participant Criteria', 'Outcomes and Endpoints', or 'Adverse Events and Publications'. The Web Search Agent uses Google Vertex AI platform to perform Grounding with Google Search with the Gemini 3 Flash model, turning the Gemini response into Observation objects. The authors note that the Web Search Agent can access sources that are not scientifically credible, controlled with a domain exclusion list, and that additional care and verification of sources is required. The Discover Agent supports predictive capabilities using graph machine learning approaches including Graph Convolutional Neural Networks, knowledge graph embedding models (RotatE, ComplEx), path-based methods (Degree-Weighted Path Count, Sörensen Similarity), and PageRank. Rankings are combined using CRank to identify predictions with strongest agreement across models. The Clinical Endpoints Agent is based on a ClickHouse database produced by the NLP Endpoints pipeline, which applies four models based on Google's BERT model to extract entities relating to drug properties and clinical trials read-outs such as RECIST, drug efficacy, drug safety, and pharmacokinetic and pharmacodynamic drug properties. The Mapping Agent extracts entities from user queries and retrieves synonyms from the BIKG node index, allowing the system to recognize that EGFR is also called ERBB or HER1. The Glossary Agent translates various terms used within AstraZeneca, such as acronyms. The Human Protein Atlas Agent obtains expression patterns of genes from the HPA REST API. The In Vivo Agent discovers preclinical data from drug experiments, including toxicity studies, animal models, compound testing, and dosing. The OFF-X Agent retrieves information on drug safety alerts from the Clarivate OFF-X API.

The Tool Agent Picker is a router module responsible for mapping the user query to sets of Tool Agents. The authors experimented with an approach using a handcrafted database of example queries embedded and stored in ChromaDB, but found it hard to maintain and dropped it. They subsequently adopted an approach using automated topic modeling on a database containing tens of thousands of real user interactions, shortlisting topics such as Chemistry Information, Biological Relationships, General Biomedical Knowledge, Drug Safety and Recent Events. At run time, the user query is automatically assigned these predefined topics and the union of the Tool Agents assigned to the selected topics is run in parallel.

For evaluations, the authors used the BioASQ Training 10b question-answer set, specifically a random subset of 100 yes/no questions to calculate balanced accuracy compared to vanilla LLMs like GPT 4. They note that improvements made based on user feedback did not translate to performance improvements on the BioASQ question set, exemplifying the orthogonality of biomedical QA benchmarks and real-life user requirements. They also used the STaRK dataset, mapping entities in PrimeKG to the internal AstraZeneca knowledge graph BIKG, to drive development of the Knowledge Graph tool agent. With the growing internal user base, they developed a simple LLM judge agent which, given a pair of real user question and system response, would assign a score and add comments about that interaction.

An example of programmatic use of Research Assistant is the CRAM Auto Tool, which supports Combination Risk Assessment for Patient Safety Scientists. This process evaluates the safety profile when two or more products are used in combination, especially in clinical trials. A critical step is determining whether modulation of a drug target is linked to a biological mechanism that could plausibly contribute to an adverse event. The CRAM Auto Tool embedded this usage pattern into the product by calling Research Assistant programmatically, retrieving relevant evidence and generating grounded outputs with citations.

The authors conclude that Research Assistant was designed around three practical priorities: data accuracy, response speed, and low operating cost. It occupies a different space from more open-ended autonomous research systems, such as Google CoScientist or Robin, and is aimed at helping scientists and clinicians quickly access grounded biomedical information in day-to-day R&D work. Research Assistant can also serve as a grounded biomedical data endpoint for larger agentic systems through its MCP and REST API endpoints. Important challenges remain, including hallucinations and limited sensitivity to biological nuance, such as distinctions between closely related gene paralogs.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems, along with the resulting capabilities:

  • Replace single-pass LLM generation with a parallel tool-agent system where multiple specialized agents (literature, knowledge graph, clinical trials, chemistry) execute concurrently

  • Implement a two-stage pipeline: (a) retrieve structured evidence from domain-specific APIs, (b) synthesize answers using a larger LLM only at the final step

  • Add a mandatory citation/observation layer that tags every generated statement with source URLs and metadata

  • Replace free-form LLM tool selection with a two-tier approach: (a) classify user query into predefined topics (e.g., biological relationships, drug safety, chemistry information) using lightweight topic modeling, (b) map topics to fixed sets of tool agents via a deterministic lookup table

  • This reduces hallucination risk and improves consistency compared to letting the LLM decide which tools to call

  • For knowledge graph access, implement a three-step process: (a) entity grounding using NER to resolve mentions to canonical IDs, (b) prototype query generation against a projected schema (not the full graph), (c) deterministic rule-based adjustment for source filtering and attribute selection

  • This prevents non-deterministic graph queries and timeouts

  • For large external records (e.g., clinical trials), fetch only the specific data blocks requested by the user (e.g., eligibility criteria, endpoints, adverse events) rather than full records

  • This reduces token consumption and latency

  • Combine semantic search scores with publication impact metrics (citation-normalized percentile, journal h-index) using a weighted sum for literature retrieval

  • This prevents low-quality but keyword-matching sources from dominating results

  • Implement a deep research mode that: (a) generates a directed acyclic graph of sub-questions, (b) validates it with a judge agent (with revision limits), (c) topologically sorts and executes sub-questions using the fast mode, (d) rewrites dependent questions using context from prior answers

  • This enables complex, multi-hop queries while maintaining bounded runtime and token budget

  • Integrate multiple graph ML models (GCN, RotatE, ComplEx, path-based methods) and combine their ranked predictions using CRank to identify high-confidence novel gene-disease-compound associations

  • Use this only in deep research mode, followed by evidence retrieval to support or refute predictions

  • Answers questions like What causes Sturge-Weber Syndrome? with citation-linked evidence from literature, knowledge graphs, and clinical trials in 10–30 seconds

  • Distinguishes between direct gene-disease associations and inferred pathways (e.g., via intermediate pathways or drugs in trial)

  • Resolves synonyms automatically (e.g., EGFR = ERBB = HER1) before searching

  • Handles multi-step queries like What classes of compounds affect genes expressed in the liver? by planning sub-questions, executing them sequentially, and synthesizing a coherent answer

  • Supports iterative analyses across multiple entities (e.g., Repeat this analysis for genes A, B, C) with automatic plan generation

  • Can generate novel hypotheses (e.g., Gene X is likely associated with Disease Y) and then retrieve supporting/contradicting evidence

  • Provides drug safety alerts from OFF-X, adverse event data, and target safety profiles with source citations

  • Supports combination risk assessment by answering questions like Does EGFR inhibition have a role in nausea? with grounded evidence

  • Exposes REST API and MCP endpoints, allowing other agentic systems to query biomedical data with grounded, citation-linked responses

  • Maintains low per-query cost (0.16) and predictable latency, making it suitable for batch processing and integration into larger workflows

  • Automatically filters out non-credible sources (e.g., forums, blogs) via domain exclusion lists

  • Uses a judge LLM to score real user interactions, enabling rapid identification of system weaknesses and targeted improvements

Abstract

We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range of data sources. The system provides a chat-style interface that brings together evidence from scientific literature, knowledge graphs, chemistry, clinical trials, safety resources, expression data, and internal experimental systems. It supports both a fast mode for direct question answering and a multi-step mode for more complex research tasks. Responses are grounded in retrieved evidence and linked back to the original sources, allowing users to review and further explore the underlying data. In this technical note, we outline the system architecture, the main design choices behind the product, and lessons learned from deploying it at scale to support day-to-day R&D workflows across AstraZeneca.

Related papers