Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models

arXiv:2608.30022 · cs.AI · Submitted 2026-08-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models".

Tom: This paper presents an end-to-end approach to automatically convert textual clinical guidelines from the UK National Institute for Health and Care Excellence (NICE) into an "executable model capable of generating…

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, to recap, we're looking at the paper "Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models," and it focuses on automating the translation of unstructured text guidelines into a usable computational model. Jane, can you explain what that means in plain language for our listeners?

Jane: Absolutely. Think about those official medical advice documents from NICE; they’re written in normal human language, which is great for doctors reading them, but it's hard for a computer to follow directly. This paper uses large language models to automatically build a formal model—a set of logical rules—that captures all the necessary logic from those guidelines.

Lu: It’s about taking something messy and converting it into something precise, like Answer Set Programming, which is a way for computers to reason under complex conditions. That formalization step seems like the core innovation here, moving beyond just simple text summarization.

Meng: I'm thinking about the authors; they used LLMs to do the heavy lifting of this transformation. My main question is about fidelity: how close is that final computational model to what a human expert actually intends when they write those guidelines?

Lalam: The implication here, in my view, is that we can start creating automated decision support tools for clinical care much faster than before, provided we can trust the translation process. This opens up possibilities for standardizing clinical reasoning across different regions or even different types of conditions.

Tom: Exactly. So it’s not just summarizing the text; they are building a structured, executable representation of the knowledge itself. We'll talk more about how they actually do this in the next part, and then we’ll look at what specific techniques they used to make that work.

The paper's summary: Jane: Well, the paper outlines a two-stage framework called Data-to-knowledge and knowledge-to-performance. Basically, the first stage takes the unstructured guidelines and converts them into an executable model using LLM prompting, which is what we’re discussing now.

Lu: The authors break it down into three specific steps for that conversion process: constant extraction, predicate generation, and rule generation. They use in-context examples to guide the model through these stages so it doesn't get lost during the translation.

Meng: That sounds like a careful decomposition strategy; I like that they aren't just hoping the LLM spits out a finished program. The idea of inspecting intermediate artifacts allows for human correction at each step, which is crucial for building something reliable.

Lalam: From a cultural perspective, this structured approach shows how we can systematically turn existing expert knowledge into machine-readable assets without having to manually rewrite every single guideline from scratch. That’s a big shift in how we manage medical documentation globally.

Tom: So they are using the LLM not just as a translator, but as an assistant that helps build the structure piece by piece—first identifying diseases and symptoms, then figuring out how those things relate to each other, and finally writing the actual logic rules.

Jane: Right. And they test this conversion process on pancreatic and lung cancer guidelines specifically to see if it works in a real-world clinical context before moving into the performance testing phase.

Lu: It’s interesting that they used a lightweight ontology aligned with SNOMED-CT concepts during the first step of constant extraction; that gives the LLM a concrete vocabulary to work with, which should help ground its understanding of medical terminology.

Meng: I’m curious about the results from that evaluation. Did they find significant issues where the logic was fundamentally broken, or were the problems more about missing small details? That tells us a lot about where we need to focus our engineering efforts.

The paper's improvements: Tom: Moving on to how they tackled potential weaknesses in this translation process, the paper suggests a few key improvements. They’re not just relying on one big prompt; they are decomposing the task and using these intermediate artifacts for inspection.

Jane: They also focus heavily on making sure the final output is actually usable by using a dedicated reasoning engine for the second stage, taking those generated models and applying them to actual patient data to produce recommendations.

Lu: That two-stage approach—D2K followed by K2P—is smart because it separates the knowledge encoding from the inference process. It’s about creating a patient-agnostic model first, and then plugging in the specific patient features later.

Meng: That decoupling is key for me; if you can separate what the guideline *says* from how it's *applied* to a specific person, you build in better checks and balances for practical deployment. It’s less prone to errors when dealing with real patient variability.

Lalam: This structure implies that the system is built with transparency in mind; because we can inspect the rules generated in the first stage, if something goes wrong later, we know exactly which part of the translation was faulty. That level of auditability is vital for any system touching patient care.

Tom: So they’re proposing a method where you build a solid logical foundation first and then layer on the specific patient context for decision support. This moves us toward models that are both general and highly personalized at the same time.

Conclusion: Jane: So to wrap up, this paper on "Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models" shows a feasible path toward taking massive amounts of unstructured medical text and turning it into structured, executable logic using LLMs as the engine for that translation.

Lu: The implication is that we can start automating the creation of computable guidelines, which means scalable knowledge representation becomes achievable without needing massive manual encoding efforts every single time.

Meng: From an engineering standpoint, the paper demonstrates a way to build a pipeline where you get intermediate checks, which significantly lowers the risk when moving from text to executable code. It shows how we can engineer trust into these translation layers before they even reach the inference stage.

Lalam: I think the cultural impact is that it democratizes access to evidence-based decision support by making complex guidelines usable by a wider range of clinical tools and researchers, not just those with deep expertise in formal logic programming.

Tom: It’s a significant step in automating guideline generation, opening the door for computable guidelines that preserve the original structure while allowing transparent inspection and modification. We’ve seen how they managed to build this bridge between natural language and executable logic here in "Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models."

Jane: It’s really exciting because it shows we can start building tools that don't just summarize but truly encode the conditional logic for clinical use. What a foundation for future work.

Ashvin Gupta, Denys Prociuk, Alessandra Russo, Brendan C. Delaney

Department of Computing, Imperial College London, London, UK · Department of Surgery and Cancer and I-X /Digital Foundry, Imperial College London, London, UK

cs.AI

Submitted: 2026-08-30

Updated: 2026-08-30

Comments: 18 pages. Published in Learning Health Systems (2026)

Journal ref: Learning Health Systems, 2026, 10(4):e70114. PubMed lists the article in volume 10, issue 4

DOI: 10.1002/lrh2.70114

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 72/100

The gist: This paper presents an end-to-end approach to automatically convert textual clinical guidelines from the UK National Institute for Health and Care Excellence (NICE) into an "executable model capable

Key concepts

NICE Guidelines
These are official medical advice documents from the UK National Institute for Health and Care Excellence. They are written in normal human language, which is difficult for computers to follow directly, making them a target for this conversion process.
Executable Computational Model
This is a formal model, like Answer Set Programming, that captures the precise logical rules from the guidelines. It moves beyond simple text summarization to create something computers can reason under complex conditions.
Data-to-Knowledge and Knowledge-to-Performance
This is a two-stage framework. The first stage converts unstructured guidelines into an executable model (Data-to-Knowledge), and the second stage applies that model to patient data to produce recommendations (Knowledge-to-Performance).
Intermediate Artifacts Inspection
The process involves breaking down the translation into steps like constant extraction, predicate generation, and rule generation. Inspecting these intermediate artifacts allows for human correction at each step, which is crucial for building a reliable system.

Terminology

Summary

This paper presents an end-to-end approach to automatically convert textual clinical guidelines from the UK National Institute for Health and Care Excellence (NICE) into an executable model capable of generating explainable patient-specific recommendations. This research is critical because while NICE guidelines provide evidence-based recommendations to support clinical care, they currently exist in unstructured natural language form, which prevents their direct use in automated clinical decision-support systems.

The Two-Stage Framework

The authors propose a framework divided into Data-to-knowledge (D2K) and knowledge-to-performance (K2P). The D2K stage utilizes LLM-based prompting to automatically convert NICE diagnostic guidelines into an executable model expressed in Answer Set Programming (ASP). ASP is a declarative programming paradigm well suited to knowledge representation and reasoning under complex, conditional logic. This process creates a patient-agnostic executable model that encodes the guideline logic.

The K2P stage then takes patient-specific features and the generated model, utilizing a reasoning engine to produce guideline-conformant recommendations. By combining these stages, the framework achieves Level 4 (executable) computable biomedical knowledge, as defined by existing biomedical knowledge frameworks.

The D2K Transformation Process

The D2K stage follows a three-step process that utilizes in-context examples to ensure semantic coherence. By decomposing the task, the researchers allow for inspection and correction at each intermediate stage through human-inspectable intermediate artifacts. The steps include:

  1. Constant extraction: The LLM identifies domain-specific constants, such as diseases and symptoms, using a lightweight ontology aligned with SNOMED-CT concepts.

  2. Predicate generation: The system identifies predicates that express relationships over the extracted constants, creating a reusable vocabulary of relations that form the backbone of the executable model.

  3. Rule generation: The LLM generates ASP rules that formalize the conditional logic expressed in the NICE guidelines, ensuring the output is a logically coherent representation.

Experimental Evaluation

The methodology was tested on pancreatic and lung cancer NICE guidelines and evaluated through both expert human review and clinical performance. During the D2K evaluation, clinicians assessed the alignment of the rules, finding strong alignment between the natural language guidelines and the generated executable models. Most discrepancies were partial omissions of specific details rather than incorrect logic, such as insufficiently specific predicates, and instances of hallucinated or fundamentally incorrect rules were rare. Inter-reviewer agreement, quantified using Cohen's κ coefficient, was found to be moderate to high.

To test the K2P stage, the researchers used 20 pancreatic cancer patient vignettes synthesized by oncologists. When executed, the resulting models produced patient-specific recommendations with an F1 score of 82.5%. The authors conclude that this work demonstrates the feasibility of automated guideline generation, which opens the door to scalable computable guidelines that preserve guideline structure and allow transparent inspection and modification.

Improvements for AI systems

1. Implementation of a Multi-Stage Neuro-Symbolic Transformation Pipeline

  • The Improvement: Replace end-to-end text-to-code generation with a decomposed, three-stage pipeline: (1) Constant Extraction (identifying domain-specific entities), (2) Predicate Schema Generation (defining relational vocabularies), and (3) Logical Rule Synthesis (encoding conditional logic).

  • What the improved system can do: It can transform highly complex, unstructured natural language manuals (medical, legal, or technical) into executable formal logic programs (such as Answer Set Programming) with high fidelity. By generating human-inspectable intermediate artifacts, the system allows for human-in-the-loop auditing, preventing the propagation of hallucinations from the initial text analysis to the final executable model.

2. Integration of Symbolic Reasoning Engines for Verifiable Decision Support

  • The Improvement: Decouple the linguistic processing from the logical reasoning by using LLMs strictly as translators into a formal logic language, while utilizing a dedicated symbolic solver (e.g., Clingo) for the actual inference.

  • What the improved system can do: It can provide decision-making recommendations that are 100% explainable and mathematically traceable. Instead of a black box prediction, the system can provide a precise logical proof showing exactly which guideline rule and which specific patient data points triggered a particular recommendation, meeting the safety and transparency requirements of high-stakes industries.

3. Structural In-Context Learning (ICL) for Rapid Domain Adaptation

  • The Improvement: Move beyond simple instruction-following to structural prompting, where the LLM is provided with multi-layered in-context examples that map [Natural Language to Constants to Predicates to ASP Rules].

  • What the improved system can do: The system can achieve rapid zero-shot or few-shot generalization across disparate domains. By simply updating the structural templates and examples, the system can pivot from encoding oncology guidelines to engineering safety protocols or financial regulatory frameworks without the need for expensive, domain-specific fine-tuning or retraining.

4. Semantic Alignment Layer for Patient-to-Predicate Mapping

  • The Improvement: Enhance the Knowledge-to-Performance (K2P) stage by implementing a semantic validation loop that forces the LLM to extract patient facts that strictly adhere to the predicate schema generated during the Data-to-Knowledge (D2K) phase.

  • What the improved system can do: It can significantly increase the precision and recall of automated decision support when operating on noisy, real-world data (such as Electronic Health Records). It prevents the logic mismatch errors where a patient's symptoms are identified but fail to trigger a rule because the extracted predicate does not perfectly match the rule's required arity or vocabulary.

Abstract

Introduction: NICE guidelines provide evidence-based recommendations for clinical care but remain largely in unstructured natural language. Existing approaches to converting them into computable representations often focus on individual diseases, require substantial manual encoding, and do not scale. Large language models (LLMs) may enable much of this translation to be automated. Methods: We present an end-to-end approach that converts textual clinical guidelines into executable models capable of generating explainable, patient-specific recommendations. A stepwise LLM-based transformation with in-context examples produces human-inspectable intermediate artifacts. We apply the approach to NICE pancreatic and lung cancer guidelines, use expert review to assess rule alignment, and evaluate the executable pancreatic cancer model on 20 patient vignettes. Results: Expert review showed strong alignment between the source guidelines and generated executable models. Most discrepancies were partial omissions rather than incorrect logic, while hallucinated or fundamentally incorrect rules were rare. On the patient vignettes, the executable model achieved an F1 score of 82.5%. Conclusion: LLMs can transform natural-language NICE guidelines into interpretable, executable models that preserve guideline structure, support transparent inspection and modification, and generate patient-specific recommendations. These findings demonstrate the feasibility of scalable automated generation of computable clinical guidelines.

Related papers