KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

summary

Video file (mp4)

The gist

This paper presents KREL (Knowledge-Guided Reasoning over Clinical Evidence with LLMs), a novel framework for Automatic Medical Coding (AMC).

In short

This episode discusses the KREL paper, which introduces a system for automatic medical coding using Large Language Models. By utilizing a three-stage process—Query Extractor, Candidate Selector, and Code Verifier—the system transforms unstructured clinical notes into standardized codes by grounding its reasoning in official medical rules and patient evidence.

Key concepts

Medical Coding
The process of converting unstructured clinical text, like doctor's notes, into standardized codes for hospital billing and research. This task is complex and requires following strict guidelines from organizations like the World Health Organization to ensure accuracy across healthcare environments.
Hierarchy-aware Beam Search
An algorithm used during the Candidate Selector stage to navigate a massive space of over seventy-two thousand possible codes. It uses a Knowledge Graph to understand how diseases are nested within each other, ensuring the system doesn't discard correct answers due to hierarchical differences.
Code Verifier
The final stage of the KREL system that acts as an audit layer. It checks candidate codes against the original clinical notes and specific coding rules to ensure every decision is grounded in actual evidence, which significantly improves the system's precision and reliability.

Terminology used across episodes

This episode discusses

The paper

KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs · Read on arXiv

Macquarie University · Beijing Intelligent Decision Medical Technology Company Limited · The University of New South Wales · University of Göttingen

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs".

Jane: The paper was written by Xubin Chen, Yipeng Zhou, Wen Sun, Chengkai Huang, Xiaoming Fu et al. from Macquarie University and Beijing Intelligent Decision Medical Technology Company Limited and The University of New South Wales and University of Göttingen.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs. That is quite a mouthful to start our show, Jane!

Jane: It definitely sounds like a lot to unpack, Tom, but the concept is actually very beautiful once you strip away all that academic jargon.

Tom: They're basically trying to solve the massive headache of medical coding, which is how hospitals turn messy doctor's notes into standardized codes for billing and research.

Jane: And it's a huge job that usually requires humans to follow incredibly complex guidelines from the World Health Organization to get it right.

Lu: What I find most exciting about this paper is how they use the term "Knowledge-Guided" in that title.

Tom: Do you think that's the real secret to their success, Lu?

Lu: Absolutely, because instead of just letting an AI guess based on patterns it saw online, they are actually feeding it official medical rules and hierarchies to guide its reasoning.

Meng: That sounds like a smart way to approach it, but I'm wondering how they prevent the model from just hallucinating when the clinical text gets confusing.

Jane: That’s exactly what the "Reasoning over Clinical Evidence" part is there for, Meng, because the system is designed to ground every single decision in actual text from the patient's file.

Meng: So it isn't just spitting out a code; it has to prove that the code is actually supported by what was written in the note?

Jane: That's exactly right, which makes it much more reliable for a real-world hospital environment.

Lalam: It also changes how we think about the relationship between technology and clinicians.

Tom: How do you see that playing out, Lalam?

Lalam: If we can automate this tedious administrative work with something this dependable, doctors can focus on being present with their patients instead of fighting with software.

Jane: It really turns the AI into a partner in care rather than just another digital chore to deal with.

Tom: It's a fascinating starting point, and I want to see how they actually built this engine to handle it.

Summary: Tom: Now that we've set the stage with KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs, let's look at the machinery inside.

Jane: They’ve broken the process down into three distinct stages: a Query Extractor, a Candidate Selector, and finally a Code Verifier.

Tom: And the first step is all about cleaning up that rambling clinical text, isn't it?

Jane: Exactly, the Query Extractor takes those long, unstructured notes and pulls out specific clinical descriptions to create very clear search queries.

Tom: Then they move into that Candidate Selector stage you mentioned earlier.

Jane: Right, and this is where they use a Knowledge Graph to navigate through more than seventy-two thousand different possible codes.

Lu: I love the way they handle that search using the Hierarchy-aware Beam Search algorithm.

Tom: Is that really different from just doing a standard keyword search, Lu?

Lu: It is much more sophisticated because it understands how diseases are nested within each other, so it doesn't accidentally discard a correct answer just because it looks a bit different at the top level of the hierarchy.

Meng: I'm thinking about the efficiency of that process, especially when you have such a massive search space to navigate.

Jane: They manage it by using an embedding model to score relevance and then focusing their search on the most promising branches of the medical tree.

Meng: That sounds like a much more scalable way to handle a huge label space than just asking an LLM to list everything it knows.

Lalam: It also ensures that the AI is following medical logic rather than just playing a game of probability.

Tom: Which brings us to the final gatekeeper, the Code Verifier.

Jane: The verifier is the part that takes those candidates and checks them against both the original note and any specific coding rules.

Tom: It's like a final audit to make sure nothing gets through that shouldn't.

Jane: It really is, but let's see if those layers of defense actually result in better performance in the real world.

Improvements: Tom: We've walked through the process, so let's get into the actual results for KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs.

Jane: The performance gains they reported are really striking, especially when you look at how they handled the full label space.

Tom: They managed to push the F1 score from zero point three two all the way up to zero point five one on the MDACE dataset!

Jane: That is a massive leap in accuracy, especially considering they aren't being given any hints about which codes might be in the note.

Lu: I was particularly impressed by how much better they performed on those tricky combination codes.

Tom: You mean the ones that require multiple pieces of evidence to be correct?

Lu: Yes, their recall for those combination codes hit zero point six one five, while the other models were struggling at less than zero point zero eight!

Meng: I also spent some time looking at their ablation studies to see if every part of the system was actually doing heavy lifting.

Jane: And they found that the verifier is absolutely critical for keeping things accurate.

Meng: Right, because when they removed the verifier, the precision dropped from zero point four nine down to a measly zero point one zero!

Jane: It basically turned into a system that just throws out random guesses without any real oversight.

Lalam: It really highlights that true intelligence in medical AI requires both the ability to find information and the discipline to verify it.

Tom: And they showed this wasn't just a fluke in one specific area, either.

Jane: They saw improvements across many different disease chapters, from infectious diseases to metabolic disorders.

Tom: It really feels like they've cracked the code on combining LLM reasoning with hard medical facts.

Conclusion: Tom: As we wrap up our discussion on KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs, it's clear this is a major milestone.

Jane: It really is, and it provides a very clear roadmap for how to use structured knowledge to make AI more reliable in high-stakes fields.

Tom: We've seen how they combined LLM reasoning with hard medical facts to solve a massive, complex problem.

Lu: I can see this evolving into even more advanced reasoning agents that might eventually help doctors identify patterns in rare diseases that are hard for humans to spot!

Meng: From an engineering standpoint, seeing this kind of structured reliability makes me much more confident about actually deploying these systems in hospital production environments.

Lalam: It's a step toward a culture where technology acts as a silent, supporting partner, allowing the human element of medicine to stay front and center.

Tom: That's a perfect way to end on, Lalam.

Jane: We've had such a great time breaking this one down with all of you today.

Tom: Thanks for listening, and we'll catch you at the next paper!

Jane: See you next time!

More episodes

← Home