paper title
summary
The gist
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-theshelf large language models (LLMs) rank the correct disease first in only 35.4% of benchmark cases.
In short
The episode explores a paper introducing liteOdyssey, an agentic AI system designed for rare disease diagnosis. It translates expert clinical reasoning into a scalable, auditable natural language policy using Policy Iteration with Human Feedback (PIHF). The system demonstrates strong performance on public benchmarks and real-world clinical data.
Key concepts
- Agentic Diagnostic System
- This is a system where the AI's reasoning procedure is defined in an explicit natural-language policy, rather than being embedded in the model's internal weights. It follows a structured eight-phase workflow to guide the AI through diagnosis.
- Policy Iteration with Human Feedback (PIHF)
- This is the method used to create the policy. Experts evaluate AI failures and then write concrete 'policy deltas'—specific rule revisions—into a written policy document, ensuring human judgment drives improvements.
- Natural-Language Policy Artifact
- The entire reasoning workflow is consolidated into a coherent, unified document rather than separate modules. This artifact is versioned and auditable, allowing the AI to execute complex task behavior using expert knowledge in context.
Terminology used across episodes
This episode discusses
- Teaching agentic AI to generalize expert diagnostic reasoning in rare diseases · Paper Radio
- A Specialized Large Language Model for Clinical Reasoning and Diagnosis in Rare Diseases
- End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning
- Language Models are Few-Shot Learners
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows
- How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- Deep reinforcement learning from human preferences
- Training language models to follow instructions with human feedback
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Inference-Time Scaling for Generalist Reward Modeling
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Reward Is Enough: LLMs Are In-Context Reinforcement Learners
- Learning by Distilling Context
- On-Policy Context Distillation for Language Models
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Fine-Tuning Language Models from Human Preferences
- Kimi K3: Open Frontier Intelligence
The paper
Teaching agentic AI to generalize expert diagnostic reasoning in rare diseases · Read on arXiv
University1 · Company2
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer. Large language models rank the correct disease first in only 35.4% of benchmark cases and often rely on learned phenotype-disease associations rather than reusable diagnostic reasoning strategies. We developed liteOdyssey through Policy Iteration with Human Feedback, a process in which model failures and expert corrections are iteratively consolidated into a clinician-gated, natural-language policy executed by a language model. Across 1,243 public benchmark cases spanning 722 rare diseases, liteOdyssey ranked the correct disease first in 59.3% of cases versus 26.5% without the policy, with comparable gains in cases involving diseases excluded from policy development. The same policy transferred across model families and sizes without retraining. Adaptation of the policy to the Undiagnosed Diseases Network (UDN) improved diagnostic accuracy among 515 UDN patients, with gains confirmed by blinded physician adjudication. These results show that expert reasoning can be externalized into an inspectable and revisable natural-language policy that generalizes across rare diseases, transfers across model backbones, and adapts to a real-world patient cohort.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "paper title".
Jane: The paper was written by author1 and author2 from University1 and Company2.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We've established that current LLMs struggle with rare disease diagnosis, but now we need to understand exactly how liteOdyssey works and why its design is so novel.
Jane: It’s important to look at the core mechanism; it isn't just another retrieval-augmented generation system that looks up a solution.
Tom: The paper describes an agentic diagnostic system built on a single principle: the reasoning procedure lives in an explicit natural-language policy, not in the model weights or orchestration code.
Lu: That distinction is crucial; it’s shifting from optimizing internal weight space to defining the behavior of executing that policy.
Meng: So, when you give a patient’s clinical features—whether they're written as free text or mapped to HPO terms—the AI doesn't just guess, it follows this guided procedure end-to-end.
Jane: And this process is what the authors call the eight-phase diagnostic workflow: pattern recognition, candidate generation, triage, deep investigation, confidence assessment, corrective search, and finally adjudication.
Tom: It’s structured but not strictly linear; the model can go back to earlier phases if new evidence changes its mind.
Lu: That iterative nature is key for capturing the nuanced way a human geneticist would revisit an initial hypothesis when they see conflicting information.
Meng: The design ensures that the tool library and the reasoning workflow are treated as components of this single readable policy artifact, not separate modules.
Jane: This means we aren't just looking at a complex pipeline; we are looking at a coherent, unified reasoning document guiding the AI.
Summary: Tom: We’ve seen how it works, but let’s dig into the specific methodology—how did they actually create this policy?
Jane: They used something called Policy Iteration with Human Feedback, or PIHF for short. This is where the human expertise comes back in.
Tom: It sounds like a reinforcement learning loop, but without relying on a machine-readable reward signal.
Lu: Exactly; you have a frozen model running the current policy on development cases, and then an expert-designed credit-assignment system evaluates those runs against the benchmark outcomes.
Meng: The critical step is that when the AI fails, the expert doesn's just correct it, they write down a concrete policy delta—a specific rule revision—into a written policy document.
Jane: That’s such an elegant way to capture human judgment; instead of learning from human preference labels, the experts author the criteria and admit every persistent revision into the policy text.
Tom: It’s not just about fixing mistakes; it's about consolidating those failures into a versioned, auditable artifact.
Lu: The correspondence to traditional RL is structural rather than formal, but we are creating a "text delta" instead of a gradient step.
Meng: This is powerful because the AI model’s weights aren’t updated; the policy document is what that becomes-and it' stays auditable and human-readable.
Improvements: Tom: So, we have this beautifully engineered policy, but did it actually perform better than other systems?
Jane: The results on the public benchmark—one thousand two hundred forty-three cases across seven hundred twenty-two rare diseases—were incredibly strong.
Tom: They achieved a Recall@one of fifty-nine point three percent in finding the correct diagnosis first, which is dramatically higher than the parametric baseline at only twenty-six point five percent.
Lu: And this gain was consistent across both the LIRICAL and PhenoPacket Store datasets, suggesting that simple case-level performance isn' a fluke.
Meng: What’s most impressive from an engineering standpoint is how well it transfers—it worked across closed and open-weight models, including a three-billion active parameter model.
Jane: It seems to generalize extremely well, too; they tested it on six hundred seventy-nine diseases that were never seen during the policy development, which proves it's not just overfitting to those fifty development cases.
Tom: The system is also highly portable because the policy is a natural language document, not tied to one specific model architecture.
Lu: This portability is huge when we consider that in the real world, we use so many different models and resources.
Meng: Plus, its footprint—just forty-six MB for everything needed to run the policy on the public benchmarks—is tiny compared to other state-of-the-art systems.
Jane: It looks like a lean, highly effective solution that allows us to move beyond just what's already encoded in model training.
Conclusion: Tom: We’ve seen how liteOdyssey works, how the policy is built using PIHF, and we have some truly exciting performance numbers from the public benchmarks.
Jane: But we also looked at a real-world clinical cohort of five hundred fifteen patients from the Undiagnosed Diseases Network.
Tom: And even there, it showed measurable gains; on that complex patient data, liteOdyssey achieved a Recall@one of twenty point four percent, which is significantly better than the parametric baseline at sixteen point seven percent.
Lu: The fact that this policy successfully adapted to the atypical and multi-system presentations of the UDN cohort suggests it' has real diagnostic power in complex clinical environments.
Meng: It also passed expert clinical review—physicians rated its differentials as being more often exact and less often unhelpful than the baseline output.
Jane: This isn't just a theoretical improvement; it’s a practical tool that helps bridge the gap between human expertise and AI capability.
Tom: The core of this paper is that by externalizing expert reasoning into a scalable, auditable policy, we have found a way to make scarce diagnostic knowledge reusable for AI.
Lu: It's moving the needle from "how can we train this model?" to "what is the optimal process for how should this system think?"
Jane: That’s the promise of teaching agentic AI to learn expert reasoning, and it looks like a huge step forward for rare disease diagnosis.
Tom: It's definitely a breakthrough that I think we can all celebrate.
Title: Tom: We've talked about the problem and now we need to understand the core of "Teaching agentic AI to learn expert reasoning for rare disease diagnosis" and its authors' vision.
Jane: It starts with this critical issue that rare diseases are individually uncommon but collectively affect millions, leading to five-year diagnostic delays for patients.
Tom: This is a massive human cost, and the authors are addressing it by showing how AI can move beyond being just "off-the-shelf" knowledge.
Jane: They're introducing liteOdyssey as their agentic solution that aims to turn that expert reasoning into a scalable AI capability through a governed learning process.
Tom: It sounds like they aren't trying to brute force the model with massive fine-tuning, which is great for resource-constrained settings.
Lu: The authors's idea of converting expert reasoning into an explicit natural-language policy suggests they are looking at the theoretical capacity of pre-trained LLMs to execute complex task behavior in context.
Meng: That ability is what makes this practical—the system can take a complex set of clinical features and follow the policy without needing a massive, costly retraining effort.
Lalam: The implication for me is that we are shifting our view of AI from seeing it as an oracle that already knows everything, to seeing it as an agent that knows *how* to find the answer.
Tom: And Jane, you mentioned earlier how difficult it is to transfer this reasoning; the authors seem confident they' have found a path where transferable and reusable will be key.
Jane: They are making it scalable and portable, which is a huge win for clinical adoption.
Summary: Tom: We have the vision, but let’s focus on what the paper summarizes in terms of their approach to solving this problem with lightOdyssey.
Jane: The core idea is that PIHF—Policy Iteration with Human Feedback—is the engine driving it.
Tom: It’s a system where expert correction and LLM critics converge, essentially building a "policy" through human input.
Lu: This process, in-context policy-learning, allows us to develop a sound procedure comprising the reasoning process that fluidly adapts to various situations.
Meng: They start with an initial policy transcribed from clinical practice and then have the models run this policy on development cases to find failures.
Jane: When a failure occurs, the experts are brought in to interpret it, and they revise the specific rules in that written policy document.
Tom: So, rather than just retrying the whole model, they localize the failure to a rule and then refine that update into a concrete policy delta.
Lu: This is highly efficient; you are pinpointing exactly where the logic failed and correcting it using an expert-authored rubric.
Meng: The process allows us to consolidate improvements into this versioned document, which is an incredibly clear way to track how the system improved over time.
Jane: It’s a structured way of saying "we learned from that failure" without relying on massive, opaque backpropagation through a reward model.
Improvements: Tom: The results are where this all come together, and the improvements here seem genuinely impressive across many metrics.
Jane: The public benchmark results for the LIRICAL and PhenoPacket Store datasets showed a clear lift in diagnostic accuracy.
Tom: We're talking about fifty-nine point three percent of cases where liteOdyssey ranked the correct disease first, which is nearly double the twenty-six point five percent baseline performance without using that policy at all.
Lu: The fact that this performance persists on cases spanning six hundred seventy-nine diseases—all of them unseen during development—is a huge testament to its generalization capability.
Meng: It’s not just working on the training data; it’ is performing robust, across diverse and novel clinical presentations.
Jane: And it's also incredibly robust across hardware; you can use the same policy with different models, from Qwen3 point 6-35B to GPT-five point four, without any modification.
Tom: It seems like they have found a way to make scarce diagnostic expertise scalable and portable for every model in any deployment environment.
Lu: This is far beyond just "prompt engineering"; the the policy itself remains an explicit, readable document that experts can inspect and revise as medical knowledge evolves.
Meng: It's a robust engineering solution, not just a clever prompt trick, that provides verifiable gains across all the benchmarks.
Conclusion: Tom: We’ve covered the mechanism of PIHF and seen how it has dramatically improved performance on public benchmarks.
Jane: Now, let's summarize what this means for our listeners, especially regarding real-world application.
Tom: The final test was the Undiagnosed Diseases Network cohort, which had five hundred fifteen patients with incredibly complex cases that traditional care could not resolve.
Lu: And even though this group is naturally more challenging than benchmarks, the policy still delivered a solid boost to diagnostic accuracy in that difficult environment.
Meng: It confirms that by creating this policy, we have provided a tool that can adapt a diagnostic strategy developed on curated public data and apply it to real-world clinical populations.
Jane: And the human feedback loop is validated, with two physicians blinded to model identity confirming the output was more accurate and less unhelpful than the baseline.
Tom: Ultimately, this paper offers a path toward scaling diagnosis for rare diseases by making expert reasoning an AI capability that is both portable and human-governable.
Lu: It’s a way to make sure that when we are building these diagnostic agents, we aren're not just copying existing knowledge, but teaching the *process* of expert how to think.
Jane: We are thrilled with these results and the confidence in the future of AI-assisted rare disease diagnosis.
Title: Tom: Let’s start by talking about "Teaching agentic AI to learn expert reasoning for rare disease diagnosis" and why this is such a big deal for us today.
Jane: The paper highlights that while genetic sequencing has improved, the skill of integrating clinical and genetic findings to make a diagnosis remains elusive, taking years in many patients.
Tom: This system solves the problem by making that expert knowledge explicit, which is something very rare in LLM development right now.
Lu: I think it's exciting because we are moving away from the idea that AI must simply memorize vast amounts of data to be effective at diagnosis.
Meng: The authors are essentially building a set of instructions—a policy—that allows us to use an off-the-shelf model and make it act like a seasoned specialist without retraining it.
Jane: It’s about giving the AI the right strategy rather than just giving it a huge knowledge base.
Tom: The authors are demonstrating that we can teach agentic AI to learn the *methodology* of diagnosis, not just the answer itself.
Summary: Tom: Now, let's summarize the core of this work—the method behind PIHF and liteOdyssey.
Jane: The paper explains that PIHF is a way to translate expert clinical practice into a concrete policy that guides the AI’s decision-making process.
Tom: It's not just one big prompt; it’s an eight-phase structured workflow, from initial pattern recognition to the final ranked differential diagnosis.
Lu: This framework allows us to capture how a human expert might handle various steps, such as triaging evidence or performing a corrective search when confidence is low.
Meng: The process starts with an initial policy derived from experts and then uses the AI's failures as feedback to refine that written policy.
Jane: It’s like a structured refinement loop where the human judgment is central to updating the policy in-context.
Tom: We are taking those failures—missed phenotypes or poor logic—and writing them back into the policy as concrete rules for future runs.
Lu: This creates a durable, versioned artifact that is much more robust than a temporary set of instructions.
Improvements: Tom: The performance gains on the public benchmarks are what I think really drive home the success of this approach.
Jane: On those one thousand two hundred forty-three cases across seven hundred twenty-two unique diseases, liteOdyssey ranked the correct diagnosis first in a remarkable fifty-nine point three percent of cases.
Tom: That’s nearly double the baseline performance at twenty-six point five percent, which is incredible for an AI system that hasn's been trained on those specific diseases.
Lu: The generalization is the most impressive part; they have shown this policy works on cases where the disease was completely unseen during development.
Meng: It’s a very efficient solution, using only forty-six MB of footprint, which is a massive improvement over other state-of-the-art systems that require gigabytes.
Jane: And Tom, it does seem to hold up across models—it works on the smallest open-weight models and the largest closed frontier models.
Tom: It’s proof that this policy isn' portable across a different model family, not just a specific prompt tailored for one model.
Conclusion: Tom: We've seen how this works, but let’s wrap up by discussing the real-world impact and what this means for our listeners.
Jane: The evaluation on the Undiagnosed Diseases Network cohort showed that the policy could adapt to highly atypical clinical presentations.
Tom: That fifty-nine point three percent success rate on public data combined with its performance in that real-world, messy patient environment is a strong indicator of reliability.
Lu: It’s a clear sign that the method of "teaching" is teaching the process, not just the knowledge, and this has profound implications for how we view AI in medicine.
Meng: We can build systems like this to be auditable and deployable without having to worry about massive model retraining or scaling laws dictating our success.
Jane: It’s providing a path toward making that scarce diagnostic expertise scalable, reusable, and accessible.
Tom: A fantastic piece of research that allows AI to learn expert reasoning for rare diseases.
Title: Tom: Let's start by talking about "Teaching agentic AI to learn expert reasoning for rare disease diagnosis" and what the authors are trying to achieve with this concept.
Jane: The central problem is that traditional LLMs are struggling, achieving only about thirty-five point four percent success rate in finding the correct diagnosis first on public benchmarks.
Tom: This paper proposes a solution through a method called Policy Iteration with Human Feedback, or PIHF, to make that agentic AI capable of expert reasoning.
Lu: I see this as a crucial shift from simply having an AI search for answers to giving it the methodology of *how* to search and evaluate those answers.
Meng: The practical application is creating a robust policy that allows us to use existing, off-the-shelf models and make them act like highly skilled diagnostic agents.
Jane: It’s about providing a structure—a set of rules—that guides the AI toward the correct diagnosis even if it has no specific training on those rare conditions.
Tom: The authors want to show that this learned policy is something that can be inspected and revised by humans, which is a big deal for trust and compliance.
Summary: Tom: Now, let's summarize the mechanism of how this works in terms of the workflow.
Jane: It’s not a single step; it’s an eight-phase process—from pattern recognition to final adjudication—that is guided by the PIHF-derived policy.
Tom: This policy acts as a set of instructions for the AI, telling it which tools to use and when to revisit earlier steps if evidence changes.
Lu: The core idea is that in every phase, we are building a case for or against the candidate diseases based on the patient’s phenotype, much like a human geneticist would.
Meng: We take those failures from development cases—where the AI made mistakes—and refine them into this written policy document.
Jane: It’s an iterative learning loop where we use expert-designed credit assignment to pinpoint exactly why the AI failed at that phase, and then correct the policy accordingly.
Tom: This ensures that we are not just fixing one failure, but building a complete, robust procedure for all of those situations.
Improvements: Tom: Let's discuss the improvements and results of this approach in detail.
Jane: The performance jump on public benchmarks is quite dramatic; the system achieved fifty-nine point three percent accuracy in finding the correct disease first, up from twenty-six point five percent.
Tom: And what’s equally impressive is that it generalizes extremely well, performing comparably on over six hundred seventy-nine diseases that were never seen during its development.
Lu: This suggests that this isn't just a collection of solutions for specific cases, but a reusable diagnostic strategy for identifying and evaluating candidates.
Meng: The engineering impact is also clear: the lightweight footprint and the ability to run it across different models make it highly deployable.
Jane: It seems like we have found an incredibly powerful way to translate human expertise into an AI that has proven reliable on both public data and real-world clinical scenarios.
Conclusion: Tom: We've covered the mechanism, the results, and now we conclude by talking about what this means for the future of medicine.
Jane: The findings in the Undiagnosed Diseases Network cohort show that this policy can adapt to complex, real-world patient presentations that are difficult to classify.
Tom: It’s not just succeeding on neat datasets; it’ is working on patients who needed help the most.
Lu: This ability to learn and be adaptable suggests a shift toward agentic AI being the standard for high-stakes reasoning tasks in specialized fields.
Meng: We are looking at a future where we can deploy scalable, auditable diagnostic agents that are far more effective than today’s models.
Jane: The paper shows how to make scarce expert knowledge reusable and accessible to provide a path forward for rare disease diagnosis.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization