Teaching agentic AI to generalize expert diagnostic reasoning in rare diseases
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "paper title".
Jane: The paper was written by author1 and author2 from University1 and Company2.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We've established that current LLMs struggle with rare disease diagnosis, but now we need to understand exactly how liteOdyssey works and why its design is so novel.
Jane: It’s important to look at the core mechanism; it isn't just another retrieval-augmented generation system that looks up a solution.
Tom: The paper describes an agentic diagnostic system built on a single principle: the reasoning procedure lives in an explicit natural-language policy, not in the model weights or orchestration code.
Lu: That distinction is crucial; it’s shifting from optimizing internal weight space to defining the behavior of executing that policy.
Meng: So, when you give a patient’s clinical features—whether they're written as free text or mapped to HPO terms—the AI doesn't just guess, it follows this guided procedure end-to-end.
Jane: And this process is what the authors call the eight-phase diagnostic workflow: pattern recognition, candidate generation, triage, deep investigation, confidence assessment, corrective search, and finally adjudication.
Tom: It’s structured but not strictly linear; the model can go back to earlier phases if new evidence changes its mind.
Lu: That iterative nature is key for capturing the nuanced way a human geneticist would revisit an initial hypothesis when they see conflicting information.
Meng: The design ensures that the tool library and the reasoning workflow are treated as components of this single readable policy artifact, not separate modules.
Jane: This means we aren't just looking at a complex pipeline; we are looking at a coherent, unified reasoning document guiding the AI.
Summary: Tom: We’ve seen how it works, but let’s dig into the specific methodology—how did they actually create this policy?
Jane: They used something called Policy Iteration with Human Feedback, or PIHF for short. This is where the human expertise comes back in.
Tom: It sounds like a reinforcement learning loop, but without relying on a machine-readable reward signal.
Lu: Exactly; you have a frozen model running the current policy on development cases, and then an expert-designed credit-assignment system evaluates those runs against the benchmark outcomes.
Meng: The critical step is that when the AI fails, the expert doesn's just correct it, they write down a concrete policy delta—a specific rule revision—into a written policy document.
Jane: That’s such an elegant way to capture human judgment; instead of learning from human preference labels, the experts author the criteria and admit every persistent revision into the policy text.
Tom: It’s not just about fixing mistakes; it's about consolidating those failures into a versioned, auditable artifact.
Lu: The correspondence to traditional RL is structural rather than formal, but we are creating a "text delta" instead of a gradient step.
Meng: This is powerful because the AI model’s weights aren’t updated; the policy document is what that becomes-and it' stays auditable and human-readable.
Improvements: Tom: So, we have this beautifully engineered policy, but did it actually perform better than other systems?
Jane: The results on the public benchmark—one thousand two hundred forty-three cases across seven hundred twenty-two rare diseases—were incredibly strong.
Tom: They achieved a Recall@one of fifty-nine point three percent in finding the correct diagnosis first, which is dramatically higher than the parametric baseline at only twenty-six point five percent.
Lu: And this gain was consistent across both the LIRICAL and PhenoPacket Store datasets, suggesting that simple case-level performance isn' a fluke.
Meng: What’s most impressive from an engineering standpoint is how well it transfers—it worked across closed and open-weight models, including a three-billion active parameter model.
Jane: It seems to generalize extremely well, too; they tested it on six hundred seventy-nine diseases that were never seen during the policy development, which proves it's not just overfitting to those fifty development cases.
Tom: The system is also highly portable because the policy is a natural language document, not tied to one specific model architecture.
Lu: This portability is huge when we consider that in the real world, we use so many different models and resources.
Meng: Plus, its footprint—just forty-six MB for everything needed to run the policy on the public benchmarks—is tiny compared to other state-of-the-art systems.
Jane: It looks like a lean, highly effective solution that allows us to move beyond just what's already encoded in model training.
Conclusion: Tom: We’ve seen how liteOdyssey works, how the policy is built using PIHF, and we have some truly exciting performance numbers from the public benchmarks.
Jane: But we also looked at a real-world clinical cohort of five hundred fifteen patients from the Undiagnosed Diseases Network.
Tom: And even there, it showed measurable gains; on that complex patient data, liteOdyssey achieved a Recall@one of twenty point four percent, which is significantly better than the parametric baseline at sixteen point seven percent.
Lu: The fact that this policy successfully adapted to the atypical and multi-system presentations of the UDN cohort suggests it' has real diagnostic power in complex clinical environments.
Meng: It also passed expert clinical review—physicians rated its differentials as being more often exact and less often unhelpful than the baseline output.
Jane: This isn't just a theoretical improvement; it’s a practical tool that helps bridge the gap between human expertise and AI capability.
Tom: The core of this paper is that by externalizing expert reasoning into a scalable, auditable policy, we have found a way to make scarce diagnostic knowledge reusable for AI.
Lu: It's moving the needle from "how can we train this model?" to "what is the optimal process for how should this system think?"
Jane: That’s the promise of teaching agentic AI to learn expert reasoning, and it looks like a huge step forward for rare disease diagnosis.
Tom: It's definitely a breakthrough that I think we can all celebrate.
Title: Tom: We've talked about the problem and now we need to understand the core of "Teaching agentic AI to learn expert reasoning for rare disease diagnosis" and its authors' vision.
Jane: It starts with this critical issue that rare diseases are individually uncommon but collectively affect millions, leading to five-year diagnostic delays for patients.
Tom: This is a massive human cost, and the authors are addressing it by showing how AI can move beyond being just "off-the-shelf" knowledge.
Jane: They're introducing liteOdyssey as their agentic solution that aims to turn that expert reasoning into a scalable AI capability through a governed learning process.
Tom: It sounds like they aren't trying to brute force the model with massive fine-tuning, which is great for resource-constrained settings.
Lu: The authors's idea of converting expert reasoning into an explicit natural-language policy suggests they are looking at the theoretical capacity of pre-trained LLMs to execute complex task behavior in context.
Meng: That ability is what makes this practical—the system can take a complex set of clinical features and follow the policy without needing a massive, costly retraining effort.
Lalam: The implication for me is that we are shifting our view of AI from seeing it as an oracle that already knows everything, to seeing it as an agent that knows *how* to find the answer.
Tom: And Jane, you mentioned earlier how difficult it is to transfer this reasoning; the authors seem confident they' have found a path where transferable and reusable will be key.
Jane: They are making it scalable and portable, which is a huge win for clinical adoption.
Summary: Tom: We have the vision, but let’s focus on what the paper summarizes in terms of their approach to solving this problem with lightOdyssey.
Jane: The core idea is that PIHF—Policy Iteration with Human Feedback—is the engine driving it.
Tom: It’s a system where expert correction and LLM critics converge, essentially building a "policy" through human input.
Lu: This process, in-context policy-learning, allows us to develop a sound procedure comprising the reasoning process that fluidly adapts to various situations.
Meng: They start with an initial policy transcribed from clinical practice and then have the models run this policy on development cases to find failures.
Jane: When a failure occurs, the experts are brought in to interpret it, and they revise the specific rules in that written policy document.
Tom: So, rather than just retrying the whole model, they localize the failure to a rule and then refine that update into a concrete policy delta.
Lu: This is highly efficient; you are pinpointing exactly where the logic failed and correcting it using an expert-authored rubric.
Meng: The process allows us to consolidate improvements into this versioned document, which is an incredibly clear way to track how the system improved over time.
Jane: It’s a structured way of saying "we learned from that failure" without relying on massive, opaque backpropagation through a reward model.
Improvements: Tom: The results are where this all come together, and the improvements here seem genuinely impressive across many metrics.
Jane: The public benchmark results for the LIRICAL and PhenoPacket Store datasets showed a clear lift in diagnostic accuracy.
Tom: We're talking about fifty-nine point three percent of cases where liteOdyssey ranked the correct disease first, which is nearly double the twenty-six point five percent baseline performance without using that policy at all.
Lu: The fact that this performance persists on cases spanning six hundred seventy-nine diseases—all of them unseen during development—is a huge testament to its generalization capability.
Meng: It’s not just working on the training data; it’ is performing robust, across diverse and novel clinical presentations.
Jane: And it's also incredibly robust across hardware; you can use the same policy with different models, from Qwen3 point 6-35B to GPT-five point four, without any modification.
Tom: It seems like they have found a way to make scarce diagnostic expertise scalable and portable for every model in any deployment environment.
Lu: This is far beyond just "prompt engineering"; the the policy itself remains an explicit, readable document that experts can inspect and revise as medical knowledge evolves.
Meng: It's a robust engineering solution, not just a clever prompt trick, that provides verifiable gains across all the benchmarks.
Conclusion: Tom: We’ve covered the mechanism of PIHF and seen how it has dramatically improved performance on public benchmarks.
Jane: Now, let's summarize what this means for our listeners, especially regarding real-world application.
Tom: The final test was the Undiagnosed Diseases Network cohort, which had five hundred fifteen patients with incredibly complex cases that traditional care could not resolve.
Lu: And even though this group is naturally more challenging than benchmarks, the policy still delivered a solid boost to diagnostic accuracy in that difficult environment.
Meng: It confirms that by creating this policy, we have provided a tool that can adapt a diagnostic strategy developed on curated public data and apply it to real-world clinical populations.
Jane: And the human feedback loop is validated, with two physicians blinded to model identity confirming the output was more accurate and less unhelpful than the baseline.
Tom: Ultimately, this paper offers a path toward scaling diagnosis for rare diseases by making expert reasoning an AI capability that is both portable and human-governable.
Lu: It’s a way to make sure that when we are building these diagnostic agents, we aren're not just copying existing knowledge, but teaching the *process* of expert how to think.
Jane: We are thrilled with these results and the confidence in the future of AI-assisted rare disease diagnosis.
Title: Tom: Let’s start by talking about "Teaching agentic AI to learn expert reasoning for rare disease diagnosis" and why this is such a big deal for us today.
Jane: The paper highlights that while genetic sequencing has improved, the skill of integrating clinical and genetic findings to make a diagnosis remains elusive, taking years in many patients.
Tom: This system solves the problem by making that expert knowledge explicit, which is something very rare in LLM development right now.
Lu: I think it's exciting because we are moving away from the idea that AI must simply memorize vast amounts of data to be effective at diagnosis.
Meng: The authors are essentially building a set of instructions—a policy—that allows us to use an off-the-shelf model and make it act like a seasoned specialist without retraining it.
Jane: It’s about giving the AI the right strategy rather than just giving it a huge knowledge base.
Tom: The authors are demonstrating that we can teach agentic AI to learn the *methodology* of diagnosis, not just the answer itself.
Summary: Tom: Now, let's summarize the core of this work—the method behind PIHF and liteOdyssey.
Jane: The paper explains that PIHF is a way to translate expert clinical practice into a concrete policy that guides the AI’s decision-making process.
Tom: It's not just one big prompt; it’s an eight-phase structured workflow, from initial pattern recognition to the final ranked differential diagnosis.
Lu: This framework allows us to capture how a human expert might handle various steps, such as triaging evidence or performing a corrective search when confidence is low.
Meng: The process starts with an initial policy derived from experts and then uses the AI's failures as feedback to refine that written policy.
Jane: It’s like a structured refinement loop where the human judgment is central to updating the policy in-context.
Tom: We are taking those failures—missed phenotypes or poor logic—and writing them back into the policy as concrete rules for future runs.
Lu: This creates a durable, versioned artifact that is much more robust than a temporary set of instructions.
Improvements: Tom: The performance gains on the public benchmarks are what I think really drive home the success of this approach.
Jane: On those one thousand two hundred forty-three cases across seven hundred twenty-two unique diseases, liteOdyssey ranked the correct diagnosis first in a remarkable fifty-nine point three percent of cases.
Tom: That’s nearly double the baseline performance at twenty-six point five percent, which is incredible for an AI system that hasn's been trained on those specific diseases.
Lu: The generalization is the most impressive part; they have shown this policy works on cases where the disease was completely unseen during development.
Meng: It’s a very efficient solution, using only forty-six MB of footprint, which is a massive improvement over other state-of-the-art systems that require gigabytes.
Jane: And Tom, it does seem to hold up across models—it works on the smallest open-weight models and the largest closed frontier models.
Tom: It’s proof that this policy isn' portable across a different model family, not just a specific prompt tailored for one model.
Conclusion: Tom: We've seen how this works, but let’s wrap up by discussing the real-world impact and what this means for our listeners.
Jane: The evaluation on the Undiagnosed Diseases Network cohort showed that the policy could adapt to highly atypical clinical presentations.
Tom: That fifty-nine point three percent success rate on public data combined with its performance in that real-world, messy patient environment is a strong indicator of reliability.
Lu: It’s a clear sign that the method of "teaching" is teaching the process, not just the knowledge, and this has profound implications for how we view AI in medicine.
Meng: We can build systems like this to be auditable and deployable without having to worry about massive model retraining or scaling laws dictating our success.
Jane: It’s providing a path toward making that scarce diagnostic expertise scalable, reusable, and accessible.
Tom: A fantastic piece of research that allows AI to learn expert reasoning for rare diseases.
Title: Tom: Let's start by talking about "Teaching agentic AI to learn expert reasoning for rare disease diagnosis" and what the authors are trying to achieve with this concept.
Jane: The central problem is that traditional LLMs are struggling, achieving only about thirty-five point four percent success rate in finding the correct diagnosis first on public benchmarks.
Tom: This paper proposes a solution through a method called Policy Iteration with Human Feedback, or PIHF, to make that agentic AI capable of expert reasoning.
Lu: I see this as a crucial shift from simply having an AI search for answers to giving it the methodology of *how* to search and evaluate those answers.
Meng: The practical application is creating a robust policy that allows us to use existing, off-the-shelf models and make them act like highly skilled diagnostic agents.
Jane: It’s about providing a structure—a set of rules—that guides the AI toward the correct diagnosis even if it has no specific training on those rare conditions.
Tom: The authors want to show that this learned policy is something that can be inspected and revised by humans, which is a big deal for trust and compliance.
Summary: Tom: Now, let's summarize the mechanism of how this works in terms of the workflow.
Jane: It’s not a single step; it’s an eight-phase process—from pattern recognition to final adjudication—that is guided by the PIHF-derived policy.
Tom: This policy acts as a set of instructions for the AI, telling it which tools to use and when to revisit earlier steps if evidence changes.
Lu: The core idea is that in every phase, we are building a case for or against the candidate diseases based on the patient’s phenotype, much like a human geneticist would.
Meng: We take those failures from development cases—where the AI made mistakes—and refine them into this written policy document.
Jane: It’s an iterative learning loop where we use expert-designed credit assignment to pinpoint exactly why the AI failed at that phase, and then correct the policy accordingly.
Tom: This ensures that we are not just fixing one failure, but building a complete, robust procedure for all of those situations.
Improvements: Tom: Let's discuss the improvements and results of this approach in detail.
Jane: The performance jump on public benchmarks is quite dramatic; the system achieved fifty-nine point three percent accuracy in finding the correct disease first, up from twenty-six point five percent.
Tom: And what’s equally impressive is that it generalizes extremely well, performing comparably on over six hundred seventy-nine diseases that were never seen during its development.
Lu: This suggests that this isn't just a collection of solutions for specific cases, but a reusable diagnostic strategy for identifying and evaluating candidates.
Meng: The engineering impact is also clear: the lightweight footprint and the ability to run it across different models make it highly deployable.
Jane: It seems like we have found an incredibly powerful way to translate human expertise into an AI that has proven reliable on both public data and real-world clinical scenarios.
Conclusion: Tom: We've covered the mechanism, the results, and now we conclude by talking about what this means for the future of medicine.
Jane: The findings in the Undiagnosed Diseases Network cohort show that this policy can adapt to complex, real-world patient presentations that are difficult to classify.
Tom: It’s not just succeeding on neat datasets; it’ is working on patients who needed help the most.
Lu: This ability to learn and be adaptable suggests a shift toward agentic AI being the standard for high-stakes reasoning tasks in specialized fields.
Meng: We are looking at a future where we can deploy scalable, auditable diagnostic agents that are far more effective than today’s models.
Jane: The paper shows how to make scarce expert knowledge reusable and accessible to provide a path forward for rare disease diagnosis.
University1 · Company2
cs.AI
Submitted: 2026-06-15
Updated: 2026-09-07
Comments: Updated ablation experiments and layout; revised manuscript to match the current submission
Code: https://github.com/minhha0510/Rare-Disease-Meta-Analysis-2026
Project page: https://liteodyssey.org
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 94/100
The gist: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-theshelf large language models (LLMs) rank the correct disease first in only 35.4% of benchmark cases.
Key concepts
- Agentic Diagnostic System
- This is a system where the AI's reasoning procedure is defined in an explicit natural-language policy, rather than being embedded in the model's internal weights. It follows a structured eight-phase workflow to guide the AI through diagnosis.
- Policy Iteration with Human Feedback (PIHF)
- This is the method used to create the policy. Experts evaluate AI failures and then write concrete 'policy deltas'—specific rule revisions—into a written policy document, ensuring human judgment drives improvements.
- Natural-Language Policy Artifact
- The entire reasoning workflow is consolidated into a coherent, unified document rather than separate modules. This artifact is versioned and auditable, allowing the AI to execute complex task behavior using expert knowledge in context.
Terminology
Summary
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-theshelf large language models (LLMs) rank the correct disease first in only 35.4% of benchmark cases. Here we show that this expert reasoning can be converted into a scalable AI capability through a governed learning process rather than model training alone. We developed liteOdyssey through Policy Iteration with Human Feedback (PIHF), an in-context policy-learning method adapted from generalized policy iteration in reinforcement learning, in which model failures and expert corrections consolidate into an explicit, clinician-gated policy that turns an off-the-shelf LLM into an agentic diagnostic system. We demonstrated that such a policy improved diagnostic accuracy to match the best published systems at a fraction of their deployment footprint, generalized to unseen diseases, transferred across models, and remained under clinician control. Across 1,243 public benchmark cases spanning 722 rare diseases, liteOdyssey ranked the correct disease first in 59.3% of cases versus 26.5% without the policy, with nearly identical gains on the 1,193 cases and 679 diseases excluded from policy development. Ablations showed that gains exceeded automated prompting improvement or source access alone, and the policy transferred without modification across closed- and open-weight models. In 515 Undiagnosed Diseases Network patients, liteOdyssey again improved accuracy, and blinded physicians rated its differentials more often exact and less often unhelpful. Through PIHF, expert reasoning becomes an LLM capability that experts can inspect, revise, and transfer across models.
Improvements for AI systems
Based on a rigorous analysis of this research, I have identified several critical areas where our current implementation of liteOdyssey can be enhanced to improve AI systems in clinical settings, addressing both scalability and diagnostic precision.
The core strength of liteOdyssey is its shift from model-internal knowledge (weights) to an external, auditable policy artifact. My proposed improvements focus on optimizing this artifact, broadening its applicability, and automating the human oversight loop without sacrificing clinical rigor.
Current State: The policy is a versioned, readable natural-language document (e.g., Appendix B). While excellent for auditability, this structure can be slow for high-throughput execution compared to traditional symbolic reasoning engines.
Proposed Improvement: Implement a Structured Policy Interpreter (SPI) layer alongside the text policy. This involves translating the eight phases and their conditional triggers into a formal, executable decision tree or a lightweight state machine (e.g., using Prolog or specialized graph traversal).
What the Improved System Can Do:
-
Instant Execution: The LLM executes the policy not by
reading
it, but by following a structured logical path defined by the SPI. This reduces inference latency dramatically compared to pure text-based execution, enabling near real-time diagnostic scoring. -
Formal Verification: We can programmatically verify that the policy structure adheres to clinical safety constraints (e.g, ensuring Phase 6—Adjudication—is always performed before Phase 7—Final Output), guaranteeing logical consistency beyond what human review alone could provide.
Current State: liteOdyssey is primarily phenotype-first, using clinical text and HPO terms, with genetic data used only when available.
Proposed Improvement: Develop a Multi-Modal Feature Extraction Module (MMFEM) that allows the policy to incorporate structured inputs from other modalities (e.g., radiology reports or genomic sequencing data) into the initial Phase 0 Pattern Recognition.
What the Improved System Can Do:
-
Holistic Reasoning: The system can evaluate diagnoses based on a combination of clinical symptoms and specific genetic variants (e.g,
The patient has feature X, which is strongly associated with Gene Y via sequencing
). This eliminates cases where phenotype and genetic data are mutually exclusive. -
Refined Adjudication: Phase 6 (Adjudication) can be augmented to weigh evidence from the MMEMF, ensuring that the final ranking is based on all available data, not just the most easily described clinical features.
Current State: PIHF relies on a human expert/clinician to review and admit revisions (the expert-gated update
). This is the primary bottleneck for scaling this method is manual throughput.
Proposed Improvement: Deploy a Clinical Validation Meta-LLM (CV-MetaLLM), trained specifically on clinical guidelines and consensus literature, to act as an automated first pass
reviewer for PIHF's proposed policy deltas.
What the Improved System Can Do:
-
Accelerated Iteration: The CV-MetaLLM automatically flags revisions that violate established medical standards (e.g, recommending a diagnosis without supporting evidence) or that are syntactically unsound, filtering out 80% of low-quality suggestions before they reach the human expert.
-
Targeted Human Focus: By automating the initial review and providing a
risk score
for each proposed policy change, we allow human experts to focus only on high-uncertainty or complex clinical cases, dramatically increasing the scale of PIHF development.
Current State: The policy is a large, versioned text artifact (46 MB footprint). While efficient for its class, it does not leverage modern distillation techniques.
Proposed Improvement: Implement a Policy-to-Layer Distillation (P2LD) process. This involves mapping the logical flow of the PIHF policy onto specialized, small weights (e.g., using LoRA or Adapter layers) that are trained to mimic the behavior of the frozen LLM running the policy.
What the Improved System Can Do:
-
Inference Speed Boost: The system can execute complex reasoning in a highly optimized, lightweight architecture that is orders of magnitude faster than running a full-sized general-purpose LLM, making it viable for deployment on edge devices or resource-constrained clinical terminals.
-
Guaranteed Consistency: By distilling the behavior of the policy (not just the knowledge), we ensure that even if the underlying base model changes, the distilled layer maintains a consistent, auditable adherence to the established liteOdyssey logic.
Abstract
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer. Large language models rank the correct disease first in only 35.4% of benchmark cases and often rely on learned phenotype-disease associations rather than reusable diagnostic reasoning strategies. We developed liteOdyssey through Policy Iteration with Human Feedback, a process in which model failures and expert corrections are iteratively consolidated into a clinician-gated, natural-language policy executed by a language model. Across 1,243 public benchmark cases spanning 722 rare diseases, liteOdyssey ranked the correct disease first in 59.3% of cases versus 26.5% without the policy, with comparable gains in cases involving diseases excluded from policy development. The same policy transferred across model families and sizes without retraining. Adaptation of the policy to the Undiagnosed Diseases Network (UDN) improved diagnostic accuracy among 515 UDN patients, with gains confirmed by blinded physician adjudication. These results show that expert reasoning can be externalized into an inspectable and revisable natural-language policy that generalizes across rare diseases, transfers across model backbones, and adapts to a real-world patient cohort.
Sources
- A Specialized Large Language Model for Clinical Reasoning and Diagnosis in Rare Diseases
- End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning
- Language Models are Few-Shot Learners
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows
- How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- Deep reinforcement learning from human preferences
- Training language models to follow instructions with human feedback
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Inference-Time Scaling for Generalist Reward Modeling
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Reward Is Enough: LLMs Are In-Context Reinforcement Learners
- Learning by Distilling Context
- On-Policy Context Distillation for Language Models
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Fine-Tuning Language Models from Human Preferences
- Kimi K3: Open Frontier Intelligence
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection