From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation

arXiv:2608.16831 · cs.AI, cs.CL · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "From Answers to Policies".

Tom: This paper introduces Policy Iteration with Human Feedback (PIHF), a framework designed to bring post-training reinforcement learning principles to in-context learning systems, specifically for complex tasks like rare disease diagnosis.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, to wrap up what we've seen so far, "From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation" introduces a framework where human experts actively guide an AI’s policy and tool set through a recurrent loop of evaluation and revision. Jane, can you distill that core idea into something really easy for our listeners to grasp in one sentence?

Jane: Absolutely, Tom. Think of it like this: instead of just giving the AI more instructions and hoping it gets better, this system lets human experts step in during the learning process. They look at the AI's work—how it reasons and what tools it uses—and they provide feedback that directly steers how the AI updates its internal plan. It’s about a continuous, guided refinement process where human judgment acts as the steering wheel for the AI’s strategy.

Lu: I find that concept fascinating because it fundamentally separates how we think about learning from scratch versus fine-tuning. By treating the policy as an external artifact that gets versioned and explicitly edited by humans, you get this very structured path to better behavior without having to retrain the whole foundation model every time. It’s like we’re not just teaching a student; we're mentoring them through a series of specific, high-quality assignments guided by an experienced teacher.

Meng: From my side, I'm thinking about the practical impact of that structure on deployment. If we can manage to refine the policy artifacts this way, it means we could potentially deploy specialized diagnostic agents that are highly accurate for rare diseases but don't require massive retraining cycles every time a new clinical scenario comes up. That level of adaptation without constant re-engineering sounds incredibly efficient for our operational needs.

Lalam: I see this as a really powerful step toward creating AI systems that are not just reactive tools but genuinely self-improving agents in their own domain. This capability to integrate structured expert review into the core learning mechanism means we can build agents that develop specialized diagnostic strategies over time, making them much more reliable and accountable for high-stakes decisions.

Tom: That’s a huge shift, Lalam! It’s moving AI from being something that just follows instructions to something that actively learns and adapts its own internal logic based on deep human insight. Jane, you mentioned the concept of process feedback earlier; how does this paper handle separating what was wrong with the reasoning steps versus what was wrong with the tools themselves?

Jane: Well, it's very deliberate about that separation. The framework allows for a critic to review specific parts of the execution trajectory to pinpoint exactly where a policy rule or a tool call failed, which is process feedback. Then, there’s another check that focuses on whether the final outcome—the actual diagnosis—is sound, which is outcome validation. This dual focus gives the system much more precision in its revisions than just looking at a final score.

Lu: That separation is what makes it so robust for complex tasks like rare disease diagnosis. If we only had outcome validation, we wouldn't know *why* the model failed; if we only had process feedback, we might correct the reasoning but miss a subtle issue in how it used a specific tool. This method lets us target those failures specifically, which is incredibly valuable when dealing with nuanced clinical data.

Meng: From an engineering standpoint, that targeted feedback loop simplifies debugging immensely. When we see a failure, instead of wading through millions of parameters to figure out what's wrong internally, the system can tell us if the issue was a flawed policy rule or if it chose the wrong tool for a given step. That localization speeds up our iteration cycle significantly.

Lalam: For our culture at the startup, this level of accountability is huge. It means we’re not just accepting an AI's output; we are actively participating in refining its internal decision-making logic based on expert critique at every stage. This fosters a much more rigorous and trustworthy development environment where the AI evolves under careful human supervision.

Tom: It sounds like this framework isn't just about making the model smarter in a general sense; it's about building a mechanism for expert collaboration that directly improves the underlying strategy of the system. Jane, what does this mean for how we approach building these systems moving forward?

Jane: It means we should start thinking less about static training and more about designing these iterative refinement loops from the very beginning. We need to build in mechanisms where human feedback is not just a final check, but an active part of the learning process itself, shaping the policy as it goes.

Lu: The implication for future research is that we can treat language models as powerful execution engines and focus our architectural energy on designing these sophisticated meta-learning and feedback loops, rather than just trying to brute-force performance on larger models. It suggests a new way to structure intelligence in AI systems.

Meng: I think the real world will see this in specialized diagnostic tools where accuracy is paramount. If we can deploy an agent that continuously improves its diagnostic policy based on expert review, it could dramatically increase the reliability of those tools for rare conditions where every bit of accuracy matters.

Lalam: Ultimately, this work points toward a future where our AI agents are not just passive responders but active collaborators in their own refinement process, developing specialized expertise under expert guidance. This is really about building AI that learns responsibly and adaptively within its domain.

Tom: So we’re looking at a framework called "From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation," which provides us with a blueprint for building AI that learns strategically from expert investigation. Jane, this has been fantastic—thanks for walking us through the concepts of policy artifacts and feedback loops!

Jane: It’s been my pleasure, Tom; it’s a really warm way to think about how AI can grow under careful human guidance.

Lu: I'm just eager to see the next set of experiments that explore how this structure handles massive amounts of diverse data while maintaining that expert-driven precision.

Meng: I’m looking forward to seeing how this concept gets implemented in production, because that’s where we need to start thinking about scaling it effectively.

Lalam: I feel really optimistic about the future of AI development with this framework, seeing systems become truly adaptable and specialized through guided iteration.

The paper's summary: Tom: So, we’ve gone over the core idea of Policy Iteration with Human Feedback, which is essentially using expert reviews to guide an AI's internal policy revisions through a structured loop. Now, what about the specific improvements this paper suggests for making these systems actually perform better? Jane, can you explain how they are sharpening the system's capabilities beyond just having that initial feedback loop?

Jane: They focus on creating a more precise way to manage two types of feedback: one that corrects the immediate reasoning steps and tool usage—that’s process-based feedback—and another that validates whether the final diagnosis or result is actually clinically correct, which is outcome validation. This separation allows for much finer tuning of what exactly needs changing in the policy artifact.

Lu: I think this dual approach is incredibly smart because it tackles the two main weaknesses in current learning systems: getting stuck on the wrong path and producing an incorrect final answer even if the path seemed plausible. By isolating those failures, you get a much clearer signal about where to apply your next revision effort.

Meng: That precision in feedback translates directly into reliability for us. If we can tell whether a failure was due to poor tool selection or flawed internal reasoning, our engineering team knows exactly where to focus their effort on refining the policy artifacts instead of guessing what the AI needs. It makes the development cycle much more focused and less wasteful.

Lalam: For our culture, this level of structured feedback is transformative because it shifts our approach from simply accepting an output to actively participating in its refinement through expert critique. We’re building AI that learns how to self-correct its strategy by receiving targeted, validated input on both its methods and its results. That builds a much stronger foundation for trustworthy systems.

Tom: It sounds like the paper is proposing a way to move past vague performance gains toward targeted, actionable improvements in the AI's behavior and tool usage. Jane, how do these specific suggestions translate into tangible benefits for complex tasks like rare disease diagnosis?

Jane: The tangible benefit is higher diagnostic accuracy, potentially with significant gains when compared to static models because the system isn't just guessing; it’s continuously evolving its strategy based on validated clinical outcomes. It means the AI learns *how* to reason effectively in a way that aligns with expert knowledge of rare diseases.

Lu: The real potential here is portability, Tom. If we can refine this policy-based approach, the resulting expertise should be transferable across different language model backbones and architectures without needing to retrain from scratch on every new model, which is a huge theoretical win for modular AI design.

Meng: Portability would be fantastic for our startup because it means we could build one core diagnostic policy and deploy it across various foundational models, saving us the massive computational cost associated with training entirely new models just to adapt them to a specific disease profile. That kind of efficiency is what we need for real-world scaling.

Lalam: I see this as paving the way for AI agents that can develop specialized, deeply ingrained clinical expertise over time through iterative human mentorship. This isn't just about better answers; it’s about building an agent with a refined, adaptable diagnostic strategy that grows smarter with every expert review it receives.

Tom: So we're looking at a framework that doesn't just give us a better starting point, but provides a mechanism for continuous, expert-guided strategic evolution of the AI's very method of operation. Jane, what’s the next big question we need to ask about these results?

Jane: We’re seeing a strong push toward building systems that learn not just from data, but from structured human collaboration during the learning phase itself, which is a really warm way to think about AI evolution.

The paper's improvements: Tom: So, to wrap up our discussion on "From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation," we've covered how human expertise can be formally integrated into an AI's learning process through a structured loop of evaluation and revision. Jane, how would you summarize the paper’s main contribution one last time for our listeners?

Jane: This work shows a formal way to integrate human expertise directly into an AI's learning process by letting experts guide persistent revisions of its policy and tool set through a structured loop. It’s about making the AI adapt its internal strategy using targeted human guidance rather than just relying on static training data.

Lu: I think the theoretical leap here is treating the policy itself as a versioned artifact that undergoes explicit external manipulation based on human review, which gives us a very concrete way to apply reinforcement learning principles to in-context learning systems. It opens up possibilities for building highly modular and specialized AI agents.

Meng: From an engineering standpoint, the practical implication is creating diagnostic agents that can continuously improve their reasoning strategy based on validated clinical feedback, which could be a real game-changer for high-stakes applications where accuracy matters most.

Lalam: For our culture, this points toward a future where AI agents are not just passive tools but active participants in their own refinement, developing specialized expertise under expert mentorship through rigorous review cycles. That fosters a much more accountable and reliable development environment.

Tom: It really is exciting to see how they’ve managed to bridge the gap between complex RL theory and the practical need for expert-driven adaptation in real-world AI systems. Jane, what are we looking at next on this topic?

Jane: We’re seeing a strong push toward building systems that learn not just from data, but from structured human collaboration during the learning phase itself, which is a really warm way to think about AI evolution.

Lu: I think the next big thing will be seeing how researchers take this framework and apply it across vastly different modalities, maybe even integrating it with visual reasoning models or physical simulation systems.

Meng: I’m watching how this translates into deployment; if we can keep the management of these versioned policies efficient, we could see specialized AI solutions deployed much faster than traditional retraining schedules allow.

Lalam: I believe the most impactful vision is an AI that develops specialized diagnostic skills through iterative human mentorship, ensuring its reasoning remains sound and clinically relevant as it evolves.

Tom: So to wrap up, "From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation" gives us a solid blueprint for building AI that learns strategically from expert investigation. Jane, this has been fantastic—thanks for walking us through the concepts of policy artifacts and feedback loops!

Jane: It’s been my pleasure, Tom; it’s a really warm way to think about how AI can grow under careful human guidance.

Conclusion: Tom: So, to wrap up our discussion on "From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation," we've covered how human expertise can be formally integrated into an AI's learning process through a structured loop of evaluation and revision. Jane, how would you summarize the paper’s main contribution one last time for our listeners?

Jane: This work shows a formal way to integrate human expertise directly into an AI's learning process by letting experts guide persistent revisions of its policy and tool set through a structured loop. It’s about making the AI adapt its internal strategy using targeted human guidance rather than just relying on static training data.

Lu: I think the theoretical leap here is treating the policy itself as a versioned artifact that undergoes explicit external manipulation based on human review, which gives us a very concrete way to apply reinforcement learning principles to in-context learning systems. It opens up possibilities for building highly modular and specialized AI agents.

Meng: From an engineering standpoint, the practical implication is creating diagnostic agents that can continuously improve their reasoning strategy based on validated clinical feedback, which could be a real game-changer for high-stakes applications where accuracy matters most.

Lalam: For our culture, this points toward a future where AI agents are not just passive tools but active participants in their own refinement, developing specialized expertise under expert mentorship through rigorous review cycles. That fosters a much more accountable and reliable development environment.

Tom: It really is exciting to see how they’ve managed to bridge the gap between complex RL theory and the practical need for expert-driven adaptation in real-world AI systems. Jane, what are we looking at next on this topic?

Jane: We’re seeing a strong push toward building systems that learn not just from data, but from structured human collaboration during the learning phase itself, which is a really warm way to think about AI evolution.

Lu: I think the next big thing will be seeing how researchers take this framework and apply it across vastly different modalities, maybe even integrating it with visual reasoning models or physical simulation systems.

Meng: I’m watching how this translates into deployment; if we can keep the management of those versioned policies efficient, we could see specialized AI solutions deployed much faster than traditional retraining schedules allow.

Lalam: I believe the most impactful vision is an AI that develops specialized diagnostic skills through iterative human mentorship, ensuring its reasoning remains sound and clinically relevant as it evolves.

Tom: So to wrap up, "From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation" gives us a solid blueprint for building AI that learns strategically from expert investigation. Jane, this has been fantastic—thanks for walking us through the concepts of policy artifacts and feedback loops!

Jane: It’s been my pleasure, Tom; it’s a really warm way to think about how AI can grow under careful human guidance.

Lu: I'm just eager to see the next set of experiments that explore how this structure handles massive amounts of diverse data while maintaining that expert-driven precision.

Meng: I’m looking forward to seeing how this concept gets implemented in production, because that’s where we need to start thinking about scaling it effectively.

Lalam: I feel really optimistic about the future of AI development with this framework, seeing systems become truly adaptable and specialized through guided iteration.

Minh-Ha Nguyen, Cathy Shyr

Department of Epidemiology, Vanderbilt University · Department of Pediatrics, Vanderbilt University Medical Center · Department of Biostatistics, Vanderbilt University Medical Center · Department of Biomedical Informatics, Vanderbilt University Medical Center

cs.AI, cs.CL

Submitted: 2026-08-17

Updated: 2026-09-29

Comments: PIHF-MCP. Formalizing and automating in-context policy development with PIHF

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: This paper introduces Policy Iteration with Human Feedback (PIHF), a framework designed to bring post-training reinforcement learning principles to in-context learning systems, specifically for

Key concepts

Policy Iteration with Human Feedback (PIHF)
A framework designed to bring post-training reinforcement learning principles to in-context learning systems. It uses human experts to guide the AI's policy and tool set through a recurrent loop of evaluation and revision, steering the AI's strategy.
Process Feedback
Feedback that corrects specific parts of an execution trajectory, focusing on where a policy rule or a tool call failed. This is separated from outcome validation to pinpoint exactly what needs changing in the AI's reasoning steps.
Outcome Validation
A check focused on whether the final result, such as a diagnosis, is actually clinically correct. Separating this from process feedback allows for finer tuning of the policy artifact based on clinical accuracy.

Terminology

Summary

This paper introduces Policy Iteration with Human Feedback (PIHF), a framework designed to bring post-training reinforcement learning principles to in-context learning systems, specifically for complex tasks like rare disease diagnosis. It addresses the need for models to adapt their behavior from instructions and demonstrations by employing a recurrent evaluate-and-improve structure where human feedback guides persistent revision of a versioned natural-language policy and tool set. This approach leverages pretrained language models as fixed execution substrates, allowing expert guidance to develop an effective diagnostic policy that can be reused across different model backbones while maintaining fixed weights.

RL Principles and Objective

The foundation of PIHF is rooted in reinforcement learning, which represents behavior by a policy and improves that policy by increasing expected return. For language-model feedback learning, the objective involves maximizing a regularized log-ratio (Equation 1):

J(θ) = E x∼D y∼πθ(·x) [R(x, y)] − β Ex∼D[KL(πθ(· x) ∥ πanc(· x))], where R is the expected evaluator score and the second term penalizes departure from the anchor policy πanc. This objective has a closed-form exponential-tilt optimizer. The core improvement mechanism is explicitly shown in Equation 2, which states that the resulting conditional distribution is:

π∗(y x) = πanc(y x) exp[R(x, y)/β] / Zβ(x). This equation shows that the anchor supplies the base response distribution, while exp[R(x, y)/β] reweights each response according to its evaluated quality.

Structural Bridge to PIHF

PIHF establishes a structural bridge between weight-space reinforcement learning and external artifact manipulation. In weight-space RL, optimization changes model parameters θ, and the correspondence is θ ←→ At = (Pt, Tt). PIHF implements this by indexing iterations with t and defining the external artifact At = (Pt, Tt), where Pt is the versioned natural-language policy and Tt is the set of available tools. The behavior-level correspondence is then πθ(y x) ←→ πM,At(τ z), where M is a frozen executor model and z is a clinical case. This separation allows for distinct functions: the critic and expert review complete-panel trajectories to localize recurrent failures, while Recall@1 and Recall@5 measure terminal diagnostic outcomes after a frozen candidate executes.

In-Context Policy Definition

The in-context policy representation is defined as the mutable artifact At = (Pt, Tt). For a clinical case z, the policy instantiates an executor prompt xt(z) = Prompt(Pt, z). A complete rollout trajectory τ is defined as (e1,..., eL, yb), where el is a recorded model output or tool result. The policy organizes execution into ordered stages Φt = (ϕ(1)t,..., ϕ(mt)t). A critique localizes the earliest stage where τ departs from Pt and attributes the departure to a policy rule, tool use, or returned evidence. Conformance asks whether τ followed Pt, and Adequacy asks whether Pt remains clinically sound and useful.

PIHF Operator and Composed Runs

The PIHF operator is defined as (P⋆, T⋆) = PIHF (M, Pinit, Tinit, Ddev), where M is the frozen executor model. The inputs include the initial policy Pinit and tools Tinit. The output (P⋆, T⋆) is the final admitted artifact. Two composed invocations are used: first for public-policy development from an initial artifact on a LIRICAL development panel (PL, TL), and second for UDN development warm-started from a frozen LIRICAL artifact (PU, TU). Development exclusion is policy-specific, meaning claims about (PL, TL) exclude Ddev L from its held-out evaluation.

Critic and Expert Proposal Formation

After executor outputs are sealed into Et, a second LLM critic reviews Et for recurrent reasoning and tool-use failures to propose an interpretation uGt. The expert then reviews this proposal against the incumbent artifact At, forming the expert-authorized proposal ut = Hform(uGt, Et, At). Admission and rollback remain expert decisions after complete-panel evaluation. Candidate formation involves applying an edit δt to produce a candidate A′t = Freeze(At ⊕ δt). The recall-preservation indicator Irecall t checks if the candidate preserves Recall@1 and Recall@5 relative to the incumbent, while Iexpert t is a qualitative expert indicator summarizing sound clinical suggestions and generalizable pattern addresses. The next incumbent is set only when both indicators equal one: At+1 = (A′t, Irecall t Iexpert t = 1, At, otherwise).

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the Policy Iteration with Human Feedback (PIHF) framework, and what those improved systems can achieve:


  1. Improve reasoning accuracy in complex, multi-step diagnostic tasks for rare diseases by replacing fixed models with versioned, expert-governed policies.

  2. Enable persistent learning and adaptation of diagnostic strategies using a recurrent evaluate-and-improve loop that leverages human expertise to localize and correct failures in policy or tool usage.

  3. Develop AI systems capable of maintaining high recall in diagnosis (Recall@1 and Recall@5) across diverse, unseen clinical cases by continuously refining their internal execution policies based on complete panel reasoning trajectories.

  4. Create portable diagnostic agents that can effectively transfer learned policies between different underlying language model backbones (e.g., GPT-5.4 to Qwen3.6-35B) while retaining expert-guided behavioral improvements, ensuring that performance gains are independent of the specific model architecture used for execution.

  5. Allow AI systems to explicitly separate and refine two distinct types of feedback: process-based feedback (correcting reasoning steps and tool use) and outcome-based validation (checking the final diagnostic accuracy), leading to more precise policy revisions.

  6. Implement a system where human experts maintain ultimate authority over the AI's evolution by reviewing, accepting, rejecting, or revising proposed policy changes derived from a language-model critic, thereby ensuring safety and clinical soundness.

This improved AI system can:

  1. Perform rare disease diagnosis with significantly higher accuracy (up to 32.7 percentage points gain reported) compared to static models by dynamically updating its execution policy based on real-world diagnostic outcomes validated against a development panel.

  2. Continuously refine its decision-making logic through iterative cycles where it identifies exactly where a failure occurred—whether in the language policy or tool use—and receives targeted, expert-approved revisions.

  3. Maintain robust diagnostic performance (Recall@1 and Recall@5) even when deployed on different foundational models, demonstrating strong portability of the learned expertise across various LLM architectures.

  4. Act as a self-improving reasoning agent that learns how to sequence complex tool calls and internal model generations (process feedback) while ensuring the final clinical conclusion remains clinically sound (outcome validation).

Abstract

Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures. We formalize Policy Iteration with Human Feedback (PIHF), which makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP. Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing tools, review and persistent inquiry records to develop reusable task policies. Across general reasoning benchmarks (BIG-Bench Extra Hard, HoVer and LiveBench-Math), PIHF-MCP improved performance of the baseline model by 16.9, 22.2 and 4.7 percentage points, respectively. With a matched baseline model, development used about 1/5 of the labelled examples and 4% of the task rollouts reported by a previous SOTA in-context optimizer, making it about 9 times faster and 3 times cheaper at comparable or higher scores. In a low-data rare-disease diagnosis setting, policies developed from previous SOTA prompt optimizers trailed a previously published PIHF-developed system on every held-out cohort (on average 16 percentage points). These findings support a route to more efficient inference-time scaling: PIHF-MCP develops reusable policies from a few examples that improve performance on unseen cases and across models. Because each policy comes from an explicit, recorded investigation, the process also keeps humans in the loop and enables ownership and learning, making it well suited to high-stakes decisions.

Sources

Related papers