TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series

arXiv:2606.09030 · cs.LG, cs.AI, cs.CL · Submitted 2026-06-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series".

Jane: Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS),

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let's talk about the title and who came up with this research. The paper is called "TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series." It’s clear they are focusing on making the reasoning part of the prediction process, which is a big step forward.

Jane: The authors include Hyeongwon Jang, Gyouk Chu, Changhun Kim, Joonhyung Park, Hangyul Yoon, and Eunho Yang from KAIST and UW-Madison. It shows this is a collaborative effort coming from some strong AI research groups in the field.

Lu: I think their focus on dialectical reasoning is the real innovation here; it’s not just about asking for an explanation after a prediction, but structuring the model to generate arguments for competing outcomes simultaneously.

Meng: It makes sense that they are focusing on dialectical reasoning because it directly attacks that risk polarization problem we discussed earlier, where models get stuck committing too early to one outcome.

Lalam: And I think the implications are huge for trust; when a clinician can actually verify the model's logic by looking at these explicit rationales, it moves the AI from being a black box suggestion to a reasoning partner.

Tom: Exactly, Lalam. This paper is showing us how to move beyond just getting an answer and start getting a verifiable explanation for that answer. It’s about making the AI's thinking transparent in a clinical context.

Jane: And they specifically target ISMTS, which is crucial because most clinical data isn't perfectly regular, and this framework is designed to handle those messy time series effectively.

Lu: They are proposing a two-stage training pipeline involving dialectical reasoning supervision followed by self-refinement through reinforcement learning using Group Relative Policy Optimization, which is a clever way to fine-tune that complex reasoning ability.

Meng: The engineering approach with the GRPO loss function, where they supervise the reasoning tokens but only apply cross-entropy loss to the final decision token, seems like a smart way to balance training both the intermediate steps and the final output.

Lalam: That sounds like a solid way to ensure that even though we are focusing on reasoning generation, we don't lose sight of actually producing a useful classification result.

The paper's summary: Tom: So, what’s the core summary of TRIAGE? Essentially, the paper introduces TRIAGE as a framework to train an LLM to generate dialectical reasoning over competing clinical outcomes by asking for outcome-specific rationales.

Jane: That means instead of one long explanation that points toward one conclusion, they are training the model to generate separate arguments for every possible clinical result. This is designed specifically to fight that risk polarization issue where the scores get skewed towards extremes.

Lu: They set up a problem setup where they treat clinical risk prediction from ISMTS as a supervised classification task, aiming for a patient-specific class distribution of probabilities across all outcomes.

Meng: The setup involves defining the input data as an ISMTS, which is a multivariate time series record containing time, value, and variable index information for different clinical measurements.

Lalam: So they are taking that complex time series data and trying to learn a predictor that outputs a probability distribution over all possible outcomes for each patient, rather than just one final guess.

Tom: And the key mechanism they use is eliciting these outcome-specific rationales from an LLM separately for each candidate outcome, making sure those rationales don't reference other outcomes.

Jane: That’s a really important constraint because if the rationale references another outcome, it becomes contrastive instead of strictly outcome-specific, which the authors found was a problem in previous methods.

Lu: They then use these collected rationales, along with ground-truth answer tokens, to form reasoning traces that guide the final risk estimation within a single LLM structure.

Meng: It shifts the paradigm from having an LLM predict first and *then* trying to explain it, to having the LLM generate the explanation *as* part of its internal decision-making process for a patient-comparable score.

Lalam: And that’s what really excites me about this summary; it fundamentally changes how we think about using LLMs in high-stakes domains like medicine—it forces structured, balanced deliberation instead of simple pattern matching.

The paper's improvements: Tom: Now, let’s talk about the actual improvements they propose in TRIAGE. The paper details a two-stage training pipeline: first, Dialectical Reasoning Supervision, and second is Self-Refinement through reinforcement learning with GRPO.

Jane: In the first stage, they are teaching the model to produce those outcome-conditioned rationales by making sure the LLM only generates evidence supporting one specific outcome at a time.

Lu: They enforce two strict constraints there: first, no reference to any alternative outcome in the rationale, and second, it must not fabricate unobserved evidence; if it can't find evidence for an outcome, it leaves the rationale blank.

Meng: That constraint about leaving the rationale blank when evidence is missing is something I’m thinking about practically; we need a robust way to handle uncertainty without just making up justifications for things that aren't there.

Lalam: It addresses the issue of hallucination in reasoning, which is a huge problem with current LLMs, because if it can’t find evidence, it stops trying to commit to a false narrative.

Tom: The second stage refines this using reinforcement learning with Group Relative Policy Optimization, and they introduce a batch-level reward function that promotes cross-patient comparability alongside the calibration benefit of on-policy training.

Jane: That reward function is interesting because it’s trying to get the model to be accurate on one patient while also being comparable across many patients simultaneously, which tackles that cross-patient comparability issue directly.

Lu: Their evaluation shows that this approach leads to significant improvements, with an average AUPRC increase of three point three percent and a calibration error reduction of eighty-one percent compared to competitive baselines on benchmarks like P12 and MIMIC-III.

Meng: An eighty-one percent reduction in calibration error is a big deal because it means the continuous risk scores they produce are much more trustworthy for actual triage decisions, which is exactly what we need when dealing with uncertain medical trajectories.

Conclusion: Tom: So, wrapping up the paper on TRIAGE: it seems the main implication is moving from brittle binary predictions to a system that yields continuous risk scores built on explicit, verifiable clinical reasoning.

Jane: That’s right, Tom; they’ve shown how dialectical reasoning can mitigate the risk polarization problem and allow a single LLM to provide patient-comparable risk grounded in clear evidence.

Lu: The future work mentioned involves evaluating this framework more broadly and seeing if these learned dialectical principles are robust enough to transfer effectively across different medical tasks or even different model architectures.

Meng: From an engineering perspective, the limitation they flag is that their evaluation is restricted to binary prediction tasks, so we need more work to see how well this holds up when we try to apply it directly for continuous risk estimation in a real-world continuous setting.

Lalam: I think the biggest takeaway for us all is that this framework provides a path toward building AI systems where the reasoning isn't hidden; it’s an explicit, balanced deliberation process that clinicians can trust and check.

Tom: It really feels like we’re getting closer to having AI tools that don't just give us a number but also show us the clinical thinking behind that number. Thanks for joining us on this deep dive into TRIAGE today!

KAIST · AITRICS

cs.LG, cs.AI, cs.CL

Submitted: 2026-06-08

Updated: 2026-10-01

Code: https://github.com/HyeongWon-Jang/TRIAGE

Importance score: 90/100

The gist: Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS), must deliver both calibrated risk

Key concepts

Risk Polarization Problem
This occurs when asking an LLM for reasoning before scoring causes the final risk score distribution to become extremely polarized. This happens because the generated text usually commits to one outcome, making the prediction nearly certain, and it rarely presents evidence supporting all possible outcomes simultaneously.
Dialectical Reasoning Supervision
This is a training stage where the model is taught to generate rationales for every potential clinical outcome separately. The key constraint is that each rationale must be strictly outcome-specific, meaning it cannot reference or contrast with any other alternative outcome, ensuring focused evidence generation.
Continuous Risk Estimation
Instead of forcing a single binary prediction, TRIAGE uses the model's implicit probability distribution at a specific answer position to derive a continuous risk score. This method preserves a graded signal that reflects the uncertainty and likelihood across all possible outcomes, rather than collapsing into an absolute yes or no.
LLM-as-a-Judge Assessment
This is a qualitative evaluation method where another LLM is used to assess the quality of clinical reasoning traces produced by TRIAGE. This helps measure how well the model's generated explanations align with actual, clinically sound justifications for patient outcomes.

Terminology

Summary

Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS), must deliver both calibrated risk scores for patient triage and interpretable rationales that clinicians can verify. This framework proposes a method to address the risk polarization problem in Large Language Models (LLMs) by training them to generate dialectical reasoning over competing clinical outcomes, enabling a single model to yield continuous risk scores grounded in explicit clinical reasoning.

The gist

TRIAGE is a framework that trains an LLM to generate dialectical reasoning over competing clinical outcomes by eliciting outcome-specific rationales, which mitigates risk polarization and enables a single LLM to yield continuous risk scores grounded in explicit clinical reasoning.

Problem Identification: Risk Polarization in LLMs

The paper identifies a risk polarization problem where eliciting natural language reasoning before risk scoring collapses the score distribution into an extreme. This occurs because (1) rationales typically pre-commit to a single outcome, making the final prediction nearly deterministic given the preceding biased text, and (2) rationales are typically one-sided, presenting only evidence that supports a committed outcome rather than evidence bearing on all possible outcomes. This issue is exacerbated in ISMTS where trajectories contain coexisting signals of deterioration and stabilization.

Proposed Solution: TRIAGE Framework

TRIAGE addresses this by grounding risk estimation in dialectical reasoning over alternative outcomes. Instead of steering reasoning toward a single outcome, TRIAGE generates a dedicated rationale for each candidate outcome and derives the final risk from the LLM’s implicit probability conditioned on these rationales. This is achieved through a two-stage training pipeline: (i) Dialectical Reasoning Supervision followed by (ii) Self-Refinement.

Dialectical Reasoning Supervision

Stage 1 focuses on teaching the model to produce outcome-conditioned rationales. This involves eliciting rationales from a strong LLM separately for each candidate outcome, adhering to two constraints: first, the LLM should not reference any alternative outcome, ensuring each rationale remains strictly outcome-specific rather than contrastive, and second, it should not fabricate unobserved evidence; in such cases, the rationale is left blank. These outcome-conditioned rationales are then concatenated with the ground-truth answer token to form reasoning traces.

Self-Refinement via Reinforcement Learning

Stage 2 refines the model through reinforcement learning using Group Relative Policy Optimization (GRPO). The objective function is structured as: L(θ) = LGRPO(θ) + λ · LCE(θ), where the GRPO objective supervises the reasoning tokens, and the Cross-Entropy (CE) loss is applied only to the final decision token. This process improves both discrimination and calibration, with a novel batch-level reward function introduced to promote cross-patient comparability alongside the calibration benefit of on-policy training.

Evaluation and Results

Evaluations were conducted on three ISMTS benchmarks: P12, P19, and MIMIC-III. TRIAGE achieves an average AUPRC improvement of 3.3% and reduces calibration error by 81% compared to competitive baselines. Furthermore, an LLM-as-a-judge assessment shows that the rationales surpass post-hoc explanations from the baseline by 20% in clinical reasoning quality, achieving a higher total IDEA score (7.744 vs. 6.474). The framework demonstrates robustness under limited information and low-resource training regimes, showing improvements over baselines even when data is subsampled to 1%.

Clinical Reasoning Quality

Qualitative results using an LLM-as-a-judge assessment show that TRIAGE produces more patient-specific and clinically grounded reasoning traces. For instance, in a case study of a survived patient, the model's rationale explicitly incorporates evidence for survival such as Glasgow Coma Score improves over time from very low values (3–6) to 9 by around 40–44 hours, while its rationale for death explicitly includes factors like advanced age (85 years) and persistently low body temperatures. This contrasts with post-hoc methods that might attribute a Glasgow Coma Score of 15 to mortality.

Inference Strategy

The framework is designed for continuous risk estimation. The final risk score is derived from the model’s implicit outcome-token distribution at a fixed answer position, calculated as: P(yk x) = exp(lk)/P k′ exp(lk′). This design preserves a graded risk signal rather than collapsing into near-deterministic outputs, and the bidirectional averaging of scores from both reasoning orderings is adopted as the default inference strategy.

Limitations

Limitations include evaluation restricted to binary prediction tasks, the increased computational cost of LLM-based pipelines compared to lightweight baselines, and reliance on LLM-as-a-judge for clinical reasoning quality assessment instead of expert assessment.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the TRIAGE framework, and what these improved systems will be capable of:


) Improved AI System Capabilities

The TRIAGE framework fundamentally shifts LLM-based clinical risk prediction from a brittle prediction/explanation trade-off to a robust system grounded in explicit, dialectical clinical reasoning. The resulting improved AI systems can perform the following specific tasks:

  1. Individualized, Cross-Patient Risk Stratification with Calibrated Confidence:

The improved system will generate a continuous, patient-specific risk score (e.g., probability of mortality) that is highly calibrated (reduced calibration error by 81% compared to baselines). Unlike current LLMs that collapse risk into overconfident binary predictions, this system can reliably distinguish between patients with slightly different risk profiles and accurately rank them across a cohort, enabling precise resource allocation and triage decisions.

  1. Clinically Verifiable Why for Every Prediction:

The system will produce explicit, outcome-specific rationales for every candidate outcome (e.g., Rationale for survival, Rationale for in-hospital death). These rationales are structured dialectically, meaning they explicitly articulate evidence supporting the predicted outcome while simultaneously detailing the counter-evidence or alternative signals that were weighed but ultimately dismissed. This allows clinicians to verify the model's logic by checking if its reasoning aligns with established clinical knowledge (as validated by an LLM-as-a-judge).

  1. Robustness to Data Sparsity and Missing Variables:

The system will maintain high predictive performance even when input data is incomplete or noisy. Through the use of set-based encoding for ISMTS and the dialectical reasoning structure, TRIAGE can effectively weigh available signals against missing information, leading to superior performance (e.g., maintaining strong AUPRC on MIMIC-III with 50% variable masking) compared to models that rely solely on post-hoc feature attribution.

  1. Discovery of Clinically Plausible Reasoning Patterns:

The system will move beyond simple pattern matching by generating reasoning traces that exhibit balanced deliberation over competing signals (P2). This allows researchers to identify clinically reasonable counterevidence—such as temporal recovery trends or the persistence of specific vital sign patterns—that standard models might discard due to confirmation bias. This capability is crucial for identifying subtle prognostic indicators missed by traditional methods.

  1. Enhanced Data Efficiency in Low-Resource Settings:

By utilizing a two-stage training pipeline (Dialectical Supervision followed by Self-Refinement), the system can learn complex reasoning structures from smaller, outcome-conditioned datasets (e.g., via rsLoRA fine-tuning). This allows clinical risk prediction to leverage large pre-trained medical priors without requiring massive amounts of task-specific labeled data, making it viable for rare diseases or novel clinical conditions.

  1. Model Agnosticism and Generalization:

The framework is designed to be backbone-agnostic (tested on Qwen3, Llama 3.2, etc.). The reasoning supervision generalizes across different model architectures, suggesting that the learned dialectical reasoning principles are robust and transferable, meaning the framework can be deployed efficiently on any state-of-the-art LLM without needing extensive retraining for each new clinical task.

Sources

Related papers