TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series
summary
The gist
Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS), must deliver both calibrated risk
In short
TRIAGE trains a Large Language Model to generate dialectical reasoning over competing clinical outcomes. By requiring the model to create separate, outcome-specific rationales for each possibility, it avoids 'risk polarization,' allowing a single model to produce continuous risk scores grounded in explicit clinical justification.
Key concepts
- Risk Polarization Problem
- This occurs when asking an LLM for reasoning before scoring causes the final risk score distribution to become extremely polarized. This happens because the generated text usually commits to one outcome, making the prediction nearly certain, and it rarely presents evidence supporting all possible outcomes simultaneously.
- Dialectical Reasoning Supervision
- This is a training stage where the model is taught to generate rationales for every potential clinical outcome separately. The key constraint is that each rationale must be strictly outcome-specific, meaning it cannot reference or contrast with any other alternative outcome, ensuring focused evidence generation.
- Continuous Risk Estimation
- Instead of forcing a single binary prediction, TRIAGE uses the model's implicit probability distribution at a specific answer position to derive a continuous risk score. This method preserves a graded signal that reflects the uncertainty and likelihood across all possible outcomes, rather than collapsing into an absolute yes or no.
- LLM-as-a-Judge Assessment
- This is a qualitative evaluation method where another LLM is used to assess the quality of clinical reasoning traces produced by TRIAGE. This helps measure how well the model's generated explanations align with actual, clinically sound justifications for patient outcomes.
Terminology used across episodes
This episode discusses
- TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series · Paper Radio
- gpt-oss-120b & gpt-oss-20b Model Card
- The Llama 3 Herd of Models · Paper Radio
- A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
- Adam: A Method for Stochastic Optimization
- OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data
- EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
- Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
- Understanding R1-Zero-Like Training: A Critical Perspective
- TimeX++: Learning Time-Series Explanations with Information Bottleneck
- Sequence Level Training with Recurrent Neural Networks
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- A Review of Deep Learning Methods for Irregularly Sampled Medical Time Series Data
- Kimi K2: Open Agentic Intelligence
- Qwen3 Technical Report
The paper
TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series · Read on arXiv
KAIST · AITRICS
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series".
Jane: Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS),
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the title and who came up with this research. The paper is called "TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series." It’s clear they are focusing on making the reasoning part of the prediction process, which is a big step forward.
Jane: The authors include Hyeongwon Jang, Gyouk Chu, Changhun Kim, Joonhyung Park, Hangyul Yoon, and Eunho Yang from KAIST and UW-Madison. It shows this is a collaborative effort coming from some strong AI research groups in the field.
Lu: I think their focus on dialectical reasoning is the real innovation here; it’s not just about asking for an explanation after a prediction, but structuring the model to generate arguments for competing outcomes simultaneously.
Meng: It makes sense that they are focusing on dialectical reasoning because it directly attacks that risk polarization problem we discussed earlier, where models get stuck committing too early to one outcome.
Lalam: And I think the implications are huge for trust; when a clinician can actually verify the model's logic by looking at these explicit rationales, it moves the AI from being a black box suggestion to a reasoning partner.
Tom: Exactly, Lalam. This paper is showing us how to move beyond just getting an answer and start getting a verifiable explanation for that answer. It’s about making the AI's thinking transparent in a clinical context.
Jane: And they specifically target ISMTS, which is crucial because most clinical data isn't perfectly regular, and this framework is designed to handle those messy time series effectively.
Lu: They are proposing a two-stage training pipeline involving dialectical reasoning supervision followed by self-refinement through reinforcement learning using Group Relative Policy Optimization, which is a clever way to fine-tune that complex reasoning ability.
Meng: The engineering approach with the GRPO loss function, where they supervise the reasoning tokens but only apply cross-entropy loss to the final decision token, seems like a smart way to balance training both the intermediate steps and the final output.
Lalam: That sounds like a solid way to ensure that even though we are focusing on reasoning generation, we don't lose sight of actually producing a useful classification result.
The paper's summary: Tom: So, what’s the core summary of TRIAGE? Essentially, the paper introduces TRIAGE as a framework to train an LLM to generate dialectical reasoning over competing clinical outcomes by asking for outcome-specific rationales.
Jane: That means instead of one long explanation that points toward one conclusion, they are training the model to generate separate arguments for every possible clinical result. This is designed specifically to fight that risk polarization issue where the scores get skewed towards extremes.
Lu: They set up a problem setup where they treat clinical risk prediction from ISMTS as a supervised classification task, aiming for a patient-specific class distribution of probabilities across all outcomes.
Meng: The setup involves defining the input data as an ISMTS, which is a multivariate time series record containing time, value, and variable index information for different clinical measurements.
Lalam: So they are taking that complex time series data and trying to learn a predictor that outputs a probability distribution over all possible outcomes for each patient, rather than just one final guess.
Tom: And the key mechanism they use is eliciting these outcome-specific rationales from an LLM separately for each candidate outcome, making sure those rationales don't reference other outcomes.
Jane: That’s a really important constraint because if the rationale references another outcome, it becomes contrastive instead of strictly outcome-specific, which the authors found was a problem in previous methods.
Lu: They then use these collected rationales, along with ground-truth answer tokens, to form reasoning traces that guide the final risk estimation within a single LLM structure.
Meng: It shifts the paradigm from having an LLM predict first and *then* trying to explain it, to having the LLM generate the explanation *as* part of its internal decision-making process for a patient-comparable score.
Lalam: And that’s what really excites me about this summary; it fundamentally changes how we think about using LLMs in high-stakes domains like medicine—it forces structured, balanced deliberation instead of simple pattern matching.
The paper's improvements: Tom: Now, let’s talk about the actual improvements they propose in TRIAGE. The paper details a two-stage training pipeline: first, Dialectical Reasoning Supervision, and second is Self-Refinement through reinforcement learning with GRPO.
Jane: In the first stage, they are teaching the model to produce those outcome-conditioned rationales by making sure the LLM only generates evidence supporting one specific outcome at a time.
Lu: They enforce two strict constraints there: first, no reference to any alternative outcome in the rationale, and second, it must not fabricate unobserved evidence; if it can't find evidence for an outcome, it leaves the rationale blank.
Meng: That constraint about leaving the rationale blank when evidence is missing is something I’m thinking about practically; we need a robust way to handle uncertainty without just making up justifications for things that aren't there.
Lalam: It addresses the issue of hallucination in reasoning, which is a huge problem with current LLMs, because if it can’t find evidence, it stops trying to commit to a false narrative.
Tom: The second stage refines this using reinforcement learning with Group Relative Policy Optimization, and they introduce a batch-level reward function that promotes cross-patient comparability alongside the calibration benefit of on-policy training.
Jane: That reward function is interesting because it’s trying to get the model to be accurate on one patient while also being comparable across many patients simultaneously, which tackles that cross-patient comparability issue directly.
Lu: Their evaluation shows that this approach leads to significant improvements, with an average AUPRC increase of three point three percent and a calibration error reduction of eighty-one percent compared to competitive baselines on benchmarks like P12 and MIMIC-III.
Meng: An eighty-one percent reduction in calibration error is a big deal because it means the continuous risk scores they produce are much more trustworthy for actual triage decisions, which is exactly what we need when dealing with uncertain medical trajectories.
Conclusion: Tom: So, wrapping up the paper on TRIAGE: it seems the main implication is moving from brittle binary predictions to a system that yields continuous risk scores built on explicit, verifiable clinical reasoning.
Jane: That’s right, Tom; they’ve shown how dialectical reasoning can mitigate the risk polarization problem and allow a single LLM to provide patient-comparable risk grounded in clear evidence.
Lu: The future work mentioned involves evaluating this framework more broadly and seeing if these learned dialectical principles are robust enough to transfer effectively across different medical tasks or even different model architectures.
Meng: From an engineering perspective, the limitation they flag is that their evaluation is restricted to binary prediction tasks, so we need more work to see how well this holds up when we try to apply it directly for continuous risk estimation in a real-world continuous setting.
Lalam: I think the biggest takeaway for us all is that this framework provides a path toward building AI systems where the reasoning isn't hidden; it’s an explicit, balanced deliberation process that clinicians can trust and check.
Tom: It really feels like we’re getting closer to having AI tools that don't just give us a number but also show us the clinical thinking behind that number. Thanks for joining us on this deep dive into TRIAGE today!
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization