CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives
Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang
Stevens Institute of Technology · Genesis Research Group · New Jersey Institute of Technology
cs.CL, cs.AI, cs.IR
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/LEAF-Lab-Stevens/TemporalAnalysis
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
The gist: CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives Abstract Summary: The paper addresses the challenge of understanding temporal progression of symptoms in clinical
Terminology
Summary
CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives
Abstract Summary:
The paper addresses the challenge of understanding temporal progression of symptoms in clinical narratives, which is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives rarely provide explicit temporal anchors. Current approaches focus predominantly on pairwise relation classification across multi-visit and timestamp-rich records, leaving reconstruction of structured symptom trajectories from individual anchor-sparse reports largely unaddressed. The authors propose CRAFT, an LLM framework that pairs a generator with a constraint-based verifier to iteratively produce and refine stage-wise symptom timelines through targeted feedback. Evaluation is conducted on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports. Experiments across four LLM backbones demonstrate that CRAFT consistently improves temporal ordering accuracy, with ablation analysis isolating the contribution of generator and verifier components across model capability levels.
Introduction:
Clinical temporal reasoning—recovering a stage-wise trajectory of clinical events from narrative text—is central to disease progression modeling, treatment outcome monitoring, and safety signal detection. However, constructing such orderings from free text remains labor-intensive. Temporal cues in these narratives are frequently implicit or expressed relative to other events rather than anchored to fixed dates, and co-mentions, restatements, and status updates further obscure the true chronological order of symptoms. A large body of work on temporal information reasoning has focused on identifying events and temporal expressions and predicting pairwise relations to derive a global ordering, with recent efforts expanding to exhaustive relation coverage and complex temporal fact extraction. In the clinical domain, progress has been bottlenecked by limited benchmark diversity, with most work concentrated on a small number of corpora and relation inventories. Recent LLM-based approaches to clinical temporal reasoning further assume multi-visit timelines, timestamp-linked supervision, or constrained data settings. As a result, standardized methods and benchmarks for ordering symptom progressions within single-report, anchor-sparse clinical narratives remain underexplored.
To address this gap, the authors propose CRAFT (Clinical Refinement with Adaptive Feedback for Temporal ordering), a generator–verifier framework that models temporal trajectory reconstruction as an iterative structured prediction task under weak temporal anchoring, where candidate trajectories are refined via constraint-based feedback. Inspired by iterative refinement paradigms such as Self-Refine, CRAFT introduces a structured trajectory representation and a task-specific verification mechanism tailored to temporal ordering, and can be instantiated with different generator and verifier configurations. In this work, CRAFT is instantiated as CRAFT-Full, which pairs a full-regeneration generator with a multi-criterion additive verifier; two baselines (PIVOT, GUIDE) and two ablations (CRAFT-G, CRAFT w/o V) are additionally defined to isolate the contribution of the generator and verifier components respectively.
To enable rigorous evaluation, the authors introduce MedTempo (Medical Temporal Ordering Benchmark), a benchmark for temporal progression reconstruction from medical free text. MedTempo contains 5,347 narrative reports spanning three vaccine types, each consisting of a single report per patient with no explicit absolute time anchor and paired with a provided symptom list. The benchmark task focuses on the 3,166 reports that exhibit temporal evidence of distinct symptom progression, for which expert-validated stage-wise ordering annotations are provided. The remaining reports contain no temporal progression and are retained in the dataset release to support future work on temporal-evidence identification, but fall outside the scope of the current benchmark evaluation.
Primary Contributions:
-
The authors propose CRAFT, a generator–verifier framework for iterative temporal reasoning refinement under weak temporal anchoring, with controlled baselines and ablations that isolate generator and verifier contributions.
-
They introduce MedTempo, an expert-annotated benchmark for evaluating structured temporal trajectories over anchor-sparse clinical narratives.
-
They conduct extensive experiments across four LLMs, revealing distinct refinement behaviors tied to model capability and verifier calibration.
Related Work:
Temporal information extraction has been shaped by the TimeML annotation framework, which introduced a markup language for events, temporal expressions, and their relations. The TempEval shared tasks established standardized evaluation for pairwise temporal relation classification, with systems progressing from rule-based and CRF/SVM approaches to neural methods. A common paradigm is to predict pairwise relations and then induce a globally consistent ordering, while more recent benchmarks emphasize richer annotation for event ordering and LLM-based strategies for temporally grounded fact extraction. Transformer-based methods have become the dominant paradigm. However, these efforts typically center on local relation correctness rather than the end-to-end reconstruction of an ordered, grouped trajectory under weak anchoring.
In the clinical domain, the i2b2 2012 challenge brought temporal relation extraction to clinical discharge summaries, followed by the Clinical TempEval shared tasks on the THYME corpus, where best-performing systems evolved from CRF/SVM classifiers to LSTMs. Surveys document persistent difficulties including implicit time anchors, inter-sentence relations, and the gap between relation-level extraction and usable patient timelines. More recent work applies neural end-to-end methods to established clinical corpora, and LLM-based approaches have begun to examine prompting and fine-tuning for clinical temporal relation extraction as well as timeline extraction from medical case reports. However, these approaches typically require longitudinal records spanning multiple visits or rely on structured temporal metadata. In contrast, CRAFT operates on single-report clinical narratives where temporal cues are sparse or implicit.
A parallel line of work applies iterative refinement to structured prediction, most notably the Self-Refine paradigm, which iterates over model outputs using self-generated feedback. In the clinical domain, Hein et al. apply iterative refinement with human-in-the-loop review cycles to improve extraction precision. However, such approaches are not suitable for systematic benchmark evaluation across model tiers, as they conflate model capability with human reviewer effort. CRAFT leverages iterative refinement for structured clinical temporal extraction in a fully automated setting, pairing a generator with a multi-criterion verifier and enabling principled comparison across frontier and open-weight models.
Dataset Creation:
MedTempo is constructed from VAERS (Vaccine Adverse Event Reporting System), managed jointly by the CDC and FDA, a passive surveillance database in which healthcare professionals, patients, and manufacturers submit reports of adverse events following immunization. Each report includes demographics, vaccination details, free-text clinical narratives, and MedDRA-coded symptom lists. The authors focus on three widely administered COVID-19 vaccines: Pfizer-BioNTech, Moderna, and Janssen, covering reports from 2021 to 2024.
Data Sampling:
Starting from the full VAERS corpus, a multi-stage filtering and stratified sampling pipeline was applied. First, non-symptom MedDRA terms were removed by building an exclusion list from three non-clinical System Organ Classes (Surgical and medical procedures, Social circumstances, Product issues) combined with human-annotated non-symptom labels, yielding 5,601 excluded terms. Reports with three or fewer distinct symptoms were discarded as they lack sufficient temporal variation. Reports of ten or fewer words (typically bare symptom lists without narrative context) were also removed. For reports between 11 and 30 words, rule-based temporal keyword filtering (relative markers such as before/after, duration terms, date patterns) was applied to retain only those with explicit temporal cues; reports of 30 or more words were retained unconditionally. Finally, stratified sampling balanced by year and report length was applied per vaccine to draw 2,000 records each, preserving the distribution of the underlying VAERS corpus.
Annotation of Temporal Sequence:
Temporal timelines were produced via a human-in-the-loop three-phase protocol: GPT-4o mini generated initial stage-ordered timelines; two annotators with a medical NLP background independently reviewed and labeled each sequence; and all disagreements and uncertain cases were resolved through collaborative adjudication. Human inter-annotator agreement (IAA) was 93%. After adjudication, 80% of LLM annotations were accepted without change, 9% were corrected, and 11% were excluded for temporal ambiguity, yielding a final corpus of 5,347 records.
Dataset Statistics:
Table 2.1 summarizes the 3,166 temporally-evident reports that form the primary benchmark subset MedTempo-T; the remaining 2,181 reports contain no temporal progression and are released separately as MedTempo-NT. Symptom count is consistent across vaccines (median 6), while narrative length varies: Pfizer-BioNTech reports are longest (median 102 words) versus Moderna (83) and Janssen (87). Because symptom count is stable regardless of length, longer reports likely contribute richer contextual cues rather than additional adverse events, providing stronger signals for temporal extraction. Figure 3.1 shows that thermoregulatory symptoms (pyrexia, chills) dominate early stages, while headache and fatigue are the most frequent overall, illustrating the diversity of symptom trajectories in the benchmark.
Method:
CRAFT operates as a fully automated iterative loop: the generator proposes a candidate temporal sequence, the verifier scores it against structural and temporal constraints and returns targeted feedback, and the loop repeats until the candidate is accepted or a fixed iteration budget is exhausted.
Problem Formulation and Temporal Representation:
For each post-vaccination report r, let xr denote its free-text clinical narrative and F(r) = f1,..., fn the provided list of adverse clinical findings (MedDRA Preferred Terms from the VAERS SYMPTOM fields). F(r) is treated as given; entity extraction or coding is not performed. The goal is to predict an explicit temporal ordering over F(r) when the narrative expresses a temporal progression (i.e., at least one new finding appears after previously mentioned findings). Concurrent onset, changes in severity or resolution without new onsets, and timelines inferable only from durations are excluded. Formally, the output is an ordered sequence of non-empty time buckets B(r) = (B1, B2,..., BK), where each finding in F(r) is assigned to exactly one bucket Bk, grouping together findings that occurred at the same point in the patient’s clinical course, and index k orders the buckets from earliest to latest. This representation is stored as a JSON list of buckets, which is the format used by both the generator and verifier. Models are evaluated solely on temporal structure (ordering and grouping); optional evidence snippets are not scored.
Generator Agent:
At iteration i, the generator LLM implements g: F(r), xr, feedback(i−1) → B̂(i)(r), where feedback(i−1) is empty at i=1 and contains verifier guidance thereafter. CRAFT-Full uses a full-regeneration generator: a single prompt template with full task instructions at each iteration, with verifier feedback appended to the context when available. The generator receives: (i) a task description and definition of temporal progression, including phenomena that should not be treated as progression; (ii) the finding list F(r), which must appear exactly as given in the output; and (iii) the free-text narrative xr. It outputs a JSON bucket sequence using only findings from F(r) and avoiding unsupported temporal inferences.
Verifier Agent:
The verifier implements v: F(r), xr, B̂(i)(r) → decision(i), feedback(i), score(i), where decision(i) ∈ ACCEPT, REVISE, feedback(i) describes issues to fix, and score(i) ∈ 0,..., 5 supports early stopping. The verifier first calls FormatTool, a deterministic helper module that normalizes the raw generator output into the JSON bucket schema without altering temporal content, then applies an additive rubric that assigns one point each for: valid JSON with earliest→latest ordering; placing non-mentioned symptoms as “none” in the final group; using each symptom exactly once; grouping symptoms that occur around the same time; and ordering groups according to temporal cues in the narrative. When the score meets threshold θ, the verifier returns ACCEPT; otherwise it returns REVISE with targeted feedback for the next iteration. If no candidate is accepted within Tmax iterations, the loop terminates and the last candidate is returned.
Experimental Setup:
Three research questions are organized: RQ1 (Method effectiveness) asks whether CRAFT-Full outperforms baselines (PIVOT, GUIDE) across model tiers; RQ2 (Model capability) asks how well different LLM backbones perform on MedTempo and whether capability ordering is stable across configurations; RQ3 (Vaccine discrepancy) asks to what extent performance and error patterns vary across vaccine types.
Dataset: MedTempo contains 5,347 vaccine adverse-event narratives across three COVID-19 vaccine types, each paired with a provided symptom list. Evaluation is on the 3,166 reports from MedTempo-T with temporal evidence of distinct symptom progression.
Models and Settings: Four LLMs are evaluated: GPT-4.1, Claude Sonnet 4.5, MedGemma-27B, and Llama-3.3-70B. Proprietary models are accessed via their respective APIs with deterministic decoding. Open-weight models run locally using Hugging Face transformers with 4-bit quantization (NF4 via bitsandbytes). Five settings are evaluated: CRAFT-Full pairs the full-regeneration generator with the additive rubric verifier; PIVOT pairs the full-regeneration generator with an anchor-based verifier inspired by the document-creation-time (DCT) anchoring tradition; GUIDE pairs an edit-conditioned generator with the same anchor-based verifier; CRAFT-G swaps in the edit-conditioned generator while holding the additive verifier fixed; CRAFT w/o V removes the verification loop entirely, running one generator pass without feedback.
Implementation Details: All runs use fixed prompt templates that enforce a strict JSON schema and require every symptom in the provided list to appear exactly once, grouping symptoms into the same stage when the narrative does not support an internal order. Deterministic decoding (no sampling) is used with max new tokens=512. For generator–verifier configurations, up to max iter=4 refinement iterations are run, and outputs are accepted when the verifier score is at least θ = 3. Open-weight experiments are run on a workstation equipped with two NVIDIA RTX A5000 GPUs (24GB each), with device sharding and 4-bit quantization.
Evaluation Metrics: Temporal sequences are evaluated as ordered lists of buckets, where each bucket is a set of items and within-bucket order is trivial. Three metrics are used: Strict Exact Match (EM), which requires the prediction to exactly reproduce the gold segmentation and inter-bucket order (within-bucket order ignored); Kendall’s τb, which measures pairwise concordance/discordance with tie handling; and Group-Aware LCCS (Longest Common Contiguous Subsequence), which treats each bucket as a token and rewards unbroken spans of perfectly matched phases.
Results:
RQ1: CRAFT-Full vs. Baselines. Against PIVOT, CRAFT-Full gains +1.0 EM points for GPT-4.1 (35.61 vs. 34.60), +0.8 for Llama-3.3-70B (28.04 vs. 27.22), +0.7 for Claude Sonnet 4.5 (37.14 vs. 36.44), and +0.1 for MedGemma-27B (20.85 vs. 20.75). Margins over GUIDE are uniformly larger (GPT-4.1: +1.6; Llama: +2.0; Claude: +1.5; MedGemma: +0.7). CRAFT-Full achieves the highest EM in every model block compared to baselines. Although baselines occasionally match or exceed CRAFT-Full on τb or LCCS, these metrics credit partial ordering agreement and can score highly even when the full trajectory structure is incorrect. EM, which requires the entire stage-wise grouping and ordering to match gold, is the most demanding metric and the one on which CRAFT-Full consistently leads.
CRAFT-Full actively uses its refinement budget: AvgIters reaches 2.99 for GPT-4.1 and 3.46 for Claude, and performance builds across iterations. For GPT-4.1, EM rises from 26.90 at i=1 to 35.61 at i=4, a gain of 8.7 points across rounds. In contrast, PIVOT and GUIDE converge after a single pass in most instances (AvgIters ≈ 1.2–2.0) because the anchor-based verifier accepts outputs quickly without substantive ordering improvement. For GPT-4.1, PIVOT peaks at EM@2=34.73 and GUIDE degrades from its peak of 34.92 at i=1. This confirms that CRAFT-Full’s additive rubric provides richer, more actionable feedback that sustains improvement across the full refinement budget, whereas the anchor-based verifier’s narrower signal cannot drive continued gains beyond the first pass.
RQ2: Overall Model Performance. The capability ordering is stable across all configurations and metrics: Claude Sonnet 4.5 > GPT-4.1 > Llama-3.3-70B > MedGemma-27B. Under CRAFT-Full, Claude achieves the highest total EM (37.14), followed by GPT-4.1 (35.61), Llama (28.04), and MedGemma (20.85). This ordering holds without exception across all three metrics and all vaccine strata. Models differ notably in how they respond to iterative refinement. GPT-4.1 benefits most from additional iterations: under CRAFT-Full, EM climbs from 26.90 at i=1 to 35.61 at i=4 (+8.7 points), with LCCS and τb following similar upward trends (47.10→61.46 and 45.71→57.77 respectively). In contrast, Claude starts strong (EM 36.57 at i=1) but gains only +0.6 points across iterations, suggesting its first-pass outputs already capture most of the temporal structure. Llama and MedGemma show similarly flat iteration curves (EM gains of +0.6 and +0.2 respectively), but for a different reason: both converge quickly (AvgIters 1.2–1.5), indicating that the verifier accepts their outputs early rather than driving further improvement. This divergence between models that saturate from high initial quality (Claude) and those that stall from limited capacity to act on feedback (Llama, MedGemma) highlights that iteration utility is tied to model capability.
RQ3: Vaccine-Stratified Performance. Performance is consistently stratified by vaccine type across all models and configurations: Moderna narratives yield the highest EM in every case, followed by Janssen, with Pfizer-BioNTech systematically lowest. Under CRAFT-Full, the Moderna–Pfizer gap is 7.4 points for GPT-4.1 (40.06 vs. 32.68), 6.5 for Claude Sonnet 4.5 (41.18 vs. 34.65), 5.1 for Llama-3.3-70B (31.90 vs. 26.77), and 7.2 for MedGemma-27B (25.99 vs. 18.80). This ordering is preserved across all settings and all four models without exception. The vaccine gap is more pronounced on EM than on LCCS and τb, suggesting that models make segmentation errors on Pfizer narratives specifically rather than systematically misranking symptoms within groups. Since EM requires an exact match on both bucket composition and inter-bucket order while LCCS rewards contiguous correct spans and τb captures global pairwise ordering, the metric divergence points to Pfizer narratives being harder to segment into correct temporal groups rather than harder to order within those groups.
Ablation Study:
Verifier contribution. Removing the verification loop causes substantial performance drops for three of four models. CRAFT-Full gains +9.3 EM over CRAFT w/o V for GPT-4.1 (35.61 vs. 26.30), +7.4 for Llama-3.3-70B (28.04 vs. 20.66), and +6.4 for MedGemma-27B (20.85 vs. 14.44). For Llama and MedGemma, virtually all gain is captured at i=1: the verifier’s first-pass feedback corrects schema violations in unverified outputs, after which the output is accepted without substantive ordering improvement. For GPT-4.1, gains accumulate steadily through i=3 (AvgIters = 2.99), confirming that the verification loop provides genuine multi-iteration value for capable models. For Claude Sonnet 4.5, CRAFT-Full (37.14) falls slightly below CRAFT w/o V (37.83): the fixed threshold θ=3 does not recognize Claude’s near-correct first-pass output as satisfactory, forcing continued revision that introduces errors rather than correcting them. This highlights the importance of threshold calibration for further improvement.
Generator contribution. Replacing the full-regeneration generator with the edit-conditioned variant (CRAFT-G) consistently degrades final performance. For GPT-4.1, CRAFT-G achieves a strong EM of 34.89 at i=1 but degrades monotonically to 31.08 by i=4 (−3.8 points), while CRAFT-Full builds to 35.61. The edit-conditioned prompt is effective for initialisation but too constrained to sustain quality under continued verifier feedback across the full budget. For MedGemma-27B, CRAFT-G is actively harmful: EM drops from 20.15 at i=1 to 18.35 by i=2, indicating the model misinterprets edit-only instructions and degrades its own output under revision. CRAFT-Full avoids this across all tiers by regenerating from full task context at every iteration.
Case Study:
A representative example shows how verifier feedback guides the generator toward the correct trajectory. For GPT-4.1 with CRAFT-Full, the symptom list is Insomnia, Muscle disorder, Herpes zoster, Memory impairment. At i=1, the model correctly orders all four symptom groups but over-merges Herpes zoster and Memory impairment into a single stage, receiving a score of 2. The verifier precisely identifies the grouping error and provides a targeted one-sentence correction: “Symptoms are grouped in temporal order and all are present, but ‘Herpes zoster’ and ‘Memory impairment’ should be in separate groups as they are not clearly described as occurring at the same time.” At i=2 the model splits the two symptoms into separate stages, exactly matching the gold standard and receiving a score of 3, which meets the acceptance threshold θ=3 and triggers early stopping. This example illustrates that a single round of targeted feedback is sufficient for a capable model to correct a local grouping error within the refinement budget.
A second case study (Claude Sonnet 4.5, GUIDE) illustrates verifier miscalibration: Claude produces the exact gold-standard output at i=1, but GUIDE’s verifier assigns a score of 2 and instructs a Merge operation because the narrative uses relative date markers rather than explicit absolute anchors. The model follows the instruction, merging all symptoms into a single stage and producing a wrong output. The verifier then issues a Split instruction, the model recovers the correct grouping at i=3, and the cycle repeats. The loop terminates at i=4 on a merge step, returning an incorrect final output despite the model having produced the correct answer twice. This illustrates how GUIDE’s verifier’s strict requirement for explicit temporal anchors is miscalibrated for narratives that express temporal order through relative date references.
Conclusion:
The authors presented CRAFT, a generator–verifier LLM framework for iterative temporal reasoning over clinical narratives, and MedTempo, an expert-annotated benchmark of 5,347 adverse-event narratives, to evaluate structured symptom trajectory construction under sparse temporal anchoring. Experiments across four LLM backbones show that CRAFT consistently improves temporal ordering accuracy. Ablation analysis confirms that both generator and verifier components contribute meaningfully, with CRAFT-Full emerging as the strongest configuration across all model tiers. Results also reveal that refinement behavior varies by model capability, suggesting that adaptive verification strategies may further improve performance. Future work will leverage the 2,181 non-temporal reports from MedTempo-NT as supervision for a learned temporal-evidence identification module, extending CRAFT into a unified end-to-end framework for automatic progression detection and stage-wise temporal ordering.
Improvements for AI systems
Based on the paper, here are specific improvements to AI systems and what the improved systems can do:
1. Implement a Generator–Verifier Iterative Refinement Architecture
-
Add a structured output generator paired with a multi-criterion verifier that provides targeted, actionable feedback (not just pass/fail)
-
Use a full-regeneration generator that re-prompts from complete task context at each iteration, rather than edit-conditioned generators that degrade under feedback
-
Include a deterministic format-normalization tool between generator and verifier to ensure schema compliance without altering content
2. Design a Constraint-Based Additive Scoring Rubric for Temporal Reasoning
-
Score outputs on five explicit criteria: valid JSON ordering, handling of non-mentioned symptoms, complete symptom coverage, correct temporal grouping, and alignment with narrative cues
-
Set an acceptance threshold calibrated per model capability (e.g., θ=3 works for GPT-4.1 but over-rejects for Claude, suggesting adaptive thresholds)
-
Return specific, one-sentence corrective feedback (e.g., “Symptoms X and Y should be in separate groups”) rather than generic instructions
3. Integrate Weak-Anchor Temporal Representation Learning
-
Represent temporal trajectories as ordered buckets of co-occurring events, allowing grouping of simultaneous symptoms
-
Handle narratives with implicit temporal cues (relative dates, duration terms) without requiring absolute timestamps
-
Distinguish true temporal progression from concurrent onset, severity changes, or duration-only inferences
4. Add Model-Capability-Aware Refinement Control
-
Detect when a model saturates early (like Claude) and stop refinement to avoid error introduction
-
Detect when a model stalls from limited capacity (like Llama/MedGemma) and either accept early or provide simpler feedback
-
Track iteration utility per model and adjust refinement budget dynamically
5. Build Vaccine/Report-Type-Specific Error Diagnostics
-
Monitor segmentation errors separately from ordering errors (EM vs. τb divergence indicates grouping difficulty)
-
Identify narrative characteristics (e.g., Pfizer reports) that systematically challenge temporal grouping and adapt prompts accordingly
-
Reconstruct stage-wise symptom timelines from single, anchor-sparse clinical narratives (e.g., vaccine adverse-event reports) without requiring multi-visit records or explicit dates
-
Iteratively refine its own temporal ordering outputs using targeted, criterion-based feedback, improving accuracy by up to +9.3 EM points over single-pass generation
-
Handle implicit temporal expressions (relative markers, duration phrases) that current systems fail on, without forcing false absolute anchors
-
Group co-occurring symptoms into correct temporal stages while preserving exact ordering between stages, achieving up to 37.14% exact-match accuracy on expert-validated benchmarks
-
Adapt refinement behavior to its own capability level: strong models benefit from multiple iterations, while weaker models avoid degradation through early acceptance
-
Identify which narratives lack temporal progression (2,181 reports in MedTempo-NT) and flag them for separate handling, enabling future end-to-end progression detection
-
Provide explainable feedback that pinpoints specific grouping or ordering errors, enabling human-in-the-loop review and model debugging
-
Generalize across vaccine types and model tiers (proprietary and open-weight), with stable performance ordering and consistent improvement over baselines
Abstract
Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors. Current approaches to temporal information reasoning focus predominantly on pairwise relation classification across multi-visit and timestamp-rich records, leaving the reconstruction of structured symptom trajectories from individual anchor-sparse reports largely unaddressed. We propose CRAFT, an LLM framework that pairs a generator with a constraint-based verifier to iteratively produce and refine stage-wise symptom timelines through targeted feedback. We conduct evaluation on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports. Experiments across four LLM backbones demonstrate that CRAFT consistently improves temporal ordering accuracy, with ablation analysis isolating the contribution of generator and verifier components across model capability levels.
Sources
- Timeline-based Sentence Decomposition with In-Context Learning for Temporal Fact Extraction
- Prompting Large Language Models for Clinical Temporal Relation Extraction
- MedGemma Technical Report
- Self-Refine: Iterative Refinement with Self-Feedback
- Transformer-Based Temporal Information Extraction and Application: A Review
- Zero-shot Temporal Relation Extraction with ChatGPT
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering