EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation

arXiv:2608.07299 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Chengyi Peng, Haoyu Yang, Meixing Shi, Yuxiang Cai, Yankai Jiang

Zhejiang University · Shanghai Artificial Intelligence Laboratory

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Minor revision: fixed author metadata rendering in the arXiv HTML version

Code: https://github.com/Maybach-dream/EliSeg

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 70/100

The gist: The paper addresses report-grounded abnormality segmentation, where a model receives a chest radiograph and its associated radiology report, and must "identify the findings that warrant spatial

Terminology

Summary

The paper addresses report-grounded abnormality segmentation, where a model receives a chest radiograph and its associated radiology report, and must identify the findings that warrant spatial localization, and generate a corresponding mask for each. Unlike target-conditioned segmentation, this setting provides no preselected finding category, referring expression, target list, point, or bounding box. The authors argue that existing segmentation methods largely bypass this ambiguity by receiving a target identity or spatial prompt before inference, which acts as a hidden target oracle.

The key challenge is that "Radiology reports are written for clinical communication rather than as segmentation prompts. They may describe findings that are present, negated, uncertain, historical, resolved, or outside the segmentation scope. Multiple valid abnormalities may also coexist and spatially overlap. Consequently, the occurrence of a disease term does not by itself define a segmentation target."

The paper focuses on seven chest radiographic abnormalities: cardiomegaly, edema, pleural effusion, atelectasis, lung opacity, pneumonia, and consolidation. A finding is eligible for segmentation only when the report asserts its current presence. Negated, prior, uncertain, and out-of-scope findings are considered ineligible.

Removing the target oracle "exposes a missing stage between report understanding and mask prediction: target construction. This stage determines whether a reported finding is eligible for segmentation, how many valid targets are present, and how the resulting semantic slots correspond to the findings. Errors therefore extend beyond inaccurate boundaries to include false targets, omitted findings, incorrect target counts, and semantic mismatches."

Formally, for each sentence r t in report R = (r1,..., r T), the model predicts a target set T t = [(f t,1, M t,1),..., (f t,K t, M t,K t)], where K t is the target cardinality, f t,k the semantic identity, and M t,k the binary mask. "The semantic category, rather than an anatomical instance or connected component, constitutes the prediction unit. Each category contributes at most one target per sentence; bilateral, multifocal, and diffuse manifestations are merged into a single category-level mask."

The sentence-level action a t ∈ SEG, REJ, SILENT summarizes target status: "a t = SEG if K t > 0, a t = REJ if no eligible target exists but explicit ineligible evidence is present, and a t = SILENT otherwise. When eligible and ineligible findings co-occur, SEG takes precedence."

EliSeg is a propose–verify–revise framework comprising three components:

The Actor comprises an image encoder E, a causal language model, a projection module g, and a mask decoder D. It emits a compact vocabulary of control tokens — an empty response for SILENT, [REJ] for rejection, and [SEG1] to [SEGK t A] for segmentation slots, with grammar-constrained decoding limiting cardinality to K max = 3 (chosen because the vast majority of eligible sentences contain no more than three targets). Each segmentation token serves as a semantic query for mask decoding: M = D(E(I), g(z t,k A)). The Actor does not emit explicit finding labels. Instead, the semantic identity of each slot is encoded by its report-conditioned representation and the canonical target ordering used during supervision.

The Actor is trained with L Actor = L CE + λ BCE L BCE + λ Dice L Dice, supervising the control sequence and masks simultaneously.

The Verifier is a frozen Qwen2.5-VL-7B model that, Given only (h t, r t), deterministically generates a structured record under a fixed few-shot prompt, without access to the radiograph or any Actor outputs. It assigns one of five textual states: positive, negated, prior, uncertain, or absent to each finding. Positively asserted findings beyond the segmentation scope are recorded separately. The verified inventory L t V is formed by arranging positive findings in canonical order, giving K t V. The Verifier operates exclusively on textual evidence and does not assess reference-mask availability or spatial properties.

Revision converts the verified inventory into a target plan. Cardinality is revised as K̃ t = min(K t V, K max). "A nonempty verified inventory yields a SEG action. When no eligible finding is verified, the action is set to REJ if the report contains explicit negated, prior, uncertain, or positive out-of-scope evidence, and to SILENT otherwise."

For a revised SEG action, "Revision constructs a canonical control sequence containing K̃ t segmentation tokens. Rather than invoking autoregressive decoding again, the complete sequence is supplied to the shared Actor through an inference-mode, teacher-forced forward pass. This introduces no gradient update or additional learned parameters. Critically, The verified finding names are not injected into the Actor. The Verifier inventory determines only the revised action and target cardinality, while the semantic correspondence of the slots is reconstructed from the report-conditioned Actor representations under the canonical ordering."

A deterministic consistency gate "compares the actions and cardinalities predicted by the Actor and Verifier. If they agree, the original Actor output is retained. Otherwise, Revision constructs a corrected control sequence and re-executes the shared Actor."

  • Evaluation sets: MIMIC-CXR-ILS (test set: 1,008 finding-level segmentation targets with valid reference masks) and 600 ineligible report mentions (balanced across negated, prior, uncertain) for rejection analysis. CheXlocalize is used for external zero-shot transfer.

  • Input settings: Under R, the model receives radiograph + unfiltered report; under R+G, the gold eligible-finding inventory is provided; under G native, each gold target is provided via the model's native interface (finding prompt, point, or bounding box).

  • Baselines: Text-based methods (GSVA, MedCLIP-SAMv2, BiomedParse, MedSAM3, CheXagent, MAIRA-2, ROSALIA), extract-then-segment cascades (CheXbert→various segmenters), and spatial-prompt methods (MedSAM, IMIS-Net).

  • Implementation: The Actor is initialized from ROSALIA-7B with rank-8 LoRA, AdamW at lr 5×10−5, batch size 2, images resized to 1024×1024.

  • Metrics: IoU, Dice, NSD, HD95. Missing, rejected, or failed predictions for eligible targets are retained and treated as empty masks. For ineligible mentions, FSR (false-segmentation rate) is reported jointly with IoU+.

EliSeg achieves the strongest Report-Inferred performance: IoU 59.4, Dice 74.5, NSD 23.2, HD95 213.9, compared to ROSALIA (40.8/57.9/18.1/351.2) and CheXbert→ROSALIA (53.9/70.0/21.9/311.6). The improvement of CheXbert→ROSALIA "confirms that target construction is a major source of error, but also exposes the limitation of a decoupled cascade: omitted findings cannot be recovered after target extraction, while spurious findings are propagated directly to mask generation."

Per-finding, EliSeg achieves the highest IoU and Dice for every finding, with the gain largest for pneumonia and consolidation and smallest for cardiomegaly. When the gold inventory is supplied (R+G), EliSeg changes only marginally, indicating most eligibility and cardinality uncertainty has already been resolved from the report.

EliSeg achieves FSR of 14.2% while retaining the highest IoU+ of 60.8. The paper notes a structural separation between eligibility control and segmentation capacity. Target-conditioned backends generally produce a mask once a finding is requested, resulting in high FSR without an eligibility front end. GSVA's lower FSR is misleading because its IoU+ of 1.4 indicates that this apparent selectivity largely reflects limited mask-generation capacity. Evidence-specific analysis shows prior mentions as the main remaining source of rejection error (29.0% FSR for prior, vs. 2.0% for negated and 11.5% for uncertain).

Without the Verifier, Revision is conditioned on the Actor's own target cardinality and therefore tends to reproduce the original target structure rather than correct it. Without Revision, the Verifier can detect errors in eligibility and cardinality, but this information remains at the decision level and cannot alter the number or content of the predicted masks. The two partial configurations perform similarly (IoU 49.6 vs 49.0), but Closing the loop improves IoU by 4.5–5.1 points and Dice by 5.9–6.4 points, while substantially reducing HD95 (from 426-442 to 229.4). Revision fires on 22% of sentences, with the largest gains from the 10% where the Actor declined to segment entirely (+0.551 Dice), while being a measured no-op (−0.001) on the 78% with correct full slot structure — evidence that the improvement comes from repairing target structure rather than from systematically redrawing masks that were already correct.

On CheXlocalize (given-target setting, no paired reports), the EliSeg Actor alone achieves the strongest performance: IoU 31.9, Dice 48.8, NSD 7.5, HD95 267.5, vs. ROSALIA (29.6/45.7/6.5/391.3). The relatively small NSD gain together with the pronounced HD95 improvement indicates that the main benefit lies in reducing severe boundary outliers rather than uniformly improving local contour agreement.

The authors note: The study covers seven chest radiographic findings and uses one category-level mask per finding, without distinguishing bilateral or multifocal instances. The fixed K max = 3 may truncate rare cases with more eligible targets, and semantic errors that preserve the target count may remain undetected since the gate compares only action and cardinality. External evaluation is limited to the Actor because CheXlocalize lacks paired reports.

The paper concludes: "We studied abnormality segmentation from radiology reports, where valid targets must be identified and delineated without prespecified target cues. We introduced EliSeg, which combines an Actor with independent textual verification and conditional re-execution to correct target eligibility and cardinality before final mask decoding. Experiments on MIMIC-CXR-ILS show that EliSeg improves segmentation across findings while suppressing masks for ineligible mentions. Ablation studies confirm the complementary roles of verification and revision, and evaluation on CheXlocalize demonstrates effective transfer of the Actor to an external dataset."

Improvements for AI systems

Based on EliSeg, here are concrete improvements to AI systems and what the improved system can do:

Improvement 1: Add a report-verification stage independent of the segmenter.

  • Use a frozen multimodal language model as a verifier that reads only the sentence and produces a structured finding inventory (positive, negated, prior, uncertain, absent).

  • The improved system can correctly decide whether a finding is currently present, how many distinct targets exist, and which findings are out of scope—without being biased by the segmenter’s own predictions or by visual patterns.

Improvement 2: Add verification-guided revision with conditional re-execution.

  • After the actor proposes segmentation slots, compare its action and cardinality against the verifier’s inventory. If they disagree, rebuild the control sequence (e.g., [SEG1], [SEG2], or [REJ]) and run the shared actor in teacher-forcing mode to decode masks again.

  • The improved system can recover omitted findings, remove spurious masks, correct target counts, and suppress masks for negated/prior/uncertain mentions—without retraining or extra parameters.

Improvement 3: Separate action control from mask decoding via grammar-constrained control tokens.

  • Use a compact vocabulary: silence for no finding, [REJ] for explicit ineligible evidence, and [SEG1]…[SEG3] for eligible segmentation slots. Constrain decoding so the actor never emits more than three targets per sentence.

  • The improved system can jointly predict report-level action and masks, while structurally preventing unbounded or malformed target sequences.

Improvement 4: Treat the semantic category as the prediction unit, not the anatomical instance.

  • Merge bilateral, multifocal, and diffuse manifestations into one category-level mask per sentence, and enforce at most one target per category per sentence.

  • The improved system avoids redundant or conflicting masks for the same finding and produces clinically consistent, category-aligned segmentations.

Improvement 5: Explicitly supervise target construction, not just mask boundaries.

  • Train the actor with cross-entropy on the control-token sequence plus BCE/Dice on masks, so the model learns when to segment, reject, or stay silent.

  • The improved system reduces false targets, missed findings, and semantic mismatches—errors that boundary-only training cannot address.

Improvement 6: Use a canonical target ordering during supervision and inference.

  • Arrange positive findings in a fixed canonical order (e.g., cardiomegaly, edema, pleural effusion, atelectasis, lung opacity, pneumonia, consolidation) so the actor’s slot representations carry semantic identity.

  • The improved system can map each segmentation mask to the correct finding without requiring explicit label injection at inference time.

Improvement 7: Replace cascaded extract-then-segment pipelines with an end-to-end propose–verify–revise loop.

  • Avoid irreversible errors from external extractors (e.g., CheXbert) by feeding the verified inventory back into the same actor for mask refinement.

  • The improved system can recover omitted findings that a decoupled cascade would lose, and can correct spurious findings that a decoupled cascade would propagate.

Improvement 8: Keep the verifier text-only and the actor vision-language, then combine them at decision time.

  • The verifier decides eligibility/cardinality from text alone; the actor handles image-conditioned mask decoding.

  • The improved system benefits from complementary signal sources: textual evidence for target existence and visual evidence for spatial delineation, leading to higher IoU/Dice and lower false-segmentation rates simultaneously.

What the improved AI system can do overall:

  • Given only a chest radiograph and its raw radiology report, it can identify all currently present abnormalities among seven findings, generate one mask per finding, and correctly ignore negated, prior, uncertain, resolved, or out-of-scope mentions.

  • It can match or exceed the segmentation quality of target-conditioned systems while also suppressing ineligible masks, with a false-segmentation rate around 14% and the highest IoU+ among report-driven methods.

  • It transfers zero-shot to external datasets where target prompts are given, improving boundary quality (e.g., reducing Hausdorff distance outliers) even without paired reports.

  • It is robust to ambiguous report language and can repair its own target structure during inference, rather than committing to an irrevocable first-pass prediction.

Abstract

Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, prior,uncertain, or irrelevant findings, while multiple valid abnormalities may coexist. Existing segmentation methods largely bypass this ambiguity by receiving a target identity or spatial prompt before inference, which acts as a hidden target oracle. We study report-grounded abnormality segmentation, where a model must determine target eligibility, cardinality, and finding-to-mask correspondence directly from an unfiltered report before delineating the corresponding regions. We propose EliSeg, an atcor--verify--revise framework that integrates target construction with mask generation. A grammar-constrained Actor proposes target slots and masks, an independent text-only Verifier reconstructs the eligible finding inventory, and Revision selectively re-executes the shared Actor when their target structures disagree. EliSeg requires no predefined target identity, finding prompt, point, or bounding box. Experiments on MIMIC-CXR-ILS show that EliSeg consistently outperforms direct segmentation methods and extract-then-segment cascades across findings, while effectively suppressing masks for ineligible report mentions. Ablation studies confirm the complementary roles of verification and revision, and evaluation on CheXlocalize demonstrates effective transfer of the EliSeg to an external dataset.Code is available at https://github.com/Maybach-dream/EliSeg.

Sources

Related papers