MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

arXiv:2608.10635 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Yuan Wang, Hualiang Wang, Yixin Chen, Songtao Jiang, Shujian Gao, Jiaming Lin, Siming Fu, Jian Wu, Zuozhu Liu

Zhejiang University · Fudan University

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 10 pages, 6 figures

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 95/100

The gist: MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models Abstract Medical Vision-Language Models (MedVLMs) excel at verbalizing visual content, yet precise visual

Terminology

Summary

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

Abstract

Medical Vision-Language Models (MedVLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modules that decouple perception from understanding, creating representation gaps for region-language alignment. We present MedUP, a Med-VLM that natively unifies perception and understanding within a shared token space. At its core lies UniMedTok, a region tokenizer that encodes masks as discrete tokens in the LLM vocabulary, enabling the model to seamlessly interleave mask tokens with text. We curate UniMed-Train, a 1.84M-instance corpus spanning text-guided segmentation, region-grounded understanding, medical VQA and CoT-based segmentation, and introduce UniMed-Bench for unified evaluation. Extensive experiments show that MedUP outperforms native, agentic, and dual-decoder Med-VLMs across all tasks while remaining competitive with specialist segmentors, demonstrating the strong potential of unified understanding and perception modeling.

Introduction

Fueled by massive medical vision-language corpora, Medical Vision-Language Models (MedVLMs) have risen to prominence through a unified paradigm: verbalizing anything they see. This formulation endows them with powerful visual understanding and versatile linguistic generation capabilities, substantially advancing tasks ranging from visual question answering and report generation to medical reasoning. Among these tasks, precise visual perception, such as grounding and segmentation, serves as an indispensable prerequisite for trustworthy medical image understanding, providing explicit localizations of pathological regions and key visual cues before decision-making.

However, the paradigm of verbalizing anything constrains how Med-VLMs achieve perception natively: existing models resort to a text-centric strategy, wherein spatial references are verbalized as discrete numerical strings, such as bounding-box coordinates and segmentation keypoints. This formulation suffers from inherent limitations: Med-VLMs lack spatial sensitivity to such coordinate strings, fundamentally disconnecting the localization semantics they encode from the visual feature space.

Alternatively, a line of work pursues an orthogonal solution: equipping Med-VLMs with external segmentation modules. Representative approaches include tool-using agent models, which orchestrate off-the-shelf models (e.g., SAM) as callable tools during inference, and dual-decoder architectures, which completely decouple the output space into an LLM branch for linguistic content understanding and a dedicated visual decoder for segmentation mask prediction. Both paradigms, however, introduce notable drawbacks: they incur substantial additional parameters and architectural complexity. More fundamentally, whether by delegating perception to external tools or routing it through a separate decoder, the decoupling of understanding and perception creates a significant representation gap between the two capabilities. As a result, this decoupled design often leads to ineffective region-language alignment, as we empirically verify in Table 4, where decoupled Med-VLMs deliver underwhelming performance on visually-grounded tasks.

These observations motivate a fundamental question: Can we natively unify understanding and perception of Med-VLMs within a shared representation space? We argue that the key insight is to tokenize regions into discrete mask tokens within the shared token space of language, i.e., region as language.

Guided by this principle, we present MedUP, a medical VLM for unified region-language modeling. At its core lies UniMedTok, a native region tokenizer that encodes medical regions as discrete mask tokens and then aligns them with their language counterparts. Consequently, MedUP can seamlessly interleave mask tokens with text in a single sequence, grounding pathological findings to precise regions, and describing arbitrary regions in natural language, thereby achieving unified perception and understanding.

We train MedUP in two stages. In Stage 1, UniMedTok is pretrained via masked region reconstruction, where it learns to encode masks into discrete tokens and decode them back into masks. In Stage 2, we align UniMedTok with the VLMs, teaching it to natively “speak” mask tokens within text sequences. To this end, we curate a large-scale training corpus, UniMed-Train, comprising 902,648 text-guided segmentation samples, 902,648 region-grounded understanding samples, 4,000 CoT-based text-to-mask reasoning samples, and 27,738 standard medical image understanding samples, 1,837,034 training instances in total.

To evaluate MedUP systematically, we further build UniMed-Bench, a unified medical region-language benchmark with three tasks: Medical VQA, Text-Guided Segmentation, and Region-Grounded Understanding. This benchmark is designed to test not only whether a model can answer questions correctly, but also whether it can associate answers with the right medical regions and operate bidirectionally between language and masks. Under this protocol, we compare general-domain VLMs, medical VLMs, adapted mask-token baselines, and MedUP. We further introduce Seg-CoT task, a segmentation-oriented chain-of-thought paradigm for text-to-mask prediction. Instead of treating mask generation as a direct decoding problem alone, Seg-CoT encourages the model to produce segmentation through intermediate reasoning about anatomy, abnormality attributes, and localization cues. This improves semantic grounding and makes mask prediction more compatible with the reasoning behavior already exhibited by large vision-language models.

Extensive experiments demonstrate that MedUP achieves strong and consistent performance across all three tasks in UniMed-Bench, outperforming native Med-VLMs, agentic Med-VLMs, and dual-decoder models, while remaining competitive with specialist medical segmentors (e.g., MedSAM1–3) on text-guided segmentation.

Our contributions are three-fold:

  • Architectural Contribution. We propose MedUP, a Med-VLM equipped with UniMedTok, a native region tokenizer that encodes masks as discrete tokens within the LLM vocabulary, unifying perception and understanding.

  • Data Contribution. We curate UniMed-Train (1.84M instances) and UniMed-Bench, providing the large-scale region-language corpus and benchmark for bidirectional medical region-language evaluation, including Seg-CoT, a new reasoning-guided segmentation paradigm.

  • Empirical Contribution. MedUP consistently outperforms native, agentic, and dual-decoder Med-VLMs across all tasks, while remaining competitive with specialist segmentors, demonstrating the strong potential of unified understanding-perception modeling.

Methods

Problem Formulation

We study grounded medical vision-language modeling under a unified region-language setting. Given a medical image I, a text instruction or question X, and an optional region mask M, the model is required to support three downstream task types: (1) Medical VQA, where the output is a free-form textual answer; (2) Text-Guided Segmentation, where the output is a segmentation mask corresponding to a language description; and (3) Region-Grounded Understanding, where the model receives a target region and generates a clinically meaningful description, label, or answer conditioned on that region. The goal is to model these tasks in one autoregressive framework rather than by coupling separate segmentation and language systems.

Formally, we introduce a mask serialization operator S(·) that maps a dense region mask into a short token span, and denote the output sequence by Y. MedUP models all tasks with a single conditional autoregressive distribution:

ptheta(Y I, X) = ∏t=1..Y ptheta(yt I, X, y<t).

The three tasks differ only in how the input-output pair is instantiated. For Medical VQA, the target sequence is a textual answer A. For Text-Guided Segmentation, the target sequence is the serialized mask span S(M). For Region-Grounded Understanding, the serialized region span S(M) is appended to the instruction as part of the conditioning context, and the model predicts the textual answer A. This formulation reduces region understanding and region generation to next-token prediction in a shared text-mask space.

Overview

MedUP is built around UniMedTok, a native mask-token interface that places medical regions in the same autoregressive space as text. The system has two stages. In Stage 1, we train a medical mask tokenizer to convert a region mask into compact discrete codes and reconstruct the mask from them. In Stage 2, we expand the VLM vocabulary with mask tokens, convert all region-related supervision into text-mask sequences, and jointly train on the four streams of UniMed-Train: Medical VQA, Text-Guided Segmentation, Region-Grounded Understanding, and Seg-CoT.

This design separates mask representation learning from region-language modeling. The tokenizer is responsible for faithful bidirectional conversion between dense masks and discrete codes, while the VLM only needs to learn how to read and generate these codes in context. As a result, UniMedTok avoids adding a trainable segmentation head inside the language model.

Stage 1: Medical Mask Tokenizer

We implement UniMedTok as an image-conditioned vector-quantized mask autoencoder. Given an image I and region mask M, the tokenizer encodes the mask into a continuous representation and then compresses it with residual vector quantization into an ordered two-code representation:

q = Q(Etok(I, M)) = [c1, c2], c1, c2 ∈ 0,..., 255,

yielding an MT256×2 tokenization scheme. The discrete code pair is then decoded, conditioned on the same image, back into a dense mask M̂ = Dtok(I, q). Because decoding remains image-conditioned, the codes act as compact region prompts rather than standalone pixel descriptions. Stage 1 is trained as a mask reconstruction objective with quantization regularization, and the tokenizer is frozen after convergence for all downstream Stage-2 data construction and inference. In practice, we use a medical SAM2-style backbone, non-shared codebooks, and image-plus-box conditioning for localization stability.

Stage 2: Mask Tokens as Language

Vocabulary expansion. After training the tokenizer, we convert each discrete code into a textual special token. Specifically, we add a start token, an end token, and 512 mask code tokens to to the VLM vocabulary. Since each mask uses two codebook levels of size 256, the first token corresponds to the first codebook and the second token corresponds to the second codebook with an offset of 256. A mask is therefore represented as:

S(M) = [ts, tc1, t256+c2, te].

This textualization turns each mask into a short, language-compatible span that can be inserted into prompts or generated as output.

Mask-as-input. For mask-grounded understanding, we encode the target region with the frozen tokenizer and insert the resulting mask-token span into the user prompt: Xreg = [X; S(M)]. The model then answers questions conditioned on both the image and the explicit region reference. This formulation allows region-level reasoning without modifying the base VLM architecture. In contrast to crop-based or overlay-based prompting, the mask is represented in a symbolic form that can be composed with arbitrary text instructions.

Mask-as-output. For text-guided segmentation, the VLM autoregressively generates a mask-token span in response to a referring instruction, q̂ = [ĉ1, ĉ2] ∼ ptheta(· I, X), where the generated tokens are parsed into a code pair before decoding. Thus, the VLM itself only generates short discrete codes, while pixel-level reconstruction is handled by the frozen tokenizer learned in Stage 1: M̂ = Dtok(I, q̂).

Unified Multi-Task Training

We train the Stage-2 VLM with mixed supervision from the four streams of UniMed-Train. All samples are converted into standard conversational sequences, so optimization remains the usual autoregressive next-token loss:

Lstage2 = −∑(I,X,Y)∈D ∑t=1..Y log ptheta(yt I, X, y<t),

where D = ∪4k=1 Dk merges Medical VQA, Text-Guided Segmentation, Region-Grounded Understanding, and Seg-CoT. This unified objective lets the model answer image-level medical questions, generate mask tokens from text, interpret mask tokens as symbolic region references, and perform reasoning-augmented text-to-mask prediction within a single training pipeline. In Stage 2, the tokenizer is frozen and only the VLM is optimized. No segmentation-specific reconstruction loss is used in this stage; cross-task transfer is induced entirely by next-token prediction over mixed text-mask sequences. For Seg-CoT specifically, the target is written as a concatenated reasoning-and-mask sequence Y = [R; S(M)], where R denotes the intermediate textual rationale.

UniMed-Train and UniMed-Bench

UniMed-Train: Stage-2 Training Corpus

Stage 2 is trained on UniMed-Train, a four-stream medical instruction corpus covering Text-Guided Segmentation, Region-Grounded Understanding, Medical VQA, and Seg-CoT. The current release contains 902,648 text-guided segmentation samples, 902,648 region-grounded understanding samples, 27,738 Medical VQA samples, and 4,000 reasoning-augmented Seg-CoT samples, for a total of 1,837,034 instances. The two mask-centric streams are constructed from 80 + 1 medical segmentation datasets using the frozen Stage-1 tokenizer: one stream trains mask-as-output generation from referring text, while the other trains mask-as-input understanding by inserting the serialized region into the question context. Medical VQA preserves image-level clinical reasoning, and Seg-CoT adds intermediate anatomical, attribute, and localization reasoning before the final mask-token span.

Because not all datasets are equally compatible with a compact two-token mask representation, we apply round-trip filtering before Stage 2: each ground-truth mask is encoded and decoded by the tokenizer, and low-fidelity datasets are downsampled according to reconstruction quality. This filtering is applied consistently to both mask-centric streams and improves training stability.

UniMed-Bench: Unified Evaluation Benchmark

We build UniMed-Bench as a held-out benchmark for unified grounded medical vision-language evaluation. It covers three tasks: Medical VQA, Text-Guided Segmentation, and Region-Grounded Understanding. Medical VQA includes 8,273 test questions. Text-Guided Segmentation contains 219,636 test samples from an 80-dataset benchmark family. Region-Grounded Understanding is built from the same 80 datasets, with 218,244 v2 tokens samples and 219,257 v1 masks samples. In the current benchmark instantiation, most region-understanding questions are concise category- or label-oriented prompts. In the main paper, we report answer accuracy for Medical VQA, Dice/IoU for Text-Guided Segmentation, and exact match for Region-Grounded Understanding, with weighted token recall used as a complementary detailed metric.

Experiments

Experimental Setup

We train MedUP on UniMed-Train and evaluate it on UniMed-Bench. Our main models are MedUP-H, built on HuluMed-4B, and MedUP-Q, built on Qwen3-VL-4B. Both variants share the same Stage-1 MT256x2 tokenizer, round-trip-filtered training data, and four-stream Stage-2 objective. Medical VQA is evaluated on SLAKE, PathVQA, and VQA-RAD with overall accuracy; Text-Guided Segmentation is evaluated on 80 datasets with mean Dice; and Region-Grounded Understanding is evaluated on the same dataset family under both v2 tokens and v1 masks with exact match. Consistent with the framing in the Introduction, we compare MedUP against three main baseline families: native medical VLMs without native mask-token interfaces, represented by HealthGPT-M3 and UniBiomed; dual-decoder or externally grounded baselines, represented by LISA++ and SAM4MLLM; and agentic grounded medical models, represented by MMedAgent.

Main Results on the Unified Benchmark

Table 1 reports the unified comparison across Medical VQA, Text-Guided Segmentation, and Region-Grounded Understanding. Across both backbones, MedUP achieves the strongest overall results among the methods included in this three-task setting, outperforming native medical VLM baselines, agentic grounded models, and dual-decoder or externally grounded baselines on the grounded tasks while remaining strong on Medical VQA. MedUP-H reaches 66.5 accuracy, 67.9 mDice, and 81.8 exact match, while MedUP-Q reaches 63.5, 64.4, and 78.5, showing that the proposed interface transfers across both medical-domain and more general multimodal foundations.

Comparison with Specialist Medical Segmentors

Table 3 compares MedUP with specialist medical segmentors on text-guided segmentation. Although specialist models receive stronger oracle visual prompts, MedUP remains competitive across modalities and substantially outperforms the grounded / VLM baselines under micro-Dice. This highlights the practical value of native text-driven segmentation: MedUP trades some oracle-prompt advantage for a much more flexible language interface while retaining strong dense localization performance.

Protocol Study for Mask-Grounded Understanding

To understand whether discrete mask tokens are an effective interface for region-grounded understanding, we compare MedUP with alternative region-presentation protocols. The v1 masks setting exposes the target region visually, whereas v2 tokens represents the region in a language-compatible token space. This experiment isolates the benefit of the interface itself. Table 4 shows a large protocol gap across both the Qwen-based and Hulu-based backbones, while Table 5 reports the per-modality weighted token recall comparison. Across all settings, v2 tokens consistently outperforms v1 masks, indicating that discrete mask tokens provide a more effective interface for region-grounded understanding. The results further suggest that the region interface is a core factor in connecting localized visual evidence with language reasoning. Under v1 masks, the model must additionally map visual overlays to textual semantics, which becomes fragile for small or anatomically ambiguous regions. In contrast, v2 tokens places both region references and linguistic context within a shared token space, reducing the representation gap between visual grounding and language reasoning and leading to substantially stronger performance.

Effect of Training Data Scale

Figure 4 shows that both grounded tasks improve as the amount of Stage-2 training data increases. For Text-Guided Segmentation, weighted Dice and weighted IoU rise consistently from 10% to 100% data scale, with a total Dice gain of 0.099. Region-Grounded Understanding follows the same trend, and v2 tokens remains stronger than v1 masks at every scale. These results indicate that the proposed interface continues to benefit from additional supervision and remains favorable throughout the tested data regime.

Effect of Round-trip Filtering

Figure 5 isolates the contribution of our round-trip filtering strategy. Without filtering, tokenizer reconstruction errors introduce noisy supervision into Stage-2 mask generation, especially on datasets with small, irregular, or semantically ambiguous regions. After filtering, both backbones improve substantially: MedUP-Q increases from 57.3 to 64.4 mDice, while MedUP-H rises from 58.6 to 67.9. The larger gain on MedUP-H suggests that stronger medical backbones can better exploit cleaner tokenized mask supervision once low-fidelity training cases are removed.

Per-Task and Protocol Analysis

Beyond the unified main table, Tables 2, 3, 4, and 5 provide task-specific views of the benchmark. Table 2 reports the closed-question Medical VQA breakdown on SLAKE, PathVQA, and VQA-RAD. Table 3 compares text-guided segmentation against specialist medical segmentors and grounded VLM baselines under per-modality micro Dice. Tables 4 and 5 analyze region-grounded understanding through protocol comparison and modality-level token recall, respectively. Figure 4 complements these tables with a training-scale analysis for the two grounded tasks.

Qualitative Analysis

We provide qualitative examples for text-guided segmentation, region-grounded understanding, and failure cases. In particular, we visualize how Seg-CoT changes generation behavior, whether the predicted masks are semantically aligned with the reasoning trace, and how the same region is handled under crop, overlay, and mask-token protocols.

Conclusion

We introduced MedUP, a family of grounded medical vision-language models built on UniMedTok, a unified mask-token interface for region-language modeling. By treating masks as discrete language-compatible tokens, MedUP unifies Medical VQA, Text-Guided Segmentation, and Region-Grounded Understanding within a single autoregressive framework. We further organized training as the four-stream UniMed-Train corpus and evaluation as the three-task UniMed-Bench, with Seg-CoT improving text-to-mask generation through intermediate anatomical and localization reasoning. These results suggest that native region-language interfaces are a promising direction for building grounded medical VLMs.

Limitations

While MedUP shows that a native mask-token interface is effective for unified medical understanding and perception, several aspects remain open for further study. Our current region representation is intentionally compact, and future work may explore whether richer tokenizations are helpful for some very small, irregular, or visually subtle structures. Our evaluation also focuses primarily on offline benchmark settings, and it would be valuable to further study behavior in more deployment-oriented scenarios such as interactive refinement, longitudinal workflows, or distribution shift. In addition, although we study two backbones under a shared interface, broader validation across model scales and training regimes would help better characterize the generality of the proposed design.

Improvements for AI systems

Based on this paper, here are the specific improvements I can make to AI systems:

Improvement: Implement a native region tokenizer (like UniMedTok) that converts segmentation masks into discrete tokens within the LLM's existing vocabulary, rather than using coordinate strings or external segmentation modules.

Capability: The AI system can now seamlessly interleave mask tokens with text in a single autoregressive sequence, enabling bidirectional region-language understanding without architectural decoupling.

Improvement: Train the system to both accept mask tokens as input context (for region-grounded understanding) and generate mask tokens as output (for text-guided segmentation) within the same framework.

Improvement: Implement a reasoning-augmented segmentation paradigm where the model produces intermediate anatomical, attribute, and localization reasoning before generating the final mask tokens.

Improvement: Apply a quality-control mechanism where ground-truth masks are encoded and decoded by the tokenizer, and low-fidelity datasets are downsampled based on reconstruction quality before training.

Improvement: Train the system jointly on four streams—Medical VQA, Text-Guided Segmentation, Region-Grounded Understanding, and Seg-CoT—using a single next-token prediction loss over mixed text-mask sequences.

Improvement: Represent target regions as symbolic mask tokens in the prompt rather than visual overlays or crops.

Improvement: Use a residual vector quantization scheme that compresses each mask into just two discrete codes (256 options each), with image-conditioned decoding.

Improvement: Structure training so that performance consistently improves with more data (verified from 10% to 100% of training corpus).

Abstract

Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modules that decouple perception from understanding, creating representation gaps for region-language alignment. We present MedUP, a Med-VLM that natively unifies perception and understanding within a shared token space. At its core lies UniMedTok, a region tokenizer that encodes masks as discrete tokens in the LLM vocabulary, enabling the model to seamlessly interleave mask tokens with text. We curate UniMed-Train, a 1.84M-instance corpus spanning text-guided segmentation, region-grounded understanding, medical VQA and CoT-based segmentation, and introduce UniMed-Bench for unified evaluation. Extensive experiments show that MedUP outperforms native, agentic, and dual-decoder Med-VLMs across all tasks while remaining competitive with specialist segmentors, demonstrating the strong potential of unified understanding and perception modeling.

Sources

Related papers