Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining

arXiv:2606.01049 · cs.CL · Submitted 2026-05-31 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining".

Tom: Biomedical figures are explained not by captions alone but by body-text passages that discuss them, and this paper introduces context-grounded reconstruction,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So we’ve been diving deep into how this paper builds a much smarter way to train medical models by focusing on linking figures and text together properly rather than just treating them as separate pieces of data.

Jane: Exactly, and I think the title itself, "Beyond Captions," really captures the core idea that we need to move past just having simple descriptions of images.

Lu: It’s fascinating because they aren't just adding more text; they are reconstructing the entire document structure to ensure that every piece of context stays attached exactly where it belongs relative to the visual information.

Meng: From a practical standpoint, I’m thinking about how much cleaner the data pipeline becomes when you have these structured sequences instead of messy, disconnected records.

Lalam: And what really stands out is their evidence-aware allocation strategy; they’re not just throwing everything at the model randomly; they’re carefully balancing visual proof against quantitative tables and structural explanations.

Tom: That careful balancing act is key, I think, because it means the resulting training signal for these new medical models will be much more robust than what we get from standard methods.

Jane: It suggests that when we train AI on this kind of context-grounded data, the models learn to reason about medical concepts in a way that’s much closer to how a human specialist thinks.

Lu: Imagine the possibilities if these models can accurately synthesize information across different evidence types, which is what they’re aiming for with that four-bucket taxonomy.

Meng: That level of contextual understanding could make AI tools in diagnostics much more reliable because they wouldn't just see a picture; they would understand the underlying mechanisms too.

Lalam: It really points toward a future where these AI systems can support complex clinical reasoning by integrating visual and textual data seamlessly, which is a huge cultural shift for how we use medical technology.

Conclusion: Tom: So, we've seen how this paper reconstructs medical data to link figures and text coherently rather than treating them as separate items for training, and now we need to talk about what that title really means for us.

Jane: I think "Beyond Captions" is a really simple way of saying they’re aiming past just reading a description next to an image; they want the AI to see the whole story.

Lu: They're essentially rebuilding the document structure so that every piece of information, especially those crucial context paragraphs, stays exactly where it belongs when paired with its corresponding figure.

Meng: From an engineering view, it’s about moving away from just feeding raw data into a model and instead giving it a highly organized sequence that tells it *how* the visual and textual parts relate to each other.

Lalam: It's about creating training material that mirrors how medical experts actually read and understand complex scientific papers, which is a massive step for improving the reliability of AI in medicine.

Tom: And when we look at who wrote this, their focus on evidence-aware allocation really shows they’re not just adding fluff; they’re strategically deciding which types of data—like visual proof versus tables—are most important for boosting model performance.

Jane: It suggests that the future of training these multimodal models isn't just about bigger datasets, but about making those datasets intelligently structured from the start.

Lu: This work opens up completely new ways to design entire training ecosystems for complex scientific areas, moving far beyond simple data aggregation and toward genuine contextual understanding.

Meng: I’m thinking this has a direct impact on how we build specialized AI applications because those applications will need to handle multi-modal inputs with high fidelity and contextual awareness, which is crucial for diagnostics.

Lalam: Ultimately, this research contributes to making AI more reliable in healthcare by ensuring the training data reflects the actual way medical knowledge is presented, rather than just isolated snippets.

Tom: It really highlights that we need to focus on the relationship structure between the visual and textual parts when we want these models to perform well in real-world clinical scenarios.

Jane: And that structural understanding is what makes this approach so effective for continuing to train these large multimodal models properly.

Hong Kong Polytechnic University

cs.CL

Submitted: 2026-05-31

Updated: 2026-10-07

Importance score: 91/100

The gist: Biomedical figures are explained not by captions alone but by body-text passages that discuss them, and this paper introduces context-grounded reconstruction, a source-grounded framework that

Key concepts

Source-Grounding
This initial step restores original captions and normalizes text while keeping track of where the information came from. It ensures that the reconstructed sequences maintain their proper provenance, which is crucial for accurately linking text to visual elements in medical documents.
Reference-Constrained Interleaving
This technique builds multi-figure sequences by carefully interleaving text and images. The key rule is that a paragraph of context can only be attached to the specific figures it explicitly mentions, preventing redundant or misplaced information.
Evidence-Aware Allocation
This strategy controls how much exposure a model gets to different types of scientific evidence (like visual data vs. tables). By choosing configurations like 'BVE-enhanced,' researchers can tailor the training signal to prioritize the most medically relevant evidence for better performance.

Terminology

Summary

Biomedical figures are explained not by captions alone but by body-text passages that discuss them, and this paper introduces context-grounded reconstruction, a source-grounded framework that converts PubMed Central Open Access (PMC-OA) records into referentially coherent interleaved sequences to create a richer training signal for generative medical MLLMs.

How it works

The framework begins with source-grounded recovery to restore provenance-linked captions and normalize source text while preserving reference anchors. This is followed by reference-constrained interleaving, which forms multi-figure sequences without duplicating shared context, ensuring that a context paragraph is attached only to the figures it explicitly references. To handle discontinuities, the framework employs a coherence-constrained curation procedure that detects discontinuities among retained source passages, preserves locally supported context, and prunes image slots that no longer have referential support.

How it works (Continued)

The process then moves to semantic curation, which filters the candidate pool for textual usability and medical relevance at full-corpus scale. Finally, evidence-aware allocation shapes the CPT learning signal by regulating the model’s exposure to complementary forms of scientific evidence. This involves a four-bucket taxonomy: Biomedical Visual Evidence (BVE), Quantitative/Table Evidence (QTE), Mechanism/Structure Evidence (MSE), and Auxiliary (AUX). The final corpus, PMC-InterCPT, is formed through a mixture configuration, such as the BVE-enhanced allocation which uses 45% BVE, 30% QTE, 20% MSE, and 5% AUX, selected for its strongest medical average.

How it works (Continued)

The quality and relevance filtering stage uses a two-stage approach. Stage 1 involves using an LLM annotation step with Gemini-3.1-pro-preview to assign labels for medical relevance (binary) and text quality (three levels: 0, 1, or 2). The quality score penalizes heavy repetition, copy-paste artifacts, severe garbling, major incoherence, and specifically penalizes malformed tables or LATEX dumps. Stage 2 applies these predicted labels to every reconstructed record; only records with a quality score of q ≥ 1 are retained for the source-pool construction.

How it works (Continued)

The evidence-aware allocation strategy controls the relative exposure to complementary forms of scientific evidence. The allocation can be set to Natural (preserving the source-pool distribution), Balanced (approximately equalizing buckets), BVE-enhanced (prioritizing biomedical visual evidence while retaining other types), or BVE-only. This allows researchers to test different mixture configurations, such as the BVE-enhanced allocation, which is presented as a medical-oriented tradeoff.

How it works (Continued)

The evaluation assesses the resulting corpus on five medical benchmarks and four general/scientific benchmarks. The main results show that PMC-InterCPT improves Qwen3.5-4B-Base by 1.46 medical average points and 3.11 general/scientific average points over a token-matched raw source control, with gains transferring to other model backbones like Qwen3.5-2B-Base and LLaVA-OneVision-1.5-4B-Base, demonstrating that context-grounded reconstruction is central to useful biomedical multimodal CPT.

How it works (Continued)

Controlled ablations confirm the necessity of the framework: article context becomes useful only after referent-aware reconstruction and curation. Breaking figure-context links decreases the medical average by 1.75 points and the general/scientific average by 3.40 points, proving that article context is only useful when it remains aligned to the figure it is meant to explain. This effect is isolated from text quantity or visual content alone.

How it works (Continued)

The final corpus composition shows that BVE receives the largest share of tokens at 45.0%, while QTE and MSE together account for half of the corpus, illustrating a medical-visual-evidence-enhanced profile. The study concludes that source-grounded curation contributes beyond simply increasing the volume of in-domain raw data.

How it works (Continued)

The evaluation protocol uses zero-shot prompting unless otherwise stated, employing deterministic decoding with temperature 0 for inference on benchmarks like MMMU and SciVQA, and LLM judge scoring for ChartQA and CharXiv-Val. The results show that the BVE-enhanced allocation provides a medical average improvement of +0.92 points over the Natural distribution in the 1B token setting, suggesting it is a "medical-oriented tradeoff.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems by implementing the methodology described in Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining (PMC-InterCPT):


The following improvements focus on moving beyond simple image-caption pairs toward building sophisticated, context-aware biomedical multimodal Large Language Models (MLLMs).

  1. Improvements to MLLM Pretraining Data Construction:

  2. Improvements to Model Robustness and Generalization:

  3. Improvements in Medical Reasoning Capabilities:

  4. Improvements in Scientific Figure Understanding:

The resulting improved AI system, powered by PMC-InterCPT data, will possess the following specific capabilities:

  1. A significantly more accurate and contextually rich understanding of biomedical figures (e.g., CT scans, histopathology slides, chemical structures).

  2. The ability to perform complex medical visual question answering (VQA) tasks with high precision on clinical and pathological images.

  3. Superior reasoning capabilities across diverse scientific domains, including complex chart interpretation and textual analysis within scientific articles.

Detailed Specific Improvements:

  1. A more robust and scalable pipeline for constructing high-quality biomedical CPT data by replacing simple image-caption pairs with context-grounded sequences derived from PubMed Central Open Access (PMC-OA) records.

  2. The ability to generate interleaved, referentially coherent sequences where surrounding body text is correctly attached to the figures it discusses, rather than being discarded or appended arbitrarily.

  3. A mechanism to automatically prune unsupported image-text attachments by verifying that every retained image slot has a traceable, article-native figure reference justifying its context.

  4. A method for repairing incoherent context by detecting and resolving discontinuities among retained source passages, ensuring that the text segments associated with a figure faithfully represent the evidence linked to it.

  5. A rigorous quality control system that filters out low-quality or corrupted textual data (e.g., heavy repetition, malformed table dumps, LaTeX artifacts) using LLM-supervised classifiers, ensuring the final training corpus is composed of clean and coherent text for pretraining.

  6. A modality-aware allocation strategy that dynamically weights the exposure to different types of scientific evidence—Biomedical Visual Evidence (BVE), Quantitative/Table Evidence (QTE), Mechanism/Structure Evidence (MSE), and Auxiliary context—allowing the model to learn from a balanced and medically prioritized distribution of information.

  7. Enhanced medical reasoning by training models on a corpus that prioritizes clinical visual data, leading to improved performance on specialized benchmarks like MMMU-Med-Test, OmniMedVQA, and PretexEval.

  8. Dramatically improved scientific figure understanding by enabling the model to correctly interpret complex relationships across multiple figures within a single source article (cross-figure grounding), which is crucial for interpreting experimental setups and comparative results.

  9. Increased generalization capability across different model backbones (e.g., Qwen3.5-4B-Base vs. LLaVA-OneVision-1.5-4B-Base), demonstrating that the corpus construction benefit is transferable to diverse MLLM architectures, not just one specific model family.

Sources

Related papers