MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports

arXiv:2605.03103 · cs.CL, cs.AI, cs.LG · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports".

Jane: The paper was written by Yingyun Li, Yu Wang and Haiyang Qian from AI Starfish, China.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: We've established what the tasks are, but to truly understand MedStruct-S, we need to look at the sheer scale and the data collection process outlined in this benchmark.

Jane: The authors created a dataset of three thousand five hundred eighty-two real-world clinical report pages. This isn's not clean EHR text; it's actual medical history that has been run through OCR, which means it carries all those imperfections and noise.

Lu: The data collection effort was significant—a five hundred sixty-person-day annotation process—to ensure that the ground truth for these complex tasks is solid, even though the input text is flawed.

Meng: I appreciate the focus on OCR noise; it’s a necessary reality. It forces us to test models in real operational environments, not just in sanitized lab conditions where everything works perfectly.

Lalam: The creation of MedStruct-S (De-ID) is also a major step, ensuring that while we preserve the structural integrity and the noise patterns of the original data, we protect patient privacy by replacing sensitive personal identifiers.

Tom: That’s a critical ethical consideration. It allows us to rigorously test models without violating compliance standards, which is a huge win for real-world testing.

Jane: The paper uses two main evaluation approaches: exact match (EM) and approximate match (AM). This lets us distinguish between needing the AI to be perfectly literal and just being semantically correct.

Lu: We are essentially quantifying the gap between demanding absolute fidelity and allowing for a semantic interpretation, which is very relevant when dealing with human language.

Meng: For me, it means that when we deploy this system, we need to know if the user requires perfect matching or if a ninety percent similar match is acceptable to deliver speed and coverage.

Lalam: It allows us to build systems that are flexible enough to handle ambiguity while still maintaining high standards of accuracy for the sake of patient care.

Improvements and Analysis: Tom: Now, we’re moving past the data itself and looking at how the models perform across these three tasks using MedStruct-S. The key finding here is a clear trade-off between model types.

Jane: We see that encoder-only models—like BERT or RoBERTa variants—are fantastic at accurate localization, which is great for Task one (Key Discovery) and Task three (Extraction).

Lu: But the results show that even when comparing models of similar size, the fine-tuned decoder-only models tend to deliver stronger overall results in specific areas. This suggests a complex synergy between generation and structure.

Meng: The performance difference is especially visible in non-null value key-conditioned QA. The encoder-only approach gives better results there, even though these models are often much smaller than the massive decoder ones we use for general text generation.

Lalam: This implies that when we need highly specific, structured knowledge extraction, rather than just fluid conversation, the targeted approach of the smaller encoder models is very effective at fulfilling a cultural need for precise information retrieval.

Tom: The authors highlight that decoder-only models are more prone to "boundary drift" and defaulting to NULL when evidence is difficult to locate. That’s a critical failure mode we need to be aware of.

Jane: And this links back to the AM vs EM discussion; if the model drifts, it’s likely failing an exact match, but maybe it’s close enough for the approximate match criterion.

Lu: The ability to identify key aliases—where multiple names for one concept exist—is a huge challenge that these models are learning how to navigate using this benchmark.

Meng: We have to consider if we should design our production system around the high precision of the encoder models or if we need the generative power of the decoder models, depending on which task is most critical for a given use case.

Lalam: The goal is not just picking one model, but using this benchmark to guide us toward finding that optimal balance—a system that respects fidelity while handling real-world messiness.

Conclusion: Tom: We've covered the mechanics and the performance metrics; now we need to wrap up our discussion by looking at the final conclusions of this study.

Jane: The researchers conclude that MedStruct-S provides a reliable, practical basis for selecting and comparing models across these different semi-structured extraction scenarios. It’s a crucial tool for model selection in AI.

Lu: They also point out that the trade-off between literal fidelity (EM) and semantic tolerance (AM) is very clear, especially as we move toward complex end-to-end extraction.

Meng: I'm glad they explicitly state that Task three is not perfectly symmetric because encoder-only models rely on a deterministic pairing heuristic, which is a design choice that the decoder models avoid by generating the pair directly.

Lalam: It’s clear this benchmark helps us move away from assuming clean data and allows us to address the real challenges of how human information—even if poorly recorded—can be translated into actionable knowledge.

Tom: We've seen how the performance is consistent whether we test against MedStruct-S or its de-identified version, which suggests these findings are robust across different scenarios.

Jane: It also highlights that while scale matters for decoder models, model family and post-training techniques like LoRA are equally important for achieving peak performance in Task three.

Lu: We can't deny the massive potential here; this is a foundation upon which we can build future systems that understand medical history on a whole person level.

Meng: It allows us to start asking the practical questions: How do we scale this kind of extraction across various hospital systems? This benchmark gives us the initial answers.

Lalam: We're looking at how AI helps bridge gaps in human knowledge, allowing us to build tools that respect both accuracy and the complex reality of messy medical documentation.

Tom: That is a fantastic way to conclude our discussion on "MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports."

Conclusion: Tom: So, as we wrap up this segment, it's clear that MedStruct-S has done a monumental job by creating a benchmark that reflects the real chaos of clinical data, which is honestly huge for testing AI systems.

Jane: It really demonstrates that simply having digital records isn't enough; we have to test how well the AI handles the actual human messiness of OCR errors and variable terminology, which is a massive win.

Lu: I think it opens up so many possibilities for future AI research because we’ aren't just training models on clean data anymore; we're teaching them to navigate ambiguity in a way that feels like genuine problem-solving.

Meng: From an engineering standpoint, this means the systems we build can actually be robust enough to handle real operational inputs without completely failing when the input isn't perfectly structured.

Lalam: And it ultimately allows us to make medical information more accessible and understandable for patients, which is a huge step toward improving how society handles health knowledge.

Tom: It’s a testament to the fact that rigorous testing is essential before deploying these advanced systems, right?

Jane: Exactly; we' need that level of accuracy when dealing with critical patient data.

Lu: I'm excited to see what new architectural breakthroughs come out of this challenge.

Meng: I just hope the practical integration into existing hospital workflows goes as smoothly as the data suggests is possible.

Lalam: It’s a tool that helps bridge the gap between all of us and with our health records, which is really impactful.

Tom: We'll carry these insights with us as we look at next paper on arXiv, so keep those eyes on your screens!

Yingyun Li, Yu Wang, Haiyang Qian

AI Starfish, China

cs.CL, cs.AI, cs.LG

Submitted: 2026-08-19

Updated: 2026-08-20

Comments: 11 pages, 5 figures. Accepted by KSEM 2026. This is the author's preprint version; the final authenticated version will be available in the Springer LNCS/LNAI proceedings

Code: https://github.com/AI-Starfish-Research/MedStruct-S

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 98/100

The gist: MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports The paper introduces MedStruct-S, a benchmark designed to evaluate

Key concepts

MedStruct-S
MedStruct-S is a benchmark created by the authors of the paper. It provides a reliable and practical basis for selecting and comparing AI models. It uses real-world clinical report pages to test how well systems perform complex tasks like Key Discovery and semi-structured extraction.
OCR Noise
This refers to the imperfections and errors present in text that has been run through Optical Character Recognition (OCR). The benchmark intentionally uses this flawed, real-world data, forcing AI models to operate in operational environments rather than sanitized lab conditions.
EM vs. AM
These are two evaluation approaches used by the researchers. Exact Match (EM) requires the AI output to be perfectly literal. Approximate Match (AM) allows for semantic correctness, allowing the system to be flexible enough to handle ambiguity while still maintaining high accuracy.
Encoder-only vs. Decoder-only Models
These are two types of AI models discussed in performance testing. Encoder-only models are noted for accurate localization and high precision in specific tasks. Decoder-only models offer strong overall results but can be prone to 'boundary drift' or defaulting to NULL when evidence is difficult to locate.

Terminology

Summary

MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports

The paper introduces MedStruct-S, a benchmark designed to evaluate semi-structured information extraction (IE) from OCR-derived clinical reports. This process is crucial to efficiently reconstruct the patients’ longitudinal medical history. In real-world practice, this scenario involves three distinct tasks: (i) field-header (key) discovery, (ii) key-conditioned question answering (QA), and (iii) end-to-end key–value pair extraction.

The authors note that existing evaluations often under-model two factors—heterogeneous and not-completely-known key representations and OCR-induced noise—making it difficult to assess model robustness in real-world settings.

Data Construction

MedStruct-S is built from 3,582 annotated real-world clinical report pages. The data collection process involved running OCR on these reports. To ensure privacy, a de-identified version, MedStruct-S (De-ID), is also released. The fidelity of this de-identification is maintained by ensuring that the consistently high page-level similarity preserves the original OCR-induced noise patterns and structural layout of the reports.

** Task Definitions and Evaluation Metrics**

The benchmark defines three tasks:

  1. Task 1 (Key Discovery): The input is an OCR-derived report page p, and the output is a predicted key set, identified without assuming a predefined key inventory.

  2. Task 2 (Key-Conditioned QA):): Takes (p, k) as input, where k is a queried key, and outputs the corresponding value.

  3. Task 3 (Semi-Structured Extraction): The input is d = (,), enables p, and the output is a predicted set of key–value pairs KV end-to-end semi-structured extraction in a single step.

To quantify robustness to OCR noise, two metrics are used: Exact Match (EM) and Approximate Match (AM). The AM metric is defined using the Levenshtein edit distance d lev(u, v), where phi(u, v) = 1 - d lev(u, v) over(u, v).

** Experimental Setup**

The study evaluates two representative paradigms: encoder-only sequence labeling and decoder-only structured generation. The models tested include:

  • Four encoder-only models (M-BERT, RoBERTa, MacBERT, and McBERT), using a BERT–BiLSTM–CRF sequence labeling model.

  • Five decoder-only models (Qwen3-0.6B, Qwen3-14B, Qwen2.5-32B, Baichuan-M2-32B, and AntAngelMed-103B), using two-shot inference or LoRA fine-tuning.

** Results and Analysis**

The results demonstrate distinct performance profiles across the tasks:

  • Task 1 (Key Discovery): Two-shot decoder-only models perform poorly on key discovery, especially at small scale. However, applying LoRA substantially improves this behavior (e.g., Qwen3-0.6B increases Ke from 0.0684 to 0.7956).The best decoder-only Task 1 results are achieved by Baichuan-M2-32B, achieving a Ke/Ka of 0.8624/0.8640 on MedStruct-S.

  • Task 2 (Key-Conditioned QA): Decoder-only models achieve high accuracy on the overall evaluation set, but their performance drops more on non-null samples. The authors observe that encoder-only models achieve best performance for non-null-value key-conditioned QA despite being substantially smaller than decoder-only models.

  • Task 3 (Semi-Structured Extraction): For fine-tuned decoder-only models, AM yields larger gains than EM. On MedStruct-S, the best result is achieved by Baichuan-M2-32B (LoRA) with a Ka Va of 0.7884.

Conclusion and Limitations

The findings illustrate that the benchmark provides a reliable and practical basis for selecting and comparing models across different semi-structured IE settings. The study concludes that end-to-end extraction remains particularly sensitive to format compliance and boundary accuracy, especially for smaller decoder-only models.

Limitations noted include: MedStruct-S is currently Chinese only, and Task 3 is not fully symmetric because encoder-only models use a deterministic nearest-neighbor pairing heuristic. The benchmark also does not yet cover the full diversity of clinical reports across institutions and layouts.

Improvements for AI systems

Based on this comprehensive set of citations—which cover advanced topics from foundational transformer architectures (14, 22) to specialized medical IE benchmarks (10, 13, 26) and modern LLM scaling/benchmarking (16, 25)—the current state-of-the-art research is highly fragmented. Specialized systems excel in narrow domains (e.g., Chinese biomedical extraction using CBLUE; PDF parsing using Omnidocbench), but they lack the seamless integration required for real-world clinical decision support, which involves heterogeneous data types (scanned forms, radiology reports, unstructured text).

The critical gap is the lack of a single, end-to-end framework that can ingest multi-modal, noisy, and structurally ambiguous documents and output a verified, semantically rich knowledge graph, while simultaneously adapting its core reasoning engine to new clinical domains without full retraining.

Here are the specific improvements I propose for an AI system:


I propose developing a novel, modular architecture called the Unified Clinical Knowledge Extraction Engine (UCKEE). This system moves beyond traditional sequence-to-sequence information extraction and instead operates as a multi-stage, self-correcting reasoning pipeline that models the entire clinical workflow from document ingestion to structured knowledge output.

1. Multi-Modal Document Ingestion & Structured Layout Understanding (Addressing [12], [16], [20]):

  • Improvement: Integrate a dedicated, pre-processing Layout Reasoning Module (LRM). This module must go beyond simple OCR to understand the functional relationship between text blocks. It will leverage advanced techniques inspired by Omnidocbench and Funsd to map document semantics (e.g., recognizing that Patient Name is always located in the top-right corner, regardless of page layout variations).

  • Mechanism: The LRM will output a preliminary structured JSON representation before the LLM sees the text, providing explicit coordinates and semantic roles for every detected entity block.

2. Context-Aware Knowledge Graph Generation (Addressing [13], [26], [28]):

  • Improvement: Replace simple relation classification with a Graph-Guided Constrained Decoding Mechanism. Instead of asking the LLM to predict a relation label (e.g., [Drug] administers to [Patient]), the system will first query a continuously updated, external knowledge graph (e.g., UMLS/SNOMED CT). The LLM’s generation process will then be constrained by the valid paths and permissible triples within this graph, significantly reducing hallucination and ensuring clinical validity.

  • Mechanism: This utilizes a Retrieval-Augmented Generation (RAG) loop where the query context is enriched by retrieving relevant graph substructures, which guides the final extraction output.

3. Domain-Adaptive Prompt Engineering & Verification Loop (Addressing [17], [28], [25]):

  • Improvement: Implement a Hierarchical Adaptive Prompting System (HAPS). This system will dynamically adjust the prompting strategy based on the detected domain and complexity level. If the input is highly technical (e.g., radiology report), it activates prompts derived from specialized benchmarks like Radgraph; if it’s cross-lingual, it activates multilingual alignment strategies inspired by Tanwar et al.

  • Mechanism: Crucially, we introduce a Self-Verification Agent. After initial extraction, the LLM is prompted to critique its own output against established clinical axioms (e.g., A drug dosage must be associated with a unit and an interval). This acts as a final, high-level filter that flags inconsistencies for human review or automated correction.

4. Unified Instruction Tuning Pipeline (Addressing [15], [21]):

  • Improvement: Develop a meta-training objective that treats all information extraction tasks (entity recognition, relation extraction, slot filling) not as separate heads, but as stages in a single Unified Structure Generation Task. This requires training the model on datasets that explicitly require the generation of nested and overlapping structures.

The UCKEE system will provide capabilities far exceeding current specialized models:

  1. End-to-End Clinical Data Synthesis: It can ingest a complex patient record comprising a scanned intake form (noisy, unstructured), an electronic lab result PDF (structured, semi-structured), and a physician's free-text consultation note (unstructured). It will automatically synthesize these three disparate sources into one coherent, verified Clinical Knowledge Graph.

  2. Automated Hypothesis Generation: By linking extracted entities across documents (e.g., linking a symptom mentioned in the note to an elevated biomarker found in the lab report), the system can proactively suggest potential differential diagnoses or necessary follow-up tests, thereby assisting clinical decision support.

  3. Zero-Shot Domain Adaptation: If presented with a document type from a novel specialty (e.g., pediatric oncology records) for which it was not explicitly trained, the system uses its foundational architectural knowledge and the external graph constraints to rapidly identify core entities and relationships with minimal fine-tuning data, achieving significantly higher robustness than current models.

  4. Explainable Extraction: For every piece of extracted knowledge (a triple in the graph), the UCKEE can provide a Source Attribution Trace, citing the exact page number, document type, and specific text span that justified that extraction, which is paramount for high-stakes medical applications.

Sources

Related papers