The COTe score: A decomposable framework for evaluating Document Layout Analysis models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "The COTe score".
Jane: Document Layout Analysis (DLA) models typically rely on general object detection metrics like IoU, F1, or mAP,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we just talked about how this paper introduces the COTe score, but let’s quickly recap what they are actually proposing as the thesis. The main argument is that traditional metrics like IoU or F1 are ill-suited for printed media because they treat the input as a simple 2D projection of three dee space rather than a natively 2D tessellation.
Jane: That’s right, Tom; they argue that this discrepancy leads to misleading interpretations of model quality when using those general object detection metrics on documents. They propose two main solutions: first, the Structural Semantic Unit, or SSU, which is a relational labelling approach focusing on semantic structure over physical position.
Lu: The SSU groups labelled regions based on sharing a common layout and semantic unit, defining what it means for two regions to belong together based on five specific conditions.
Meng: I see the SSU as this necessary step to solve the incommensurability problem they identified, which arises because different labelling schemas often overlap but aren't fully interoperable with each other.
Lalam: And then on top of that, they introduce the COTe score, which is a decomposable metric that measures parsing quality through four components: Coverage, Overlap, Trespass, and Excess.
Tom: That’s right; the COTe score is designed to be natively 2D and breaks down performance into these four distinct parts to give a more robust evaluation of DLA model quality.
Jane: So in simple terms, the paper claims that by using SSU and COTe, we can get a much more comprehensive and informative evaluation of how well an AI model actually parses a page.
Lu: The core contribution is this shift from physical position to semantic structure as the primary evaluation lens for DLA models.
Meng: From my perspective, this framework offers a practical path forward because it acknowledges the messiness of real-world labelling and provides a metric that isn't overly sensitive to those arbitrary choices.
Lalam: It’s about building tools that reflect how we actually work with documents, not just theoretical projections from other domains.
Tom: So, before we move into the conclusion, let’s see what this all means when you take the title and authors of "The COTe score: a decomposable framework for evaluating Document Layout Analysis models" into account.
Conclusion: Jane: Looking at the title and authors of "The COTe score: a decomposable framework for evaluating Document Layout Analysis models," it really emphasizes that this work is about creating a new way to evaluate document layout analysis.
Lu: It signals that the authors are not just tweaking old metrics but are introducing an entirely new conceptual structure—the SSU and COTe score—to tackle the fundamental evaluation challenge.
Meng: The implication for us is that we might need to adjust how we benchmark our DLA models because relying solely on general object detection metrics isn't sufficient for this domain.
Lalam: This work suggests that future advancements in document AI should prioritize developing evaluation frameworks that respect the unique 2D nature of printed media directly.
Tom: Exactly, and the authors show that when you evaluate five common DLA models across three different datasets, the COTe score reveals distinct failure modes depending on the model architecture used.
Jane: That means we can start using this framework not just to get a score, but to actively debug and understand the weaknesses in our current AI systems.
Lu: The future work hinted at involves extending this approach by looking at how these structural units interact with larger document-level structures, which could unlock even deeper levels of analysis.
Meng: From an engineering standpoint, I see this as a way to guide future development toward building models that are inherently more aware of the semantic organization they are trying to parse.
Lalam: This is really encouraging because it suggests we can develop AI that doesn't just find pixels but actually understands the layout and meaning of the content it sees.
University College London
cs.CV
Submitted: 2026-03-13
Updated: 2026-10-01
Comments: 10000 words, 5 Figures, 19 Tables,
Code: https://github.com/JonnoB/cotescore
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 84/100
The gist: Document Layout Analysis (DLA) models typically rely on general object detection metrics like IoU, F1, or mAP, which are ill-suited for printed media because they treat images as 2D projections of 3D
Key concepts
- Structural Semantic Unit (SSU)
- The SSU is a way to group text regions based on their common layout and meaning. It moves beyond simple physical location by defining units that share the same class, structural unit, and semantic unit. This helps capture how text is organized structurally, regardless of the exact labeling granularity used during training.
- Coverage (C)
- This measures how much of the ground truth is successfully covered by a model's prediction. It is calculated as the normalized area where a prediction overlaps with an SSU mask. High coverage indicates that the model has successfully identified most of the target text regions.
- Overlap (O)
- Overlap measures instances where a model predicts repeated phrases or sentences, often indicating 'stacked predictions.' This metric penalizes redundant or overly dense predictions, helping to identify when a model is generating unnecessary redundancy in its output.
- Trespass (T)
- Trespass quantifies errors where a prediction covers an SSU that does not belong to its assigned ground truth SSU. It specifically measures breaches of tessellation logic, identifying instances where predictions incorrectly cross semantic boundaries.
Terminology
Summary
Document Layout Analysis (DLA) models typically rely on general object detection metrics like IoU, F1, or mAP, which are ill-suited for printed media because they treat images as 2D projections of 3D space rather than natively 2D tessellations. This discrepancy leads to misleading performance interpretations; therefore, the paper introduces the Structural Semantic Unit (SSU) and the Coverage, Overlap, Trespass, and Excess (COTe) score as a decomposable framework to provide a more robust and informative evaluation of DLA model quality.
The Core Problem: Limitations of Traditional Metrics
Traditional metrics are often inadequate because they fail to capture the specific challenges inherent in parsing printed media. For instance, IoU is sensitive to the arbitrary choice of granularity at which text is labelled (e.g., line or paragraph), resulting in misleading interpretations of model quality.
Furthermore, these metrics do not account for critical failure modes such as whether predictions overlap or trespass across semantic boundaries. This lack of comprehensive evaluation frameworks leaves practitioners with generalist Object Detection methods that cannot provide meaningful insight into model performance, leading to a gap between prediction and actual task-relevant outcomes.
The Structural Semantic Unit (SSU)
The SSU is introduced as a relational labelling approach
designed to address the practical manifestation of incommensurability
arising from differing text-labelling granularities. The SSU shifts the focus from the physical position of text to its meaning by grouping labels into semantic and layout units. A single SSU is defined as any number of labelled text regions which share a common layout and semantic unit.
For two regions to belong to the same SSU, they must satisfy five conditions: (1) same class, (2) same structural unit, (3) same semantic unit, (4) adjacent in reading order, and (5) adjacency in reading order. This approach provides robustness to the level at which the text itself was labelled.
The COTe Score: A Decomposable Metric
The COTe score is a decomposable metric for measuring page parsing quality
that is natively 2D.
It measures performance across four distinct elements:
-
Coverage (C): Measures how much of the ground truth is covered by predictions, defined as the normalized area of intersection between prediction and SSU masks.
-
Overlap (O): Concerns repeated phrases and sentences, penalizing
stacked prediction
areas. -
Trespass (T): Penalizes breaches of tessellation logic by measuring how much a prediction covers an SSU that does not belong to its assigned ground truth SSU.
-
Excess (E): Acts as a support metric, measuring the area outside the bounds of any SSU, contextualizing core metrics and showing
how well the predictions fit the SSUs.
The overall COTe score is calculated additively: COTe score = C − O − T.
This additive structure means that 1 is a perfect score, and negative scores indicate that the sum of Trespass and Overlap exceeds coverage.
Granularity Robustness and Pragmatic Competence
A key finding is the granularity robustness
of the COTe score, which largely holds even without explicit SSU labelling.
This suggests that practitioners can gain much of the benefit without retraining models on custom SSU labels. The paper demonstrates this by showing that when ground truth is at line level but predictions are at paragraph level (or vice versa), traditional metrics yield poor results, whereas the COTe score remains highly informative and reliable. This robustness stems from the metric's pragmatic competence,
which allows it to track task-relevant outcomes across a wider range of granularities and schemas without requiring model retraining.
Model Evaluation and Diagnostics
The authors evaluate five common DLA models (DocLayout-YOLO, Heron, PPDoclayout) on three datasets (NCSEV2, HNLA2013, DocLayNet). The results show that no single model is universally best; performance varies by dataset. For example, DocLayout-YOLO showed high coverage but also extensive Overlap and Trespass in one case. Crucially, the COTe score reveals distinct failure modes across models—such as breaching semantic boundaries or repeatedly parsing the same region
—allowing practitioners to debug models through visual inspection and identify specific weaknesses related to training data or architecture. The paper concludes that the COTe framework provides a more insightful evaluation than IoU or F1, as it penalizes errors in spatial and semantic structure.
Practical Release
To support community adoption, the authors release an SSU labelled dataset
and a Python library called cotescore for applying the COTe framework in DLA projects. This allows researchers to use the Coverage, Overlap, Trespass, and Excess to identify system weaknesses leading to more robust and accurate models.
Improvements for AI systems
Here are the specific improvements that an AI system can achieve by implementing the concepts from this paper:
-
Improve model evaluation robustness against annotation granularity mismatch: The system will no longer rely solely on IoU/F1/mAP, which are sensitive to differing label levels (e.g., line vs. paragraph). Instead, it will use the deomposable COTe score to quantify performance across different structural and semantic units (SSUs), providing a more reliable measure of true parsing quality, especially when ground truth is loosely labeled or predictions are at a finer granularity than expected.
-
Diagnose specific model failure modes: The COTe score's decomposed metrics (Coverage, Overlap, Trespass) allow the system to pinpoint exactly why a model fails. It can distinguish between:
-
Identify and quantify semantic boundary breaches (high Trespass).
-
Detect redundant phrase parsing (high Overlap).
-
Determine poor content identification in sparse areas (low Coverage/Excess).
-
Improve post-processing strategies: By identifying high Trespass or Overlap, the system can trigger targeted mitigation steps, such as re-evaluating adjacent regions or applying context-aware filtering before final text extraction.
-
Enhance model selection and training feedback: The paper shows how different model architectures (e.g., DocLayout-YOLO vs. Heron) exhibit distinct failure modes across datasets (NCSE, HNLA2013). This diagnostic capability allows researchers to select models based on the specific error profile needed for their application or to guide targeted fine-tuning efforts toward correcting identified weaknesses (e.g., if a model consistently shows high Overlap, it suggests issues with phrase segmentation).
-
Create a more
pragmatically competent
evaluation framework: The COTe score shifts the focus from the physical projection to the semantic tessellation of printed media, making the metric inherently more relevant to human reading and document understanding tasks, thus reducing the performance interpretation gap relative to traditional metrics. -
Develop a versatile, community-supported toolkit: By releasing an open-source Python library (cotescore), practitioners can integrate this robust framework into their existing DLA pipelines immediately, lowering the barrier to entry for advanced evaluation without requiring complete retraining of models.
-
Facilitate cross-dataset generalization: The SSU concept provides a relational structure that is less dependent on exact label consistency across datasets, allowing for better comparison and transfer of knowledge between documents with different annotation schemas.
Abstract
Document Layout Analysis (DLA) is the process by which a page is parsed into meaningful elements, often using machine learning models. Typically, the quality of a model is judged using general machine vision metrics such as IoU, F1 or mAP. However, these metrics are designed for images that are 2D projections of 3D space, not for the natively 2D imagery of printed media. This discrepancy can result in misleading or uninformative interpretation of model performance. To encourage more robust, comparable, and nuanced DLA, we introduce: The Structural Semantic Unit (SSU), a relational labelling approach that shifts the focus from the physical to the semantic structure of the content; and the Coverage, Overlap, Trespass, and Excess (COTe) score, a decomposable metric for measuring page parsing quality. We demonstrate the value of these methods through case studies and by evaluating 5 common DLA models on 3 DLA datasets. We show that the COTe score is more informative than traditional metrics and reveals distinct failure modes across models, such as breaching semantic boundaries or repeatedly parsing the same region. We find that, under granularity differences between model and ground truth, the COTe score is substantially more robust than the F1. Even in the worst case, comparing character-level predictions against paragraph-level ground truth with otherwise perfect parsing, COTe returns 0.68 where F1 returns 0. Notably, we find that, on real datasets, the COTe's granularity robustness largely holds even without explicit SSU labelling, reducing the barrier to entry. Finally, we release an SSU labelled dataset and a Python library for applying COTe in DLA projects.
Sources
- LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
- FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents
- ICDAR 2023 Competition on Robust Layout Segmentation in Corporate Documents
- M$^{6}$Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis
- PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction
- ICDAR 2025 Competition on FEw-Shot Text line segmentation of ancient handwritten documents (FEST)
- Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
- Are we done with ImageNet?
- Automatic universal taxonomies for multi-domain semantic segmentation
- Detection Hub: Unifying Object Detection Datasets via Query Adaptation on Language Embedding
- Label Unification for Cross-Dataset Generalization in Cybersecurity NER
- The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models
- DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
- Advanced Layout Analysis Models for Docling
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models