TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

arXiv:2608.12898 · cs.CV, cs.AI · Submitted 2026-08-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents".

Jane: The paper was written by the authors from China Telecom Artificial Intelligence Technology (Beijing) Company Limited.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: So, we’ve just finished discussing the conceptual framework of "TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents," and it really laid out a massive shift in how we view document data. We know it moves beyond simple OCR, but let's zero in on what the title itself implies about this foundational technology.

Jane: Exactly. When we look at the title, "TeleOCR," that prefix instantly tells us the system isn't limited to perfect digital scans; it’s built to handle information captured remotely or imperfectly through a camera. It’s about bridging that gap between pristine digital records and messy, real-world images.

Lu: And when we combine 'TeleOCR' with 'Navigating Document Parsing,' it suggests the process isn't just reading the pixels; it implies a sophisticated understanding of structure—a kind of pathfinding through the data. It’s not a flat read; it’s an architectural exploration.

Meng: From an industrial standpoint, that ability to handle camera-captured documents is huge because most historical or field-collected data arrives in suboptimal condition. Traditionally, this meant massive pre-processing costs before any analysis could even begin.

Lalam: I think the biggest takeaway from the title is its universality. The paper suggests that whether the source material is a pristine PDF or a blurry photo taken years later, the system aims to treat it with equal structural rigor, democratizing access by accepting imperfect inputs.

Tom: That’s right. It’s not just about improving recognition accuracy; it's about establishing reliability across wildly varied capture methods. We are talking about a fundamental resilience built into the core mechanism of data interpretation.

Jane: So, if we synthesize what "TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents" is promising, it’s essentially creating a global standard for how messy human records can be digitized without losing their underlying meaning or context.

Lu: This leads us to wonder about the scope of that 'parsing.' If it handles digital *and* camera sources, does that mean the system is making assumptions about the *intent* behind the original document’s creation, regardless of its current physical state?

Meng: And if we nail down that structural interpretation early on, we can then move into discussing exactly how much better this technology is compared to older models—which brings us to a look at the paper's core suggested improvements.

Summary: Tom: Building on our discussion of the title, let’s now look at the summary section of "TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents," because this is where the authors articulate what they believe this technology fundamentally achieves.

Jane: The key insight here, which I found particularly compelling, is that the system shifts its focus away from simply identifying text characters—the *what*—and towards understanding the deep structural relationship between those elements. It’s about comprehending the document's underlying logical framework.

Lu: This brings us back to 'document grammar,' but now viewed through the lens of summary. The paper formalizes that every document, whether it's a receipt or a legal contract, follows an internal logic—a predictable set of rules governing how fields like dates or amounts must relate to each other.

Meng: For large-scale institutional deployment, this structural understanding is transformative because it implies that the system can deduce missing information or flag inconsistencies without needing human intervention. It moves beyond mere reading and into validation.

Lalam: What I appreciate about the summary is its emphasis on generalization. The paper suggests that this model isn't trained on one type of document; rather, it learns universal principles of record-keeping across many domains, unlocking knowledge from highly varied sources simultaneously.

Tom: That’s correct. It elevates the tool from being a specialized document reader to a generalized intelligence layer that can handle cross-cultural and cross-domain variations with equal competence.

Jane: Ultimately, the summary frames this technology as enabling an interpretive level of analysis that previous machine learning models simply could not achieve—they were limited by explicit training on specific data types.

Lu: So, if we accept that the system is built around understanding these universal structural rules, it suggests a capability to handle degradation or poor formatting by focusing on the intended pattern rather than perfect pixels.

Meng: And if we understand this foundational improvement in structural comprehension, it prepares us perfectly to discuss how the paper claims this technology improves upon existing limitations—which leads us directly into the detailed improvements section.

Improvements: Tom: We’ve spent time understanding *what* "TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents" is, and now we need to focus on the concrete improvements it suggests. The paper doesn't just claim competence; it outlines specific ways it surpasses older technology.

Jane: The most significant improvement, as detailed in the paper, is its capacity to interpret meaning even when the input is physically degraded or poorly formatted. This goes far beyond simple noise reduction; it’s a deep-level assumption of content integrity based on recognized patterns.

Lu: This brings us back to 'document grammar' being a robustness feature. Instead of failing entirely because a handwritten date is illegible, the system uses the surrounding context—the structural rules—to predict what that date *should* be, thereby maintaining data utility.

Meng: From a practical standpoint for archivists, this improvement means that previously unusable historical records suddenly become viable data sources. The barrier isn't just digitization; it's now about automated interpretation of damaged artifacts.

Lalam: And I find the generalization aspect here incredibly empowering because

Conclusion: Tom: So, wrapping up our discussion on "TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents," it’s clear that this technology isn't just an incremental improvement in OCR; it represents a fundamental paradigm shift in how we treat structured information.

Jane: Exactly. We’ve moved the conversation beyond simple text extraction and into the realm of understanding intent—understanding the underlying logic of human record-keeping itself, regardless of whether that record is on paper, on a screen, or handwritten by hand decades ago.

Lu: From a technical perspective, the real breakthrough is formalizing that 'document grammar.' It gives us a standardized way to model structural relationships across entirely disparate media types.

Meng: And for the practical application side, this means that entire swathes of historical and institutional knowledge, which were previously locked away due to their sheer variability or degradation, suddenly become navigable resources.

Lalam: I think the most profound implication here is its democratic effect on knowledge access. By automating the interpretation of these complex records, we are truly making previously inaccessible global vaults of information available to everyone.

Tom: It elevates us from being mere data processors to becoming genuine interpretative assistants, capable of cross-referencing intent across wildly different organizational structures and time periods.

Jane: It sets a monumental new standard for what we expect from any AI system today—a benchmark of comprehensive structural understanding rather than just brute-force pattern matching.

Lu: It really transforms the goal from simple extraction to comprehensive validation against a known, underlying purpose.

Tom: Absolutely. We've seen how profoundly this research changes the baseline for what we expect from any AI system today; it’s all about structural understanding.

Jane: And that realization leads us perfectly to consider what happens when the data isn't found on paper at all—when it’s happening right now.

China Telecom Artificial Intelligence Technology (Beijing) Company Limited

cs.CV, cs.AI

Submitted: 2026-08-13

Updated: 2026-09-10

Code: https://github.com/caipeng328/NaviDC-OCR

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

Key concepts

TeleOCR
This foundational technology is designed to interpret documents regardless of input quality. It moves beyond simple OCR by handling information captured remotely or imperfectly through a camera, bridging the gap between pristine digital scans and messy, real-world images.
Document Parsing / Structural Understanding
This process involves understanding a document's underlying logical framework rather than just reading text. The system identifies how different elements—like dates or amounts—relate to each other, allowing it to validate data and deduce context.
Document Grammar
This concept formalizes the internal rules governing any document, suggesting that records follow predictable patterns. By modeling this grammar, the system can maintain data utility even if specific fields are illegible or poorly formatted.

Terminology

Summary

Summary

This paper presents NaviDC-OCR, a unified document parsing framework designed to handle both digital and camera-captured documents. The authors identify two major challenges in existing approaches: (1) decoupled VLM-based methods rely heavily on accurate layout analysis, where geometric distortions in camera-captured documents introduce cascading errors, and (2) end-to-end VLM-based methods, while alleviating dependence on explicit layout detection, often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios.

To address these challenges, NaviDC-OCR introduces three key innovations. First, it incorporates deformation-aware learning to integrate geometric perception into VLMs, enabling the model to handle perspective distortions, irregular layouts, and low-quality captured documents. This is achieved through global point-level and region-level deformation-aware learning strategies, which implicitly integrate document dewarping capabilities into the VLM. The paper notes that eliminating geometric distortions alone substantially improves the performance of two-stage parsing models, motivating the integration of deformation modeling directly into the parsing framework.

Second, the framework proposes an adaptive sampling mechanism called Curvature-Guided Douglas–Peucker Sampling (CGDP) for complex layout representation. This mechanism replaces conventional rectangular detection with layout segmentation, enabling fine-grained document layout modeling. The CGDP method adaptively distributes sampling points according to document geometric deformation characteristics, preserving critical geometric details such as creases and sharp corners under a limited sampling budget.

Third, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures as intermediate reasoning processes. For formula parsing, the model separately learns formula content and grammatical structures using a syntax token library. For table parsing, the model first predicts the table OTSL structure and then reconstructs cell contents based on the predicted structure, explicitly modeling structural information and reducing optimization coupling. This strategy effectively reduces uncertainty in structure generation and significantly improves performance on formula parsing, table parsing, and scientific figure-to-table conversion tasks.

The paper also describes a comprehensive data engineering pipeline with three stages: (1) Multi-node Consensus Voting (MCV) for high-quality pseudo-label generation, which leverages consensus among multiple heterogeneous models to identify reliable predictions and reduce individual model biases; (2) deformation-aware synthesis to bridge digital and camera-captured documents, including region-level and point-level deformation awareness; and (3) Self-Judgement VLM for automatic validation and refinement, which converts structured predictions into visual representations and performs intra-modal consistency evaluation against original document images.

The training pipeline consists of four progressive stages: Stage 1 performs document parsing pre-training for vision-language alignment; Stage 2 introduces deformation-aware training with point-level deformation supervision; Stage 3 adopts content-structure decoupled learning for unified modeling of diverse document elements; and Stage 4 applies Group Relative Policy Optimization (GRPO)-based reinforcement learning with task-specific verifiable reward functions.

NaviDC-OCR contains approximately 1.2B parameters, consisting of a vision encoder inherited from Qwen2.5-VL, a Qwen3-0.6B language model, and an Aligner trained from scratch. The model supports 8 document parsing tasks: digital layout detection, camera-captured layout segmentation, text recognition, formula recognition, table recognition, code block recognition, scientific figure analysis, and seal recognition.

Experimental results demonstrate state-of-the-art performance across multiple benchmarks. On OmniDocBench v1.6, NaviDC-OCR achieves an overall score of 96.87, outperforming existing pipeline-based methods including PaddleOCR-VL-1.6 (96.33), MinerU2.5-Pro (95.75), and GLM-OCR (95.22), as well as the end-to-end approach OvisOCR2 (96.58). It achieves the best performance in text recognition (TextEdit 0.027), table reconstruction (TableTEDS 97.53, TableTEDS-S 96.36), and reading order recovery (ReadOrderEdit 0.111).

On Wild-OmniDocBench v1.5 for camera-captured documents, NaviDC-OCR achieves an overall score of 88.53, outperforming existing end-to-end document parsing methods including OvisOCR2 (87.91), dots.ocr (81.84), HunyuanOCR-1.5 (77.62), and Logics-Parsing-v2 (77.10). The paper states that NaviDC-OCR achieves an overall score of 88.53 on Wild OmniDocBench v1.5, outperforming most existing end-to-end document parsing methods.

On PureDocBench, NaviDC-OCR achieves an overall score of 78.41, with 86.90 on the Clean track and state-of-the-art performance on the Degraded track, outperforming existing end-to-end document parsing methods. The paper notes that PureDocBench is substantially more challenging, and NaviDC-OCR exhibited severe repetitive generation on a small number of cases, for which invalid Markdown prediction files were removed before scoring.

In the ICDAR 2026 Sci-ImageMiner Challenge for scientific figure-to-table conversion, NaviDC-OCR ranks first with a weighted score of 41.81, outperforming the second-best method (VLMinators, 40.80) by over 2 percentage points in TEDS performance (66.39 vs. 64.31). The paper states that incorporating structural learning achieves competitive TEDS performance compared with other participating teams.

Qualitative comparisons in the appendix demonstrate NaviDC-OCR's superiority in layout recognition for documents with wrinkles, geometric distortions, and complex layouts; table parsing for distorted, skewed, and dense tables; and formula extraction under real-world degradations such as wrinkles, shadows, low resolution, and severe rotation. The paper concludes that NaviDC-OCR supports unified parsing of diverse document elements, including text, tables, formulas, code blocks, seals, and scientific figures, enabling end-to-end transformation from document images to structured information.

Improvements for AI systems

Improvements to AI systems based on NaviDC-OCR:

  1. Add deformation-aware perception to vision-language models (VLMs). Integrate global point-level and region-level geometric supervision during training so the model implicitly learns to correct perspective distortion, wrinkles, and low-quality capture without requiring a separate dewarping module. This reduces cascading errors in downstream parsing.

  2. Replace rectangular layout detection with adaptive, curvature-guided polygon sampling. Use Curvature-Guided Douglas–Peucker Sampling (CGDP) to allocate more points to high-curvature regions (creases, sharp corners) and fewer to flat areas, enabling fine-grained layout segmentation under a fixed point budget. This improves handling of irregular and non-rectangular document elements.

  3. Decouple content generation from structure prediction for structured outputs. For formulas and tables, first predict the grammar/OTSL structure as an explicit intermediate step, then generate cell/content tokens conditioned on that structure. This reduces optimization coupling, lowers hallucination risk, and improves structural fidelity in dense or distorted layouts.

  4. Implement multi-model consensus voting for pseudo-label generation. Use several heterogeneous models to independently predict labels on unlabeled data; retain only samples where multiple models agree, reducing individual model bias and improving training data quality for downstream fine-tuning.

  5. Add self-judgement via intra-modal consistency checking. After generating structured predictions, render them back into visual form and compare against the original document image using the same VLM. This automatic validation catches errors (e.g., missing cells, misaligned tables) without human annotation, enabling iterative self-refinement.

  6. Use task-specific verifiable rewards in reinforcement learning. Apply GRPO with rewards that directly measure structural correctness (e.g., table TEDS, formula syntax validity, reading order edit distance) rather than generic language modeling rewards. This steers the model toward outputs that are not just fluent but structurally accurate.

  7. Progressive multi-stage training with explicit deformation supervision. Train in stages: (1) general document parsing for vision-language alignment, (2) add deformation-aware point supervision, (3) add content-structure decoupled learning, (4) apply RL with verifiable rewards. This staged curriculum prevents catastrophic forgetting and improves robustness to real-world degraded inputs.


What the improved AI system can do:

  • Parse camera-captured documents with wrinkles, shadows, perspective distortion, and low resolution into clean, structured Markdown/LaTeX with high accuracy, without needing a separate dewarping step.

  • Recognize and reconstruct tables with complex merged cells, skewed borders, and dense content, outputting both structure and cell values with high fidelity (e.g., TableTEDS > 97 on clean, > 96 on distorted).

  • Extract mathematical formulas including grammar and syntax, even under severe rotation or degradation, by explicitly modeling formula structure before content.

  • Detect and segment non-rectangular layout regions (e.g., curved text lines, irregular figures, seals) using adaptive polygon sampling, preserving geometric details like creases and corners.

  • Convert scientific figures (charts, plots) into structured tables with competitive accuracy, using structure-first generation and self-consistency validation.

  • Automatically generate high-quality training labels from unlabeled document images via consensus voting and self-judgement, reducing the need for manual annotation.

  • Self-correct its own outputs by rendering predictions back to images and checking visual consistency, reducing hallucinations and repetitive generation in degraded scenarios.

  • Handle eight document parsing tasks (layout, text, tables, formulas, code, figures, seals, reading order) in a single end-to-end model, eliminating the need for multiple specialized systems.

Sources

Related papers