Tables Decoded: DELTA for Structure, TARQA for Understanding
cs.CV, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026
DOI: 10.1109/WACV61042.2026.00272
Code: https://github.com/Tihiitborg/Tables-Decoded
License: http://creativecommons.org/licenses/by/4.0/
The gist: Table understanding is a core task in document intelligence, encompassing two key subtasks: table reconstruction and table visual question answering (TabVQA).
Terminology
Abstract
Table understanding is a core task in document intelligence, encompassing two key subtasks: table reconstruction and table visual question answering (TabVQA). While recent approaches predominantly rely on vision- language models (VLMs) operating on table images, we propose a more scalable and effective alternative based on structured textual representations. These representations are easier to process, align more naturally with LLMs, and eliminate the need for language-specific visual encoders, making them particularly suitable for multilingual documents. We present DELTA, which separates physical structure recognition, logical structure recognition, and OCR to extract both layout and content accurately. DELTA outputs tables in Optimised Table Structure Language (OTSL), a compact and unified format that encodes cell arrangements and textual content. On table structure recognition (TSR), DELTA achieves TEDS- Structure scores comparable with state-of-the-art methods across FinTabNet, PubTabNet, and PubTables-1M. We further establish its robustness on non-English tables through our curated Hindi benchmark, TORQUE. Building on this, we introduce TARQA, an LLM fine-tuned on OTSL sequences. Our approach yields gains of 9.3 p.p. on WTQ (TabQA) and 9.2 p.p. on FinTabNetQA (TabVQA), respectively. On TORQUE, our method ranks second among all VLMs and DELTA + LLM variants. We release our code, models, and benchmark at: https://github.com/Tihiitborg/Tables-Decoded
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- The Llama 3 Herd of Models
- Mask R-CNN
- CogAgent: A Visual Language Model for GUI Agents
- mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
- Mistral 7B
- Mixtral of Experts
- OCR-free Document Understanding Transformer
- TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
- SPRINT: Script-agnostic Structure Recognition in Tables
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
- Improved Baselines with Visual Instruction Tuning
- An End-to-End Multi-Task Learning Model for Image-based Table Recognition
- SmolVLM: Redefining small and efficient multimodal models
- TableFormer: Table Structure Understanding with Transformers
- SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models