A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Answering
cs.AI
Submitted: 2026-08-27
Updated: 2026-08-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: Answering questions over real-world documents requires processing long inputs that interleave text with tables.
Terminology
Abstract
Answering questions over real-world documents requires processing long inputs that interleave text with tables. Optical context compression, which represents context as images, promises to reduce token cost, but its effect on table understanding remains unclear. We study pixel-level table compression for question answering over documents with multiple tables, evaluating five VLMs across two benchmarks and five visual-token budgets. Representing tables as images at native resolution matches text in both performance and efficiency, but downscaling them makes models compensate the loss in readability with longer, less effective reasoning traces that cancel the expected savings. Highly downscaled tables, however, preserve enough signal to identify whether they are relevant to a question. We exploit this asymmetry with a training-free, two-step method: the model first identifies the tables needed to answer a question from a pixel-compressed context, and then reasons over those at native resolution. On long documents, our method saves 41% of total tokens and gains 7 accuracy points over single-step QA with native resolution tables. It also uses 15% fewer tokens than the most efficient single-step compressed configuration, with no accuracy loss.
Sources
- An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
- Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
- DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
- TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
- ColPali: Efficient Document Retrieval with Vision Language Models
- Qwen3-VL Technical Report
- Improving Language Understanding from Screenshots
- Gemma 4 Technical Report
- Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
- Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
- TABQAWORLD: Optimizing Multimodal Reasoning for Multi-Turn Table Question Answering
- A Comprehensive Survey on Long Context Language Modeling
- Language Modelling with Pixels
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
- NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
- Document-Level Numerical Reasoning across Single and Multiple Tables in Financial Reports
- DeepSeek-OCR: Contexts Optical Compression
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection