From Pixels to Pairs: A Comprehensive Benchmark of LLM-Driven Key-Value Extraction in Noisy Document Settings
cs.CL, cs.CV
Submitted: 2026-07-14
Updated: 2026-09-28
Comments: 25 pages, 20 tables, 5 figures
Code: https://github.com/JaidedAI/EasyOCR
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Llama 3 Herd of Models
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- PP-OCR: A Practical Ultra Lightweight OCR System
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
- Mistral 7B
- StructuralLM: Structural Pre-training for Form Understanding
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
- Quantum RNG Integration in an NG-PON2 Transceiver
- Qwen2.5 Technical Report
- DocVLM: Make Your VLM an Efficient Reader
- Achieving the Asymptotically Optimal Sample Complexity of Offline Reinforcement Learning: A DRO-Based Approach
- Unifying Vision, Text, and Layout for Universal Document Processing
- Gemma: Open Models Based on Gemini Research and Technology
- InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction
- Emergent Abilities of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering