OCR-MetaReasoning Benchmark: Evaluating the Meta-Reasoning Ability of MLLMs in Text-Rich Image Understanding
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/gengxuli/OCR-MetaReasoning
Terminology
Sources
- Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
- Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
- Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents
- Towards Visual Text Grounding of Multimodal Large Language Model
- GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
- Spatial Dual-Modality Graph Reasoning for Key Information Extraction
- Qwen3-VL Technical Report
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
- MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
- MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering