CapGeo-Bench: Decoupling Visual Perception from Reasoning and Evaluating Geometric Understanding
cs.CV, cs.AI, cs.CL
Submitted: 2025-10-10
Updated: 2026-09-13
Comments: 20 pages
Code: https://github.com/YuYingLi0/CapGeo
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen2.5-VL Technical Report
- Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
- Winning Gold at IMO 2025 with a Model-Agnostic Verification-and-Refinement Pipeline
- GPT-4o System Card
- GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning
- Benchmarking and Improving Detail Image Caption
- OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
- RAGAR, Your Falsehood Radar: RAG-Augmented Reasoning for Political Fact-Checking using Multimodal Large Language Models
- G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
- Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
- Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- Enhancing Advanced Visual Reasoning Ability of Large Language Models
- MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
- LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
- Visual-RFT: Visual Reinforcement Fine-Tuning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models