GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
cs.CV, cs.AI, cs.CL
Submitted: 2025-09-29
Updated: 2026-09-04
Comments: 59 pages, 7 figures, Project Page: https://zju-real.github.io/GSM8K-V Code: https://github.com/ZJU-REAL/GSM8K-V Datasets: https://huggingface.co/datasets/ZJU-REAL/GSM8K-V Accepted at EMNLP 2026 Main Conference. Updated to the camera-ready version with additional experiments, analyses, and revisions
Code: https://github.com/ZJU-REAL/GSM8K-V
Project page: https://zju-real.github.io/GSM8K-V
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mathematical reasoning is a key capability for vision-language models (VLMs), yet current benchmarks mainly evaluate text-based or explicitly symbolic visual inputs.
Terminology
Abstract
Mathematical reasoning is a key capability for vision-language models (VLMs), yet current benchmarks mainly evaluate text-based or explicitly symbolic visual inputs. It remains unclear whether VLMs can reason mathematically when information must be perceived and inferred from images rather than read from explicit symbols. We introduce GSM8K-V, a benchmark transforming GSM8K into multi-image sequences with semantic equivalence preserved. By mapping text-based problems into visual form via an automated pipeline and human verification, we curate 1,319 high-quality samples. In GSM8K-V, quantities must be extracted through visual perception, and reasoning chains must be reconstructed by integrating implicit cues across scenes. Evaluation of 34 VLMs reveals a striking modality gap: while most models exceed 90% on text, the best model achieves only 59% on GSM8K-V, far below the 91% human accuracy. Notably, models enhanced for visual math reasoning show no improvement on GSM8K-V despite large gains on existing benchmarks, confirming that it evaluates a distinct capability. Error analysis shows that the primary bottleneck lies in Implicit Visual Inference Error (IVIE), where models fail to recover visual semantics that are implied rather than explicitly stated. Our code and data are released at https://github.com/ZJU-REAL/GSM8K-V.
Sources
- Qwen2.5-VL Technical Report
- GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning
- Training Verifiers to Solve Math Word Problems
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
- Stable Reinforcement Learning for Efficient Reasoning
- Kimi-VL Technical Report
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Mathematical Problem Solving With the MATH Dataset
- Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Ovis2.5 Technical Report
- GPT-4o System Card
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models