Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models
cs.CL
Submitted: 2026-09-13
Updated: 2026-10-05
Comments: Accepted to EMNLP 2026 (2026 Conference on Empirical Methods in Natural Language Processing)
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophisticated synthesis of perceptual grounding and
Terminology
Abstract
Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophisticated synthesis of perceptual grounding and symbolic logic. However, in the realm of mathematical functions, our investigation reveals a critical modality interference phenomenon: even advanced models, while performing textual computational reasoning, tend to disregard or misinterpret essential visual cues. To address this challenge, we propose Func-R1, which synergistically harmonizes precise visual perception and rigorous logical reasoning. Concretely, built upon an explicitly decoupled architecture, we employ a hierarchical post-training framework to progressively identify critical visual evidence and conduct in-depth theoretical reasoning. Furthermore, the Perception-Aligned Theoretic Optimization (PATO) strategy is proposed to steer policy updating towards internalizing fundamental theoretical properties while dynamically rectifying heterogeneous visual information throughout the reasoning process. Extensive experiments across diverse benchmarks demonstrate that Func-R1 delivers the optimal performance among open-source MLLMs, even surpassing GPT-5 with an 8.4% improvement on MathVerse's function-oriented tasks.
Sources
- GPT-4 Technical Report
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen3-VL Technical Report
- Soft Adaptive Policy Optimization
- CogVLM2: Visual Language Models for Image and Video Understanding
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Gemma 3 Technical Report
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
- MinerU: An Open-Source Solution for Precise Document Content Extraction
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Generalizable Geometric Image Caption Synthesis
- Qwen3 Technical Report
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering