SciMIF: Understanding Multimodal Instruction Following in Scientific Domains
cs.AI, cs.LG
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 21pages, 9 figures, 16 tables
Code: https://github.com/shenye7436/SciMIF
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields.
Terminology
Abstract
Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. In this work, we introduce SciMIF, a novel benchmark designed to evaluate the capability of MLLMs in following complex scientific instructions. Specifically, based on an extensive analysis of 22 distinct tasks across 5 representative scientific disciplines, we propose a comprehensive taxonomy comprising 10 constraint groups that captures both general functional requirements and discipline-specific characteristics. Guided by this taxonomy, we develop a high-fidelity instruction injection pipeline to systematically augment existing scientific datasets. We conduct comprehensive experiments on multiple state-of-the-art closed-source and open-source MLLMs. Our findings reveal significant performance disparities across different scientific disciplines, with chemistry posing greater challenges for current MLLMs. Furthermore, we observe that increasing the model scale does not yield corresponding improvements in constraint adherence, and current models still struggle severely with fine-grained constraints and instructions requiring the deep application of disciplinary knowledge. SciMIF fills the current void in evaluating multimodal instruction adherence within scientific domains, laying a crucial foundation for future enhancements of MLLMs in rigorous scientific applications. Data and code will be released at https://github.com/shenye7436/SciMIF.
Sources
- MM-IFEngine: Towards Multimodal Instruction Following
- ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
- Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research
- Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
- Gemini: A Family of Highly Capable Multimodal Models
- PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
- EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
- UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
- MaScQA: A Question Answering Dataset for Investigating Materials Science Knowledge of Large Language Models
- MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
- Instruction-Following Evaluation for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection