OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/ECNU-RAIL/OmniPhys-EMNLP2026
Terminology
Sources
- Phi-4 Technical Report
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- Qwen3-VL Technical Report
- Hallucination of Multimodal Large Language Models: A Survey
- InternLM2 Technical Report
- TheoremQA: A Theorem-driven Question Answering dataset
- PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
- Using Large Language Model to Solve and Explain Physics Word Problems Approaching Human Level
- The Llama 3 Herd of Models
- RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
- Measuring Massive Multitask Language Understanding
- OpenAI o1 System Card
- Improving Physics Reasoning in Large Language Models Using Mixture of Refinement Agents
- K12Vista: Exploring the Boundaries of MLLMs in K-12 Education
- DeepSeek-V3 Technical Report
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Multi-Physics: A Comprehensive Benchmark for Multimodal LLMs Reasoning on Chinese Multi-Subject Physics Problems
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering