EngIntervene: Benchmarking Multimodal Engineering State Understanding and Design Intervention Reasoning
cs.CV, cs.SE
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/changcv2021/EngIntervene
Terminology
Sources
- CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
- Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
- MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
- Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generation
- AnalogCoder-Pro: Unifying Analog Circuit Generation and Optimization via Multi-modal LLMs
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
- FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
- BikeBench: A Bicycle Design Benchmark for Generative Models with Objectives and Constraints
- ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
- BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models