OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/SKYLENAGE-AI/D3OmniFrameworkhttps:
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic
Terminology
Abstract
Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated by this, we introduce D3-Omni, a balanced and decoupled benchmark for diagnosing fine-grained multimodal understanding, covering 53 orthogonal binary dimensions (17/22/14) and 10,671 samples (3,526/1,998/5,147) across the three tasks. Rather than re-generating outputs, which may leak information across dimensions, we fix verified fully positive seeds and derive negatives through controlled prompt rewriting and atomic, dimension-isolating perturbations. The resulting D3 design is Dual-balanced, which helps alleviate negative-sample scarcity and per-dimension label imbalance; Decoupled, so that each error is attributable to a single capability; and Dynamic, steering construction toward under-represented regions of the label distribution as generative models improve.The suite reaches near 1:1 per-dimension parity and a uniform distribution over all total-score levels. Under this balanced view, even strong OmniJudges tend to struggle on modality-related dimensions, to confirm satisfied requirements far more reliably than they detect violated ones, and to treat nominally distinct attributes as largely a single decision, suggesting that aggregate accuracy may hide systematic blind spots that a balanced and decoupled lens can help expose and, in turn, address.
Sources
- OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
- Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
- Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
- Towards A Better Metric for Text-to-Video Generation
- CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
- Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
- GPT-4o System Card
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Qwen3-Omni Technical Report
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
- Baichuan-Omni Technical Report
- MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection