HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
cs.CV, cs.AI
Submitted: 2025-05-16
Updated: 2026-09-07
Code: https://github.com/VectorInstitute/HumaniBench
Terminology
Sources
- MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
- The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence
- Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
- VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
- Red Teaming Visual Language Models
- Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
- MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
- HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
- AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
- MVP-Bench: Can Large Vision--Language Models Conduct Multi-level Visual Perception Like Humans?
- SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
- GPT-4o System Card
- Qwen2.5-VL Technical Report
- Phi-4 Technical Report
- Gemma 3 Technical Report
- CogVLM2: Visual Language Models for Image and Video Understanding
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- The Llama 3 Herd of Models
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models