CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models
cs.CV
Submitted: 2026-09-25
Updated: 2026-10-08
Terminology
Sources
- Pixtral 12B
- Concrete Problems in AI Safety
- Qwen2.5-VL Technical Report
- Seeing to Generalize: How Visual Data Corrects Binding Shortcuts
- PaLM-E: An Embodied Multimodal Language Model
- CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
- GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
- MiniMax Sparse Attention
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- Probing Visual Language Priors in VLMs
- Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance Graphs
- InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
- Shortcut Learning Susceptibility in Vision Classifiers
- Gemma 3 Technical Report
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- CCR-Bench: A Comprehensive Benchmark for Evaluating LLMs on Complex Constraints, Control Flows, and Real-World Cases
- When and why vision-language models behave like bags-of-words, and what to do about it?
- Causal Parrots: Large Language Models May Talk Causality But Are Not Causal
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models