WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
cs.CV, cs.AI
Submitted: 2026-03-18
Updated: 2026-03-18
Journal ref: European Conference on Computer Vision (ECCV), 2026
Code: https://github.com/EvolvingLMMs-Lab/openr1-multimodal
License: http://creativecommons.org/licenses/by/4.0/
The gist: Existing vision-language models (VLMs) have demonstrated impressive performance in reasoning-based segmentation.
Terminology
Abstract
Existing vision-language models (VLMs) have demonstrated impressive performance in reasoning-based segmentation. However, current benchmarks are primarily constructed from high-quality images captured under idealized conditions. This raises a critical question: when visual cues are severely degraded by adverse weather conditions such as rain, snow, or fog, can VLMs sustain reliable reasoning segmentation capabilities? In response to this challenge, we introduce WeatherReasonSeg, a benchmark designed to evaluate VLM performance in reasoning-based segmentation under adverse weather conditions. It consists of two complementary components. First, we construct a controllable reasoning dataset by applying synthetic weather with varying severity levels to existing segmentation datasets, enabling fine-grained robustness analysis. Second, to capture real-world complexity, we curate a real-world adverse-weather reasoning segmentation dataset with semantically consistent queries generated via mask-guided LLM prompting. We further broaden the evaluation scope across five reasoning dimensions, including functionality, application scenarios, structural attributes, interactions, and requirement matching. Extensive experiments across diverse VLMs reveal two key findings: (1) VLM performance degrades monotonically with increasing weather severity, and (2) different weather types induce distinct vulnerability patterns. We hope WeatherReasonSeg will serve as a foundation for advancing robust, weather-aware reasoning.
Sources
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
- HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
- Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Training Language Models to Self-Correct via Reinforcement Learning
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Bridging Day and Night: Target-Class Hallucination Suppression in Unpaired Image Translation
- MIMT: Multi-Illuminant Color Constancy via Multi-Task Local Surface and Light Color Learning
- Seeing Beyond Haze: Generative Nighttime Image Dehazing
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming
- SAM 2: Segment Anything in Images and Videos
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- RaindropGS: A Benchmark for 3D Gaussian Splatting under Raindrop Conditions
- Solving math word problems with process- and outcome-based feedback
- PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models