SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions
cs.CV, cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Qwen2.5-VL Technical Report
- SMPLitex: A Generative Model and Dataset for 3D Human Texture Estimation from Single Image
- PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- YOLOv11: An Overview of the Key Architectural Enhancements
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
- Small Language Models: Survey, Measurements, and Insights
- SmolVLM: Redefining small and efficient multimodal models
- Contrast Limited Adaptive Histogram Equalization (CLAHE) Approach for Enhancement of the Microstructures of Friction Stir Welded Joints
- CinePile: A Long Video Question Answering Dataset and Benchmark
- Gemma 3 Technical Report
- Boundless: Generating Photorealistic Synthetic Data for Object Detection in Urban Streetscapes
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Synscapes: A Photorealistic Synthetic Dataset for Street Scene Parsing
- LLaVAction: evaluating and training multi-modal large language models for action understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models