Understanding the Effects of Distractors on Reasoning Vision-Language Models
cs.CV, cs.AI, cs.CL, cs.LG
Submitted: 2025-11-26
Updated: 2026-09-14
Comments: EMNLP 2026
Code: https://github.com/luca-medeiros/langsegment-anything
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
- The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
- Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
- Intern-S1: A Scientific Multimodal Foundation Model
- O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
- DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
- Qwen3-VL Technical Report
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- Demystifying Long Chain-of-Thought Reasoning in LLMs
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models