DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning
cs.CV
Submitted: 2026-06-06
Updated: 2026-09-27
Code: https://github.com/Sammy20207109/DyCo-RL
Terminology
Sources
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Soft Adaptive Policy Optimization
- Spotlight on Token Perception for Multimodal Reinforcement Learning
- Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
- Perception-Aware Policy Optimization for Multimodal Reasoning
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Token-level Direct Preference Optimization
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models