Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?
cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- Qwen2.5-VL Technical Report
- Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Training Verifiers to Solve Math Word Problems
- Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
- Towards a Unified View of Parameter-Efficient Transfer Learning
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- Editing Models with Task Arithmetic
- When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Steering Llama 2 via Contrastive Activation Addition
- Function Vectors in Large Language Models
- Steering Language Models With Activation Engineering
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
- Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
- Multimodal Chain-of-Thought Reasoning in Language Models
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection