CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning
cs.CL
Submitted: 2026-06-01
Updated: 2026-08-28
Terminology
Sources
- MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
- When Continue Learning Meets Multimodal Large Language Model: A Survey
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
- OASIS: Online Sample Selection for Continual Visual Instruction Tuning
- Large Continual Instruction Assistant
- Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
- LLaVA-c: Continual Improved Visual Instruction Tuning
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- A Practitioner's Guide to Continual Multimodal Pretraining
- IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
- Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning
- PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
- LoKI: Low-damage Knowledge Implanting of Large Language Models
- LLaVA-CMoE: Towards Continual Mixture of Experts for Large Vision-Language Models
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering