SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
cs.LG, cs.AI
Submitted: 2026-02-02
Updated: 2026-08-26
Code: https://github.com/LAMDA-CL/Prism
Terminology
Sources
- Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
- CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
- LLaVA-c: Continual Improved Visual Instruction Tuning
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
- OASIS: Online Sample Selection for Continual Visual Instruction Tuning
- Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning
- LLaMA: Open and Efficient Foundation Language Models
- Large Continual Instruction Assistant
- LoKI: Low-damage Knowledge Implanting of Large Language Models
- HERMAN: Hierarchical Representation Matching for CLIP-based Class-Incremental Learning
- Cross-Sample Relational Fusion: Unifying Domain Generalization and Class-Incremental Learning
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- How to Teach Large Multimodal Models New Skills
- PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
- LLaVA-CMoE: Towards Continual Mixture of Experts for Large Vision-Language Models
- Instruction-Following Evaluation for Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks