Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/damianomarsili/ExpertLens
Terminology
Sources
- Qwen3-VL Technical Report
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- MoE Lens -- An Expert Is All You Need
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Towards Automated Circuit Discovery for Mechanistic Interpretability
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- Steering MoE LLMs via Expert (De)Activation
- Gemma 4 Technical Report
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- What Gets Activated: Uncovering Domain and Driver Experts in MoE Language Models
- Mixtral of Experts
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- PubMedQA: A Dataset for Biomedical Research Question Answering
- GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning
- Kimi K3: Open Frontier Intelligence
- Kimi-VL Technical Report
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models