MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations
cs.CV, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: EMNLP 2026 Main Conference
Code: https://github.com/LijieFan/LaCLIP
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen3-VL Technical Report
- Microsoft COCO Captions: Data Collection and Evaluation Server
- ComKD-CLIP: Comprehensive Knowledge Distillation for Contrastive Language-Image Pre-traning Model
- The Llama 3 Herd of Models
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- What's "up" with vision-language models? Investigating their struggle with spatial reasoning
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- Image Retrieval from Contextual Descriptions
- Modeling Caption Diversity in Contrastive Vision-Language Pretraining
- Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
- VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Towards reporting bias in visual-language datasets: bimodal augmentation by decoupling object-attribute association
- LLAVADI: What Matters For Multimodal Large Language Models Distillation
- A Survey on Knowledge Distillation of Large Language Models
- Qwen3 Technical Report
- When and why vision-language models behave like bags-of-words, and what to do about it?
- VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models