PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection
cs.CV, cs.AI, cs.CL
Submitted: 2025-02-17
Updated: 2026-08-28
Comments: Accepted to EMNLP 2026 and selected for the ACL 2026 Best Paper Consideration
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Representation Degeneration Problem in Training Natural Language Generation Models
- VizWiz Grand Challenge: Answering Visual Questions from Blind People
- LoRA: Low-Rank Adaptation of Large Language Models
- CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
- Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
- Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning
- Understanding Dimensional Collapse in Contrastive Self-supervised Learning
- Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection
- Concept-skill Transferability-based Data Selection for Large Vision-Language Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- The Double-Ellipsoid Geometry of CLIP
- LLaVA-OneVision: Easy Visual Task Transfer
- Evaluating Object Hallucination in Large Vision-Language Models
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- Graph is a Substrate Across Data Modalities
- Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- Improved Baselines with Visual Instruction Tuning
- MMBench: Is Your Multi-modal Model an All-around Player?
- Less is More: High-value Data Selection for Visual Instruction Tuning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models