MemeLens: Multilingual Multitask VLMs for Memes
cs.AI, cs.CL
Submitted: 2026-01-18
Updated: 2026-09-18
Comments: disinformation, misinformation, factuality, harmfulness, fake news, propaganda, hateful meme, multimodality, text, images
Code: https://github.com/MohamedBayan/MemeLens
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Memes are a dominant medium for online communication and manipulation because meaning emerges from interactions between embedded text, imagery, and cultural context.
Terminology
Abstract
Memes are a dominant medium for online communication and manipulation because meaning emerges from interactions between embedded text, imagery, and cultural context. Existing meme research is distributed across tasks (e.g., hate, misogyny, propaganda, sentiment, humour) and languages, which limits cross-domain generalization. To address this gap, we propose MemeLens, a unified multilingual, multitask explanation-enhanced Vision-Language Model (VLM) for meme understanding. We consolidate 38 public meme datasets, filter and map dataset-specific labels into a shared taxonomy of 20 tasks spanning harm, targets, figurative/pragmatic intent, and affect. We present a comprehensive empirical analysis across modeling paradigms, task categories, and datasets. Our findings suggest that robust meme understanding requires multimodal training, varies substantially across semantic categories, and remains sensitive to over-specialization when models are fine-tuned on individual datasets rather than trained in a unified setting. We make the experimental resources (https://github.com/MohamedBayan/MemeLens), model (https://huggingface.co/QCRI/MemeLens-VLM) and datasets (https://huggingface.co/datasets/QCRI/MemeLens) publicly available to the community.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- On VLMs for Diverse Tasks in Multimodal Meme Classification
- From Recognition to Reasoning: Advancing Multimodal Harmful Meme Detection via Chain-of-Thought Alignment
- LoRA: Low-Rank Adaptation of Large Language Models
- MIMIC: Multimodal Islamophobic Meme Identification and Classification
- ExPO-HM: Learning to Explain-then-Detect for Hateful Meme Detection
- GPT-4 Technical Report
- RoMemes: A multimodal meme corpus for the Romanian language
- MedGemma Technical Report
- OpenAI GPT-5 System Card
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qwen3-Omni Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection