Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
cs.CL, cs.AI, cs.MM
Submitted: 2026-07-21
Updated: 2026-09-01
Terminology
Sources
- MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
- Qwen3-VL Technical Report
- StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos
- I Spy a Metaphor: Large Language Models and Diffusion Models Co-Create Visual Metaphors
- Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
- Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- A Survey of Multimodal Sarcasm Detection
- From Recognition to Reasoning: Advancing Multimodal Harmful Meme Detection via Chain-of-Thought Alignment
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- TextMI: Textualize Multimodal Information for Integrating Non-verbal Cues in Pre-trained Language Models
- Chumor 1.0: A Truly Funny and Challenging Chinese Humor Understanding Dataset from Ruo Zhi Ba
- Language Models, Agent Models, and World Models: The LAW for Machine Reasoning and Planning
- MemeCap: A Dataset for Captioning and Interpreting Memes
- BottleHumor: Self-Informed Humor Explanation using the Information Bottleneck Principle
- MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
- Meme-ingful Analysis: Enhanced Understanding of Cyberbullying in Memes Through Multimodal Explanations
- D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
- When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering