PEST: Parameter Efficient Steering of Blackbox VLMs via Agentic Few-shot Alignment for Hateful Meme Moderation

arXiv:2601.04692 · cs.CL, cs.CV · Submitted 2026-01-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "PEST: Parameter Efficient Steering of Blackbox VLMs via Agentic Few-shot Alignment for Hateful Meme Moderation".

Tom: In this work, a novel framework is proposed to simultaneously address hateful meme moderation by combining classification, explanation, and intervention using task-specific generative AI agents.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, wrapping up this discussion on "PEST: Parameter Efficient Steering of Blackbox VLMs via Agentic Few-shot Alignment for Hateful Meme Moderation," the authors are proposing a framework that integrates classification, explanation, and intervention using task-specific generative AI agents through a few-shot prompting style.

Jane: Essentially, they are showing how you can use these agents to generate synthetic training data to guide large multimodal models to perform all three moderation tasks at once, which solves the problem of having those components studied in isolation.

Lu: The real implication here is that they're providing a generalizable solution for content moderation under conditions where creating massive annotated datasets is just not practical.

Meng: From an engineering standpoint, this parameter efficient method combined with agentic steering offers a viable path toward more robust and less resource-intensive moderation pipelines.

Lalam: I think the focus on making the explanation and intervention capabilities coherent means that users get much better feedback when content is flagged, which could positively impact how people interact online.

Tom: Exactly, so the authors are really pushing for this simultaneous capability because they feel it better reflects how content moderation actually needs to operate on social media platforms.

Jane: The title itself suggests that the steering is parameter efficient, meaning they're achieving good results without needing to train enormous models from scratch, which is a big consideration.

Lu: If this approach proves effective across different meme styles and content types, it sets a precedent for how we can adapt these agentic steering techniques to other complex AI moderation challenges.

Meng: It’s about making the AI models more adaptable without needing a complete overhaul of their architecture, which is something we need when dealing with evolving internet content.

Lalam: Ultimately, this work suggests that the future of moderation might involve systems that are not just detectors but active participants in helping users navigate harmful content and understand the reasoning behind those decisions.

Conclusion: Tom: So, we've been diving deep into this paper about PEST, which is all about steering those big blackbox VLMs using agentic few-shot alignment for meme moderation.

Jane: It’s really fascinating how they manage to weave together classification, explanation, and intervention into one unified system without needing a massive amount of data to train everything separately.

Lu: The authors are showing that you can use these task-specific agents—the caption generator, the explainability generator, and the intervention generator—to create synthetic training material for the main models.

Meng: From an engineering standpoint, achieving parameter efficiency while maintaining this level of multi-task alignment is a significant hurdle they’ve managed to clear.

Lalam: What's really striking is how the system handles that gap where explanation and intervention are usually studied in isolation from detection, offering a holistic solution instead.

Tom: Exactly, it’s not just another detector; it’s a system that actually tells you *why* something is flagged and suggests how to handle it immediately.

Jane: And looking at the title, "Parameter Efficient Steering," they're pointing toward a more accessible way to use these powerful multimodal models for real-world moderation tasks.

Lu: I think the implication is that we don't always need petabytes of perfectly labeled data to get sophisticated, multi-faceted AI systems working effectively across different domains.

Meng: So, if this method works on memes and hate speech, it suggests a scalable architecture for applying this steering technique to other nuanced content moderation challenges.

Lalam: I see a future where our AI systems become much more transparent and proactive in how they handle harmful content, which could really improve the culture online.

Tom: Indeed, it feels like we’re moving toward AI that doesn't just spot problems but understands and responds to them intelligently.

Jane: It’s a powerful concept because it tackles the complexity of human language and intent within visual content in a very direct way.

Lu: The authors’ work really pushes us to think about how we can design these agentic loops to be more adaptable across diverse linguistic and visual contexts.

Meng: I'm curious what kind of constraints they found when testing this on different model architectures beyond GPT-4o, because that will tell us a lot about its practical deployment limits.

Lalam: That’s a great point, Meng; understanding those limits is crucial for building systems that are both powerful and reliable in production environments.

Naquee Rizwan, Subhankar Swain, Paramananda Bhaskar, Gagan Aryan, Shehryaar Shah Khan, Animesh Mukherjee

Indian Institute of Technology (IIT), Kharagpur

cs.CL, cs.CV

Submitted: 2026-01-08

Updated: 2026-09-28

Code: https://github.com/doccano/doccano

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: In this work, a novel framework is proposed to simultaneously address hateful meme moderation by combining classification, explanation, and intervention using task-specific generative AI agents.

Key concepts

Task-Specific Generative AI Agents
These are specialized AI models trained for distinct moderation subtasks: generating meme captions, creating label-aware explanations, and drafting interventions. They work together in a few-shot prompting process to enrich data.
Silver Training Data Generation
This refers to creating high-quality training examples for larger multimodal models using smaller, task-specific agents. The agents generate contextually rich captions and explanations that serve as synthetic ground truth, solving the problem of lacking sufficient labeled data.
Few-Shot Prompting Paradigm
Instead of needing thousands of examples, this technique uses a small number of high-quality examples to guide the AI. The framework selects relevant examples based on similarity (SigLIP embeddings) and then enriches them with agent outputs to teach the final model how to predict labels, explanations, and interventions.
Explanatory vs. Intervention Analysis
Analysis reveals differences between explanations and interventions: explanations are shorter but more diverse in vocabulary, focusing on descriptive terms like 'mocks.' Interventions are longer and more repetitive, using action-oriented words like 'promote' or 'avoid'.

Terminology

Summary

In this work, a novel framework is proposed to simultaneously address hateful meme moderation by combining classification, explanation, and intervention using task-specific generative AI agents. This approach tackles the critical research gap where explanation and intervention have been studied separately from detection, offering a generalizable solution for content moderation under limited data conditions.

How it works

The core of the proposed framework is built on a few-shot prompting paradigm utilizing task-specific agentic models to generate silver training data for larger multimodal models. This addresses the lack of coherent training samples that can act as ground truth annotations for end-to-end explainable classification. The process involves three distinct task-specific agents:

  1. Meme oriented caption generation (AC): Trained on the MemeCap dataset to provide meme-oriented captions for a given input, noting that Memes significantly depart compared to general images as their contextualized meaning is portrayed as a combination of embedded text and the image itself.

  2. Label aware explainability generator (AE): Trained on HatReDAug, which combines the HatReD dataset with manually annotated non-hateful memes from FHM, to generate very good silver training data for the train split of unseen datasets.

  3. Intervention generator for hateful memes (AI): Trained on the MemeSense dataset, which provides ground-truth causal reasoning and interventions for toxic memes, to express why a sample is classified as hateful.

The Few-Shot Framework

The few-shot prompting process involves two major steps: exemplar selection and enrichment of the selected exemplars. Exemplar selection is performed by choosing few-shot examples that have the highest cosine similarity with the t in terms of the SigLIP embeddings. For each selected exemplar, it is enriched by passing it through the task-specific agents. Specifically, AC returns a caption, AE returns an explanation for hateful or non-hateful memes based on their label awareness, and AI generates an intervention only if the meme is predicted as hateful. This enrichment allows the test sample to predict the label, the explanation and the intervention (if t is predicted as hateful).

Data Curation and Evaluation

The paper addresses data incoherence by utilizing prior datasets for different agent training purposes: HatReD for explanation generation, MemeCap and MemeSense for captioning and intervention generation. For evaluation, a novel approach is used where the test splits of FHM and MAMI datasets are extended with explanation and intervention annotations to create a dataset that includes the classification label alongside these new components. The final unified framework is evaluated using large multimodal models like GPT-4o, Intern-VL3, and Pixtral.

Key Findings

The experiments demonstrate significant performance gains across all tasks compared to existing baselines. GPT-4o attained macro-F1 scores of 80.25% and 89.07% on FHM and MAMI datasets respectively, outperforming all previous methods, including Vlm-Lim and LoReHM in classification accuracy. For explanation, GPT-4o provided significantly better results compared to other open source models and HatReD framework. Conversely, for intervention generation, Intern-VL3 and Pixtral were found to be the best models when compared against GPT-4o. Analysis of textual properties revealed that while explanations are generally shorter but more diverse in vocabulary (higher type-token ratio), intervention texts are longer but more repetitive with higher unigram perplexity. Furthermore, word shift graph analysis showed that explanations for hateful predictions contain words like ‘mocks’, ‘stereotype’, and ‘dehumanizing,’ whereas interventions lean toward words like ‘promote’ and ‘avoid.’

Limitations and Ethics

The current dataset is primarily in English, limiting the immediate generalizability of the helper agents to other languages. Future work could involve comparing large fine-tuned models against the few-shot setup across all three axes. The authors emphasize that they have taken proper care of privacy by extending previous datasets without adding new samples and running experiments with both open-source and proprietary models. They plan to release their dataset upon paper acceptance while ensuring no privacy leakage occurs.

The gist: A novel framework is proposed to simultaneously address hateful meme moderation by combining classification, explanation, and intervention using task-specific generative AI agents. This approach tackles the critical research gap where explanation and intervention have been studied separately from detection, offering a generalizable solution for content moderation under limited data conditions.

  1. Meme oriented caption generation (AC): Trained on the MemeCap dataset to provide meme-oriented captions for a given input.

  2. Label aware explainability generator (AE): Trained on HatReDAug, which combines the HatReD dataset with manually annotated non-hateful memes from FHM, to generate very good silver training data for the train split of unseen datasets.

Improvements for AI systems

As a fastidious researcher, I see several high-leverage areas where this framework offers significant improvement over existing methods in hateful meme moderation. The core innovation lies in moving from isolated components (classification, explanation, intervention) to a unified, few-shot agentic system.

Here are the specific improvements and what the improved AI system can achieve:


)Improved AI System Capabilities: Unified Hateful Meme Moderation Agent

The proposed framework enables an end-to-end multimodal agent capable of simultaneously performing three critical tasks on any given hateful meme input, leveraging few-shot learning to adapt to new meme types without requiring massive, task-specific fine-tuning datasets.

Specific Improvements:

  1. Simultaneous Tripartite Reasoning (The Core Leap):

Hateful memes are no longer just classified; the system can immediately generate a rationale and a suggested moderation action in a single inference pass.

  1. Few-Shot Generalizability via Task-Specific Agents:

Instead of needing millions of labeled examples, the system leverages task-specific silver data generation agents (Captioning Agent, Label-Aware Explanation Agent, Intervention Agent) trained on small datasets (HatReDAug, MemeSense). This allows the main Large Multimodal Model (LMM) to perform complex reasoning tasks with only a few high-quality exemplar shots.

  1. Adaptive Exemplar Selection via Multimodal Embeddings:

The system utilizes task-agnostic multimodal embedding retrievers (SigLIP, CLIP, BLIP) to select the most relevant visual and textual examples for in-context learning, ensuring the prompt context is maximally informative for the specific classification/explanation/intervention needed.

  1. Nuanced Textual Output Control:

The system provides fine-grained control over the generated output based on the confidence of its prediction (e.g., generating detailed explanations only when correctly classified as hateful, or generating interventions that align with specific sentiment goals).

Specific System Functionality Details:

Component Input Required Output Generated Specific Improvement Over Baselines

:---:---:---:---

  1. Classification Agent (AC) & LMM (GPT-4o/Intern-VL3/Pixtral) Meme Image, Caption, OCR Text. Few-shot Exemplars. Binary Label: Hateful / Non-Hateful. Achieves state-of-the-art macro F1 scores on classification by leveraging few-shot context superior to fine-tuning methods (e.g., Vlm-Lim).

  2. Label-Aware Explanation Agent (AE) Meme Image, Caption, OCR Text + Ground Truth Label. A detailed explanation justifying the label based on definitions (30 words max). Generates explanations with higher semantic coherence and better linguistic diversity than models trained only on hateful data (HatReD). Better at explaining wrong predictions (wp cases).

  3. Intervention Agent (AI) Meme Image, Caption, OCR Text + Ground Truth Label. A specific moderation action/advice (e.g., Do not post publicly). Generates interventions with higher semantic coherence and better alignment with safety principles than models trained on general toxic data (MemeSense). Superior to GPT-4o in generating negative/restrictive intervention text for open models.

  4. Multi-Task Inference Engine Test Sample + Selected Exemplars (F) Integrated Triple Output: Label, Explanation, Intervention (if Hateful). The key novelty: Simultaneously performs C+E+I, resulting in a holistic moderation decision that is more robust than sequential or isolated systems.

In summary, the improved system moves beyond simple detection to provide a complete Detect-Explain-Intervene pipeline for hateful memes, offering superior performance and robustness across classification accuracy (via GPT-4o) and qualitative reasoning (via specialized agents).

Abstract

In this work, we examine hateful memes from three complementary angles - how to detect them, how to explain their content and how to intervene them before being posted - by applying a range of strategies built on top of generative AI models. To the best of our knowledge, explanation and intervention have typically been studied separately from detection, which does not reflect real-world conditions. Further, since curating large annotated datasets for meme moderation is prohibitively expensive, we propose a novel framework - PEST - that leverages task-specific generative VLMs and the few-shot adaptability of large VLMs to cater to different types of memes. We believe this is the first work focused on generalizable hateful meme moderation under limited data conditions, and has strong potential for deployment in real-world production scenarios. Warning: Contains potentially toxic contents.

Sources

Related papers