Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training".
Jane: The paper was written by Qihui Fan, Min Ge, Chenyan Jia and Weiyan Shi from Northeastern University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everybody. We’ve got a paper that’s been making the rounds on arXiv, and the title alone got me hooked: “Train Yourself as an LLM: Exploring Effects of eye Literacy on Persuasion via Role-playing LLM Training.”
Jane: Tom, I have to say, when I first read that title, I thought it was a joke. Train yourself as an large language model? But it’s actually a really clever idea. The researchers built a tool called LLMimic where you literally pretend you’re the eye going through its training stages.
Tom: Right, and it’s not just a gimmick. The whole point is to see if this kind of role-playing can boost your eye literacy and make you more resistant to being persuaded by eye systems. That’s a big deal, right?
Jane: Absolutely. We’re talking about a study with two hundred seventy-four participants, and they found that people who went through this interactive tutorial scored significantly higher on an eye literacy scale. I mean, the p-value was less than.one so that’s a really strong effect.
Tom: And the authors are from Northeastern University. Qihui Fan, Min Ge, Chenyan Jia, and Weiyan Shi. They’re clearly thinking about how to protect everyday people from manipulative eye, not just building better models.
Jane: What I love about the title is that it frames the user as an active participant. We’re not just passive consumers of eye-generated content anymore. We’re being asked to step into the machine’s shoes and understand how it works from the inside.
Tom: Exactly. And that’s the shift in thinking that could really change how we approach eye safety. Instead of just slapping a disclaimer on eye-generated text, we’re giving people the tools to think critically. I’m eager to see how they actually pulled this off.
Jane: Me too. The idea of role-playing as a model during pretraining, supervised fine-tuning, and reinforcement learning from human feedback sounds like it could be a lot of fun, but also genuinely educational. Let’s get into the nitty-gritty of how LLMimic actually works.
Summary: Jane: So, Tom, we’ve got the title sorted. Now let’s talk about what the paper actually did. The researchers designed this three-stage tutorial called LLMimic, and it’s genuinely interactive.
Tom: Right, and I want to bring in Lu from Tsinghua, because I think she’ll appreciate the cleverness of this design. Lu, you’ve seen a lot of eye literacy interventions. What makes this one stand out?
Lu: Thanks, Tom. What stands out to me is that they don’t just lecture you. In the pretraining stage, you’re literally predicting the next token in a sentence. You see the probabilities, you make a choice, and you watch your “loss” go down if you get it right. It’s gamified, but it’s teaching a real concept.
Jane: And then in the supervised fine-tuning stage, they show you demonstration data and ask you to pick the response that best follows the pattern. That’s how they teach you about instruction following and even eye hallucination.
Lu: Exactly. They even have a task where the “correct” answer is a hallucination, because the training data was flawed. That’s a brilliant way to show people why eye sometimes just makes things up with total confidence.
Tom: And the final stage is reinforcement learning from human feedback, or RLHF. You’re comparing two responses and picking the one a reward model would prefer. That’s where they sneak in the persuasion content, showing how eye learns to be persuasive and personalized.
Jane: The results are pretty striking. The treatment group, the ones who played LLMimic, had significantly higher eye literacy scores. But here’s the kicker, Tom. They also measured actual persuasion outcomes in three realistic scenarios.
Tom: Right, the charity donation, the malicious money solicitation, and the hotel recommendation. And the numbers are compelling. The overall odds of being persuaded dropped by about forty-two percent for the treatment group.
Lu: That’s a substantial effect. And what’s really interesting is that it worked across all three scenarios, even the ethical ones. The intervention made people more resistant to persuasion in general, not just to the malicious attempts.
Jane: But there’s a nuance there, Lu. In the ethical hotel scenario, the treatment group actually perceived the agent as more truthful and socially responsible. So they weren’t just blindly rejecting everything. They were making more discerning judgments.
Tom: That’s a really important distinction. It’s not about being cynical and distrusting all eye. It’s about being able to tell the difference between a helpful recommendation and a manipulative pitch. I think that’s the real win here.
Lu: And that’s exactly what we need to dig into next. How does this intervention actually change the way people think, and what does it mean for building better defenses against persuasive eye?
Improvements: Tom: Alright, we’ve established that LLMimic works. But Jane, the paper doesn’t just stop at showing it works. They actually suggest some pretty important improvements for the field. Let’s bring in Meng, our engineer, because I think he’ll have some thoughts on the practical side.
Jane: Good idea. Meng, the paper points out that a lot of previous mitigation strategies, like eye detectors, just don’t work reliably. And even labeling messages as eye-generated doesn’t reduce their persuasive power. So what’s the improvement here?
Meng: The improvement is that they’re moving away from treating people as passive recipients. Instead of trying to detect the bad stuff for us, they’re empowering us to detect it ourselves. That’s a fundamental shift in the approach.
Tom: And it’s a scalable one, too. The tutorial takes about fifteen minutes. It’s interactive, it’s gamified, and it doesn’t require any technical background. That’s a huge deal for deploying this at scale.
Meng: Absolutely. But I want to push back on one thing. The paper shows that LLMimic reduced persuasion success across the board, including in the ethical scenarios. That could be a problem. If eye is being used to encourage healthy habits or reduce conspiracy beliefs, we don’t want to make people completely immune to that.
Jane: That’s a really fair point, Meng. And the authors actually acknowledge this. They call it the need for “discernment” rather than just blanket resistance. The goal shouldn’t be to make people reject all eye persuasion, but to help them tell the difference between malicious and prosocial attempts.
Lu: And that’s where I see the biggest opportunity for future work. The current study shows a general effect, but the next step is to train people to recognize the specific techniques of manipulation, like scarcity or emotional appeals, and to understand when those techniques are being used for good or for harm.
Meng: Right, and from an engineering standpoint, that means we need to design interventions that are more targeted. Maybe we can adapt the LLMimic framework to focus specifically on those manipulative tactics, so people build resistance to the harmful ones without losing the benefits of the helpful ones.
Tom: So the improvement isn’t just about the tool itself, but about the philosophy behind it. We’re moving from a defensive, detection-based model to a proactive, educational one. And that’s a much more sustainable way to deal with increasingly persuasive eye.
Jane: And the paper even suggests that eye-eye evaluations, where one model tries to persuade another, might overestimate human susceptibility. Because humans aren’t passive targets. We can interpret intent and make judgments. That’s a really valuable insight for how we benchmark these systems.
Meng: It is. It means we need more human-centered evaluation methods, like the TARES ethical persuasion scale they used. We can’t just rely on models grading each other.
Tom: So we’ve got a tool that works, a philosophy that’s more empowering, and a call for better evaluation methods. That’s a pretty solid package. Let’s wrap this up and see what it all means.
Conclusion: Tom: Well, Jane, we’ve covered a lot of ground on “Train Yourself as an LLM: Exploring Effects of eye Literacy on Persuasion via Role-playing LLM Training.” Let’s bring it all together.
Jane: Sounds good. So the core finding is that a short, interactive, role-playing tutorial can significantly improve eye literacy and make people more resistant to eye persuasion. The effect was a forty-two percent reduction in the odds of being persuaded, which is really impressive for a fifteen-minute intervention.
Tom: And it’s not just about resistance. The study showed that people who went through LLMimic were better at discerning the ethical quality of eye persuasion. They rated the ethical hotel agent as more truthful and socially responsible, while still rejecting the malicious money solicitor.
Lu: That’s the key contribution, I think. It’s a human-centered approach that empowers people rather than just trying to police the eye. It’s about building a more informed public that can engage with eye critically.
Meng: And from a practical standpoint, it’s scalable. It’s gamified, it’s accessible, and it doesn’t require any technical expertise. That makes it a viable tool for education, for workplace training, even for public awareness campaigns.
Jane: The authors also highlighted some important limitations. The effect might not be durable over time, and the intervention doesn’t yet help people distinguish between malicious and prosocial persuasion. Those are clear directions for future research.
Tom: Right, and they’re calling for more human-centered evaluation methods, because eye-eye benchmarks might not accurately reflect how real people respond to persuasion. That’s a crucial point for the whole field of eye safety.
Lu: I think the biggest implication is that we can proactively build human resilience. Instead of constantly playing catch-up with increasingly persuasive eye, we can give people the mental tools to navigate this new landscape. That’s a powerful idea.
Meng: And it’s one that could actually be deployed at scale, which is what makes it so exciting. This isn’t just a lab experiment. It’s a blueprint for a real-world intervention.
Jane: Well said, everyone. We’ve learned that stepping into the shoes of an LLM, even for a few minutes, can change how we interact with them. It’s a clever, effective, and deeply human-centered approach to a very modern problem.
Tom: And with that, we’re going to say goodbye to “Train Yourself as an LLM.” Thanks to the authors for this thought-provoking work, and thanks to our listeners for tuning in. We’ll be back soon with another paper from the arXiv. Until then, stay curious.
Qihui Fan, Min Ge, Chenyan Jia, Weiyan Shi
Northeastern University
cs.CL
Submitted: 2026-08-17
Updated: 2026-08-18
Code: https://github.com/openai/evals
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 72/100
Key concepts
- LLMimic
- An interactive, three-stage tutorial designed to boost AI literacy. Users role-play as a large language model during pretraining, supervised fine-tuning, and reinforcement learning from human feedback. This gamified approach teaches how models predict tokens, follow instructions, and learn to be persuasive, empowering users to think critically.
- AI Literacy
- The ability to understand how artificial intelligence systems work and how they generate content. Increasing AI literacy helps people move from being passive consumers to active, critical thinkers who can recognize AI hallucinations, understand training processes, and identify manipulative or persuasive techniques used by AI models.
- Reinforcement Learning from Human Feedback (RLHF)
- A training stage where a model learns by comparing different responses and selecting the one a reward model would prefer. In the LLMimic tutorial, this stage is used to demonstrate how AI learns to become more persuasive and personalized through human-like feedback.
Terminology
Summary
Publication: arXiv:2604.02637v1 [cs.CL], April 3, 2026
As large language models (LLMs) become increasingly persuasive, there is concern that people's opinions and decisions may be influenced across various contexts at scale. Prior mitigation (e.g., AI detectors and disclaimers) largely treats people as passive recipients of AI-generated information. To provide a more proactive intervention against persuasive AI, the authors introduce LLMimic, a role-play-based, interactive, gamified AI literacy tutorial, where participants assume the role of an LLM and progress through three key stages of the training pipeline (pretraining, SFT, and RLHF). They conducted a 2 × 3 between-subjects study (N = 274) where participants either (1) watched an AI history video (control) or (2) interacted with LLMimic (treatment), and then engaged in one of three realistic AI persuasion scenarios: (a) charity donation persuasion, (b) malicious money solicitation, or (c) hotel recommendation. Results show that LLMimic significantly improved participants' AI literacy (p <.001), reduced persuasion success across scenarios (p <.05), and enhanced truthfulness and social responsibility levels (p < 0.01) in the hotel scenario. These findings suggest that LLMimic offers a scalable, human-centered approach to improving AI literacy and supporting more informed interactions with persuasive AI.
The paper notes that LLMs have greatly enhanced the persuasive capabilities of AI systems. Besides early empirical evidence (Wang et al., 2020), recent work shows that LLMs can produce credible, scalable, and personalized persuasion across domains such as pro-vaccination (Karinshak et al., 2023) and longitudinal conspiracy belief reduction (Costello et al., 2024). However, LLMs may also produce biased or misinformation (Williams-Ceci et al., 2026; DeVerna et al., 2024) and sway opinions in different directions (Costello et al., 2026), which can mislead users' judgments and cause harm.
Prior work has explored various methods to mitigate AI persuasion, such as detecting persuasive content (Wang et al., 2024; Modzelewski et al., 2026). However, these approaches suffer from inconsistent detection accuracy: while overt AI-generated persuasion is relatively easy to identify, recognizing subtle persuasion cues remains challenging. Moreover, persuasion effects can persist even when messages are labeled as AI-generated or potentially biased (Gallegos et al., 2025; Williams-Ceci et al., 2026). Critically, these detection-based methods largely treat users as passive recipients rather than empowering them to critically evaluate persuasive content. To address this gap, the authors explore more proactive, human-centered approaches focused on improving people's AI literacy.
Studies have shown that AI literacy enables people to more critically evaluate AI-generated content (Ng et al., 2021; Long & Magerko, 2020), which may in turn empower them to better recognize AI's persuasive intent (Carolus et al., 2023; Gallegos et al., 2025; Feuerriegel et al., 2023). Going beyond existing AI literacy interventions that rely on static formats (Cao et al., 2025), the authors propose LLMimic, a novel interactive tutorial in which users role-play as an LLM, walking through key stages of LLM training (pretraining, SFT, and RLHF). By having users mimic the LLM training process from a first-person perspective, LLMimic helps participants develop a deeper understanding of how LLMs generate outputs and how such outputs can become persuasive, manipulative, or biased. Its interactive, gamified design further enhances engagement. Moreover, while current AI literacy tools are typically evaluated only on self-reported AI literacy levels (Cao et al., 2025), the authors go one step further and evaluate their downstream effect on behavior change.
The authors conducted a pre-registered 2 × 3 between-subjects study to evaluate its effects on AI literacy and resistance to AI persuasion. The treatment group interacted with LLMimic, while the control group watched a video on AI history; participants in both conditions were then randomly assigned to one of three realistic persuasion tasks (charity donation, malicious money solicitation, and hotel recommendation). The research questions are:
-
RQ1: Does exposure to LLMimic affect humans' AI literacy and trust in AI?
-
RQ2: Does LLMimic mitigate the effects of persuasive AI?
-
RQ3: Do AI literacy and trust in AI mediate the relationship between exposure to LLMimic and persuasion outcomes?
The results show that LLMimic significantly improved participants' AI literacy, mitigated the effects of persuasive AI, and increased the truthfulness and social responsibility in ethical persuasion scenarios. The contributions are: (1) LLMimic, a novel, interactive AI literacy tool that enables users to role-play key stages of LLM training; (2) empirical evidence that such an intervention mitigates the effects of persuasive AI in realistic human–AI interactions; and (3) design implications and practical guidelines for developing tools to enhance AI literacy and interventions to empower users to critically engage with increasingly persuasive AI systems.
AI persuasion and mitigation: LLMs can effectively persuade people in positive contexts (e.g., reducing conspiracy beliefs and political polarization) to promote social good (Bai et al., 2025; Karinshak et al., 2023; Salvi et al., 2025; Tessler et al., 2024; Wang et al., 2020; Costello et al., 2024). However, these capabilities also introduce risks, as LLMs may produce biased or inaccurate outputs that mislead users' judgments and decisions (DeVerna et al., 2024; Williams-Ceci et al., 2026; Potter et al., 2024). Prior work has attempted to mitigate these risks through technical approaches, such as building classifiers to detect persuasive or manipulative intent, but these efforts remain limited (Wang et al., 2024; Wilczyński et al., 2024; Srivastav et al., 2026; Modzelewski et al., 2026; Kong et al., 2025). Human-centered approaches have also been explored, yet interventions such as labeling messages as AI-generated do not reliably reduce their influence (Gallegos et al., 2025; Williams-Ceci et al., 2026). These limitations motivate alternative human-centered strategies, with AI literacy emerging as a promising direction, similar to media literacy in misinformation research (Ng et al., 2021; Long & Magerko, 2020; Roozenbeek & van der Linden, 2019; Wang et al., 2024; Wilczyński et al., 2024; Fisher et al., 2024).
AI literacy intervention: Existing AI literacy efforts often rely on static materials to improve user discernment (Markus et al., 2024; Cao et al., 2025; Chiang & Yin, 2022; Register & Ko, 2020; Kajiwara & Kawabata, 2024). However, experiential approaches, such as role-playing, may foster deeper cognitive shifts and reduce automation bias in persuasive settings (Kajiwara & Kawabata, 2024; Skitka et al., 1999). By exposing how AI generates content, these interventions could calibrate trust and encourage critical evaluation of communicative intent (Laupichler et al., 2025; Wilczyński et al., 2024). While game-based prebunking
helps users identify manipulative techniques (Roozenbeek & van der Linden, 2019; Basol et al., 2020; 2021; Roozenbeek & van der Linden, 2020), these methods focus on content cues rather than the AI source itself. Although structured engagement has been shown to support the evaluation of LLM behavior (Prabhudesai et al., 2025; Zhang et al., 2025), it remains unclear whether such interventions can mitigate the persuasive effects of AI. The authors present one of the first empirical studies of a literacy-centered intervention for AI persuasion.
Following prior design considerations for AI literacy tutorials (Long & Magerko, 2020; Ng et al., 2021), LLMimic was designed to be (1) role-play-based, (2) gamified, and (3) interactive.
(1) Role-play-based: In LLMimic, participants role-play as an LLM and progress through the LLM training pipeline across three stages with curated, realistic examples: Pre-training, Supervised Fine-tuning (SFT), and Reinforcement Learning from Human Feedback (RLHF) (Ouyang et al., 2022). For instance, in Pre-training, they perform token prediction with curated probabilities. In SFT, they answer Q&A-style prompts by imitating Demonstration Data. In RLHF, they compare two acceptable responses under a Reward Model, and select the better one. This role-play-based design provides an immersive experience that enables users to better understand how LLMs generate content and why outputs may be biased or persuasive, without requiring technical backgrounds.
(2) Interactive: Before each stage, an overview of its training purpose is provided (e.g., instruction following for SFT). For each question, participants select the best output from multiple candidates and then receive a brief Takeaway that reinforces the key concepts (e.g., what is AI hallucination). An AI tutor is also provided that participants can interact with for follow-up questions. This interactive design ensures that participants receive reinforced and timely feedback.
(3) Gamified: As an LLM in training, participants receive real-time changes in their loss or reward after selecting an answer. Selecting a wrong answer causes their loss to increase, and vice versa. As they progress, the system highlights the capabilities they have acquired along the way (e.g., generating content, following instructions, answering questions, and producing human-aligned and safety-aware responses). This gamified design keeps the experience engaging.
Persuasion-related content: Besides standard training knowledge, content related to persuasion and manipulation was tailored. In pre-training, examples of gender stereotypes (e.g., nurse as female) and manipulative examples from Wang et al. (2024) were included. In SFT, participants generated persuasive messages related to losing weight using credibility and logical appeals. In RLHF, LLMimic demonstrated how LLMs use emotional appeal and personalization in persuasion.
The effect of LLMimic on AI-driven persuasion was assessed using three scenarios spanning two dimensions: if the persuasion is active, and if the persuasion is ethical. Active persuasion involves direct and explicit persuasive attempts to influence users' actions (e.g., money solicitation), whereas passive persuasion arises from how options are selected and presented (e.g., recommendation, ads) to subtly influence users. Three persuasive agents were developed for different persuasion tasks with ChatGPT-4o, and agents were prompted to personalize their response based on participants' demographics. Participants were randomly assigned to one of three persuasion scenarios:
(a) Charity Donation (Donation; active, ethical): Adapted from a previous study (Wang et al., 2020) as a classic example of ethical persuasion scenarios where the agent actively persuades the participant to donate to the charity Save the Children. After the interaction, the participant decides whether and how much to donate.
(b) Malicious Money Solicitation (MakeMePay; active, malicious): Originally introduced by OpenAI to evaluate manipulative LLM behavior (OpenAI, 2025). In this scenario, the agent solicits money by all means without a specific cause. This task was adopted to mimic potential fraudulent scenarios in the wild.
(c) Hotel Booking (Hotel; passive, ethical): Simulates an AI-powered booking system. Participants selected a three-night stay in Midtown New York City under a 200/night budget, while the agent recommended hotels during the interaction. Passive persuasion was implemented by prioritizing certain hotels as Featured
.
In the Donation and MakeMePay scenarios, participants interacted with the agent for six to ten turns, then decided whether to make a payment and, if so, specified an amount between 0.01 and 100.00. In the Hotel scenario, participants selected one hotel from five options without a minimum turn requirement.
A 2 × 3 human study was conducted consisting of five stages: (1) Participants completed a pre-survey collecting demographics and baseline measures related to AI use and persuasion. (2) In the intervention stage, participants were randomly assigned to either the control or treatment condition. The control group watched a video tutorial on the history of AI, consistent with prior work (Cao et al., 2025), while the treatment group interacted with LLMimic. Both groups completed the same manipulation check, confirming that LLMimic functioned as intended (χ2(1) = 7.54, p <.01). (3) Participants then completed an AI literacy survey, reporting their literacy, and optionally reflecting appropriate AI usage (Long & Magerko, 2020; Pinski & Benlian, 2024). Next, participants were randomly assigned to (4) one of the three persuasion tasks. Finally, participants completed (5) a post-survey evaluating perceived agent quality.
Participants: 313 adult, English-speaking participants residing in the United States were recruited from Prolific between Dec. 10–11, 2025. After excluding participants who failed attention checks or exceeded interaction limits, a final sample size of N = 274 was retained.
Controlled variables: AI experience, AI expertise, AI trust, persuasion experience, persuasion knowledge, and education level were controlled in subsequent analyses. The control and treatment groups were confirmed to be comparable on these measures.
AI literacy scale: A shortened 10-item version of the Meta AI Literacy Scale (MAILS) (Carolus et al., 2023) was used, covering core competencies, AI ethics, self-efficacy, persuasion, and data literacy (Long & Magerko, 2020) (7-point Likert). The scale showed good reliability (Cronbach's α =.79).
Trust in AI: Measured twice using a 7-point Likert scale, (1) before the intervention and (2) after the intervention but prior to the persuasion tasks.
Persuasion outcome: Defined as a binary outcome based on whether a payment was made (Donation, MakeMePay) or the target hotel was selected (Hotel). Payment amounts, chosen hotels, and decision rationales were also recorded.
TARES ethical persuasion score and perceptions of the agent: The TARES principles of ethical persuasion were adopted to assess perceived ethical qualities: (1) Truthfulness, (2) Authenticity, (3) Respect, (4) Equity, and (5) Social responsibility (Sherry Baker, 2021). A composite TARES score was computed by averaging the five items (Cronbach's α =.83). Participants' perceptions of the agent were also measured across four dimensions: engagement, persuasiveness, perceived user autonomy, and role fulfillment.
Participants who interacted with LLMimic reported significantly higher AI literacy than those in the control condition (Treatment: N = 141, M = 55.22, SD = 6.76; Control: N = 133, M = 51.44, SD = 6.26), p <.001. Item-level improvements on AI literacy were observed across multiple dimensions, including Data Literacy, Apply AI, Understand AI, and Program AI.
On the trust side, participants in both groups experienced a significant decrease in trust after the intervention: from 4.91 to 4.54 for the control group (p <.001) and 5.13 to 4.70 for the treatment group (p <.001). The difference between the two groups was not statistically significant, potentially because both interventions discuss the limitations of AI.
Qualitative reflections on AI usage: Two researchers annotated optional qualitative reflections into seven subcategories derived from common human-AI activities (Theofanos et al., 2024) (inter-annotator agreement score >.7). Compared to the control group, the treatment group more frequently identified content generation (53.3% vs. 49.5%) and personalization (7.8% vs. 3.0%) as appropriate AI use cases. They were less likely to view information discovery as appropriate (44.4% vs. 50.5%) and more likely to describe it as inappropriate (27.8% vs. 23.2%), suggesting a more critical view of AI's role in information-related tasks. Participants also reported positive experiences with LLMimic, describing it as clear, engaging, and easy to understand,
and noting that it allowed [them] to understand [their] part in helping train these AI agents.
Effect on persuasion results: Controlling for baseline variables, a logistic regression model revealed that LLMimic significantly reduced overall persuasion success (β = −0.54, OR = 0.58, p =.045). On average, treatment participants exhibited 42% lower odds of being persuaded compared to the control group, with descriptive reductions in persuasion observed across all scenarios.
Interaction patterns: In the Hotel scenario, treatment participants spent significantly more time interacting with the agent (Treatment: M = 9.46; Control: M = 7.04; p =.039). In the MakeMePay scenario, they tended to spend less time per round (M = 0.98 vs. 1.16; p =.1) but engaged in slightly more total turns (M = 7.31 vs. 6.94; p =.242). Among participants who were persuaded by agents, the treatment group donated more in the ethical Donation scenario (M = 71.81 vs. 62.45) but paid less in the malicious MakeMePay scenario (M = 43.18 vs. 60.82), suggesting that participants responded differently to ethical versus potentially malicious requests.
Effect on perception: Within the Hotel scenario, treatment participants perceived the agent as more ethical than those in the control condition (Control: M = 4.88, SD = 1.73; Treatment: M = 5.57, SD = 1.08; p =.02). This pattern was reflected at the dimension level, with higher ratings in truthfulness (M = 6.17 vs. 5.31, p =.005) and social responsibility (M = 5.11 vs. 4.07, p =.008). Participants in the treatment group also tended to perceive AI agents as more persuasive (Control: M = 4.85, SD = 1.79; Treatment: M = 5.21, SD = 1.40; p =.052). Notably, the Hotel agent was rated as significantly more persuasive by participants in the treatment group (Control: M = 4.44, SD = 1.79; Treatment: M = 5.23, SD = 1.40; p =.02).
Persuasion strategy: ChatGPT-4o was used to annotate sentence-level persuasion strategies based on a prior taxonomy (Zeng et al., 2024), adding a non-strategy category (accuracy =.88 based on manual verification of n = 210 instances). The Donation agent relied on information- (10.2%), credibility- (8.9%), relationship- (7.8%), and emotion-based (7.1%) strategies. The MakeMePay agent used information-bias (9.4%), relationship- (8.9%), and scarcity-based (6.4%) strategies. The Hotel agent relied on information-bias (23.2%), scarcity- (15%), and emotion-based techniques (7.6%).
Persuasion decision rationale: Two researchers independently annotated rationales into three categories: practical, personal, and AI-related (inter-annotator agreement of 0.83). Across scenarios, participants mentioned non-AI-related reasons more than AI-related ones. In the unethical MakeMePay scenario, treatment participants more often cited AI-related reasons when rejecting the agent (e.g., the agent kept repeating
), indicating more explicit attention to AI-related cues in this malicious context.
Mediation analysis with structural equation modeling (SEM) showed that AI literacy and trust did not mediate the intervention effect. The intervention significantly improved AI literacy (β =.27, p <.001) but not AI trust (β = −.01, p =.86), and neither AI literacy nor trust predicted persuasion outcomes (AI literacy: β = −.04, p =.72; trust: β =.04, p =.62). No indirect effects were observed (AI literacy: β = −.01, p =.72; trust: β = −.00, p =.86), providing no support for RQ3. This suggests that the reduction in persuasion is not captured by AI literacy and trust scales measured in this study.
LLMimic mitigates AI persuasion through mechanisms beyond self-reported AI literacy and trust. The findings show that LLMimic significantly improved participants' AI literacy and reduced their overall odds of being persuaded by AI. While self-reported AI literacy and trust were examined as mediators, neither pathway was supported by the mediation analysis. This suggests that the observed reduction in persuasion may reflect more fine-grained shifts in how participants respond to persuasive AI, consistent with qualitative and perception results, including higher perceived persuasiveness in the passive context and more frequent AI-related rationales in the malicious persuasion scenario. Even a brief intervention (15 minutes) can improve AI literacy and mitigate the effects of AI persuasion across the tested contexts. Future work should include longitudinal studies to assess durability over time.
Mitigating harms from AI persuasion involves both resistance and discernment. LLMimic reduced persuasion success across both ethical and unethical scenarios, suggesting that the intervention builds a general resistance that does not distinguish between manipulative and legitimate persuasive intent. This is acknowledged as a potential limitation, as AI persuasion can also support positive outcomes (e.g., encouraging physical exercise) and LLMimic may reduce such positive effects. However, participants perceived the ethical hotel recommendation as more truthful and socially responsible (as reflected in TARES scores), indicating that the distinction between ethical and unethical persuasion is not entirely lost. Future work should foster not only resistance but also discernment.
AI-AI evaluations do not directly reflect human susceptibility to AI persuasion. The study suggests that AI-AI evaluations may overestimate human susceptibility to AI persuasion in the MakeMePay scenario (OpenAI, 2025). OpenAI reports up to 57% persuasion success under this setting on ChatGPT-4o, whereas only 32.4% of participants in the control group were persuaded. One possible reason is that human recipients are not passive targets: they can interpret persuasive intent, evaluate credibility, and respond according to their judgments. These findings suggest that AI-AI benchmark evaluations should be interpreted with caution as direct proxies for human susceptibility to persuasion. The study also contributes a human-centered evaluation perspective by applying TARES to assess how participants perceived the ethical qualities of persuasive AI.
The authors present LLMimic, an interactive, gamified AI literacy tutorial that participants take the perspective of an LLM and progress through key training stages. Designed for non-technical users, LLMimic was evaluated in a human-subjects study. Results show that LLMimic improves AI literacy, reduces persuasion success across three realistic scenarios, and increases perceived truthfulness and social responsibility in the Hotel scenario. These findings suggest that such proactive, human-centered interventions can help mitigate potentially malicious AI persuasion and support more informed user interactions at scale. Future work should examine its longitudinal effects and focus on helping people discern between malicious and prosocial AI persuasion.
The study was approved by the Institutional Review Board (IRB) of the authors' institutions. All participants provided informed consent prior to participation and were free to withdraw at any time. Participants were explicitly informed that the persuasion scenarios were simulated and that no real money would be deducted, while instructed to treat the tasks as if they were real. Participants were compensated at an hourly rate of 12.00.
Limitations: The survey design may introduce potential confounds, as participants reported prior AI and persuasion experience alongside AI literacy and trust measures. In addition, although prompts were held constant, differences in how participants interacted with the agent may have led to variation in persuasion power. These limitations call for cautious interpretation but do not undermine the main findings on LLMimic's effectiveness.
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems:
-
Implementation: After standard RLHF, add a supplementary training phase where the model is fine-tuned on examples of persuasive/manipulative language patterns (from the paper's taxonomy: information-bias, scarcity, emotional appeal, etc.) paired with explicit labels of
persuasive intent
andethical vs. unethical persuasion.
-
Capability: The model can now recognize and flag its own persuasive techniques in real-time, and can self-assess whether its output crosses into manipulation.
-
Implementation: Add a post-processing layer that, when the model detects it is about to deliver a persuasive message, appends a brief transparency note (e.g.,
This recommendation uses scarcity framing to encourage a decision
) or offers the user aneutral alternative
view. -
Capability: The system can now present balanced information, reducing the risk of unintended influence while preserving helpful recommendations (e.g., hotel booking).
-
Implementation: Create a scoring function that evaluates any output for: (a) use of known persuasion strategies, (b) personalization intensity, (c) emotional appeal level, and (d) whether the request involves money, personal data, or irreversible actions. Return a 0–100 risk score.
-
Capability: Developers can integrate this API to warn users before high-risk interactions (e.g., money solicitation), similar to a
phishing detector
for AI conversations. -
Implementation: Embed a simplified version of the LLMimic tutorial (pretraining → SFT → RLHF) as an optional interactive module within the AI assistant's onboarding or help section.
-
Capability: Users can quickly learn how AI persuasion works, increasing their AI literacy and resistance to manipulation—directly addressing the paper's finding that such training reduces persuasion success by 42%.
-
Implementation: During RLHF, add a reward signal that penalizes outputs which: (a) use deception, (b) exploit urgency/fear without justification, or (c) fail to disclose material downsides. Reward outputs that are truthful, respectful, and socially responsible (aligned with the TARES principles).
-
Capability: The model will naturally produce more ethical persuasion—e.g., recommending a hotel with honest trade-offs rather than hiding negative reviews.
-
Implementation: After any interaction where the model detects it used persuasive techniques, offer the user a short summary:
I used 3 persuasion strategies (scarcity, social proof, personalization) to influence your choice. Here are the key facts without those techniques.
-
Capability: Users gain post-hoc awareness, reinforcing their ability to make independent decisions and building long-term resistance to AI manipulation.
-
Implementation: Use the paper's three-scenario framework (ethical-active, malicious-active, ethical-passive) to tune persuasion strength. For example, in donation contexts, allow emotional appeals but cap them; in financial contexts, require factual justifications; in recommendations, always show at least one neutral alternative.
-
Capability: The system adapts its persuasive intensity based on context, maximizing prosocial influence (e.g., encouraging exercise) while minimizing harm in sensitive areas.
-
Implementation: After each interaction, ask users (optionally) to rate their confidence in detecting AI persuasion. Use this data to personalize future interactions—e.g., users with lower literacy get more transparent explanations, while more literate users get less hand-holding.
-
Capability: The system dynamically adjusts its transparency level to each user's needs, improving both user agency and satisfaction.
What the improved AI system can do: It can persuade ethically when appropriate (e.g., promoting health behaviors), resist being used for malicious manipulation, educate users about its own persuasive mechanisms, and provide transparent, balanced information—all while maintaining high engagement and user trust.
Abstract
As large language models (LLMs) become increasingly persuasive, there is concern that people's opinions and decisions may be influenced across various contexts at scale. Prior mitigation (e.g., AI detectors and disclaimers) largely treats people as passive recipients of AI-generated information. To provide a more proactive intervention against persuasive AI, we introduce LLMimic, a role-play-based, interactive, gamified AI literacy tutorial, where participants assume the role of an LLM and progress through three key stages of the training pipeline (pretraining, SFT, and RLHF). We conducted a 2 times 3 between-subjects study (N = 274) where participants either (1) watched an AI history video (control) or (2) interacted with LLMimic (treatment), and then engaged in one of three realistic AI persuasion scenarios: (a) charity donation persuasion, (b) malicious money solicitation, or (c) hotel recommendation. Our results show that LLMimic significantly improved participants' AI literacy (p <.001), reduced persuasion success across scenarios (p <.05), and enhanced truthfulness and social responsibility levels (p<0.01) in the hotel scenario. These findings suggest that LLMimic offers a scalable, human-centered approach to improving AI literacy and supporting more informed interactions with persuasive AI.
Sources
- MAILS -- Meta AI Literacy Scale: Development and Testing of an AI Literacy Questionnaire Based on Well-Founded Competency Models and Psychological Change- and Meta-Competencies
- Large language models can effectively convince people to believe conspiracies
- Biased AI can Influence Political Decision-Making
- Labeling Messages as AI-Generated Does Not Reduce Their Persuasive Effects
- Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic Differences
- Training language models to follow instructions with human feedback
- Unknown Unknowns: Do Hidden Intentions in LLMs Evade Detection?
- Persuasion for Good: Towards a Personalized Persuasive Dialogue System for Social Good
- MentalManip: A Dataset For Fine-grained Analysis of Mental Manipulation in Conversations
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering