Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training

summary

Video file (mp4)

In short

Researchers developed LLMimic, an interactive tutorial where users role-play as an LLM through pretraining, fine-tuning, and RLHF stages. The study found that this training significantly boosts AI literacy and reduces susceptibility to AI persuasion by 42%, helping users better discern between helpful recommendations and manipulative content.

Key concepts

LLMimic
An interactive, three-stage tutorial designed to boost AI literacy. Users role-play as a large language model during pretraining, supervised fine-tuning, and reinforcement learning from human feedback. This gamified approach teaches how models predict tokens, follow instructions, and learn to be persuasive, empowering users to think critically.
AI Literacy
The ability to understand how artificial intelligence systems work and how they generate content. Increasing AI literacy helps people move from being passive consumers to active, critical thinkers who can recognize AI hallucinations, understand training processes, and identify manipulative or persuasive techniques used by AI models.
Reinforcement Learning from Human Feedback (RLHF)
A training stage where a model learns by comparing different responses and selecting the one a reward model would prefer. In the LLMimic tutorial, this stage is used to demonstrate how AI learns to become more persuasive and personalized through human-like feedback.

Terminology used across episodes

This episode discusses

The paper

Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training · Read on arXiv

Qihui Fan, Min Ge, Chenyan Jia, Weiyan Shi

Northeastern University

As large language models (LLMs) become increasingly persuasive, there is concern that people's opinions and decisions may be influenced across various contexts at scale. Prior mitigation (e.g., AI detectors and disclaimers) largely treats people as passive recipients of AI-generated information. To provide a more proactive intervention against persuasive AI, we introduce LLMimic, a role-play-based, interactive, gamified AI literacy tutorial, where participants assume the role of an LLM and progress through three key stages of the training pipeline (pretraining, SFT, and RLHF). We conducted a 2 times 3 between-subjects study (N = 274) where participants either (1) watched an AI history video (control) or (2) interacted with LLMimic (treatment), and then engaged in one of three realistic AI persuasion scenarios: (a) charity donation persuasion, (b) malicious money solicitation, or (c) hotel recommendation. Our results show that LLMimic significantly improved participants' AI literacy (p <.001), reduced persuasion success across scenarios (p <.05), and enhanced truthfulness and social responsibility levels (p<0.01) in the hotel scenario. These findings suggest that LLMimic offers a scalable, human-centered approach to improving AI literacy and supporting more informed interactions with persuasive AI.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training".

Jane: The paper was written by Qihui Fan, Min Ge, Chenyan Jia and Weiyan Shi from Northeastern University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. We’ve got a paper that’s been making the rounds on arXiv, and the title alone got me hooked: “Train Yourself as an LLM: Exploring Effects of eye Literacy on Persuasion via Role-playing LLM Training.”

Jane: Tom, I have to say, when I first read that title, I thought it was a joke. Train yourself as an large language model? But it’s actually a really clever idea. The researchers built a tool called LLMimic where you literally pretend you’re the eye going through its training stages.

Tom: Right, and it’s not just a gimmick. The whole point is to see if this kind of role-playing can boost your eye literacy and make you more resistant to being persuaded by eye systems. That’s a big deal, right?

Jane: Absolutely. We’re talking about a study with two hundred seventy-four participants, and they found that people who went through this interactive tutorial scored significantly higher on an eye literacy scale. I mean, the p-value was less than.one so that’s a really strong effect.

Tom: And the authors are from Northeastern University. Qihui Fan, Min Ge, Chenyan Jia, and Weiyan Shi. They’re clearly thinking about how to protect everyday people from manipulative eye, not just building better models.

Jane: What I love about the title is that it frames the user as an active participant. We’re not just passive consumers of eye-generated content anymore. We’re being asked to step into the machine’s shoes and understand how it works from the inside.

Tom: Exactly. And that’s the shift in thinking that could really change how we approach eye safety. Instead of just slapping a disclaimer on eye-generated text, we’re giving people the tools to think critically. I’m eager to see how they actually pulled this off.

Jane: Me too. The idea of role-playing as a model during pretraining, supervised fine-tuning, and reinforcement learning from human feedback sounds like it could be a lot of fun, but also genuinely educational. Let’s get into the nitty-gritty of how LLMimic actually works.

Summary: Jane: So, Tom, we’ve got the title sorted. Now let’s talk about what the paper actually did. The researchers designed this three-stage tutorial called LLMimic, and it’s genuinely interactive.

Tom: Right, and I want to bring in Lu from Tsinghua, because I think she’ll appreciate the cleverness of this design. Lu, you’ve seen a lot of eye literacy interventions. What makes this one stand out?

Lu: Thanks, Tom. What stands out to me is that they don’t just lecture you. In the pretraining stage, you’re literally predicting the next token in a sentence. You see the probabilities, you make a choice, and you watch your “loss” go down if you get it right. It’s gamified, but it’s teaching a real concept.

Jane: And then in the supervised fine-tuning stage, they show you demonstration data and ask you to pick the response that best follows the pattern. That’s how they teach you about instruction following and even eye hallucination.

Lu: Exactly. They even have a task where the “correct” answer is a hallucination, because the training data was flawed. That’s a brilliant way to show people why eye sometimes just makes things up with total confidence.

Tom: And the final stage is reinforcement learning from human feedback, or RLHF. You’re comparing two responses and picking the one a reward model would prefer. That’s where they sneak in the persuasion content, showing how eye learns to be persuasive and personalized.

Jane: The results are pretty striking. The treatment group, the ones who played LLMimic, had significantly higher eye literacy scores. But here’s the kicker, Tom. They also measured actual persuasion outcomes in three realistic scenarios.

Tom: Right, the charity donation, the malicious money solicitation, and the hotel recommendation. And the numbers are compelling. The overall odds of being persuaded dropped by about forty-two percent for the treatment group.

Lu: That’s a substantial effect. And what’s really interesting is that it worked across all three scenarios, even the ethical ones. The intervention made people more resistant to persuasion in general, not just to the malicious attempts.

Jane: But there’s a nuance there, Lu. In the ethical hotel scenario, the treatment group actually perceived the agent as more truthful and socially responsible. So they weren’t just blindly rejecting everything. They were making more discerning judgments.

Tom: That’s a really important distinction. It’s not about being cynical and distrusting all eye. It’s about being able to tell the difference between a helpful recommendation and a manipulative pitch. I think that’s the real win here.

Lu: And that’s exactly what we need to dig into next. How does this intervention actually change the way people think, and what does it mean for building better defenses against persuasive eye?

Improvements: Tom: Alright, we’ve established that LLMimic works. But Jane, the paper doesn’t just stop at showing it works. They actually suggest some pretty important improvements for the field. Let’s bring in Meng, our engineer, because I think he’ll have some thoughts on the practical side.

Jane: Good idea. Meng, the paper points out that a lot of previous mitigation strategies, like eye detectors, just don’t work reliably. And even labeling messages as eye-generated doesn’t reduce their persuasive power. So what’s the improvement here?

Meng: The improvement is that they’re moving away from treating people as passive recipients. Instead of trying to detect the bad stuff for us, they’re empowering us to detect it ourselves. That’s a fundamental shift in the approach.

Tom: And it’s a scalable one, too. The tutorial takes about fifteen minutes. It’s interactive, it’s gamified, and it doesn’t require any technical background. That’s a huge deal for deploying this at scale.

Meng: Absolutely. But I want to push back on one thing. The paper shows that LLMimic reduced persuasion success across the board, including in the ethical scenarios. That could be a problem. If eye is being used to encourage healthy habits or reduce conspiracy beliefs, we don’t want to make people completely immune to that.

Jane: That’s a really fair point, Meng. And the authors actually acknowledge this. They call it the need for “discernment” rather than just blanket resistance. The goal shouldn’t be to make people reject all eye persuasion, but to help them tell the difference between malicious and prosocial attempts.

Lu: And that’s where I see the biggest opportunity for future work. The current study shows a general effect, but the next step is to train people to recognize the specific techniques of manipulation, like scarcity or emotional appeals, and to understand when those techniques are being used for good or for harm.

Meng: Right, and from an engineering standpoint, that means we need to design interventions that are more targeted. Maybe we can adapt the LLMimic framework to focus specifically on those manipulative tactics, so people build resistance to the harmful ones without losing the benefits of the helpful ones.

Tom: So the improvement isn’t just about the tool itself, but about the philosophy behind it. We’re moving from a defensive, detection-based model to a proactive, educational one. And that’s a much more sustainable way to deal with increasingly persuasive eye.

Jane: And the paper even suggests that eye-eye evaluations, where one model tries to persuade another, might overestimate human susceptibility. Because humans aren’t passive targets. We can interpret intent and make judgments. That’s a really valuable insight for how we benchmark these systems.

Meng: It is. It means we need more human-centered evaluation methods, like the TARES ethical persuasion scale they used. We can’t just rely on models grading each other.

Tom: So we’ve got a tool that works, a philosophy that’s more empowering, and a call for better evaluation methods. That’s a pretty solid package. Let’s wrap this up and see what it all means.

Conclusion: Tom: Well, Jane, we’ve covered a lot of ground on “Train Yourself as an LLM: Exploring Effects of eye Literacy on Persuasion via Role-playing LLM Training.” Let’s bring it all together.

Jane: Sounds good. So the core finding is that a short, interactive, role-playing tutorial can significantly improve eye literacy and make people more resistant to eye persuasion. The effect was a forty-two percent reduction in the odds of being persuaded, which is really impressive for a fifteen-minute intervention.

Tom: And it’s not just about resistance. The study showed that people who went through LLMimic were better at discerning the ethical quality of eye persuasion. They rated the ethical hotel agent as more truthful and socially responsible, while still rejecting the malicious money solicitor.

Lu: That’s the key contribution, I think. It’s a human-centered approach that empowers people rather than just trying to police the eye. It’s about building a more informed public that can engage with eye critically.

Meng: And from a practical standpoint, it’s scalable. It’s gamified, it’s accessible, and it doesn’t require any technical expertise. That makes it a viable tool for education, for workplace training, even for public awareness campaigns.

Jane: The authors also highlighted some important limitations. The effect might not be durable over time, and the intervention doesn’t yet help people distinguish between malicious and prosocial persuasion. Those are clear directions for future research.

Tom: Right, and they’re calling for more human-centered evaluation methods, because eye-eye benchmarks might not accurately reflect how real people respond to persuasion. That’s a crucial point for the whole field of eye safety.

Lu: I think the biggest implication is that we can proactively build human resilience. Instead of constantly playing catch-up with increasingly persuasive eye, we can give people the mental tools to navigate this new landscape. That’s a powerful idea.

Meng: And it’s one that could actually be deployed at scale, which is what makes it so exciting. This isn’t just a lab experiment. It’s a blueprint for a real-world intervention.

Jane: Well said, everyone. We’ve learned that stepping into the shoes of an LLM, even for a few minutes, can change how we interact with them. It’s a clever, effective, and deeply human-centered approach to a very modern problem.

Tom: And with that, we’re going to say goodbye to “Train Yourself as an LLM.” Thanks to the authors for this thought-provoking work, and thanks to our listeners for tuning in. We’ll be back soon with another paper from the arXiv. Until then, stay curious.

More episodes

← Home