Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation
summary
In short
The episode discusses a paper showing LLM agents can predict social media reactions with up to 97% accuracy when given detailed profiles, driven by psychological information, not just demographics. Hosts discuss implications for testing recommendation systems and potential risks like creating realistic fake accounts. The key finding is that accuracy drops significantly when profile details are removed.
Key concepts
- Profile-Consistent Reaction Prediction
- The ability of an AI agent to accurately guess whether a specific person would like or dislike a social media post based on a detailed description of that person's profile.
- Zero-Shot Generalization
- The capability of Large Language Models to handle new posts they have never seen before without any prior training, unlike traditional machine learning methods which can only memorize patterns.
- Behavioral Validity Standards
- Suggested governance improvements requiring minimum accuracy thresholds and heterogeneity requirements when using LLM simulations for real-world decisions affecting users.
Terminology used across episodes
This episode discusses
- Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation · Paper Radio
- Project Sid: Many-agent simulations toward AI civilization
- Embodied LLM Agents Learn to Cooperate in Organized Teams
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension
- Simulating Social Media Using Large Language Models to Evaluate Alternative News Feed Algorithms
- Ethical and social risks of harm from Language Models
- OASIS: Open Agent Social Interaction Simulations with One Million Agents
The paper
Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation · Read on arXiv
Ljubiša Bojić, Ljiljana Matić, Jörg Matthes, Milan Čabarkapa, Bojana Dinić, Jue Wang
Institute for Artificial Intelligence Research and Development of Serbia · University of Belgrade · Complexity Science Hub · University of Kragujevac · University of Vienna · University of Novi Sad · Nanyang Technological University
Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. A central open question is whether persona-prompted LLMs can simulate individual-level social media reactions with sufficient accuracy to support either application, and how accuracy depends on profile completeness, model selection, and the generalization challenge posed by novel post content. This study benchmarks twelve LLM configurations on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts under three profile conditions, with leave-post-out machine learning classifiers as baselines. Across full-profile conditions, accuracy ranges from 75.54% to 96.68%, with a 30-point spread attributable primarily to model selection and confirmed by paired McNemar tests with agent-level bootstrap intervals. GPT-5.5 Pro accuracy degrades monotonically from 96.68% under a full profile to 62.32% under a reduced profile and to 51.00% with demographics alone, the last indistinguishable from the majority-class baseline, which confirms that demographic inference provides negligible predictive signal. Supervised classifiers collapse to 15.4% under leave-post-out, while LLMs sustain genuine zero-shot generalization unavailable to trained methods. Adaptive reasoning improves accuracy substantially for some models. Inter-model agreement is nearly double for posts with direct profile anchors (mean = 0.44) than for posts without them (= 0.23), and the least heterogeneous configuration homogenizes 34% of simulated population reactions. Results validate LLM-based simulation for recommender system stress-testing while documenting the behavioral accuracy that makes large-scale synthetic agent swarms a credible threat to public opinion.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation".
Jane: The paper was written by Ljubiša Bojić, Ljiljana Matić, Jörg Matthes, Milan Čabarkapa, Bojana Dinić et al. from Institute for Artificial Intelligence Research and Development of Serbia and University of Belgrade and Complexity Science Hub and University of Kragujevac and University of Vienna and University of Novi Sad and Nanyang Technological University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everybody. Today we're digging into a paper with a title that really says it all — "Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation." Jane, I have to say, that title is doing a lot of work.
Jane: It really is, Tom. And honestly, the claim in that title is bold. Near-perfect prediction of how people react to social media posts, just from knowing their profile? That's a big statement. The authors are saying that if you give an AI agent a detailed enough description of a person, it can guess whether that person would like or dislike a specific post with up to ninety-seven percent accuracy.
Tom: Ninety-seven percent. That's not just good, that's basically reading minds. And the team behind this is impressive — researchers from Serbia, Austria, Singapore, all working together. Ljubiša Bojić leads it, and they've got Jörg Matthes from Vienna, who's a big name in communication research.
Jane: What I love about this title is how direct it is. "Knowing you is everything" — that's the whole thesis in four words. The more they know about a person, the better they can simulate that person's behavior. And the paper backs that up with a really clean experiment where they strip away information layer by layer and watch the accuracy collapse.
Tom: Right, and that's the part that should make us all sit up and pay attention. When they gave the AI the full profile, it hit nearly ninety-seven percent. When they removed the attitudinal stuff and kept just demographics, it dropped to fifty-one percent — basically a coin flip. So the title isn't hype. It's literally what the data shows.
Jane: And the implications are huge, Tom. If you can simulate individual human reactions this accurately, you can build a synthetic population that behaves like a real one. That's powerful for testing recommendation algorithms before they go live. But it's also a warning about what bad actors could do with that same capability.
Tom: Exactly. So let's hold onto that tension as we go through the paper — the constructive uses and the scary ones. Because this title is really a promise, and the paper delivers on it in ways that are both exciting and unsettling.
Jane: And we're just getting started. Next we'll look at what the abstract actually promises and whether the full paper backs it up.
Summary: Tom: So we've established the title is a promise. Now let's look at what the abstract actually delivers in "Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation." Jane, what's the core claim here?
Jane: The core claim is that they tested twelve different AI model configurations on a really specific task: predicting whether someone would like or dislike a social media post. They had two hundred ninety-six real people fill out surveys, built profiles from those surveys, and then asked the AI agents to predict reactions. The best model hit ninety-six point six eight percent accuracy.
Tom: And that's with the full profile. But here's the kicker — when they stripped the profile down to just demographics, that same model dropped to fifty-one percent. That's literally the same as guessing. So the abstract is telling us that the predictive power comes from the detailed psychological and attitudinal information, not from knowing someone's age or where they live.
Jane: Right. And the abstract also makes a really important comparison. They ran traditional machine learning classifiers on the same data, and those collapsed to fifteen percent accuracy when they had to generalize to new posts. But the language models kept working on posts they'd never seen before. That's zero-shot generalization — the AI can handle novel content without any training on it.
Tom: That's the part that blew my mind, honestly. Traditional methods can memorize patterns but they can't transfer them. The LLMs actually reason about what a person would do. And the abstract frames this as a dual-use finding — great for stress-testing recommender systems before deployment, but also a credible threat if someone wants to run a swarm of fake accounts that look genuinely human.
Jane: And there's a subtle point in there about inter-model agreement. When posts had a direct anchor in the survey data, the models tended to agree with each other. When posts were about contested topics without that anchor, agreement dropped way down. So the models are not just parroting each other — they're genuinely interpreting the profile.
Tom: So the abstract is really setting up a roadmap: accuracy numbers, degradation patterns, generalization capabilities, and the societal stakes. And I think the most important takeaway is that this isn't theoretical. This is measured, reproducible, and the numbers are stark.
Jane: And that's what makes it so important to understand. Next we should look at what the paper suggests we actually do with this capability — the improvements and applications they're proposing.
Improvements: Tom: We've covered the headline numbers from "Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation." Now let's talk about what the authors think we should actually do with this. Jane, what improvements are they suggesting?
Jane: The biggest one is validation. They're saying that before you deploy an LLM-based simulation to test a recommender system change, you need to check that the agent population actually behaves like the real population. They found that some models homogenize reactions — meaning a third of the simulated population gives the same answer on certain posts. That's a red flag if you're trying to model genuine diversity.
Tom: Right, and that connects to their heterogeneity analysis. The best model, GPT-five point five Pro, produced zero homogenized posts across all fifty-six. But Claude Haiku — a smaller, cheaper model — homogenized thirty-four percent of posts. So the paper is really saying: don't just pick any model and assume it works. You have to measure the behavioral spread.
Jane: Exactly. And they also suggest that future work should replicate this in other countries and languages. The study was done in Serbia, with Serbian participants and Serbian posts. The geopolitical context matters — attitudes toward Russia and the EU are specific to that region. Would the same accuracy hold in the US or Germany? We don't know yet.
Tom: There's also a really interesting suggestion about prompt engineering and retrieval-augmented generation. They're saying the current approach just gives the model a static profile, but you could imagine giving the agent access to additional information on demand — pulling in relevant facts about a topic before asking for a reaction. That could push accuracy even higher.
Jane: And that's the constructive side. But the paper also suggests improvements on the governance side. They're calling for behavioral validity standards — minimum accuracy thresholds, heterogeneity requirements, documented handling of contested political content. If you're going to use these simulations to make decisions that affect real users, you need standards.
Tom: So the improvements are really twofold: technical improvements to make the simulations better, and governance improvements to make sure they're used responsibly. And I think that balance is what makes this paper stand out.
Jane: It does. And now we should step back and look at the actual first page of the paper — how they set up the problem and why it matters.
First Page: Tom: So we've talked about the title, the abstract, and the proposed improvements. Now let's actually look at the opening of "Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation." Jane, how do they frame the problem on page one?
Jane: They start with a really sobering observation. Social media is now the primary way billions of people encounter political information and form opinions. And that system is no longer purely human — algorithmic systems decide what content reaches which users, and AI-generated actors are already participating in online discourse, posting and reacting in ways that are indistinguishable from humans.
Tom: And they cite some serious evidence. There's a study showing that changing the feed algorithm on X can shift political attitudes across entire user populations. And another one showing that coordinated AI agent swarms can manufacture the appearance of grassroots consensus. So the stakes are real, and they're happening right now.
Jane: The first page also introduces the central question: how faithfully can AI agents replicate the reactions of the specific people they're meant to represent? That's what they call individual behavioral fidelity. And they trace the history — from the early generative agents that just seemed socially credible, to more recent work where agents grounded in interviews with real people could reproduce their survey responses.
Tom: And there's a really important distinction they draw early on. Some researchers think LLMs are just doing demographic stereotyping — if you tell the model someone is from a certain group, it defaults to the stereotype for that group. But this paper is testing something more specific: can the model integrate the actual attitudinal information in a profile, not just the demographic labels?
Jane: Right, and that's why their degradation experiment matters so much. When they removed the attitudinal information and kept only demographics, accuracy collapsed to chance. That's direct evidence that the models are not just stereotyping — they're genuinely using the detailed psychological profile.
Tom: So the first page is really setting up the intellectual stakes: this is about whether LLMs can simulate individuals, not just groups. And that's a much harder problem.
Jane: It is. And it's also the problem that makes this work so consequential. Now let's wrap up and think about what this all means.
Conclusion: Tom: We've spent this whole episode on "Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation," and I think we need to take a moment to really absorb what we've learned. Jane, how would you sum it up?
Jane: I'd sum it up this way: we now have empirical evidence that LLM agents can predict individual social media reactions with up to ninety-seven percent accuracy when given a detailed profile. That accuracy is driven by the attitudinal and psychological information in the profile, not by demographics. And the models can generalize to new posts they've never seen, which traditional machine learning cannot do.
Tom: And the flip side is that this capability is dual-use. It's genuinely valuable for testing recommender systems before they affect real users. But it's also exactly what you'd need to run a network of synthetic accounts that look like real, diverse human beings. The paper doesn't shy away from that tension.
Jane: No, it doesn't. And I think the most important contribution is the degradation ladder. Watching accuracy fall from ninety-seven percent to sixty-two percent to fifty-one percent as you strip away profile information is the clearest possible demonstration that the models are doing real inference, not stereotyping. That's a finding that should shape how we think about AI simulation for years.
Tom: And the heterogeneity finding — that some models homogenize a third of the population's reactions — is a warning that not all models are suitable for population-level simulation. You can't just pick a cheap model and assume it works.
Jane: Exactly. So we're saying goodbye to this paper with a clear picture: the capability is real, it's measurable, and it has consequences. We should be both excited about the research possibilities and careful about the risks.
Tom: Well said, Jane. That's a wrap on "Knowing You Is Everything." Next up, we've got another paper on our list, and I can't wait to see what it brings. Thanks for listening, everyone.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization