LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
summary
The gist
LUNAR is the first benchmark for evaluating cross-domain behavioral personalization in large language models (LLMs), where responses must be grounded in heterogeneous daily-life activities.
In short
The LUNAR paper benchmarks how well Large Language Models personalize answers using real user behavior logs across domains like travel and food. Findings show models struggle with evidence retrieval, meaning they cannot effectively find relevant data. The research concludes that improving retrieval methods is key to personalization, while also highlighting the necessary balance between helpfulness and user privacy.
Key concepts
- LUNAR Benchmark
- A standardized test designed to measure how well Large Language Models personalize their responses based on a user's actual, messy behavioral data. It forces the model to act like a detective by sifting through real-world logs rather than relying on stated preferences.
- Universal User Behavior Logs
- The raw, comprehensive data used for testing, including app interactions, purchases, travel bookings, and housing choices. These logs represent a user's actual actions across various domains to provide evidence for personalization.
- Evidence Retrieval Challenge
- The difficulty models face in identifying the specific, relevant pieces of data from a large set of logs. The paper found that models often fail at 'finding' the right information, making retrieval a critical bottleneck before they can reason over it.
- Personalization vs. Privacy
- A conflict where higher levels of personalization often lead to lower privacy protection. Models might become too eager to be helpful, resulting in the over-exposure of irrelevant or sensitive behavioral details.
Terminology used across episodes
This episode discusses
- LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs · Paper Radio
- Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces
- KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation
- POPI: Personalizing LLMs via Optimized Natural Language Preference Inference · Paper Radio
- LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Towards Natural Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions · Paper Radio
- OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents · Paper Radio
- Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
- Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
- PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI
- PersonaVLM: Long-Term Personalized Multimodal LLMs
- On Memory Construction and Retrieval for Personalized Conversational Agents
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
- DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings
- MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
The paper
LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs · Read on arXiv
Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang
Shenzhen University of Advanced Technology · Ant Group
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs".
Jane: The paper was written by Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang et al. from Shenzhen University of Advanced Technology and Ant Group.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, listeners, welcome back. We've got a fresh one from the arXiv this week, and it's called "LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs." Jane, I gotta say, that acronym is doing some heavy lifting.
Jane: It really is, Tom. But the idea behind it is actually pretty straightforward once you unpack it. We've got all these AI assistants now, right? And they're supposed to know us. But the question is, how do we actually test whether they know us, or whether they're just giving generic advice to everyone?
Tom: Right, and that's exactly what this paper is tackling. They've built a benchmark, which is basically a standardized test, to see how well these models can personalize their answers based on your actual behavior.
Jane: And not just your stated preferences, like "I like Italian food." They're using your actual behavior logs. We're talking about your app interactions, your purchases, your travel bookings, all that messy, real-world data.
Tom: Lu, you've been staring at this paper all morning. What's the big deal from a research perspective?
Lu: The big deal is that most existing benchmarks give the model a persona, a little text blurb that says "this user is a budget-conscious traveler." But that's cheating. LUNAR forces the model to figure that out for itself by looking at a messy pile of behavioral data.
Jane: Exactly. It's the difference between someone telling you they're a neat person and you looking at their desk. One is a claim, the other is evidence.
Tom: So they're testing whether the model can be a detective, essentially. Digging through the logs to find the clues that actually matter for the question being asked.
Lu: Precisely. And the key word in the title is "UNiversal." They're not just looking at one type of behavior. They're looking across clothing, food, housing, and mobility. So the model has to integrate evidence from all these different domains to answer a single question.
Meng: From an engineering standpoint, that's a nightmare. You're not just retrieving one relevant document; you're trying to find a signal in a haystack of different data types and then synthesize it into a coherent, personalized answer.
Tom: And that's the real challenge, right? It's not just about having the data; it's about knowing which piece of data is relevant to this specific person, at this specific time, for this specific question.
Jane: It's a much harder test than anything we've seen before. And the name, LUNAR, it kind of fits. It's a benchmark that's trying to get us to the moon of personalization, which is a long way off.
Tom: Well, they've built the test. The next question is, how did the models actually do on it? That's where things get really interesting.
Summary: Jane: So, Tom, we've established that LUNAR is a tough test. But the results from the paper are what really got me excited. They tested nineteen different mainstream LLMs, and the findings are a bit of a wake-up call.
Tom: A wake-up call is an understatement. The paper basically says that just giving a model a user's behavioral logs is necessary, but it's not enough for deep personalization. The average scores across all models were pretty low, even in the best-case scenario.
Lu: Right, the best model only hit an average score of three point nine zero out of five point zero in the full context setting. And that's with the entire behavioral history. It shows that we're a long way from these models truly understanding us.
Tom: And here's the kicker, Jane. They found that bigger models don't necessarily mean better personalization. The paper shows that a 32B model from the Qwen3 series actually scored lower than its smaller 14B and 8B counterparts.
Jane: That's so counterintuitive. You'd think throwing more parameters at a problem would always help. But it seems like personalization is a different kind of skill. It's not just about raw reasoning power.
Meng: It's more about information filtering and knowing what to ignore. A bigger model might have more knowledge, but if it can't figure out which of the one hundred forty-three thousand behavioral records in the log are relevant, it's just going to get overwhelmed by the noise.
Tom: And that's exactly what they found. When they gave the models a "curated context," which is just the relevant evidence pre-selected for them, the scores went up significantly. The best model jumped to four point zero seven.
Jane: So the bottleneck isn't just about having the information, it's about the retrieval process. The models are struggling to find the needle in the haystack on their own.
Lu: That's a critical insight. It separates the problem into two parts: finding the relevant evidence and then reasoning over it. LUNAR shows that the "finding" part is a major challenge for current models.
Tom: And the paper digs even deeper into that. They compared a simple retrieval system, RAG, against a more complex agentic memory system. And guess what? The simpler one won.
Meng: Yeah, that makes sense to me. The agentic memory compresses the data into summaries, which can lose important details. RAG just pulls the raw records, so the model gets the fine-grained information it needs.
Jane: So the lesson is, don't over-process the data before the model gets to see it. Let the model do the reasoning on the original, unadulterated facts.
Tom: It's a fantastic result because it gives us a clear direction for improvement. We don't need bigger models; we need better ways to get the right information to the models we already have.
Jane: And that's a much more actionable takeaway. But there's another layer to this that we haven't even touched on yet, and it's about how the models handle the privacy implications of all this personal data.
Improvements: Tom: We're back, and we've just seen that LUNAR shows us models struggle to find the right evidence. But the paper goes further. It suggests some concrete ways to improve, and it also highlights a pretty serious trade-off.
Jane: Right, the privacy angle. The paper shows that as personalization scores go up, privacy protection often goes down. Some models were so eager to be helpful that they started listing irrelevant behavioral details, which just feels creepy.
Meng: It's the difference between saying "your flight is at eight:thirty-five you should head to the gate" and saying "I see you bought a sunscreen spray at six:eighteen and you stayed at a family hotel, so you probably have kids with you." One is helpful; the other is surveillance.
Lu: Exactly. And the paper quantifies this. They found a clear Pareto frontier, where some models achieve high personalization with good privacy, like Kimi-K2 point 6, while others, like Gemini Flash, are more aggressive and over-expose behavioral details.
Tom: So the improvement isn't just about making models smarter at retrieval. It's about teaching them restraint. They need to learn not just what to use, but how to use it without making the user feel like they're being watched.
Jane: That's a really important distinction. The paper calls it "evidence selection" and "evidence expression control." It's not just about picking the right data; it's about deciding how explicitly to reveal that you're using it.
Meng: And that's a hard engineering problem. You can't just add a line to the prompt saying "be private." You need to build in a mechanism that evaluates the sensitivity of the information before it's included in the response.
Lu: The paper also suggests that the way we evaluate these models needs to change. They use a "retroductive" method, which means they start from a generic answer and work backward to figure out what personalized information could improve it. This helps create a much more grounded rubric for evaluation.
Tom: So instead of just asking "is this a good answer?", they're asking "did this answer use the user's data to go beyond what any generic assistant could have said?"
Jane: And that's the core of what makes LUNAR so valuable. It's not just a test; it's a diagnostic tool. It tells us exactly where the models are failing, whether it's in retrieval, reasoning, or privacy control.
Tom: So the path forward is clear. We need better retrieval systems, we need models that can reason over cross-domain evidence, and we need them to do it all while respecting user privacy.
Meng: And we need to be able to measure all of that. LUNAR gives us the yardstick to do it.
Jane: It really does. And I think this is going to push the whole field forward. It's a benchmark that actually captures the complexity of real-world personalization.
Conclusion: Tom: Well, that brings us to the end of our look at "LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs." Jane, what's the big picture here?
Jane: The big picture is that personalization is a much harder problem than we thought. This paper proves that simply feeding a model a user's data isn't enough. The models need to learn how to be detectives, finding the relevant clues, and then how to be diplomats, using that information without being intrusive.
Lu: And it gives us a clear roadmap. We need to focus on improving evidence retrieval and cross-domain reasoning, and we need to bake privacy awareness into the model's behavior, not just as an afterthought.
Meng: From my side, it's a validation that simpler retrieval methods can be more effective than complex memory systems. And it gives us a benchmark to measure our own progress against.
Tom: So, a tough test, some surprising results, and a clear direction for future work. That's a pretty solid paper in my book.
Jane: Absolutely. And the fact that they built this benchmark using a synthesis pipeline grounded in real-world data makes the results even more credible. It's not a toy problem; it's a realistic simulation of what these assistants will face in the wild.
Tom: So as we say goodbye to LUNAR, we're not just closing a chapter on a paper. We're opening a conversation about what it really means for an AI to know a person.
Jane: And that's a conversation we'll definitely be continuing. Thanks for listening, and we'll see you on the next one.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language