LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

arXiv:2608.05246 · cs.AI · Submitted 2026-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs".

Jane: The paper was written by Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang et al. from Shenzhen University of Advanced Technology and Ant Group.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, listeners, welcome back. We've got a fresh one from the arXiv this week, and it's called "LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs." Jane, I gotta say, that acronym is doing some heavy lifting.

Jane: It really is, Tom. But the idea behind it is actually pretty straightforward once you unpack it. We've got all these AI assistants now, right? And they're supposed to know us. But the question is, how do we actually test whether they know us, or whether they're just giving generic advice to everyone?

Tom: Right, and that's exactly what this paper is tackling. They've built a benchmark, which is basically a standardized test, to see how well these models can personalize their answers based on your actual behavior.

Jane: And not just your stated preferences, like "I like Italian food." They're using your actual behavior logs. We're talking about your app interactions, your purchases, your travel bookings, all that messy, real-world data.

Tom: Lu, you've been staring at this paper all morning. What's the big deal from a research perspective?

Lu: The big deal is that most existing benchmarks give the model a persona, a little text blurb that says "this user is a budget-conscious traveler." But that's cheating. LUNAR forces the model to figure that out for itself by looking at a messy pile of behavioral data.

Jane: Exactly. It's the difference between someone telling you they're a neat person and you looking at their desk. One is a claim, the other is evidence.

Tom: So they're testing whether the model can be a detective, essentially. Digging through the logs to find the clues that actually matter for the question being asked.

Lu: Precisely. And the key word in the title is "UNiversal." They're not just looking at one type of behavior. They're looking across clothing, food, housing, and mobility. So the model has to integrate evidence from all these different domains to answer a single question.

Meng: From an engineering standpoint, that's a nightmare. You're not just retrieving one relevant document; you're trying to find a signal in a haystack of different data types and then synthesize it into a coherent, personalized answer.

Tom: And that's the real challenge, right? It's not just about having the data; it's about knowing which piece of data is relevant to this specific person, at this specific time, for this specific question.

Jane: It's a much harder test than anything we've seen before. And the name, LUNAR, it kind of fits. It's a benchmark that's trying to get us to the moon of personalization, which is a long way off.

Tom: Well, they've built the test. The next question is, how did the models actually do on it? That's where things get really interesting.

Summary: Jane: So, Tom, we've established that LUNAR is a tough test. But the results from the paper are what really got me excited. They tested nineteen different mainstream LLMs, and the findings are a bit of a wake-up call.

Tom: A wake-up call is an understatement. The paper basically says that just giving a model a user's behavioral logs is necessary, but it's not enough for deep personalization. The average scores across all models were pretty low, even in the best-case scenario.

Lu: Right, the best model only hit an average score of three point nine zero out of five point zero in the full context setting. And that's with the entire behavioral history. It shows that we're a long way from these models truly understanding us.

Tom: And here's the kicker, Jane. They found that bigger models don't necessarily mean better personalization. The paper shows that a 32B model from the Qwen3 series actually scored lower than its smaller 14B and 8B counterparts.

Jane: That's so counterintuitive. You'd think throwing more parameters at a problem would always help. But it seems like personalization is a different kind of skill. It's not just about raw reasoning power.

Meng: It's more about information filtering and knowing what to ignore. A bigger model might have more knowledge, but if it can't figure out which of the one hundred forty-three thousand behavioral records in the log are relevant, it's just going to get overwhelmed by the noise.

Tom: And that's exactly what they found. When they gave the models a "curated context," which is just the relevant evidence pre-selected for them, the scores went up significantly. The best model jumped to four point zero seven.

Jane: So the bottleneck isn't just about having the information, it's about the retrieval process. The models are struggling to find the needle in the haystack on their own.

Lu: That's a critical insight. It separates the problem into two parts: finding the relevant evidence and then reasoning over it. LUNAR shows that the "finding" part is a major challenge for current models.

Tom: And the paper digs even deeper into that. They compared a simple retrieval system, RAG, against a more complex agentic memory system. And guess what? The simpler one won.

Meng: Yeah, that makes sense to me. The agentic memory compresses the data into summaries, which can lose important details. RAG just pulls the raw records, so the model gets the fine-grained information it needs.

Jane: So the lesson is, don't over-process the data before the model gets to see it. Let the model do the reasoning on the original, unadulterated facts.

Tom: It's a fantastic result because it gives us a clear direction for improvement. We don't need bigger models; we need better ways to get the right information to the models we already have.

Jane: And that's a much more actionable takeaway. But there's another layer to this that we haven't even touched on yet, and it's about how the models handle the privacy implications of all this personal data.

Improvements: Tom: We're back, and we've just seen that LUNAR shows us models struggle to find the right evidence. But the paper goes further. It suggests some concrete ways to improve, and it also highlights a pretty serious trade-off.

Jane: Right, the privacy angle. The paper shows that as personalization scores go up, privacy protection often goes down. Some models were so eager to be helpful that they started listing irrelevant behavioral details, which just feels creepy.

Meng: It's the difference between saying "your flight is at eight:thirty-five you should head to the gate" and saying "I see you bought a sunscreen spray at six:eighteen and you stayed at a family hotel, so you probably have kids with you." One is helpful; the other is surveillance.

Lu: Exactly. And the paper quantifies this. They found a clear Pareto frontier, where some models achieve high personalization with good privacy, like Kimi-K2 point 6, while others, like Gemini Flash, are more aggressive and over-expose behavioral details.

Tom: So the improvement isn't just about making models smarter at retrieval. It's about teaching them restraint. They need to learn not just what to use, but how to use it without making the user feel like they're being watched.

Jane: That's a really important distinction. The paper calls it "evidence selection" and "evidence expression control." It's not just about picking the right data; it's about deciding how explicitly to reveal that you're using it.

Meng: And that's a hard engineering problem. You can't just add a line to the prompt saying "be private." You need to build in a mechanism that evaluates the sensitivity of the information before it's included in the response.

Lu: The paper also suggests that the way we evaluate these models needs to change. They use a "retroductive" method, which means they start from a generic answer and work backward to figure out what personalized information could improve it. This helps create a much more grounded rubric for evaluation.

Tom: So instead of just asking "is this a good answer?", they're asking "did this answer use the user's data to go beyond what any generic assistant could have said?"

Jane: And that's the core of what makes LUNAR so valuable. It's not just a test; it's a diagnostic tool. It tells us exactly where the models are failing, whether it's in retrieval, reasoning, or privacy control.

Tom: So the path forward is clear. We need better retrieval systems, we need models that can reason over cross-domain evidence, and we need them to do it all while respecting user privacy.

Meng: And we need to be able to measure all of that. LUNAR gives us the yardstick to do it.

Jane: It really does. And I think this is going to push the whole field forward. It's a benchmark that actually captures the complexity of real-world personalization.

Conclusion: Tom: Well, that brings us to the end of our look at "LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs." Jane, what's the big picture here?

Jane: The big picture is that personalization is a much harder problem than we thought. This paper proves that simply feeding a model a user's data isn't enough. The models need to learn how to be detectives, finding the relevant clues, and then how to be diplomats, using that information without being intrusive.

Lu: And it gives us a clear roadmap. We need to focus on improving evidence retrieval and cross-domain reasoning, and we need to bake privacy awareness into the model's behavior, not just as an afterthought.

Meng: From my side, it's a validation that simpler retrieval methods can be more effective than complex memory systems. And it gives us a benchmark to measure our own progress against.

Tom: So, a tough test, some surprising results, and a clear direction for future work. That's a pretty solid paper in my book.

Jane: Absolutely. And the fact that they built this benchmark using a synthesis pipeline grounded in real-world data makes the results even more credible. It's not a toy problem; it's a realistic simulation of what these assistants will face in the wild.

Tom: So as we say goodbye to LUNAR, we're not just closing a chapter on a paper. We're opening a conversation about what it really means for an AI to know a person.

Jane: And that's a conversation we'll definitely be continuing. Thanks for listening, and we'll see you on the next one.

Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang

Shenzhen University of Advanced Technology · Ant Group

cs.AI

Submitted: 2026-08-14

Updated: 2026-08-17

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 67/100

The gist: LUNAR is the first benchmark for evaluating cross-domain behavioral personalization in large language models (LLMs), where responses must be grounded in heterogeneous daily-life activities.

Key concepts

LUNAR Benchmark
A standardized test designed to measure how well Large Language Models personalize their responses based on a user's actual, messy behavioral data. It forces the model to act like a detective by sifting through real-world logs rather than relying on stated preferences.
Universal User Behavior Logs
The raw, comprehensive data used for testing, including app interactions, purchases, travel bookings, and housing choices. These logs represent a user's actual actions across various domains to provide evidence for personalization.
Evidence Retrieval Challenge
The difficulty models face in identifying the specific, relevant pieces of data from a large set of logs. The paper found that models often fail at 'finding' the right information, making retrieval a critical bottleneck before they can reason over it.
Personalization vs. Privacy
A conflict where higher levels of personalization often lead to lower privacy protection. Models might become too eager to be helpful, resulting in the over-exposure of irrelevant or sensitive behavioral details.

Terminology

Summary

LUNAR is the first benchmark for evaluating cross-domain behavioral personalization in large language models (LLMs), where responses must be grounded in heterogeneous daily-life activities. It addresses the limitations of existing personalized LLM benchmarks that primarily rely on textual personas or isolated behavioral signals.

The benchmark represents users through longitudinal behavior histories of universal daily-life behaviors—spanning clothing, food, housing, and mobility—and requires models to generate personalized responses by selectively integrating heterogeneous evidence among these behavioral domains. To support scalable benchmark construction while mitigating data sparsity and privacy concerns, LUNAR uses a multi-stage coarse-to-fine synthesis pipeline grounded in real-world behavioral patterns. The pipeline consists of three stages: reality-anchored data generation, retroductive evidence and rubric generation, and evidence-grounded personalization evaluation.

The reality-anchored data generation pipeline uses anonymized behavioral logs and user requests from a large-scale commercial service platform as anchors, decomposing behavior generation into three levels: user profiles, long-span trajectories, and concrete behavior logs. This design progressively narrows the synthesis space, maintains long-term behavioral coherence, and keeps logs close to real-world patterns. The retroductive evidence and rubric generation process starts from a generic response without behavioral context, identifies what additional user-specific information could improve the answer, and works backward to locate the behavioral domains and logs that provide such information. The evidence-grounded personalization evaluation uses a referenced comparison framework where a personalized response (with behavioral logs) is compared against a generic response (without behavioral context) along two dimensions: Personalization Coverage (PC), measuring whether the response covers user-specific needs and relevant evidence, and Personalization Depth (PD), measuring whether the response turns evidence into meaningful personalization rather than generic advice.

LUNAR comprises 150 users and 300 evaluation queries, with each user associated with a 12-week behavioral history. Of these queries, 112 are single-domain, whereas the remaining 188 require evidence from at least two domains. Overall, the benchmark contains 143,008 behavioral records spanning a wide range of behavior types. The behavioral logs span four core daily-life domains (clothing, food, housing, mobility), while the queries span six representative life-service scenes (travel, hotel, affairs, bus, train, plane).

Fidelity analyses show closer alignment with real behavioral distributions than other synthetic benchmarks. The reality-anchored generation reduces Behavior-Amount JSD from 0.330 to 0.105 (a 68.2% relative reduction) and Behavior-Time JSD from 0.078 to 0.060 (a 23.1% relative reduction). External comparison shows LUNAR achieves the lowest Cond-Type-Time JSD of 0.028, reducing divergence from real users by 37.8% compared to DynamicMem and 67.1% compared to Mem-PAL.

Experiments on 19 mainstream LLMs reveal several key findings. First, personalized tasks grounded in behavioral data remain a critical bottleneck: even under the most favorable contextual conditions, the highest average score across 19 models reaches merely 3.90 (Gemini Flash) under Full Context, and the best performance under Curated Context is only 4.07 (Kimi-K2.6), with over half of the models still falling below 3.5. Second, parameter scale is not the decisive factor governing personalization capability—the Qwen3-32B dense model achieves an average score of only 2.42 under full context, lower than its 14B (2.60) and 8B (2.56) counterparts. Third, fine-grained retrieval (RAG) consistently outperforms compressed memory (Agentic Memory) across all 19 models, with an average Avg. score gap of 0.21 points.

Cross-domain analysis shows that cross-domain information integration is the true bottleneck beyond retrieval. Reducing evidence from all relevant domains (Full-E) to only one sampled relevant domain (Single-E) lowers performance for every model, with drops ranging from 0.39 for Qwen3-0.6B to 1.03 for Kimi-K2.6. Strong models achieve substantially higher scores under Full-E than their single-domain performance, while weak models score lower under Full-E than their single-domain performance, suggesting that additional domains serve as noise rather than supplementation for them. Evidence scaling experiments show that cross-domain evidence adds value monotonically, but its marginal return and ceiling are capability-dependent: top-tier models exhibit classic diminishing returns, mid-tier models achieve lower total gains with unstable trajectories, and bottom-tier models show near-flat curves.

Privacy-personalization trade-off analysis reveals that models with lower personalization scores tend to preserve privacy better, while highly personalized models often fall into the aggressive region with lower privacy protection. Six models fall into the Balanced region (high personalization, high privacy), including Kimi-K2.6 and Qwen3.6-35B-A3B, demonstrating that strong personalization does not inevitably compromise privacy. The Pareto front consists of only four models—Kimi-K2.6, Gemini Flash, GPT-4.1-mini, and GPT-4o-mini—with Kimi-K2.6 achieving the best balance among Pareto-optimal points.

Human-LLM agreement validation shows that the LLM judge achieves agreement comparable to that among human annotators, with human–LLM agreement reaching 76.6% for personalization score (avg.) and 71.1% for privacy protection score.

These findings identify evidence selection, cross-domain integration, and privacy control as key challenges for personalized LLMs. The paper concludes that behavioral logs are necessary but not sufficient for deep personalization: neither more context nor larger models guarantees better performance, and effective personalization depends on selecting and integrating relevant evidence across domains.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

Improvement: Add a dedicated module that explicitly identifies which behavioral domains (clothing, food, housing, mobility) are relevant to a query before generating a response. This module filters out irrelevant logs and prioritizes cross-domain evidence.

What the improved system can do: When a user asks What's the deal with boarding?, the system will automatically detect that mobility (flight booking), clothing (sunscreen spray purchase), and housing (hotel checkout) are relevant, while ignoring unrelated food logs. It will integrate these three domains into a single, coherent, personalized answer—rather than producing generic advice or overloading the user with irrelevant data.

Improvement: Replace agentic memory compression (e.g., Mem0) with direct retrieval of serialized behavioral records (RAG-style). Use fine-grained event-level retrieval rather than summarizing histories into atomic facts.

Improvement: Add a post-generation privacy filter that scores the response for three offense types: (1) profiling/labeling, (2) surveillance/tracking, and (3) condescension. The system then rewrites the response to remove or soften these elements while preserving personalization depth.

Improvement: Implement a capability-aware evidence budget. For weak models (e.g., <8B parameters), limit the number of behavioral records provided to 1–2 per domain, since additional evidence acts as noise. For strong models, provide full cross-domain evidence to exploit diminishing-returns saturation.

Improvement: Add a temporal reasoning layer that explicitly uses query timestamps and event timestamps to infer urgency, sequences, and conflicts.

Improvement: Before finalizing a response, the system performs an internal swap test—it replaces the user's data with a different user's data and checks whether the advice would change. If the advice remains the same, the system regenerates with deeper personalization.

Improvement: Before generating, the system scores each behavioral domain's relevance to the query on a 0–1 scale. Only domains above a threshold (e.g., 0.6) are included in the context. This prevents the noise effect observed in weak models.

Improvement: After generating a response, the system evaluates itself against a structured rubric (Personalization Coverage and Depth) using the same criteria as the paper's LLM judge. If the self-score is below 4.0, it iterates once with explicit instructions to address missing rubric items.

Sources

Related papers