Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

arXiv:2608.09507 · cs.CL, cs.AI · Submitted 2026-08-12 · Read on arXiv

Yuting Liu, Wei Wu, Jianzhe Zhao, Guibing Guo

Software College, Northeastern University · Ant International

cs.CL, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/AntResearchNLP/AlignX-Family

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: The paper addresses a central challenge in LLM personalization: task-specific preference adaptation.

Terminology

Summary

The paper addresses a central challenge in LLM personalization: task-specific preference adaptation. The authors note that "universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary therefore wastes context capacity and introduces cross-task distraction, while manually designing task-specific preference views is difficult to scale."

The key observation is illustrated in Figure 1: Given a query, only a small portion of the universal profile P is relevant to the task, while the remaining information may act as distractors and lead to suboptimal personalization. In their example, only 1% of characters in a 7,675-character universal profile are task-relevant for a specific query about staying active without straining legs. The universal profile contains dominant distractor information (e.g., community-oriented approach) that overwhelms the task-relevant clue (past college soccer injury), causing the downstream model to select an incorrect response.

The authors propose AlignXada, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones. The framework has two key desiderata for adapted representations:

  1. Sufficiency: preserving all preference information necessary for downstream personalization

  2. Compactness: removing redundant or task-irrelevant information to reduce noise and improve context efficiency

AlignXada consists of two stages:

Stage 1: Few-shot policy induction — "the framework leverages a small support set of task-specific demonstrations, each consisting of a universal preference summary, a user query, and the corresponding user response, and iteratively refines the rewriting policy generated by a meta learner using natural language feedback."

Stage 2: Policy deployment — the selected policy is consumed by a refiner to produce a refined profile for task adaptation. The refined profile is then provided to downstream models for personalized inference.

Critically, "all models—including the meta learner, the refiner, and the downstream model—remain frozen. In the following, Section 3.1 formalizes the learning problem, and Section 3.2 presents the policy induction procedure based on verbal reinforcement learning."

Given a user u with universal preference summary Pu, and a task tau, the objective is to derive a task-adapted profile P̃u such that downstream model performance improves. The support set is defined as:


S(tau) = (Pui, xi, yi) i=1 b

where Pui is the universal profile, xi is an input prompt, and yi is the corresponding response.

The learning objective is to maximize:


E(Pui,xi,yi)∼S(tau)[Score tau(M(xi, P̃ui), yi)]

The induction procedure (Algorithm 1) works as follows:

  1. Rollout: At each round t, the current policy phit is evaluated on the support set by applying the refiner R, querying the downstream model M, and computing task-specific scores.

  2. Development evaluation: To reduce overfitting, AlignXada evaluates each policy on a development set D(tau) disjoint from S(tau). The policy utility is:


J D(tau)(phit) = (1/D(tau)) Σ i=1 D(tau) si,t D

  1. Structured feedback: The rollout records are aggregated into textual diagnostic feedback including scalar signals, such as the average support-set score; instance-level signals, such as predictions and scores; and representative failure patterns derived from low-scoring examples.

  2. Policy update: The frozen meta learner produces a revised policy: phit+1 = pimeta(phit, Et). The update prompt instructs the meta learner to follow the predefined policy schema, make targeted revisions based on the feedback, and avoid example-specific rules that merely memorize support instances.

  3. Best-epoch selection: After T rounds, AlignXada returns the policy with the highest development-set performance: phi(tau) = argmax J D(tau)(phit).

The policy schema includes: the refinement goal, the types of user evidence to preserve, the compression strategy, the error patterns to avoid, the output style, or the priority order among these instructions.

The authors construct a composite benchmark by integrating PersonaMem-v2 and MemoryCD:

  • We first derive task-agnostic user summaries from both datasets and extract semantic signatures capturing stable interests, preferences, aversions, and contextual constraints.

  • We then match users one-to-one based on semantic compatibility, excluding pairs with explicit preference conflicts.

  • For each matched pair, we construct a shared universal profile by interleaving and summarizing their PersonaMem and MemoryCD histories.

The benchmark spans 13 tasks: nine PersonaMem-v2 conversational tasks and four MemoryCD tasks (item ranking, rating prediction, review-title generation, review generation).

  • Downstream models: Qwen3-8B, DeepSeek-V4-Flash, and GPT-5-mini

  • Meta learner and refiner: Gemini-2.5-Pro (default)

  • Baselines: Raw universal preference (no refinement) and RAG using BM25 with top-8 chunks within 768-token budget

  • Four-way choice questions: exact accuracy

  • Item ranking: Hit@1

  • Rating prediction: Rating-Score = max(0, 1 − y−ŷ/4)

  • Generation tasks: ROUGE-L

  • Context compression: Token ratio (TR) = average refined-profile length / average universal-profile length

Across 39 task–model cells, AlignXada achieves:

  • Improves 33 cells with an average gain of +3.82 points over raw universal preference

  • Retains only 22.8% of original profile tokens

  • Outperforms RAG in 36 cells with an average margin of +7.28 points

  • RAG reduces raw-profile score by 3.46 points on average

Per-model gains:

  • Qwen3-8B: +2.27 points

  • DeepSeek-V4-Flash: +7.00 points

  • GPT-5-mini: +2.19 points

The authors note: performance gains are nearly uncorrelated with token ratio (r = 0.06). Thus, AlignXada improves the utility of retained evidence rather than relying on longer profiles.

When evaluated separately on the original benchmarks with cleaner, domain-coherent profiles:

  • PersonaMem-v2: +1.30 points average improvement, outperforming raw on 7/9 tasks, token ratio 40.8%

  • MemoryCD: +1.69 points average improvement, improving all 4 tasks, token ratio 48.0%

  • AlignXada outperforms RAG on every task in both benchmarks

This demonstrates AlignXada goes beyond removing cross-domain noise by using support-set feedback to induce preference representations better aligned with downstream requirements.

Support batch size (Figure 3): Increasing b from 5 to 20 improves mean primary-metric change from −0.33 to +2.58 points, median from −0.91 to +1.94 points, and improved tasks from 4/13 to 11/13. Larger batches provide broader and more balanced feedback, leading to more stable policies while also producing the lowest token ratio (20.5%).

Update rounds (Figure 4): Eight of the thirteen task-specific policies emerge by the second round, and twelve within five rounds. Most tasks converge early; additional rounds mainly benefit challenging tasks like rating prediction (whose final policy appears at round 9).

The authors propose adaptive diagnostic sampling which selects support examples from four outcome-transition states:

  • improved: raw incorrect, refined correct

  • regressed: raw correct, refined incorrect

  • stable-success: both correct

  • persistent-failure: both incorrect

Compared to fixed sampling, adaptive sampling improves mean task-level gain from +1.22 to +2.58 points, increases improved tasks from 7/13 to 11/13, and increases context-token reduction from 68.3% to 72.7%.

Using DeepSeek-V4-Flash for all model roles (AlignXada-D):

  • Improves 12/13 tasks with average gain of +4.47 points

  • Improves 8/9 PersonaMem-v2 tasks (+4.22 points) and all 4 MemoryCD tasks (+5.04 points)

  • Token ratio of 23.95%

  • Outperforms RAG on all 13 tasks by 9.33 points on average

The authors conclude: AlignXada transfers across model families, but the choice of policy-induction and rewriting models still affects how effectively decision-relevant evidence is identified and retained.

Three preference sources tested on PersonaMem-v2:

  • Source A (GPT-5 compact persona): +0.61% improvement after refinement

  • Source B (raw structured JSON): −0.84% decrease (format mismatch)

  • Source C (Gemini-2.5-Pro comprehensive summary): +1.30% improvement

The authors note: AlignXada works best as a task-specific refiner for natural-language preference profiles, rather than as a replacement for upstream profile construction. Source C is the best match because it is sufficiently evidence-rich for personalization while remaining in natural language.

The audit uses DeepSeek-V4-Pro as judge model to evaluate:

Claim-level faithfulness (97.5% support, 97.9% non-hallucination, 99.8% non-contradiction):

  • 97.5% of extracted claims are supported by the original profile

  • the non-hallucination rate is 97.9%, and the non-contradiction rate reaches 99.8%

Preference-evidence retention:

  • Only 34.0% of benchmark preference evidence is recoverable from original source profiles

  • Refined profiles retain 30.2% overall

  • Conditional retention: Conditional on the original profile containing the golden preference, AlignXada preserves it in 83.3% of cases, with a critical missing rate of 15.3%

The authors conclude: AlignXada can faithfully compress and adapt the profile it receives, but its performance ceiling depends on whether the universal profile contains the cross-topic evidence needed for the decision.

  1. Problem formalization: "We formalize the task-specific preference adaptation problem, which aims to adapt universal user preferences to downstream tasks by removing redundant or task-irrelevant information, thereby paving the way for lifelong personalized agents equipped with memory."

  2. Framework introduction: "We introduce AlignXada, a meta-learning framework that induces natural language refinement policies from a small set of task-specific demonstrations. The policy is iteratively optimized via verbal reinforcement learning, making AlignXada compatible with both open-source and proprietary LLMs."

  3. Extensive evaluation: "We conduct extensive evaluations across thirteen tasks and three downstream models. The results show that AlignXada consistently achieves a favorable trade-off between task performance and context-token usage, owing to the faithfulness and compactness of the refined preferences."

The authors acknowledge two main limitations:

  1. Both benchmarks are constructed or curated evaluation settings and may not fully capture the evolving, noisy interactions of real-world lifelong agents.

  2. "AlignXada assumes that the support set used for policy induction is representative of the target task distribution. When the support set is biased or insufficiently diverse, the induced refinement policy may not generalize well."

The paper concludes: "We propose AlignXada, which induces a reusable natural-language refinement policy through verbal reinforcement learning, improving the performance–budget trade-off while preserving source-supported claims and critical evidence. By enabling effective use of universal preferences, our work offers a new perspective on memory adaptation for lifelong personalized agents."

Improvements for AI systems

Based on this paper, here are the specific improvements I can make to AI systems:

Improvement: Instead of feeding an entire user profile into the model for every query, I can dynamically extract only the preference-relevant subset for the current task.

What the improved system can do: When a user asks What's a good workout given my knee injury?, I can automatically identify and surface only the knee-related constraints and exercise preferences from a large profile—ignoring irrelevant details about food preferences, travel habits, or social style. This reduces context noise and improves answer accuracy by up to +3.82 points across tasks.

Improvement: I can learn reusable text-refinement policies through iterative natural-language feedback, without fine-tuning any model weights.

What the improved system can do: Given a small set of task demonstrations (e.g., 20 examples of product recommendation queries with user profiles and ideal responses), I can generate a policy like: Keep budget constraints, brand preferences, and past purchase context; remove demographic info and generic interests. This policy then applies to all future product-recommendation queries, improving accuracy while compressing context to 23% of original size.

Improvement: When selecting examples for policy learning, I can prioritize cases based on outcome-transition states (improved, regressed, stable-success, persistent-failure) rather than random sampling.

What the improved system can do: I can focus learning on regressed cases—where the raw profile gave a correct answer but my refinement made it wrong—to fix specific policy errors. This adaptive approach improves task-level gains from +1.22 to +2.58 points and increases the number of improved tasks from 7/13 to 11/13.

Improvement: I can compress user profiles while maintaining verifiable fidelity to the source information.

What the improved system can do: When condensing a 7,675-character profile to 1,750 characters, I can ensure that 97.5% of extracted claims are supported by the original, with a 99.8% non-contradiction rate. Crucially, when the source profile contains critical preference evidence, I preserve it 83.3% of the time—minimizing the risk of losing decision-relevant information during compression.

Improvement: I can identify which refinement policies generalize across similar tasks and reuse them, rather than inducing new policies from scratch.

What the improved system can do: After learning a policy for movie recommendation, I can apply it to book recommendation with minimal adjustment, since both tasks share preference types (genre, style, pacing). This reduces the support-set requirement and speeds up adaptation to new tasks.

Improvement: I can apply the same refinement framework regardless of the underlying downstream model architecture.

What the improved system can do: Whether the downstream model is Qwen3-8B, DeepSeek-V4-Flash, or GPT-5-mini, I can generate task-specific profiles that improve performance across all of them—with gains ranging from +2.19 to +7.00 points depending on the model. This makes the personalization layer independent of the model choice.

Improvement: I can evaluate each candidate refinement policy on a held-out development set before deployment, selecting the policy with the best generalization rather than the best support-set performance.

What the improved system can do: Instead of memorizing specific examples from the support set, I can choose policies that generalize to unseen queries. This prevents example-specific rules and ensures the refinement policy works on new user data, not just the training demonstrations.

Improvement: I can explicitly optimize for the trade-off between context-token usage and task performance.

What the improved system can do: For a task like rating prediction where longer profiles help, I can retain more evidence (higher token ratio). For item ranking where noise hurts more, I can compress aggressively. The system can adapt its compression strategy per task, achieving a 22.8% average token retention while still improving performance—and the performance gains are uncorrelated with token ratio (r=0.06), meaning the system learns to keep the right evidence, not just more evidence.

Abstract

Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary therefore wastes context capacity and introduces cross-task distraction, while manually designing task-specific preference views is difficult to scale. In this work, we study task-specific preference adaptation: given a universal user preference summary and a downstream task, derive a task-conditioned representation that preserves sufficient decision-relevant evidence while removing redundant context. To this end, we propose AlignXada, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones. The refinement policy is iteratively optimized by a meta learner through verbal reinforcement learning. Across 13 tasks and three downstream models (39 task--model cells), AlignXada achieves an average gain of 3.82 points, improving 33 cells while retaining only 22.8% of the original profile tokens and outperforming RAG in 36 cells. An extended faithfulness analysis further shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.

Sources

Related papers