From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation
Alireza S. Ziabari, Kat Ellis, Colleen Chan, Ding Tong
Netflix
cs.AI, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale.
Terminology
Summary
Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact same item with identical evidence, highlighting the unreliability of off-the-shelf LLMs in predicting user engagement. To resolve this, we develop and apply a sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales. Evaluated on real-world homepage interaction logs, this aligned reasoning approach achieves a 32.19% lift in MacroF1 score over the zero-shot baseline and matches the production feature-engineered baseline. The results demonstrate that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pipeline overhead.
The paper characterizes bidirectional rationalization as a personalization-specific failure mode of LLM judges that is structurally distinct from hallucination. This failure traces to foundational recommender-system trade-offs (e.g., short-term vs. long-term, accuracy vs. diversity, novelty vs. popularity, exploration vs. exploitation) each of which admits multiple defensible reasoning pathways that an unaligned model can elaborate into a fluent argument in either direction. After filtering rationale pairs for unfactual claims, the bidirectional disagreement patterns persist, showing that the failure is not reducible to fabrication. Specifically, 77.0% (960 of 1,246) of the balanced pairs survived the factuality filter, and among those, 95.4% (916) exhibit contrasts that fit one or a combination of the four trade-offs.
Through a systematic comparison of prompting strategies, the paper finds that reasoning-based prediction and the inclusion of immediate session context are the only consistent contributors to LLM judge accuracy. Session context, such as the current time, proved to be among the most critical factors to include in the prompt. Prompting the model to reason before outputting a final label yielded a noticeable improvement in predictive performance compared to direct label prediction, increasing Macro-F1 score 4.21% over the baseline non-reasoning prompt in the zero-shot setting. However, prompt engineering alone is insufficient to close the gap between a zero-shot LLM judge and a heavily feature-engineered baseline, motivating the move to parameter-level adaptation.
The paper proposes an alignment recipe that first applies SFT on reasoning traces, and then applies offline preference optimization over paired correct and counterfactual reasonings grounded in true engagement outcomes. This recipe closes the remaining performance gap: the resulting text-based LLM evaluator matches the feature-engineered baseline on Netflix homepage engagement prediction without any manual feature engineering, while producing human-interpretable reasoning traces that reveal the user-history signals driving each prediction.
Supervised fine-tuning provided a significant boost to predictive performance, increasing Macro-F1 Score by 12.23% over the zero-shot baseline in the simple prompt setup. Training the model with DPO proved more effective than SFT alone, leading to an improvement of 28.37% in Macro-F1 score over the zero-shot baseline with the simple prompt setup. The best overall results were achieved by sequentially chaining the two methods: first utilizing SFT to establish the domain vocabulary and formatting, followed by DPO to optimize the decision-making process based on accurate versus inaccurate reasoning. The SFT+DPO combination with reasoning-based prediction yields a 32.19% Macro-F1 lift over the zero-shot baseline—the best result, closing the gap to the feature-engineered production baseline by reaching statistical parity in Macro-F1 score (difference < 0.1%).
The paper concludes that personalized LLM judges fail not because they fabricate evidence, but because they apply underspecified, competing framings of identical evidence. Behavioral alignment via paired-rationale preference is key to making such judges reliable. Future work will explore intermediate user profile generation to address the diminishing returns observed when extending raw user histories beyond 50 events.
Improvements for AI systems
Improvements to AI systems:
-
Add a bidirectional rationalization detector and mitigation layer to LLM-based judges/evaluators. Before trusting any prediction, the system actively generates both a positive and negative rationale for the same input, checks if both are fluent and factually grounded, and flags the case as
underspecified
if both pass. This prevents overconfident, contradictory outputs in personalized decision tasks. -
Implement a two-stage alignment pipeline (SFT → DPO) with paired counterfactual rationales for any LLM used in user engagement prediction. Stage 1: fine-tune on correct reasoning traces to learn domain vocabulary and formatting. Stage 2: apply Direct Preference Optimization (DPO) using pairs of correct vs. counterfactual rationales that are grounded in real outcomes, not just syntactically valid. This yields a 32.19% Macro-F1 improvement over zero-shot and matches production feature-engineered baselines.
-
Integrate session context (e.g., current time, immediate user actions) as mandatory prompt features for LLM judges. The paper shows this is among the most critical contributors to accuracy. The improved system will automatically extract and inject such context into prompts, rather than relying on static user profiles.
-
Add a factuality filter for rationale generation that removes fabricated evidence before preference optimization. The system will verify claims against the actual user history log, discarding any rationale that references non-existent events, ensuring the model learns from true behavioral signals only.
-
Enforce reasoning-before-labeling in the output format. The system will be trained to produce a structured reasoning trace first, then a final engagement label. This alone improves Macro-F1 by 4.21% in zero-shot settings and is a prerequisite for the alignment gains.
What the improved AI system can do:
-
Predict user engagement (e.g., click, watch, like) from raw text interaction logs with accuracy statistically indistinguishable from a heavily feature-engineered production system, but without any manual pipeline maintenance.
-
Generate human-interpretable reasoning traces that explicitly cite which user-history signals (e.g.,
user watched 3 action movies in the last hour,
current time is 11 PM
) drove the prediction, enabling debugging and trust. -
Detect when its own prediction is unreliable due to underspecified trade-offs (e.g., novelty vs. popularity) and flag that case for human review or additional context, rather than emitting a confident but arbitrary answer.
-
Avoid contradictory outputs on identical evidence—if the system predicts
likely to engage
for a user, it will not also produce a fluent, equally plausibleunlikely to engage
argument for the same item, unless the underlying trade-off is genuinely ambiguous (in which case it flags it). -
Scale to new domains or platforms by fine-tuning on paired rationales from that domain’s logs, without needing to re-engineer features, because the model learns to reason directly from text.
Sources
- Reasoning Models Don't Always Say What They Think
- OneRec-Think: In-Text Reasoning for Generative Recommendation
- When Two LLMs Debate, Both Think They'll Win
- Towards Understanding Sycophancy in Language Models
- Reason-to-Recommend: Using Interaction-of-Thought Reasoning to Enhance LLM Recommendation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection