Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Consensus vs. Dissent".
Jane: This paper investigates how Large Language Models (LLMs) can dynamically select an optimal social choice-based aggregation strategy for Group Recommender Systems by mimicking nuanced human perceptions of fairness, satisfaction,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Alright everyone, let's talk about this paper from arXiv. It's called "Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders." This work is pretty interesting because it looks at how large language models can actually mimic what people think about fairness, satisfaction, and consensus when making group recommendations <ref:2607.10235#pg0>.
Jane: Exactly. The core idea here is that instead of just picking one recommendation strategy and sticking with it forever, this paper explores how an AI can dynamically choose the best strategy as the group changes or as new preferences come in <ref:2607.10235#pg1>. It claims LLMs can model these nuanced human perceptions to pick a better outcome <ref:2607.10235#pg1>.
Lu: What's really cool about this is that they didn't just use a standard setup; they fine-tuned two specific models, Judgmental Llama and Judgmental OLMo, using some synthetic data generated by another big model to help them learn how to give those human-like judgments <ref:2607.10235#pg4>.
Meng: So you're saying they built these AI judges on top of existing human feedback, but they had to use a lot of extra steps, like using DeepSeek-V3 point 1 for that reasoning data <ref:2607.10235#pg0>? How practical is that process for a real recommendation system?
Lalam: From my side, I see this as a huge step toward making AI recommendations feel more tailored to the actual group dynamics rather than just following a fixed rule <ref:2607.10235#pg4>. The paper shows they used "Knowledge Distillation via Synthetic Chain-of-Thought (CoT) Generation" to fill in some of those gaps in the original survey data <ref:2607.10235#pg3>.
Tom: Right, so they used that synthetic reasoning to create an augmented dataset suitable for fine-tuning, and then they applied ParameterEfficient Fine-Tuning using LoRA to get Judgmental Llama and Judgmental OLMo <ref:2607.10235#pg4>. That seems like a pretty solid engineering approach for training the models.
Jane: It's solid, but it leads into their pipeline where they test six different social choice-based aggregation strategies, like Additive Utilitarian and Fairness Voting <ref:2607.10235#pg2>. They don't just pick one; they calculate outcomes for all of them.
Lu: And then the system doesn't stop there; it evaluates each of those recommendations five separate times, simulating five different survey participants for each strategy <ref:2607.10235#pg5>. That’s what lets them get those fairness, satisfaction, and consensus scores across the board <ref:2607.10235#pg5>.
Meng: Five evaluations per strategy sounds intensive computationally; how much do you have to run that pipeline before you settle on the best one? I need to know if this is something that runs fast enough in a live setting.
Tom: The process involves calculating the average score across those five scores for each of those three metrics—fairness, satisfaction, and consensus—to determine the final score for each strategy <ref:2607.10235#pg5>. They rank these strategies and then present the group with the item based on that top-ranked strategy <ref:2607.10235#pg5>.
Jane: And they even add a final layer by accounting for prior group history, suggesting that the fourth-ranked item might be recommended as a final output for the group <ref:2607.10235#pg5>. That shows they're thinking about context beyond just the immediate choice.
Lalam: What I find really important is how they use these fine-tuned LLMs not just to generate static scores, but to dynamically select the one that maximizes those predicted human-like evaluations <ref:2607.10235#pg1>. It moves away from just applying a fixed rule.
Tom: They did compare their models against base models using the Wasserstein Distance metric, where a smaller distance means better performance in terms of judgment quality <ref:2607.10235#pg4>. And they found that Judgmental Olmo emerged as the best-performing model overall in terms of those distances <ref:2607.10235#pg4>.
Lu: The correlation analysis on that model was interesting too; Pearson correlation coefficients between fairness, satisfaction, and consensus scores generated by Judgmental Olmo ranged between zero point six four and zero point seven eight <ref:2607.10235#pg4>. That's consistent with the original human ground truth they started with <ref:2607.10235#pg4>.
Jane: So, when you put that all together, the user study involved two hundred eighty-four participants who tested this new LLM-based method against traditional static strategies <ref:2607.10235#pg6>. The results showed that this dynamic approach tends to outperform static social choice-based aggregation <ref:2607.10235#pg7>.
Meng: Outperforming a fixed strategy is good, but what about the caveats? If you're running this in a real environment, what are the limitations they pointed out? What doesn't this system do well?
Tom: The paper does flag that there are statistically significant interactions with group configuration for both fairness and satisfaction <ref:2607.10235#pg7>. This means the dynamic approach is really good at adapting to how different groups are structured internally <ref:2607.10235#pg7>.
Lu: For example, in coalitional groups, the LLM system actually outperformed strategies like Approval Voting and Additive Utilitarian <ref:2607.10235#pg8>. That's a specific demonstration of where the dynamic modeling really shines.
Jane: So what does this mean for people who just listen to the show? It means we can build recommendation systems that adapt their decision-making based on how complex or polarized the group of users is <ref:2607.10235#pg9>.
Tom: The overall implication is that this gives us a blueprint for Group Recommender Systems that can account for the nuanced way humans interpret what a recommendation outcome means, while still using those familiar social choice strategies <ref:2607.10235#pg9>.
Jane: We've got the full discussion on how AI can model subjective human preferences in this paper, "Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders." That’s all for this episode.
Conclusion: Tom: So we're wrapping up this look at "Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders." Basically, this paper shows how you can use large language models to pick the best recommendation strategy on the fly based on what a group actually wants.
Jane: It’s about moving past picking just one fixed way to decide what's good for everyone and instead using that AI to test different fairness and satisfaction ideas against each other dynamically.
Lu: The authors did some pretty heavy lifting here, fine-tuning these models using synthetic reasoning data to get them ready to judge recommendations in a human-like way.
Meng: From an engineering standpoint, the idea of testing six different social choice strategies—like additive utilitarian or approval voting—and then letting the AI pick the best one based on five simulated participant scores is a lot of computation.
Lalam: But what makes it interesting is that this dynamic selection actually leads to better results than just sticking with one strategy, especially when you look at how well it handles different group settings.
Tom: Right, and the main point they drive home is that this approach gives us a scalable blueprint for recommender systems that can handle the messy reality of human preference without being stuck in just one simple rule.
Jane: It means we can build recommendation tools that adapt to how different user groups actually behave and what kind of consensus they value most.
Lu: And I think the way they showed those strong correlations between the AI's scores and actual human surveys is really telling, it shows the model is learning something meaningful about what people consider fair or satisfying.
Meng: The caveat they point out is that you still have to be careful because there are significant interactions with group structure; it’s not a magic fix for every single situation.
Tom: True, so while the AI is smart at picking the right tool for a given group configuration, we still need to watch those interaction effects when deploying this in the real world.
Jane: Exactly. So what does this mean for us as people who just listen to the show? It means systems that make choices feel more thoughtful and less blindly fixed.
Lu: It suggests that the future of recommendation isn't just about finding the single right answer, but about exploring a whole landscape of possibilities based on group context.
Tom: We’ll get into how those specific model performance metrics compare next.
Cedric Waterschoot, Nava Tintarev, Francesco Barile
Maastricht University · ACM Reference Format, ACM Conference 17, Washington, DC, USA (implied venue)
cs.CL, cs.IR
Submitted: 2026-07-11
Updated: 2026-07-11
Comments: Full paper accepted at the 20th ACM Conference on Recommender Systems (RecSys 2026)
Journal ref: Cedric Waterschoot, Nava Tintarev, and Francesco Barile. 2026. Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders. In Proceedings of the 20th ACM Conference on Recommender Systems (RecSys 2026)
Code: https://github.com/Cwaterschoot/Consensusvs-dissent-2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 78/100
The gist: This paper investigates how Large Language Models (LLMs) can dynamically select an optimal social choice-based aggregation strategy for Group Recommender Systems by mimicking nuanced human
Key concepts
- Social Choice-Based Aggregation Strategies
- These are mathematical methods used to combine individual preferences within a group to reach a single decision. The paper tested six strategies like Additive Utilitarian and Majority Voting, allowing the system to explore different ways of making a recommendation.
- LLM Fine-Tuning for Judgment
- The researchers trained Large Language Models using survey data and synthetic reasoning to make them act as evaluators. This process taught the LLMs to generate human-like assessments of recommendations based on perceived fairness, satisfaction, and consensus.
- Dynamic Strategy Selection
- Instead of picking one fixed strategy, the system evaluates multiple aggregation methods repeatedly. It uses fine-tuned LLMs to judge these outcomes against human-like metrics. The final recommendation is the one that maximizes these predicted human evaluations for a specific group.
Terminology
Summary
This paper investigates how Large Language Models (LLMs) can dynamically select an optimal social choice-based aggregation strategy for Group Recommender Systems by mimicking nuanced human perceptions of fairness, satisfaction, and consensus. The gist: Our pipeline successfully generates multiple recommendation candidates based on social choice-based aggregation strategies and dynamically selects the one that maximizes these predicted human-like evaluations <ref:2607.10235#pg1>.
LLM Fine-Tuning for Judgment
The researchers fine-tuned Large Language Models (LLMs) to serve as judgmental models capable of delivering human-like evaluations by utilizing survey data. This process involved several key steps:
-
Survey Data: The starting point was
human ground truth presented in Barile [4], Barile et al. [5]
where 288 participants rated four scenarios, resulting in a final dataset containing 1, 152 assessments <ref:2607.10235#pg3>. -
Chain-of-Thought Generation: To address the
limited scale of the original human-annotated dataset,
they employed "Knowledge Distillation via Synthetic Chain-of-Thought (CoT) Generation [18] <ref:2607.10235#pg3>. -
Data Augmentation: They used a teacher model,
DeepSeek-V3 (671B) (via the Ollama cloud API2),
to generate synthetic reasoning for existing human ground truth, resulting in anaugmented dataset suitable for fine-tuning LLMs to generate both reasoning and assessment scores for group recommendations
<ref:2607.10235#pg4>. -
Fine-Tuning: They fine-tuned two open-source LLMs,
Llama3.1-8B-Instruct and OLMo-3-7B-Instruct-SFT,
using ParameterEfficient Fine-Tuning (PEFT) via LowRank Adaptation (LoRA) <ref:2607.10235#pg4>. Acustom label masking function
was applied to the loss function to ensure models learned to replicate the reasoning process and scores rather than the prompt structure and group ratings <ref:2607.10235#pg4>.
Group Recommender System Pipeline
The GRS pipeline is designed around the repeated evaluation of aggregation strategies to determine the most suitable outcome for a given group
<ref:2607.10235#pg5>. The system operates as follows:
-
Strategy Calculation: Starting from the group’s preferences, the system calculates outcomes for six distinct social choice-based aggregation strategies:
Additive Utilitarian (ADD), Approval Voting (APP), Fairness (FAI), Least Misery (LMS), Majority Voting (MAJ), and Most Pleasure (MPL)
<ref:2607.10235#pg5>. -
Strategy Evaluation: The recommendations derived by each strategy are
evaluated five separate times each, simulating the assessment of five different survey participants
<ref:2607.10235#pg5>. -
Scoring: The final score for each strategy is determined by
calculating the average across the five scores for each of the three metrics (fairness, satisfaction, consensus)
<ref:2607.10235#pg5>. -
Selection: The strategies are ranked and
the group is presented with the recommendation based on the top-ranked strategy
<ref:2607.10235#pg5>. A final step involves accounting for prior group history by recommendingthe fourth-ranked item as the final output for the group
<ref:2607.10235#pg5>.
LLM-as-Judge and Evaluation
The fine-tuned LLMs are used to generate human-like assessments of potential recommendations, moving beyond static aggregation <ref:2607.10235#pg1>.
-
LLM Generation: The system generates
multiple recommendation candidates based on social choice-based aggregation strategies and dynamically selects the one that maximizes these predicted human-like evaluations
<ref:2607.10235#pg1>. -
Model Performance: Evaluation compared the fine-tuned models against base models using the Wasserstein Distance (WD) metric, where
A smaller distance (from now on referred to as WD) is better
<ref:2607.10235#pg4>. TheJudgmental Olmo emerged as the best-performing model, achieving the lowest WD scores across all three dimensions
<ref:2607.10235#pg4>. -
Correlation Analysis: Pearson correlation coefficients between fairness, satisfaction, and consensus scores generated by Judgmental Olmo showed
relatively strong correlations (ranging between 0.64 and 0.78), consistent with the strong correlations found in the original human ground truth
<ref:2607.10235#pg4>.
User Study Validation
A user study with n=284 participants was conducted to evaluate the effectiveness of the dynamic, LLM-based method against traditional static strategies <ref:2607.10235#pg6>.
-
LLM Selection: The pipeline generated
five assessments (committee size) per potential recommendation
and averaged them to obtain final scores <ref:2607.10235#pg7>. -
Performance: The LLM-based aggregation strategy selection
tends to outperformed static social choice-based aggregation (Table 3), achieving the highest mean scores for perceived fairness (0.82) and consensus (0.81)
<ref:2607.10235#pg7>. -
Interaction Effects: The analysis showed
statistically significant interactions with group configuration
for both fairness and satisfaction, indicating that the dynamic approach is effective in adapting to specific within-group preference distributions <ref:2607.10235#pg7>. For instance, in coalitional groups, the LLM outperformed strategies like APP and ADD <ref:2607.10235#pg8>.
The study concludes that the dynamic approach provides a scalable blueprint for GRS that accounts for the nuanced human-like interpretation of a recommendation outcome
while retaining transparency through social choice-based aggregation strategies <ref:2607.10235#pg9>.
Improvements for AI systems
-
Fine-tuning LLMs for Human-Aligned Judgment The system can be fine-tuned to generate
human-like evaluations
by using adistilled reasoning dataset... grounded in human assessments of fairness, satisfaction, and consensus.
This allows the LLM to move beyond base model limitations by replicating the diversity found in human responses (sigma about 1.6). -
Dynamic Aggregation Strategy Selection The improved system can
dynamically select the one that maximizes these predicted human-like evaluations
by generating multiple recommendation candidates based on various social choice-based aggregation strategies and then selecting the outcome that achieves the highest average scores forperceived fairness and consensus.
-
Contextual Adaptation to Group Configurations The pipeline can exhibit adaptive behavior based on group dynamics, as demonstrated by finding
statistically significant interactions with group configuration (e.g., minority or coalition),
allowing it to choose strategies based on whether the group is uniform, divergent, minority, or coalitional. -
Robustness Across Aggregation Strategies For complex scenarios like coalitional groups where static strategies might fail, the system can
correctly identified that recommendations generated by MPL might be useful for obtaining the most positive assessments for uniform groups,
showing it can deviate from globally strong strategies when context demands it.
Abstract
Previous work in group recommender systems has demonstrated a sensitivity to the distribution of preferences within a group. Specifically, the selection of the preference aggregation strategy benefits from considering such group configurations. In this paper, we study whether LLMs are able to mimic this sensitivity and to select the ideal aggregation strategy (and corresponding recommendation) according to nuanced human perceptions of fairness, satisfaction, and consensus. We do this by fine-tuning Large Language Models (LLMs) on human survey data to serve as real-time judgmental models within the recommendation pipeline. Using a reasoning dataset distilled from DeepSeek-V3.1 and human ground truth assessments, we develop Judgmental Llama and Judgmental OLMo to simulate group assessments. Our pipeline successfully generates multiple recommendation candidates based on social choice-based aggregation strategies and dynamically selects the one that maximizes these predicted human-like evaluations. We further validate these suggestions in a user study (n=284) and find that our methodology achieved the highest scores for satisfaction and group consensus. Furthermore, we find that LLM judgments are most aligned with human perceptions of fairness, satisfaction and consensus when we also consider interaction effects between our LLM-based method and group configuration (e.g., minority or coalition). These findings give further support for dynamically adapting aggregation strategies to specific within-group preference distributions, and highlight the advantage of using LLMs for an adaptation that is aligned with subjective human judgments.
Sources
- Chat-REC: Towards Interactive and Explainable LLMs-Augmented Recommender System
- Self-Preference Bias in LLM-as-a-Judge
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering