Balancing Optimality and Diversity: Human-Centered Decision Making through Generative Curation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Balancing Optimality and Diversity: Human-Centered Decision Making through Generative Curation".
Jane: The paper was written by Michael Lingzhi Li and Shixiang Zhu from Harvard Business School and Carnegie Mellon University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the decision-making world, called “Balancing Optimality and Diversity: Human-Centered Decision Making through Generative Curation.” Jane, I have to say, just the title alone got me excited.
Jane: Oh, absolutely, Tom. And I love that it’s from Michael Lingzhi Li at Harvard Business School and Shixiang Zhu at Carnegie Mellon. These are folks who really understand how decisions actually get made in the real world, not just in theory.
Tom: Right, and that’s the key thing here. This paper is about a fundamental shift in how we think about algorithms that help people make decisions. Instead of just spitting out one “best” answer, this framework generates a whole set of options.
Jane: Exactly. And the reason that matters is because in so many real situations, the algorithm doesn’t have the full picture. There are always those unspoken, hard-to-quantify factors that a human decision-maker brings to the table.
Tom: Like what? Give me a concrete example.
Jane: Think about school bus routing. An algorithm can optimize for time and cost, but the school administrator knows that a certain route might cut through a neighborhood where parents are worried about traffic. That’s a qualitative factor the algorithm just can’t see.
Tom: So the algorithm gives you a bunch of good options, and the human picks the one that feels right for reasons the computer can’t understand. That’s the whole premise.
Jane: You got it. And the paper formalizes this beautifully. They call it “generative curation.” The idea is you’re not just finding one optimal solution, you’re curating a small portfolio of solutions that are both quantitatively good and qualitatively diverse.
Tom: And that’s the balancing act, right? You don’t want to give someone a hundred options, because that’s overwhelming. But you also don’t want to give them five options that are all basically the same.
Jane: Precisely. The paper is all about finding that sweet spot. And the authors have built a mathematical framework to figure out exactly what that sweet spot looks like, depending on how much we trust the quantitative model versus how much we think the human’s gut feeling will matter.
Tom: So it’s not just a nice idea, it’s a rigorous framework. I’m really curious to see how they actually model that unknown human preference. That’s got to be the hardest part.
Jane: It is, and that’s exactly what we’re going to dig into next. They use some clever statistics to represent that unknown “gut feeling,” and the results are pretty surprising.
Tom: Can’t wait. Stick around, folks, because we’re just getting to the good stuff.
Summary of the Paper: Jane: So, Tom, we just set the stage. Now let’s talk about the core of “Balancing Optimality and Diversity.” The authors tackle that hard problem I mentioned: how do you model the unknown human preference?
Tom: Right, the “gut feeling” part. I’m all ears.
Jane: They assume that this unobservable qualitative desirability follows something called a Gaussian Process. Now, don’t let the name scare you. Think of it as a very flexible statistical tool that describes a smooth, continuous landscape of unknown preferences across all possible actions.
Tom: So instead of saying “the human likes this one thing,” it says “the human’s preference is a smooth surface, and actions that are close together probably have similar hidden appeal.”
Jane: Exactly. And that assumption is powerful because it lets them do some serious math. They can actually derive a formula for the objective, which is the expected desirability of the best option in a set of recommendations.
Tom: And what does that formula tell us? I bet it’s not just “pick the best ones.”
Jane: No, it’s not. It reveals a fundamental trade-off. On one side, you have the quantitative quality, which is the average score of the actions you’re recommending. On the other side, you have something they call a diversity metric.
Tom: And that diversity metric is the key, right? It’s not just about being different for the sake of being different.
Jane: Precisely. The metric is based on the correlation between the hidden preferences of different actions. If two actions are very similar, their hidden preferences are highly correlated. That means if the human doesn’t like one, they probably won’t like the other either.
Tom: So you want to pick actions that are not just far apart in the obvious, measurable way, but also in this hidden preference space.
Jane: You hit the nail on the head. And the paper shows that the optimal strategy is to maximize the expected quantitative score plus a bonus that’s proportional to this diversity. The more uncertain you are about the human’s hidden preferences, the bigger that diversity bonus becomes.
Tom: So the algorithm is essentially saying, “I’m not sure what you really want, so I’m going to give you options that cover a lot of different potential preferences.”
Jane: Exactly. And there’s a really cool result about this. They show that a naive approach to diversity, like just maximizing the pairwise distance between solutions, actually leads to a bad outcome. It collapses to just two extreme options.
Tom: Wait, really? That seems counterintuitive.
Jane: It is, but it makes sense. If you only care about distance, the best you can do is pick the two farthest points. But that gives the human almost no choice in the middle. The Gaussian Process approach is smarter because it balances the quantitative quality with the qualitative spread in a principled way.
Tom: So it’s not just about being different, it’s about being different in a way that’s likely to match some unknown preference. That’s a huge insight. I’m really curious to see how they actually build a system that does this.
Jane: And that’s our next topic. They have two different ways to implement this, and they’re both pretty clever.
Improvements Suggested by the Paper: Tom: Alright, Jane, so we’ve got this great theory. But how do you actually use it? How do you make a computer generate these curated sets of options?
Jane: That’s the million-dollar question, and the paper offers two distinct solutions. The first is a deep generative approach. They basically train a neural network to learn the optimal distribution of actions.
Tom: So instead of the network learning to predict a single answer, it learns to produce a whole distribution of answers that maximizes that objective we talked about.
Jane: Exactly. You feed it random noise, and it outputs a set of actions. The network is trained to balance the quantitative score and the diversity metric. It’s a really elegant way to use modern AI.
Tom: But I’m guessing that doesn’t work for everything. What about problems where the action space is discrete, like a routing problem where you have to pick a sequence of stops?
Jane: You’re right, and that’s where the second method comes in. It’s called Diversified Iterative Search. Instead of learning a whole distribution at once, it builds the set of solutions one at a time.
Tom: Like a greedy algorithm?
Jane: Sort of. At each step, it looks at the solutions it has already picked and finds a new one that maximizes the quantitative score plus a bonus for being diverse from the existing set. It’s a sequential approach that can be plugged into any existing optimization solver.
Tom: So you can take a classic integer programming solver, which is great at finding one optimal solution, and modify it to find a diverse set of near-optimal ones.
Jane: Exactly. And that makes the framework incredibly versatile. They tested it on a real-world problem: redesigning police zones in Atlanta. The goal was to balance workload across zones, but the police also had to consider things like highway access and neighborhood integrity, which are hard to model.
Tom: And what happened? Did the algorithm’s recommendations make sense?
Jane: They did. The paper shows that the plans generated by their method were not only good on the quantitative metric of workload variance, but they also looked reasonable. In fact, one of the generated plans closely resembled the plan that the Atlanta Police Department actually adopted in two thousand nineteen.
Tom: That’s a fantastic validation. The algorithm, without being told about highways or neighborhood character, produced a plan that a human team ultimately found acceptable.
Jane: It’s a strong signal that the diversity metric is capturing something real. It’s not just a mathematical trick; it’s encoding the idea that a good set of options should be robust to the factors we can’t easily define.
Tom: So we have a framework, we have two implementations, and we have a real-world success story. What’s the big takeaway for the world?
Jane: I think the big takeaway is that we can build decision-support tools that truly complement human judgment, rather than trying to replace it. And that’s a huge deal for high-stakes fields like healthcare, public policy, and logistics.
Conclusion: Tom: Well, Jane, we’ve covered a lot of ground on “Balancing Optimality and Diversity.” Let’s wrap it up for our listeners.
Jane: Absolutely. We started with the core problem: algorithms often don’t know the full picture of what a human decision-maker values. This paper provides a formal way to handle that by generating a curated set of options, not just a single answer.
Tom: And the key was that diversity metric. It’s not about being different for the sake of it. It’s about covering the space of possible hidden preferences so that the human is more likely to find an option they truly like.
Jane: Right. And they showed you can implement this with a neural network for continuous problems, or with an iterative search for discrete problems like police districting. The Atlanta case study was really compelling.
Tom: It really was. It showed the framework isn’t just theoretical. It can produce plans that are both quantitatively sound and qualitatively acceptable to the people who have to live with them.
Jane: And that’s the ultimate goal, isn’t it? To build AI systems that work with humans, not against them. This paper gives us a principled way to do that, ensuring that our recommendations are useful even when our objectives are incomplete.
Tom: It’s a powerful vision for the future of human-centered AI. We’re sad to see this paper go, but we’re excited to see what comes next.
Jane: For sure. A big thank you to Michael Lingzhi Li and Shixiang Zhu for this work. And thank you, our listeners, for joining us today.
Tom: We’ll be back soon with another paper. Until then, keep asking the big questions.
Michael Lingzhi Li, Shixiang Zhu
Harvard Business School · Carnegie Mellon University
cs.LG, cs.HC, math.OC
Submitted: 2026-08-14
Updated: 2026-08-18
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 70/100
Key concepts
- Generative Curation
- A framework that moves beyond finding a single optimal solution. Instead, it generates a small portfolio of diverse options designed to be both quantitatively good and qualitatively varied for human consideration.
- Gaussian Process
- A flexible statistical tool used in the paper to model unknown human preferences. It assumes that the human's unobservable 'gut feeling' forms a smooth, continuous landscape across all possible actions.
- Diversity Metric
- A measure of how different recommended options are, not just physically, but also in their hidden preference space. It is based on the correlation between the unknown preferences of different actions.
- Diversified Iterative Search
- A versatile method for implementing the framework in discrete problems. It builds a set of solutions sequentially by finding a new option that maximizes both quantitative score and diversity from previously selected options.
Terminology
Summary
Summary
This paper introduces a framework called generative curation for human-centered decision-making, where algorithms generate a portfolio of candidate solutions (e.g., treatment plans, delivery routes, policy options) and a human decision-maker retains final authority to select the most desirable option. The authors argue that in many real-world settings, decision quality depends not on a single algorithmic optimum
but on whether the recommendation set contains at least one option the human ultimately deems desirable, especially when desirability depends on both observable objectives and unobserved qualitative considerations.
The paper formalizes the problem as follows: the true human desirability of an action a is decomposed into a quantitative component Y(a) (based on observable features) and a qualitative component U(a) (based on unobservable factors such as stakeholder preferences, political feasibility, or community acceptance). The goal is to find a generative distribution pi over actions such that, when m actions are sampled from pi, the expected maximum desirability is maximized:
[
pi E [A 1,,A m about pi (A i)] = pi E [A 1,,A m about pi (Y(A i) + U(A i))].
]
Under the assumption that U(a) follows a zero-mean stationary Gaussian process with covariance function k(a,a'), the authors derive tight upper and lower bounds for this objective (Proposition 1). The bounds reveal a fundamental trade-off between quantitative optimality and qualitative diversity, formalized through a novel diversity metric rho[pi], defined as the expected correlation between qualitative desirability components of sampled actions:
[
rho[pi] = 1 over sigma squared E[k(A i, A j)].
]
The bounds take the form:
[
pi E [A 1,,A m about pi (A i)] pi (E [(Y(A 1),,Y(A m))] + sigma sqrt 1-rho[pi] E m),
]
[
pi E [A 1,,A m about pi (A i)] pi (E A about pi[Y(A)] + sigma sqrt 1-rho[pi] E m),
]
where E m = integral-infinity infinity x d over dx[(x)] m dx and is the standard normal CDF.
The paper provides several theoretical results (Proposition 2, Corollary 1) characterizing the optimal policy pi* as a function of sigma (the variance of qualitative desirability) and m (the number of recommendations). Key findings include: (i) as sigma increases, the optimal policy becomes more diverse and less quantitatively optimal; (ii) as sigma to infinity, the optimal policy either degenerates to a point distribution or satisfies an integral equation involving the kernel; (iii) for a white noise kernel, the optimal asymptotic distribution is uniform; (iv) for a Gaussian kernel, the optimal distribution tends to form multiple clumps.
The authors also show that maximizing pairwise distance (a common heuristic for diversity) leads to a degenerate solution concentrating on just two points (Proposition 3), highlighting the value of their principled diversity metric. Additionally, they extend the framework to incorporate human preference feedback (Proposition 4), showing how binary preference signals (e.g., I prefer action 1 over action 2
) can be used to update the posterior distribution of U(a), progressively refining the model of qualitative desirability.
Two implementations are proposed: (1) Neural Net Generative Curation (NN-GC) using a generative neural network with the reparameterization trick to directly parameterize pi, and (2) Diversified Iterative Search for Generative Curation (DIS-GC) using a sequential optimization approach that iteratively adds solutions to the portfolio, suitable for discrete/combinatorial problems.
The paper validates the framework on three synthetic settings (1D Gaussian, 2D Ackley, and Knapsack) and a real-world police redistricting problem in Atlanta. Results show that both NN-GC and DIS-GC consistently reduce expected regret compared to baselines (Random, Quantitative Optimizer, Noisy QO, and Iterative Search maximizing pairwise distance). For example, in the 1D Gaussian setting, NN-GC achieves an expected regret of 0.005 versus 0.261 for QO and 0.215 for Random. In the police redistricting problem, DIS-GC achieves the lowest regret when sigma in [10, 1000], and one of the generated plans closely resembles the plan actually adopted by the Atlanta Police Department in 2019.
The paper concludes by discussing limitations, including the stationarity and Gaussian assumptions on qualitative desirability, and suggests future work on relaxing these assumptions and considering scenarios where humans partially deviate from algorithmic recommendations.
Improvements for AI systems
Based on the paper, here are specific improvements to AI systems and what the improved systems can do:
1. Replace single-optimal-output AI with generative curation for human-in-the-loop decisions
-
Improvement: Instead of outputting one
best
action, the AI learns a probability distribution over actions (π) and samples m diverse, near-optimal candidates. The objective is to maximize the expected desirability of the best option in the set, not the average. -
What the improved AI can do: In clinical decision support, instead of recommending one treatment plan, it generates 5-10 plans that are quantitatively good but qualitatively diverse (e.g., different drug combinations, different dosing schedules). The physician selects the one that best fits unmodeled patient preferences, allergies, or lifestyle constraints. This reduces the risk of recommending a plan that is optimal on paper but unacceptable in practice.
2. Incorporate a principled diversity metric derived from Gaussian process assumptions
-
Improvement: The AI uses a new diversity term, σ√(1−ρ[π])E m, where ρ[π] is the expected correlation between qualitative desirability of sampled actions, σ is the variance of unobserved qualitative factors, and E m is the expected maximum of m standard normals. This replaces ad-hoc pairwise-distance or entropy-based diversity measures.
-
What the improved AI can do: In logistics route planning, the AI generates multiple delivery routes that are not just spatially spread out but are uncorrelated in unobserved factors (e.g., traffic unpredictability, driver familiarity, road construction). The metric ensures that if one route fails due to an unmodeled issue, another route is likely to succeed, rather than all routes failing together.
3. Calibrate diversity based on the uncertainty of qualitative factors (σ)
-
Improvement: The AI explicitly estimates or tunes σ (the variance of unobserved qualitative desirability). When σ is high (qualitative factors dominate), the AI increases diversity; when σ is low, it focuses on quantitative optimality. The paper proves that the optimal policy's diversity is monotonically decreasing in σ.
-
What the improved AI can do: In public policy (e.g., COVID-19 interventions), when the AI is uncertain about political feasibility or public acceptance (high σ), it generates a wider range of policy options (e.g., different levels of lockdown, different economic support packages). When σ is low (e.g., well-understood epidemiological factors), it focuses on the quantitatively optimal policy. This prevents over-diversification when it's unnecessary and under-diversification when it's risky.
4. Use the iterative curation algorithm (DIS-GC) for combinatorial problems
-
Improvement: For discrete optimization (e.g., knapsack, police districting, school bus routing), the AI uses a sequential approach: at each step, it adds a new solution that maximizes the current objective plus the diversity term, considering already-selected solutions. This is compatible with existing integer programming solvers.
-
What the improved AI can do: In police districting, the AI generates 5-10 alternative zone configurations that are all near-optimal in workload balance but differ in shape, highway access, and neighborhood integrity. The police chief selects the one that best fits community feedback and operational constraints, which are hard to quantify. The paper shows this method outperforms single-solution optimization and random search in real Atlanta police data.
5. Implement a neural network generative policy (NN-GC) for continuous action spaces
-
Improvement: The AI uses a deep generative model (e.g., a neural network with reparameterization) to directly parameterize the distribution π. The network is trained by gradient descent on the lower-bound objective: E[Y(A)] + σ√(1−ρ[π])E m.
-
What the improved AI can do: In continuous control or treatment dosing, the AI learns a smooth distribution over actions (e.g., drug dosages, robot trajectories) that balances quantitative performance with robustness to unmodeled human preferences. It can generate a batch of 20 diverse yet high-quality options in milliseconds, allowing rapid iteration if the human rejects the first batch.
6. Adaptively refine recommendations based on human preference feedback
-
Improvement: The AI uses a Bayesian update on the qualitative desirability function U(a) after each human choice. Given a preference like
option A is better than option B,
the AI updates the posterior distribution of U(a) using the closed-form conditional distribution derived in Proposition 4. -
What the improved AI can do: In a multi-stage decision process (e.g., treatment planning over weeks), the AI learns from the physician's choices which qualitative factors matter (e.g., avoiding certain side effects, preferring oral over injectable). After a few rounds, it generates recommendations that are increasingly aligned with the physician's implicit preferences, reducing the number of rejected batches and improving long-term decision quality.
7. Avoid the pairwise distance maximization
trap
-
Improvement: The AI does not use simple pairwise distance (e.g., L2 norm) as a diversity objective. The paper proves that maximizing pairwise distance leads to a degenerate distribution that concentrates on just two extreme solutions, which is not truly diverse.
-
What the improved AI can do: In recommendation systems for complex tasks (e.g., designing a new product, choosing a marketing strategy), the AI avoids the common mistake of generating two extreme, opposite solutions. Instead, it generates a set of solutions that are spread across multiple
clumps
(as shown for Gaussian kernels), ensuring that the human has genuinely different options that cover a range of plausible preferences, not just the two ends of a spectrum.
8. Provide a theoretical guarantee on the regret bound
-
Improvement: The AI's objective is directly tied to minimizing expected regret: R(π) = l(a*) − max i l(a i). The paper shows that the proposed framework's lower bound is tight when Y is constant or m=1, and the diversity term σ√(1−ρ[π])E m is a principled approximation of the qualitative gain.
-
What the improved AI can do: In high-stakes decisions (e.g., emergency response planning), the AI can provide a quantitative estimate of the worst-case regret (95% upper bound) for a given set of recommendations. This allows the human decision-maker to know how much they are sacrificing in quantitative terms to gain qualitative robustness, enabling more informed trade-offs.
9. Handle the unknown σ
case gracefully
-
Improvement: The AI can be run with a range of σ values (e.g., 0, 1, 10, 100, 1000) to produce a family of recommendation sets, from purely quantitative to highly diverse. The paper's experiments show that the true regret is minimized when σ is set close to the true qualitative variance.
-
What the improved AI can do: In a new domain where the importance of qualitative factors is unknown, the AI presents the human with multiple
tiers
of recommendations: a conservative set (low σ, high quantitative quality), a balanced set (medium σ), and an exploratory set (high σ, high diversity). The human can then see which tier best matches their needs, effectively calibrating the AI's behavior without requiring them to articulate their preferences upfront.
10. Integrate with existing optimization solvers without changing their core
-
Improvement: The DIS-GC algorithm only requires adding a diversity term to the objective function at each iteration, making it a drop-in enhancement for any existing solver (e.g., Gurobi, CPLEX, simulated annealing).
-
What the improved AI can do: In school bus routing or hospital scheduling, the AI can be deployed without replacing the existing optimization engine. It simply modifies the objective to include the diversity term, generating a set of m solutions instead of one. This minimizes implementation cost and disruption, making the human-centered approach practical for real-world operational teams.
Sources
- Active Preference-Based Gaussian Process Regression for Reward Learning
- Learning to Make Adherence-Aware Advice
- Auto-Encoding Variational Bayes
- Diffusion Models for Black-Box Optimization
- Pareto Set Learning for Neural Multi-objective Combinatorial Optimization
- Language Models are Few-Shot Learners
- Black-Box Optimization with Implicit Constraints for Public Policy
- Data-Driven Optimization for Police Beat Design in South Fulton, Georgia
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks