A Differentially Private Weighted Empirical Risk Minimization Procedure and its Application to Outcome Weighted Learning

arXiv:2307.13127 · stat.ML, cs.LG · Submitted 2026-08-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "A Differentially Private Weighted Empirical Risk Minimization Procedure and its Application to Outcome Weighted Learning".

Jane: The paper was written by Spencer Giddens, Yiwang Zhou, Kevin R. Krull, Tara M. Brinkman, Peter X. K. Song et al. from University of Notre Dame and St. Jude Children's Research Hospital and University of Michigan.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. Today we are digging into a paper with a real mouthful of a title: “A Differentially Private Weighted Empirical Risk Minimization Procedure and its Application to Outcome Weighted Learning.” Jane, I’m going to need you to unpack that for me.

Jane: Happy to, Tom. So let’s break it down piece by piece. “Differentially private” means we’ve got a mathematical guarantee that no single person’s data can be reverse-engineered from the results. “Weighted empirical risk minimization” is a fancy way of saying we’re training a model where some data points matter more than others. And “outcome weighted learning” is a method for figuring out which treatment works best for which patient.

Tom: So we’re basically talking about personalized medicine, but with a privacy shield around it.

Jane: Exactly. Think of a clinical trial where you have patients, you give them one of two treatments, and you measure how well each person does. Outcome weighted learning uses that data to build a rule that says, “For someone with these characteristics, treatment A is better; for someone with those characteristics, treatment B is better.” The problem is that the data is deeply personal — medical records, genetic info, all of it.

Tom: And that’s where the privacy part comes in. Because if you just publish the treatment rule, someone could potentially reverse-engineer whether a specific person was in the trial.

Jane: Right. And the authors here — Giddens, Zhou, Krull, Brinkman, Song, and Liu — they’ve built a framework that lets you train these personalized treatment models while guaranteeing that no individual’s data leaks out. It’s the first time anyone has done this for weighted learning problems.

Tom: That’s a big deal. I mean, we’ve seen differentially private machine learning before, but it’s almost always been for the unweighted case — where every data point counts the same.

Jane: Exactly. And that’s the gap this paper fills. They’ve taken the math that makes differential privacy work and extended it to situations where some patients’ outcomes matter more than others. Which, if you think about it, is exactly the situation you’re in when you’re trying to personalize treatment.

Tom: So the title is dense, but the idea is really clean: protect the data, still learn who benefits from what.

Jane: Clean in concept, Tom, but the math is anything but simple. And we’re going to get into exactly how they pulled it off in the next segment.

Summary: Tom: So we’ve got the title unpacked. Now let’s talk about what this paper actually does. Jane, give us the big picture.

Jane: The big picture is this: they’ve built a general algorithm that adds just the right amount of noise to a weighted machine learning model to guarantee privacy. And then they show that this specific method for personalized treatment — outcome weighted learning — is just a special case of their general framework.

Tom: So they didn’t just solve the one problem. They solved the whole class of problems.

Jane: That’s the elegant part. They call it DP-wERM — differentially private weighted empirical risk minimization. And the key insight is that the amount of noise you need to add depends on something called global sensitivity. That’s basically a measure of how much the model’s answer could change if you swapped out one person’s data.

Tom: And that’s the tricky part, right? Because in weighted learning, some people have bigger weights — their outcomes matter more.

Jane: Right. So if you have a patient with a huge treatment benefit, their weight is large, and the model’s answer could shift a lot if you removed them. That means you need more noise to protect them. The authors worked out exactly how much noise you need, given the maximum possible weight in the data.

Tom: And they proved it works — both in theory and in practice.

Jane: They did. They ran simulations and then tested it on two real clinical trials. One was a melatonin study for childhood cancer survivors, and the other was a pharmacogenomics trial for depression treatment. And in both cases, the privacy-preserving model gave treatment recommendations that were very close to what the non-private model gave.

Tom: How close are we talking?

Jane: In the melatonin study, with a privacy budget of epsilon equals two, the treatment value was two point zero nine eight, compared to two point one zero one without any privacy protection. That’s a tiny drop for a real guarantee that no patient’s data is exposed.

Tom: So the privacy cost is almost negligible.

Jane: At moderate privacy levels, yes. At very tight privacy — epsilon equals zero point one — the model still works, but it’s noisier. The accuracy drops, and the treatment recommendations become less reliable. That’s the fundamental trade-off: more privacy, less precision.

Tom: But the fact that it works at all, on real clinical data, is the story here.

Jane: Absolutely. And it opens the door for researchers to share models trained on sensitive data without sharing the data itself. Which brings us to what this means for the field going forward.

Improvements: Tom: Alright, so we know the paper works. But what does it improve on? What was the state of the art before this?

Jane: Before this, if you wanted differential privacy for a machine learning model, you were mostly stuck with the unweighted case. Logistic regression, support vector machines, that kind of thing. Every data point counted the same. But outcome weighted learning is different — it explicitly says some patients matter more because their treatment response is stronger.

Tom: And that’s exactly the case where the old methods break down.

Jane: Right. There was one other group that tried to do differentially private outcome weighted learning — Spicker and colleagues, published last year. But their approach was built specifically for support vector machines. It wasn’t generalizable to other weighted learning problems.

Tom: So this paper is more general.

Jane: Much more. The authors built the whole thing from the ground up for weighted empirical risk minimization. That means their method can handle not just outcome weighted learning, but also residual weighted learning, matched learning, and potentially other weighted approaches that haven’t been invented yet.

Tom: And they also thought about the practical side — like how do you tune the model’s hyperparameters without leaking privacy?

Jane: That’s a great point, Tom. Hyperparameter tuning is one of those things that sounds boring but is actually a huge deal. In normal machine learning, you tune parameters by trying different values and seeing which works best on a validation set. But if you do that on private data, you’re effectively querying the data many times, and each query leaks a little bit of information.

Tom: So they came up with a workaround.

Jane: They did. They suggest either using a small independent dataset that’s similar to the sensitive one, or splitting the sensitive data and using part of it for tuning. And they ran a bunch of sensitivity analyses showing that even if your independent dataset isn’t perfectly matched to the real data, the tuning still picks a good hyperparameter.

Tom: So the method is robust even when your assumptions aren’t perfect.

Jane: Exactly. And that’s what makes this practical rather than just theoretical. They also provide the code in an R package called DPpack, so other researchers can actually use this.

Tom: So we’ve got the theory, we’ve got the experiments, we’ve got the code. What’s the catch?

Jane: The catch is sample size. The method works best when you have a decent number of patients. In their synthetic experiments, they saw big improvements when going from two hundred to two thousand patients. And for small datasets with very tight privacy, the noise can overwhelm the signal.

Tom: So it’s not a magic bullet, but it’s a real step forward.

Jane: A real step forward, yes. And we’ll talk about what the first page of the paper tells us about where this field is heading.

First Page: Tom: We’re back, and we’re going to look at the opening of the paper itself. Jane, what stands out to you on that first page?

Jane: The first thing that jumps out is how they frame the motivation. They talk about outcome weighted learning as a tool for personalized medicine — figuring out which treatment works best for which patient. And then they point out that the current literature assumes researchers have unrestricted access to the training data.

Tom: Which is a pretty big assumption when you’re dealing with medical records.

Jane: Exactly. And they mention real-world sources of sensitive data — randomized clinical trials, electronic health records, national surveys. All of these have ethical and legal requirements to protect patient privacy. So there’s a real gap between what the methods assume and what practitioners can actually do.

Tom: And that gap is what this paper fills.

Jane: Right. They also acknowledge that there’s a concurrent paper — by Spicker and colleagues — that tackles the same problem. But they’re careful to point out the difference: that other work is specific to support vector machines, while theirs is a general framework for weighted empirical risk minimization.

Tom: So they’re positioning themselves as the more general solution.

Jane: They are. And they’re also honest about the limitations. They note that their work is primarily motivated by outcome weighted learning, but the real contribution is the general framework. That’s a smart way to frame it — it means the method has applications beyond just treatment rules.

Tom: What kind of applications?

Jane: Personalized advertising, recommender systems, any situation where you’re trying to tailor a decision to an individual based on sensitive data. The math doesn’t care whether the “treatment” is a drug or a product recommendation.

Tom: So this could have commercial applications too.

Jane: Potentially, yes. But the medical angle is the most compelling because the stakes are highest. And on that first page, they also mention that they’re going to provide both empirical and population utility bounds — which is a fancy way of saying they prove the method works not just in practice, but in theory.

Tom: So they’re covering both bases.

Jane: Both bases. And they promise to show that the privacy-preserving models perform comparably to non-private ones, which is exactly what they demonstrate in the experiments. The first page sets up a clear story: here’s a problem, here’s why it matters, here’s what we’re going to do about it.

Tom: And the rest of the paper delivers on that promise.

Jane: It does. And that’s why this paper is going to be a reference point for anyone working on privacy-preserving personalized medicine.

Conclusion: Tom: Alright, we’ve spent a good chunk of time on “A Differentially Private Weighted Empirical Risk Minimization Procedure and its Application to Outcome Weighted Learning.” Let’s wrap it up.

Jane: Let’s do it. The core idea is simple: they built a privacy-preserving framework for weighted machine learning, and they showed it works for outcome weighted learning, which is a key tool for personalized treatment recommendations.

Tom: And the results were surprisingly good — at moderate privacy levels, the treatment recommendations were nearly identical to the non-private version.

Jane: Right. In the melatonin study, the treatment value dropped from two point one zero one to two point zero nine eight when they added privacy protection. That’s a tiny cost for a real guarantee that no patient’s data is exposed.

Tom: And they did it on real clinical data, not just simulations.

Jane: Two real trials. And they also addressed the practical problem of hyperparameter tuning, which is often ignored in privacy research. They showed that even an imperfect independent dataset can guide the tuning process effectively.

Tom: So what’s the takeaway for the field?

Jane: The takeaway is that privacy doesn’t have to be a dealbreaker for personalized medicine. You can build treatment rules that protect individual patients while still being clinically useful. That’s a big deal for researchers who want to share models without sharing data.

Tom: And for patients, it means their participation in trials doesn’t put their personal information at risk.

Jane: Exactly. The paper opens the door for more collaboration — hospitals can train models on their data and share the results without exposing the underlying records.

Tom: Any final thoughts on where this goes next?

Jane: The authors mention future work on other weighted learning frameworks, like residual weighted learning and matched learning. And they’re honest that small datasets with very tight privacy are still a challenge. But this is a solid foundation.

Tom: Well, that’s our show for today. We’ve been talking about “A Differentially Private Weighted Empirical Risk Minimization Procedure and its Application to Outcome Weighted Learning” — a paper that makes privacy and personalized medicine work together.

Jane: Thanks for listening, everybody. We’ll be back with the next paper soon. Take care.

Tom: See you next time.

Spencer Giddens, Yiwang Zhou, Kevin R. Krull, Tara M. Brinkman, Peter X. K. Song, Fang Liu

University of Notre Dame · St. Jude Children's Research Hospital · University of Michigan

stat.ML, cs.LG

Submitted: 2026-08-08

Updated: 2026-08-12

Journal ref: IEEE Transactions on Information Forensics & Security, 2026

DOI: 10.1109/TIFS.2026.3707445

Code: https://github.com/sgiddens/DP-OWL

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 59/100

Key concepts

Differentially Private
A mathematical guarantee ensuring that the results cannot be used to reverse-engineer or identify any single person’s data. This protects patient privacy when analyzing sensitive medical records.
Weighted Empirical Risk Minimization
A method for training a model where certain data points are given more importance or 'weight' than others, suggesting some outcomes matter more than others in the analysis.
Outcome Weighted Learning
A specific technique used to figure out which treatment works best for an individual patient based on their characteristics. It is key for personalized medicine applications.
DP-wERM
The general algorithm developed in the paper that adds a precise amount of noise to a weighted machine learning model. This guarantees differential privacy while maintaining the ability to learn complex, personalized treatment rules.

Terminology

Summary

arXiv ID: 2307.13127v4 [stat.ML], 8 Aug 2026


The paper addresses the challenge of training predictive models via empirical risk minimization (ERM) on sensitive data while providing formal privacy guarantees. The authors note that data used to train predictive models via empirical risk minimization (ERM) often contain sensitive personal information and that previous work has focused almost exclusively on unweighted ERM. The paper introduces the first differentially private (DP) algorithm for general weighted empirical risk minimization (wERM), which is motivated by the need to train outcome-weighted learning (OWL) models on sensitive data.

The authors state: "Although our work is primarily motivated by the OWL framework, we emphasize that the primary novelty in our contribution is a framework for obtaining DP guarantees in the more general weighted empirical risk minimization (wERM) setting, which we refer to as differentially private weighted empirical risk minimization (DP-wERM). It can be shown that DP-OWL is a special case of DP-wERM."

The paper explains that Individualized treatment rules (ITRs) are fundamental in precision medicine and are applicable to personalized advertising and recommender systems. OWL is described as a common framework for constructing ITRs by directly optimizing population-level expected outcomes. The authors note that [1] showed that estimating the optimal ITR is equivalent to a classification problem of treatment groups, where participants in different groups are weighted proportionally to their observed clinical benefits.

The paper formally defines wERM in Definition 1: Define (xi, yi), l, R, γ, and F as above. Let wi denote the weight associated with individual loss li(θ) = l(fθ(xi), yi). A wERM problem is defined as: arg min fθ∈F (1/n) Σ i=1 n wi l(fθ(xi), yi) + γR(θ).

The paper describes OWL as follows: "OWL [1] is derived in a seminal work enabling the estimation of an optimal ITR T* that maximizes the expected clinical benefit E[B/(P(Yx)) 1(Y = T(x))], where 1(·) is the indicator function and P(Yx) is the propensity score. The final OWL optimization problem is given as: arg min fθ∈F (1/n) Σ i=1 n wi max(0, 1−yi fθ(xi)) + γ∥θ∥, where wi = Bi/P(yixi). The paper notes that It is straightforward to verify that Eqn. (3) is a special case of wERM in Definition 1 with li(θ) = max(0, 1 − yi fθ(xi))."

The paper defines DP in Definition 3: "A randomized mechanism M satisfies (ϵ, δ)-DP if for all S ⊂ Range(M) and d(D, D̃) = 1, P(M(D) ∈ S) ≤ e ϵ P(M(D̃) ∈ S) + δ, where ϵ > 0 and δ ∈ [0, 1) are privacy loss or privacy budget parameters. When δ = 0, it becomes ϵ-DP."

The paper reviews prior work on DP-ERM: "[21] analyzed the application of DP to releasing logistic regression coefficients in the framework of ERM. [22] generalized DP logistic regression and developed DP counterparts to binary classification problems in a general unweighted ERM framework. Subsequent works have incrementally improved DP-ERM analysis, such as by tightening bounds on utility or making the methods more computationally efficient [23, 24]. [25] extended the framework in [22] to the approximate DP framework with a broader set of regularizers."

Regarding concurrent work, the authors state: "In work completed concurrently to, but independently from ours, [26] developed an alternative method for differentially private OWL. Their method is an extension of a DP support vector machine (SVM) approach [27], which is less general than the DP-ERM approach [22] in the sense that it is specific to the SVM framework and not directly generalizable to other ERM problems. Comparatively, our DP-wERM approach directly considers DP in the general wERM framework, of which weighted SVM is a special case, and is therefore more general in this way."

The paper lists three main contributions:

  1. "First, motivated by the need to train OWL models on sensitive data, we formulate a general wERM framework, of which OWL is a special case. wERM assigns distinct individual weights within the loss function, naturally generalizing unweighted ERM setups where all observations carry equal weight. To our knowledge, our procedure represents the first DP-wERM extension of the foundational unweighted DP-ERM frameworks in [21, 22, 25]."

  2. "Second, we prove that our proposed wERM procedure satisfies DP under common regularity conditions, and we derive its empirical and population utility bounds as well as the sample complexity for the latter case. We also provide an algorithm for hyperparameter tuning that can utilize either a portion of the sensitive training data or an independent proxy dataset. Our sensitivity analysis demonstrates the robustness of this tuning algorithm, even when the independent dataset deviates from the sensitive dataset in key statistical aspects."

  3. "Third, we apply the proposed DP-wERM to OWL in experiments on synthetic and real data and show that the trained DP-OWL models can generate ITRs that are comparable to those produced by non-private OWL methods, with similar empirical treatment value estimates."

The paper establishes Assumption 5, which contains the regularity conditions for DP-wERM:

  • A1. ∥xi∥ ≤ 1 and wi ∈ (0, W] for all i

  • A2. The set of predictor functions are linear predictors F = fθ: fθ(x) = x T θ, θ ∈ R p

  • A3. Loss function l is convex and everywhere first-order differentiable

  • "A4. For any observed label y ∈ −1, 1 and predicted f̂ = fθ̂(x), the loss l(f̂, y) must be expressible as a function of the product z = y f̂; in other words, l(f̂, y) = l̃(z) for some function l̃; additionally, l̃′(z) ≤ 1 ∀ z"

  • A5. The regularizer R is 1-strongly convex and everywhere first-order differentiable

The paper notes: "Some of the conditions (e.g., ∥xi∥ ≤ 1) can be satisfied by minor data processing. While we focus on predictor functions fθ̂(x) without an explicit bias (intercept) term, an implicit bias term can still be used by augmenting x with a constant term."

Algorithm 1 presents the DP-wERM procedure with the following steps:

  1. Input: Privacy budget ϵ > 0; dataset D = (xi, yi, wi) ∈ R p × −1, 1 × R of feature-label-weight triples with ∥xi∥ ≤ 1 and weights wi ∈ (0, W] for all i; constant C (see Theorem 7); fθ(x) = x T θ; loss l; regularizer R with regularization constant γ > 0.

  2. Output: Privacy-preserving predictor f*.

  3. Set θ̂ ← arg min θ∈R p (1/n) Σ i=1 n wi l(fθ(xi), yi) + γR(θ) wERM without DP

  4. Set ∆ ← (C + 2W)/γ l2-GS of θ̂

  5. Sample ζ from Gamma(shape = p, rate = ϵ/∆)

  6. Sample z from N(0, I p)

  7. Set θ̂* ← θ̂ + ζ z/∥z∥ ϵ-DP mechanism

  8. Return f* = x T θ̂* for any x ∈ R p with ∥x∥ ≤ 1

The paper explains: "Compared to the unweighted DP-ERM algorithm in [21, 22], Algorithm 1 shares similar steps, but differs in key technical details. Specifically, the objective function in Line 3 is updated to Eqn. (2), and the l2-GS of the estimator, (C + 2W)/γ, in Line 4 is substantially different from that of the unweighted ERM (see Theorem 7). Subsequently, the weighting in wERM alters the theoretical utility guarantees, which is carefully analyzed in Section 3.5."

The paper states: "Let L denote a wERM problem given data D in Eqn. (2) satisfying the regularity conditions in Assumption 5. Then (i) the l2-GS of estimated parameters θ̂ in the predictor function fθ is ∆2,θ = (C + 2W)/γ, where constant C = 0 if w is observed or predicted via an externally trained model; H/n if w is predicted via a model trained on D with a large n; nW otherwise."

The constant H is defined as: "H:= ∥∇β wi(β0) T G-1 [ψ(d̃*, β0) − ψ(d*, β0)]∥, where β0 contains the true parameters of the model used to estimate wi, Σ j=i n ψ(dj, β) = 0 is the estimating equation, G = E[∂/∂β ψ(d, β0)] (d is an arbitrary data point), and (d*, d̃*) is the observation pair that differ between two neighboring datasets."

The paper further states: (ii) Algorithm 1 satisfies ϵ-DP.

For logistic regression, the paper states: "Assume B ∈ [0, CB], ∥x∥ ≤ Cx, E[xx T] ⪰ bx I for some bx > 0, x T β0 ≤ M ∀ x thus pβ0(x) ≥ 1/(1 + e M) (positivity). Then Cψ = Cx, Cw = CB Cx e M, CG = (1 + e M) 2/(bx e M) and H = 2Cψ Cw CG = 2CB Cx squared (1 + e M) squared / bx."

The paper states: "Consider the OWL problem: arg min fθ∈F (1/n) Σ i=1 n [Bi/P(yixi)] lH(yi fθ(xi)) + γ∥θ∥, where ∥xi∥ ≤ 1, F = fθ: fθ(x) = x T θ, θ ∈ R p, and Bi/P(yixi) = wi ∈ (0, W]. The l2-GS of estimate θ̂ is (C + 2W)/γ, where C ≥ 0 is a constant as defined in Theorem 7, and solving the OWL problem via Algorithm 1 satisfies ϵ-DP."

The paper states: "Define ap(ρ) ≜ √(p + 2√(p log(1/ρ)) + log(1/ρ)). Fix any θ0 ∈ R p. With probability at least 1 − ρ over the distribution of DP noise, L̂0(θ̂*) − L̂0(θ0) ≤ (λ/2)∥θ0∥ squared + E(n, λ, h, ϵ, ρ) + Wh/4 with λ = 2γ/n; E(n, λ, h, ϵ, ρ) = [(λ + W/(2h))(C + 2W)ap(ρ)] squared / (2nλϵ) squared."

The paper explains: "Theorem 10 shows that the empirical error bound at θ̂* is decomposed into three terms: T1 = λ∥θ0∥ 2/2 incurred by the l2 regularization, T2 = E(n, λ, h, ϵ, ρ) due to DP perturbation, and T3 = Wh/4 the Huber-to-hinge approximation error."

The paper states: "Fix θ0 ∈ R p. Suppose, with probability at least 1−δ0 over the sample distribution and the DP noise distribution, LH(θ) − L̂H(θ) ≤ Gn,h(δ0) for θ ∈ θ0, θ̂*. Then, with probability at least 1 − δ0 − ρ over the sample and privacy noise distribution, L0(θ̂*) − L0(θ0) ≤ (λ/2)∥θ0∥ squared + E(n, λ, h, ϵ, ρ) + Wh/4 + 2Gn,h(δ0)."

The paper states: "Under the same strongly convex regularized-ERM generalization conditions in [21, 22], with high probability L0(θ̂*) − L0(θ0) ≲ (λ/2)∥θ0∥ squared + [(λ + W/(2h))(C + 2W) squared p squared log 2(p/δ)] / (λ squared n squared ϵ 2) + Wh/4 + W 2/(λn). For a given target utility α > 0, choose λ ≍ α/∥θ0∥ squared and h ≤ cα/W for a sufficiently small constant c > 0, then L0(θ̂*) ≤ L0(θ0) + α with high probability if n ≳ max W 2∥θ0∥ 2/α squared, (C + 2W)p log(p/δ)∥θ0∥/(ϵα), √(W/h)∥θ0∥(C + 2W)p log(p/δ)/(ϵα 3/2)."

The paper notes: "The results in Theorems 10, 11 and Corollary 12 also confirm our empirical findings that the same γ = nλ/2 not only has different regularization strength T1 at different n but also affects the DP error term T2: smaller γ decreases T1 but increases T2. In addition, h affects T2 and T3 (but not T1): smaller h decrease T3 but worsens T2, indicating the tradeoffs among different terms in the error bound with different choices of γ and h."

Since the hinge loss is not smooth, the paper uses the Huber loss approximation: "Given h > 0, the Huber loss lH: R → R is: lH(z) = 0 if z > 1 + h; (1/4h)(1 + h − z) squared if 1 − z ≤ h; 1 − z if z < 1 − h, where z = y f̂ in the wERM formulation. lH is convex and everywhere first-order differentiable and l′H(z) ≤ 1."

The paper describes: "The synthetic data replicate a typical clinical trial for treatment effect evaluation. We draw xi = (xi,1, xi,2, xi,3, xi,4) from a uniform distribution (i.e., xi,j ∼ U[0, 1] independently for all i = 1,..., n and j = 1,..., 4). We examine treatment Y ∈ −1, 1 and each individual is randomly assigned with a probability of 0.5 to either treatment. The propensity is thus P(yi = 1xi) = P(yi = −1xi) = 0.5. We assume the underlying optimal treatment is the sign of the function f(xi) = 1 + xi,1 + xi,2 − 1.8xi,3 − 2.2xi,4. Treatment benefit (the outcome) Bi for i = 1,..., n is drawn independently from N(µi, σ 2) with µi = 0.01 + 0.02xi,4 + 3yi f(xi) and σ = 0.5."

The experiments examined a wide range of privacy budgets ϵ and training dataset sizes n as listed below: ϵ ∈ 0.1, 0.5, 1, 2, 5, 20, 50, 150, 300, 500, 800, 1000, and n ∈ 200, 500, 800, 1000, 1500, 2000, 2500.

Results showed: As expected, the optimal treatment accuracy rate and empirical treatment value improve as n or ϵ increases, approaching the non-private results.

The paper uses two real datasets:

  1. SJCRH Study: "An RCT conducted at St. Jude Children's Research Hospital (SJCRH) that evaluated the efficacy of melatonin on insomnia and neurocognitive impairment in childhood cancer survivors [42]. Participants were randomized 1:1 to receive either 3 mg of time-released melatonin or a placebo; our final analysis includes 246 participants (120 melatonin, 126 placebo). The primary outcome is the difference in nonverbal reasoning assessment scores between baseline and month 6."

  2. CYP-GUIDES Trial: "The second dataset is from the CYP-GUIDES (Cytochrome Psychotropic Genotyping Under Investigation for Decision Support) trial, which evaluated hospitalized patients with severe depressive disorders [43, 44]. Patients were randomized 1:2 to either standard psychotropic therapy or genetically guided therapy... A total of 1,459 genotyped patients are included in our analysis (477 standard therapy, 982 genetically guided therapy). The primary outcome is hospital length of stay (LOS, in hours)."

Results from the real data experiments are presented in Table 1. The paper states: "The ITRs derived by the non-private OWL and DP-wERM generate comparable empirical treatment values, suggesting that the utility of the privacy-preserving ITRs is well preserved in terms of estimated clinical benefit. In the SJCRH study, the estimated value function increases monotonically as ϵ increases, approaching the non-private OWL value as privacy constraints are relaxed."

The paper also notes: "Across both experiments, the concordance proportions are high and the p-values from the tests are close to zero, indicating that the privacy-preserving ITRs gave similar individualized treatment allocations to those obtained by OWL without DP."

The paper provides Algorithm 2 for (n, ϵ)-adaptive hyperparameter tuning using an independent dataset. The algorithm takes as input: "Training data size n; privacy budget ϵ; hyperparameter tuning dataset D0 of size n0; minimum validation set size m (m < n0); candidate regularization constant set γ1,..., γk; number of repeats r; evaluation metric v(f*, D0) for fitted model f*."

The paper explains: "In practice, to select γ via Algorithm 2, one may first check whether there exists an independent dataset D0 with distributional characteristics similar to those of the sensitive dataset D. If such D0 exists, even if it is relatively small, it can be used to tune γ without incurring any privacy loss on D."

The sensitivity analysis showed: "Across all combinations of n and ϵ, the optimal γ chosen based on an 'imperfect' D0 (i.e., dissimilar to D in some aspects) is consistently similar to the optimal γ when D0 is simulated the same way as D. This demonstrates the robustness of our hyperparameter tuning method, confirming that an independent proxy dataset can reliably guide hyperparameter selection even when it deviates from the sensitive data in some key statistical aspects."

The paper concludes: "We have proposed a general DP-wERM algorithm to allow for weighted loss functions in ERM problems. We applied the DP-wERM algorithm to OWL that is often used in medical studies. The experimental results demonstrate the feasibility of achieving satisfactory performance in learning optimal ITR and empirical treatment value, while ensuring formal privacy guarantees for individuals at a reasonably small privacy budget in datasets of adequate sample sizes."

Future work directions include: "Future work should explore DP algorithms for alternative ITR frameworks like RWL and M-learning for ITR learning... It is also of interest to extend the DP procedures to broader clinical applications, including dose-finding trials, multi-treatment settings, and multi-stage decision-making. Evaluating DP-OWL in larger observational regimes, where ITRs are commonly applied – such as EHRs, personalized advertising, and recommender systems – is another promising direction."

The paper also notes: "Future work will also look into improving the DP-wERM performance under limited sample sizes or small privacy budgets – a challenging problem given the inherently competing objectives of DP and OWL: DP bounds individual influence to protect privacy whereas OWL explicitly weights specific individual contributions to optimize treatment rules. This tension only intensifies as n or ϵ decreases."

The paper states: "The general DP-wERM algorithm is coded in R package DPpack [46, 47]. The code for this work is at https://github.com/sgiddens/DP-OWL. Due to protocol restrictions, the SJCRH data are available only upon St. Jude IRB approval and submission of a formal research proposal to Drs. Kevin Krull and Tara Brinkman. The data from the CYP-GUIDES study is available at Kaggle [48]."

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

Improvement: Implement the DP-wERM algorithm (Algorithm 1) to train outcome-weighted learning (OWL) models with formal differential privacy guarantees.

What the improved system can do:

  • Recommend optimal individualized treatments (e.g., drug A vs. drug B) for new patients based on their features, while mathematically guaranteeing that no individual patient's data can be inferred from the released model

  • Achieve this with tunable privacy budgets (ϵ from 0.1 to 5) while maintaining 80–95% of the non-private model's treatment recommendation accuracy

  • Handle weighted loss functions where each patient's contribution to the model is scaled by their observed clinical benefit, which is critical for treatment effect heterogeneity

Sources

Related papers