Measuring the Prevalence of Policy Violating Content with ML Assisted Sampling and LLM Labeling

arXiv:2602.18518 · cs.LG, stat.ME, stat.ML · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Measuring the Prevalence of Policy Violating Content with ML Assisted Sampling and LLM Labeling".

Jane: The paper was written by Attila Dobi, Aravindh Manickavasagam, Benjamin Thompson, Xiaohan Yang and Faisal Farooq from Pinterest.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, folks. Today we’re digging into a paper that’s been making waves in the content safety world, titled “Measuring the Prevalence of Policy-Violating Content with ML-Assisted Sampling and LLM Labeling.” Jane, I gotta say, the title alone tells you this is a team that cares about precision.

Jane: Absolutely, Tom. And it’s from a team at Pinterest — Attila Dobi, Aravindh Manickavasagam, Benjamin Thompson, Xiaohan Yang, and Faisal Farooq. What I love is that this isn’t some abstract academic exercise. These are people who run content safety at a massive scale, with billions of impressions every day.

Tom: Right, and that’s the thing. When you hear “policy-violating content,” you might think of obvious stuff like spam or hate speech. But the paper’s core idea is about measuring *prevalence* — the fraction of user views that actually landed on violating content. Not just what got reported.

Jane: And that’s a huge distinction. Reports only capture what users bother to flag, and a lot of harmful content goes unreported. So the team wanted a metric that reflects the real user experience. They call it exposure-weighted prevalence, which is a fancy way of saying: out of every hundred impressions, how many showed something that breaks the rules?

Tom: Exactly. And the authors are quick to point out that this is complementary to reporting. Reports tell you what people noticed and cared enough to flag. Prevalence tells you what actually happened across the whole platform, even the stuff nobody reported.

Jane: The cool part is the team’s background. You’ve got folks who clearly live and breathe both statistics and engineering. They’re not just proposing a metric; they built a daily pipeline that runs in production. That’s the kind of work that moves the needle.

Tom: And it’s published at KDD two thousand twenty-six which is a big deal. Peer-reviewed, real methodology, real results. This isn’t a blog post — it’s a rigorous system with confidence intervals and everything.

Jane: Right, and that rigor matters because content safety decisions have real consequences. If you’re going to tell a company “your platform is showing X percent violative content,” you better have the math to back it up.

Tom: So we’ve got the title, the authors, and the big idea. But the real meat is in how they actually pull this off — sampling millions of items, labeling them with eye, and doing it all daily. That’s our next segment.

Jane: And trust me, the methodology is clever. They’re not just throwing money at the problem; they’re being smart about where they spend their label budget. Stick around.

Summary and Core Idea: Tom: So Jane, we’re back with “Measuring the Prevalence of Policy-Violating Content with ML-Assisted Sampling and LLM Labeling.” Let’s get into the summary, because the approach is genuinely clever.

Jane: It really is. The core problem is that violative content is rare — maybe half a percent of impressions. If you sample randomly, you’d need a massive sample just to find enough violations to get a stable estimate. That’s expensive and slow.

Tom: And that’s where the ML-assisted sampling comes in. Instead of uniform random sampling, they assign each piece of content a weight based on two things: how many impressions it got, and a risk score from their existing safety models. High-exposure, high-risk content gets sampled more often.

Jane: But here’s the kicker — they don’t just use those risk scores as labels. The scores only influence *which* items get reviewed. The actual estimator reweights everything by the known sampling probabilities, so the final prevalence estimate stays unbiased.

Tom: Right, that’s the design-based approach. They’re not letting the model’s opinion contaminate the measurement. The model just helps you look in the right places, but the math corrects for the fact that you looked there more often.

Jane: And then the labeling side. They use a multimodal LLM — think GPT-four class — to actually decide whether each sampled item violates the policy. The prompt includes the policy text, the image or video keyframes, and metadata. The LLM returns a structured label and a brief rationale.

Tom: Which is wild. A year ago, this would’ve been human reviewers looking at thousands of images a day. Now they’re doing over a million units per day with the LLM, and they claim about a hundred times the throughput at a fraction of the cost.

Jane: But they’re not naive about it. They validate the LLM against gold sets labeled by subject matter experts. They found the LLM’s accuracy was within about three percent of human performance, and F1 within half a percent. That’s basically on par.

Tom: And the whole thing runs daily. That’s the real breakthrough. Previously, prevalence studies were monthly or quarterly because human labeling was so slow. Now they get a fresh estimate every single day, with confidence intervals, and they can drill down by surface, geography, content age — all from one global sample.

Jane: That single-sample drill-down is a beautiful piece of design. Instead of running separate studies for each segment, they store the impression counts per segment for each sampled item. So you can compute prevalence for homefeed, for search, for any country — all from the same data.

Tom: So the summary is: smart sampling to find rare events, LLM labeling to make it fast and cheap, and careful statistics to keep it honest. That’s the whole package.

Jane: And the implications are huge. This isn’t just for Pinterest. Any platform with user-generated content — social media, marketplaces, forums — could use this approach. But before we get too excited, we should talk about what this actually improves and where the limits are.

Tom: Good point. The improvements are real, but there are trade-offs. Let’s dig into that next.

Improvements and Implications: Tom: We’re back with “Measuring the Prevalence of Policy-Violating Content with ML-Assisted Sampling and LLM Labeling.” Jane, let’s talk about what this paper actually improves over the status quo.

Jane: The biggest improvement is speed. The paper reports that LLM labeling reduces end-to-end latency by about fifteen times compared to human review. And cost drops by more than ten times. That’s not incremental — that’s a step change.

Tom: And that speed enables daily measurement. The paper shows they run this pipeline daily for over a year across multiple policy areas. That means they can catch emerging harms within days, not months.

Jane: But there’s another improvement that’s subtler and maybe more important: the metric itself. Prevalence is exposure-weighted, meaning it measures what users actually saw. That’s a much better reflection of harm than just counting how many violative items exist.

Tom: Right, because a piece of content that gets a million views is a bigger problem than one that gets ten views, even if they’re both violations. The weighting captures that.

Jane: And the paper shows this metric correlates only weakly with user reports. That’s a finding in itself — it means reports are missing a lot of what’s actually happening. Prevalence gives you visibility into the unreported harm.

Tom: Now, the improvements also come with practical considerations. The paper is honest about the trade-offs. For example, the ML-assisted sampling can concentrate the sample on high-risk items, but if the risk scores are spiky, you get more variance in your weights, which can actually widen your confidence intervals.

Jane: That’s why they have tunable parameters. You can dial back the risk-score weighting if it’s hurting precision. It’s a balancing act between finding rare violations and keeping the estimate stable.

Tom: And there’s the label error question. LLMs aren’t perfect. The paper includes an optional correction for false positives and false negatives using the Rogan-Gladen method. But they prefer to keep the core metric simple and just monitor label quality over time.

Jane: The other big improvement is the configurable workflow. Each policy area gets its own prompt, its own gold set, its own sampling parameters. That means they can stand up a new measurement for a new policy in days, not weeks.

Tom: So what does this mean for the world? Honestly, this could become a standard for content safety across the industry. Regulators want platforms to measure harm, not just respond to reports. This gives them a defensible, statistically sound way to do that.

Jane: And it opens the door for experimentation. The paper mentions using calibrated risk scores as surrogates for large-scale A/B testing. That means you could test a new intervention and measure its effect on prevalence within days, not months.

Tom: So the improvements are clear: speed, cost, visibility, and configurability. But there are still open questions — like how this handles very rare categories, or how it deals with labeler drift over time.

Jane: Right, and those are the limitations we should talk about. But honestly, for a production system, this is remarkably thorough. Let’s wrap up with our final thoughts.

Conclusion: Tom: Alright, we’re closing out our discussion of “Measuring the Prevalence of Policy-Violating Content with ML-Assisted Sampling and LLM Labeling.” Jane, give us the final summary.

Jane: This paper shows that you can measure the prevalence of policy-violating content daily, at scale, with statistical rigor. The combination of ML-assisted sampling to find rare events and LLM labeling to make it affordable is a genuine breakthrough.

Tom: And the key is that they never sacrifice unbiasedness. The sampling weights are known, the estimator corrects for them, and the LLM labels are validated against human gold sets. It’s a system built on trust.

Jane: The implications are broad. Any platform with user-generated content can adopt this. It gives you a metric that reflects real user experience, supports drill-downs, and enables faster intervention evaluation. That’s a big deal for safety teams everywhere.

Tom: And it’s not just about catching bad content. It’s about measuring whether your interventions actually work. That’s the kind of feedback loop that makes platforms safer over time.

Jane: There are limitations, of course. Very rare categories still need bigger samples or weekly pooling. LLM drift requires constant monitoring. And the paper is honest about all of that.

Tom: But the fact that this runs in production at Pinterest, daily, for over a year — that’s proof it works. It’s not a toy. It’s a real system with real results.

Jane: So we’ll say goodbye to this paper. It’s a strong contribution to content safety, and we’re excited to see how it gets adopted and extended.

Tom: Thanks for listening, folks. Next up, we’ve got a paper on something completely different — I won’t spoil it, but let’s just say it involves a lot of math and a little bit of magic. See you then.

Jane: Take care, everyone.

Attila Dobi, Aravindh Manickavasagam, Benjamin Thompson, Xiaohan Yang, Faisal Farooq

Pinterest

cs.LG, stat.ME, stat.ML

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: 8 pages

Code: https://github.com/facebookarchive/ml_sampler

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

Key concepts

Exposure-weighted prevalence
A metric measuring the fraction of user views that land on violating content. Unlike simple counts, it accounts for how many impressions a piece of content receives, providing a more accurate reflection of the actual harm experienced by users across a platform.
ML-assisted sampling
A method to find rare violations efficiently by weighting content based on impression counts and existing risk scores. While this focuses sampling on high-risk items, mathematical reweighting is applied to ensure the final prevalence estimate remains unbiased and accurate.
LLM labeling
Using multimodal large language models to identify policy violations in sampled content. By analyzing images, video keyframes, and metadata, LLMs offer much higher throughput and lower costs than human reviewers, with accuracy levels nearly on par with human experts.

Terminology

Summary

Summary

This paper presents a production system at Pinterest for measuring the daily prevalence of policy-violating content on the platform. The authors define prevalence as the fraction of user views (impressions) that went to content violating a given policy on a given day, formally expressed as:

[

theta d = sum j C jd Y jd over sum j C jd

]

where C jd is the total impressions of content unit j on day d, and Y jd in 0,1 indicates whether that content violates the policy. The population is the set of content units (e.g., Pins) that received at least one impression on day d. The metric is exposure-weighted, emphasizing user experience: a single item can be posted once but receive many impressions, or none.

The system has three main components. First, ML-assisted probability sampling: each content unit receives an auxiliary risk score s jd > 0 from production safety models, and a sampling weight is defined as:

[

w jd = C nu jd (s jd + epsilon) gamma, nu, gamma 0, epsilon > 0

]

The tunable exponents control efficiency trade-offs: " gamma = 0 yields impression-weighted sampling; gamma = nu = 0 yields (approximately) uniform sampling. Intuitively, gamma > 0 shifts probability mass toward model-high-risk items (increasing the sample positive rate), but can reduce effective sample size when score distributions are spiky. The sampling distribution is p jd = w jd / sum k w kd. Crucially, s jd is used only to prioritize review and does not become a label; the estimator reweights by 1/p jd so prevalence reflects impressions, not model thresholds." Implementation uses weighted reservoir sampling via the Efraimidis–Spirakis key trick, keeping memory O(m) and avoiding materializing the full population.

Second, LLM labeling: sampled items are labeled with a multimodal LLM using the content (image/video keyframes, title, text metadata), an SME-reviewed policy description, and an optimized prompt that returns a structured label and brief rationale. The authors report that the LLM-based labeler achieved accuracy within about 3.2% (relative difference) and an F1 score within about 0.5% (relative difference) of human agent performance on SME-validated gold sets. The system processes over 1M units per day, representing a throughput increase of approximately 100 times that of a human baseline, with labeling latency under 24 hours and cost reduction of about 95% per 1k labels.

Third, design-consistent estimation: the paper uses the Hansen–Hurwitz ratio estimator:

[

d = d over d = 1 over m sum i=1 m z id/p J i d over 1 over m sum i=1 m x id/p J i d

]

where z id = C J i d Y J i d and x id = C J i d. Variance is estimated via Taylor linearization with residuals r id = (1/p J i d)(z id - d x id), giving:

[

(d) about 1 over d squared times 1 over m(m-1) sum i=1 m (r id - d) squared

]

with 95% CIs reported as d plus or minus 1.96 sqrt (d).

A key design goal is one global sample with many pivots. For any segment g (e.g., surface, viewer country, content age), segment prevalence is:

[

dg = sum i=1 m z idg/p J i d over sum i=1 m x idg/p J i d

]

with z idg = C J i d g Y J i d and x idg = C J i d g. The system stores C jdg g in G for sampled items so dashboards can compute (10) or (11) on demand without re-sampling.

The paper also covers uncertainty quantification with effective sample size diagnostics (Kish ESS proxy), optional Rogan–Gladen correction for label error, and sensitivity analysis. For operational alerting, the system uses a 7-day moving average and computes minimum detectable effects (MDEs) via:

[

MDE rel about z 1-alpha/2 + z 1-beta over z 1-alpha/2 sqrt 2 over 7 h d over theta 0

]

where h d is the daily CI half-width. Production results show ML-assisted sampling increases the sample positive fraction by 6–11 times across policy areas relative to impression-only PPS at fixed label budget. The system runs daily for over a year across multiple policy areas, supporting daily monitoring, OKR-style goal tracking, and automated alerting.

The paper includes a fully specified synthetic simulation appendix (Algorithm 1, Table 5) with parameters such as population size N=300, 000, base positive rate p=0.005, heavy-tailed impressions via Pareto(alpha=1.4), and Beta-distributed scores. The simulation demonstrates that ML-assisted PPS yields materially tighter CIs than uniform sampling and PPS-by-impressions at realistic sample sizes.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems, along with what the improved system can do:

Implementation: Add a weighted reservoir sampling layer (Efraimidis–Spirakis algorithm) that uses auxiliary risk scores and impression counts to draw probability samples. The weights follow w = C ν * (s + ε) γ with tunable ν and γ.

What the improved system can do:

  • Draw fixed-size, unbiased samples from massive streams (billions of impressions) in O(m) memory

  • Concentrate label budget on high-risk, high-exposure items, achieving 6–11× higher positive-rate lift than impression-only sampling

  • Maintain design consistency even when auxiliary model thresholds drift, because reweighting corrects for sampling lensing

Implementation: Replace naive prevalence calculations with a Hansen–Hurwitz ratio estimator that reweights each sampled item by its inverse inclusion probability. Store per-item segment counts (surface, geography, content age) so any pivot can be computed without re-sampling.

Implementation: Integrate a multimodal LLM labeling pipeline with: (a) SME-reviewed policy prompts returning structured JSON (label, rationale, confidence), (b) gold-set validation before launch, (c) continuous monitoring via random human validation and periodic gold-set re-evaluation.

Implementation: Add variance estimation using residuals from the ratio estimator, plus a Rogan–Gladen correction for label error (using sensitivity and false-positive rate from validation sets). Provide a sensitivity formula linking CI half-width to ESS and base rate.

Implementation: Build a config-driven pipeline where each policy metric specifies: taxonomy, data sources, sampling parameters, LLM prompt/model, gold sets, quality thresholds, and output schemas. Persist all lineage (sampled IDs, weights, segment counts, LLM outputs, prompt versions, token usage).

Implementation: Add a 7-day moving average for daily prevalence, with MDE-based alerting that accounts for autocorrelation in the series. Use historical residuals to validate MDE thresholds empirically.


The improved AI system can measure the true user-facing prevalence of policy-violating content daily, with:

  • Unbiased estimates despite rare events and heavy-tailed exposure

  • 6–11× sampling efficiency gains over naive approaches

  • 100× labeling throughput at 95% cost reduction

  • Full drill-down capability from a single sample

  • Rigorous uncertainty quantification and drift detection

  • Configurable, auditable, and scalable across multiple policy areas

This transforms content safety from infrequent, expensive audits into a high-frequency, statistically grounded feedback loop that can detect emerging harms, evaluate interventions, and support goal-setting—all without relying on user reports, which are known to be incomplete.

Abstract

Content safety teams need metrics that reflect what users actually experience, not only what is reported. We study prevalence: the fraction of user views (impressions) that went to content violating a given policy on a given day. Accurate prevalence measurement is challenging because violations are often rare and human labeling is costly, making frequent, platform-representative studies slow. We present a design-based measurement system that (i) draws daily probability samples from the impression stream using ML-assisted weights to concentrate label budget on high-exposure and high-risk content while preserving unbiasedness, (ii) labels sampled items with a multimodal LLM governed by policy prompts and gold-set validation, and (iii) produces design-consistent prevalence estimates with confidence intervals and dashboard drilldowns. A key design goal is one global sample with many pivots: the same daily sample supports prevalence by surface, viewer geography, content age, and other segments through post-stratified estimation. We describe the statistical estimators, variance and confidence interval construction, label-quality monitoring, and an engineering workflow that makes the system configurable across policies.

Related papers