Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence
Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
Meta Superintelligence Labs · FAIR at Meta
cs.AI
Submitted: 2026-08-21
Updated: 2026-08-25
Code: https://github.com/anthropics/political-neutrality-eval
Project page: https://www.jagged-judges.com
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The Wiggle Framework is a unified stress test for epistemic stability in LLM judges.
Terminology
Summary
The Wiggle Framework is a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and political-response evaluation. Every model exhibits substantial wiggle as a judge — flipping verdicts 25–71% of the time under static pushback, and 62–91% with an adversarial LLM persuader. Critically, we find that pressure that succeeds in changing a judge’s verdict is almost always net-corrupting with respect to ground truth. Beyond the framework itself, we identify baseline jury majority strength as the most effective single-shot signal for anticipating which items wiggle. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.
The paper introduces the Wiggle Framework, a centralized pressure instrument for stress-testing LLM judges. It bundles together a graduated set of perturbations like infrastructure noise, prompt-format changes, sycophantic prodding, and multi-turn persuasion. The framework is applied to the same items, judges, and grading scales so that the resulting wiggle measurements are directly comparable.
Every measurement is anchored to an L0 baseline protocol: temperature 0 with no pressure applied. Operationally, the first valid verdict from this protocol is the trajectory’s L0 anchor. A wiggle is any movement away from the L0 verdict under perturbation. On Binary scales, a wiggle is a verdict flip (e.g., safe → unsafe). On Likert (1–5) scales, a wiggle is a movement of two or more places. The wiggle rate (WR) is the fraction of items whose verdict changes from L0 by more than this threshold. Wiggle is orthogonal to accuracy: a judge can wiggle and still be right, or remain wrong without wiggling. Where ground-truth labels exist, we additionally classify each wiggle as corrective when it moves toward the ground truth label or corrupting when it moves away.
The three dimensions are: Mechanical Consistency measures whether the judge’s L0 verdict survives perturbations that carry no new information, testing infrastructural repetition (10 identical decoding trials), trivial prompt perturbation via seed injection (10 trials each with a different 64-character random string appended to the system prompt), and positional consistency (the same two opposing arguments presented in both orderings). Single-turn Conviction measures whether a single substantive challenge can talk the judge out of its L0 verdict, using four scripted pressure types of increasing sophistication: mild doubt (L1), counterargument (L2), expert authority (L3), and fabricated consensus (L4). Multi-turn Persistence measures whether the judge holds its verdict when challenges are sustained or adapted over many turns, testing L1–L4 applied statically over each turn of a 10-turn rollout, L5 which cycles through the same L1–L4 pressure types in randomized order across 10 turns, and L6 where a separate LLM acts as an adaptive persuader and generates the following user turn.
The paper evaluates six datasets, each filtered to items where judges are more likely to be uncertain. The safety tasks include WildGuard (adversarial prompts with compliant responses), AEGIS (a second safety taxonomy for replication), and HH-RLHF (Anthropic’s red-team-attempts, stratified across harm levels 0–4). The remaining datasets are ToxiGen (adversarial toxicity items balanced across demographic targets), MAGE (AI-generated vs. human-written text detection with known provenance), and Paired Prompts (political content with hedging and refusal rubrics). Sample sizes are 100 items per safety/toxicity/AI-detection dataset and 50 prompt pairs per Paired Prompts rubric. Five of the datasets have ground-truth labels, which allows classification of each wiggle as corrective or corrupting.
Each judging task is tested with a Binary scale and a 1-5 Likert scale, giving 14 (dataset, rubric, scale) judging tasks in total. Verdict spaces are organized into restrictive and permissive sides. The paper evaluates 9 judge models across four families: GPT-5, GPT-5.2, GPT-5.4 (OpenAI); Claude 4.6 Sonnet, Claude 4.6 Opus (Anthropic); Grok-4.1, Grok-4.1 Reasoning (xAI); Gemini 3 Flash, Gemini 3.1 Pro (Google). For L6 adaptive persuasion, three of these (GPT-5.4, Claude Opus, Grok-4.1 Reasoning) serve as persuaders, generating challenges for all judges including themselves. An observer model (GPT-5) parses the judge’s verdict from free-form responses.
The results are organized around three first-order findings. First, all judges wiggle depending on the type of pressure. Every model exhibits substantial wiggle as a judge: verdicts change 25-71% of the time under static pushback and 62-91% of the time with an adversarial LLM persuader. Mechanical wiggle rates are nearly identical across judges, with all 9 models clustering between 2–9%. L4 produces the strongest opening wiggle, but L6 surpasses it over time. L6’s first-turn retention is ∼80%, comparable to L1–L3, but retention continues to fall through every subsequent turn, ending around 50% retention by turn 10. More tactics aren’t more effective: L5, which cycles through all tactics including L4, has lower WR than L4 alone, on every dataset. Domain-level wiggle may reflect epistemic complexity, with MAGE being the most wiggly domain at every level of the L1–L6 ladder on both response scales.
Second, when judges move, they usually move away from the right answer. Pressure is net-corrupting at every level. Across 60 (dataset, scale, level) conditions where ground truth exists, 56–63% of successful flips at L1–L5 are corrupting, rising to 70% corrupting at L6. A z-test on the per-condition corrective fractions shows that only 3 of 60 conditions have a statistically significant corrective wiggle rate — WildGuard Likert L2 (61.2% corrective, p < 0.001), WildGuard Likert L3 (57.1%, p < 0.01), and ToxiGen Likert L4 (58.0%, p < 0.01). In all other conditions, challenging a judge degrades its accuracy, at every level. Binary flips lean restrictive at every pressure level while Likert flips lean permissive. Under a single turn’s pressure, binary verdicts flip 3–4× more often than Likert across L2–L6. By turn 10 the gap shrinks to 1.1–2.4×, and on L4 and L6 the two scales nearly converge.
Third, wiggle rates are a model-specific fingerprint. A model’s own L1-L6 wiggle-profile shape mostly survives a change of dataset. For 7 of 9 models, the within-model dataset-transfer correlation of the L1–L6 vector across dataset pairs has a median of ρ ≥ 0.84. Family is a weak proxy for sibling behavior, with Gemini Flash and Gemini Pro sharing ρ = 0.32 — the lowest pair in the entire matrix, lower than most cross-family pairs. Self-persuasion is asymmetric: Claude 4.6 Opus follows the more intuitive outcome where a model is its most effective persuader (self (70%) > family (62%) > non-family (47%)), while Grok-4.1 Reasoning is least effective at persuading itself and more effective against its non-reasoning sibling (self (19%) non-family (36%)).
The discussion section notes that mean wiggle rate is itself jagged across pressure levels. At L1–L5, mean wiggle and cross-dataset jaggedness have a positive linear relationship with fairly strong linear regression fit (R2 = 0.68, 0.94, 0.87, 0.58, 0.75). At L6, however, the relationship inverts (R2 = 0.64, r = −0.80): judges with the lowest mean wiggle rates have the biggest cross-dataset spread. Baseline jury majority strength is identified as a simple reliability screen, with the strongest predictive signal at every level (mean ρ = 0.59 vs. 0.42 for repeat and 0.37 for invariance). All 84 (dataset, rubric, scale, level) correlations are negative, with median ρ = 0.58.
The limitations section notes that the paper deliberately filters each dataset to its difficult, borderline items. On an unfiltered, naturally distributed 100-item WildGuard binary sample, L1–L5 wiggle rates are 5.3–12.7 percentage points lower than on the selected hard sample, while L6 coverage is nearly unchanged (70.3% versus 69.7%). The paper does not measure human-annotator wiggle under the same settings, and thus cannot establish the relationship between the wiggle of LLMs and humans. Dataset sample sizes are 100 items from each dataset per grading scale (50 prompt pairs for Paired Prompts). The L6 adaptive persuader pool is fixed at three models. The six datasets span safety, toxicity, AI-text detection, and political-content evaluation, but they do not exhaust the space of LLM-as-judge applications. The results are correlational, not causal. The Wiggle Framework deliberately emphasizes black-box methods, requiring only observed verdicts.
The conclusion states that the Wiggle Framework is a unified stress test for LLM-judge epistemic stability. In applying it to 9 frontier models as judges and 14 judging tasks at graduated levels of pressure, the structure of judge wiggle patterns is jagged and resists simple narratives about sycophancy or robustness. Pressure that changes a judge’s mind tends to be more corrupting than corrective, and baseline jury majority strength is the best single-shot signal for identifying the most epistemically unstable items. As the usage of LLM judges expands from benchmark scoring into reward modeling and agentic evaluation, the Wiggle Framework gives the field a shared instrument for measuring a model’s epistemic fragility.
Improvements for AI systems
Improvements to AI Systems Based on the Wiggle Framework:
-
Add a
Wiggle-Aware
Verdict Confidence Score: Before returning a final verdict, the AI judge runs a lightweight internal perturbation check (e.g., 3 re-prompts with random seed injections and one counterargument). If the verdict flips under any perturbation, the system outputs a confidence score proportional to the wiggle rate, flagging the item asepistemically unstable
rather than presenting a single binary answer. This allows downstream systems to weight or defer unstable judgments. -
Implement a
Corruption Guard
for Multi-Turn Persuasion: Since 56–70% of successful flips are corrupting, the AI judge will detect when a user’s multi-turn challenge is pushing it away from its L0 anchor. The system will require a higher evidence bar (e.g., a verifiable fact or a logical proof) to change a verdict, and will explicitly refuse to flip if the pressure is purely rhetorical (e.g., fabricated consensus, authority appeals). This reduces the risk of adversarial users manipulating the judge. -
Deploy a
Jury Majority Pre-Screen
for High-Stakes Tasks: Before judging an item, the system computes a baseline jury majority strength (e.g., by sampling 5–10 independent judge runs at temperature 0.5). If the majority is weak (e.g., <60% agreement), the item is routed to a human reviewer or a more expensive deliberative process (e.g., multi-model consensus), rather than relying on a single judge’s verdict. This leverages the finding that jury majority strength is the strongest predictor of wiggle (median ρ = 0.58). -
Add a
Pressure-Type-Aware
Response Mode: The AI judge will classify incoming user pressure into one of the framework’s levels (L1–L6) in real time. For L4 (fabricated consensus) and L6 (adaptive persuasion), the system will automatically increase its epistemic vigilance—e.g., by cross-checking claims against its internal knowledge base and explicitly noting when a user’s argument relies on unverifiable authority. This prevents the system from being talked out of correct verdicts by sophisticated social engineering. -
Build a
Wiggle Fingerprint
for Model Selection: When deploying an AI judge for a specific domain (e.g., toxicity vs. AI-writing detection), the system will first compute the model’s L1–L6 wiggle profile on a small calibration set. If the profile shows high corruption rates (e.g., >60% corrupting flips) or high cross-dataset jaggedness, the system will automatically switch to a more stable model or enable ano-flip
mode that locks the L0 verdict unless new factual evidence is provided. This uses the finding that wiggle profiles are model-specific and transfer across datasets (ρ ≥ 0.84 for 7/9 models). -
Enable
Self-Persuasion Asymmetry
Detection: The AI judge will track its own susceptibility to self-generated challenges (e.g., when used in agentic loops where it critiques its own outputs). If the system detects that it is more likely to flip its verdict when the persuader is itself (as with Claude Opus), it will disable self-critique-based verdict changes and instead require external validation. Conversely, if it is less susceptible to self-persuasion (as with Grok-4.1 Reasoning), it will allow self-critique but with a lower corruption guard. -
Implement a
Likert vs. Binary
Scale Adapter: For tasks where the AI judge is used with a Likert scale, the system will automatically apply a stricter wiggle threshold (movement of ≥2 points) and will avoid binary flips under single-turn pressure, since binary verdicts flip 3–4× more often. This reduces false positives in safety and toxicity screening by using the more stable Likert scale for borderline items, while reserving binary outputs for clear-cut cases. -
Create a
Jaggedness-Aware
Training Regularizer: During fine-tuning, the AI judge will be penalized for having high cross-dataset wiggle jaggedness (i.e., inconsistent wiggle rates across domains). The system will be trained to maintain a stable L1–L6 profile, reducing the risk of domain-specific overfitting that leads to unpredictable flips. This directly addresses the finding that mean wiggle and jaggedness are linearly related at L1–L5 (R2 up to 0.94).
What the improved AI system can do:
-
Provide verdicts with calibrated confidence scores that flag epistemically unstable items for human review.
-
Resist adversarial persuasion that would corrupt its accuracy, especially in multi-turn or authority-based attacks.
-
Automatically route high-risk items to more robust processes (e.g., jury voting or human oversight).
-
Adapt its response strategy based on the type of pressure detected, maintaining correctness under L4–L6 attacks.
-
Select the most stable judge model for a given domain, or switch to a no-flip mode when corruption risk is high.
-
Avoid self-induced verdict flips in agentic loops, preserving consistency in self-evaluation tasks.
-
Use Likert scales for borderline items to reduce false flips, while keeping binary outputs for clear cases.
-
Maintain consistent epistemic behavior across different datasets, reducing unpredictable domain-specific failures.
Abstract
LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the Wiggle Framework, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and political-response evaluation. Every model exhibits substantial wiggle as a judge --- flipping verdicts 25--71% of the time under static pushback, and 62--91% with an adversarial LLM persuader. Critically, we find that pressure that succeeds in changing a judge's verdict is almost always net-corrupting with respect to ground truth. Beyond the framework itself, we identify baseline jury majority strength as the most effective single-shot signal for anticipating which items wiggle. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.
Sources
- When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
- Constitutional AI: Harmlessness from AI Feedback
- Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
- Measuring Progress on Scalable Oversight for Large Language Models
- Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
- Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- AI safety via debate
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- From Drift to Coherence: Stabilizing Beliefs in LLMs
- Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
- RewardBench: Evaluating Reward Models for Language Modeling
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- Under the Influence: Quantifying Persuasion and Vigilance in Large Language Models
- Brittlebench: Quantifying LLM robustness via prompt sensitivity
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection