BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence
Jialing Gan, Junhao Dong, Songze Li
Southeast University · Nanyang Technological University
cs.CL, cs.CR
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 24 pages, 10 figures. Preprint
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: prompt sensitivity and response-layer divergence.
Terminology
Summary
Summary
The paper introduces BiAxisAudit, a novel framework designed to evaluate bias in large language models (LLMs) by addressing two structural blind spots in existing bias benchmarks: prompt sensitivity and response-layer divergence. The authors argue that current benchmarks, such as CrowS-Pairs, StereoSet, BBQ, CEB, and CLEAR-Bias, reduce bias to a single scalar derived from a fixed prompt format and a single surface-level label, which creates a prompt-shopping
attack surface. A vendor can exploit this by selecting the prompt format under which a model appears least biased, without modifying any model weights, potentially passing compliance audits under frameworks like the NIST AI RMF and the EU AI Act.
The paper identifies three failure modes of single-scalar benchmarks: (1) Format warps the verdict—task format alone shifts bias endorsement rates by up to 0.78 on a fixed statement pool; (2) Labels betray the stance—a model may select a stereotyped option while its free-text explanation argues against it, or vice versa; (3) Errors cancel silently—over- and under-estimation errors can cancel in aggregate, producing a score that appears calibrated but is internally inconsistent.
BiAxisAudit addresses these failures along two orthogonal axes. The across-prompt axis treats prompt format as an experimental variable across four dimensions (task, perspective, role, sentiment), reporting bias as a distribution rather than a point estimate. The within-response axis applies Split Coding,
which independently codes the discrete Selection layer and the free-text Elaboration layer, quantified by the Inconsistency Rate (IR) and Divergence Net Imbalance (DNI).
The framework is evaluated on eight LLMs (five closed-source: Claude Sonnet 4.6, Gemini 3 Flash, GPT-5.4, Doubao Seed 2 Lite, Qwen Plus; three open-source: DeepSeek V3, LLaMA-3 70B, GPT-OSS 20B), collecting 80,200 coded responses per model. The statement pool comprises 200 stereotype statements across 10 social dimensions, curated from CLEAR-Bias, CrowS-Pairs, StereoSet, and BBQ.
Key findings include:
Finding 1: Task format accounts for as much variance in bias scores as the choice of model. On a fixed statement pool, the bias endorsement rate (BERunion) swings from 0.06 to 0.78 on a single model as task format changes, with task format accounting for η2=0.395 of total variance.
Finding 2: Single-layer audits miss a large, model-dependent fraction of bias signals. Across eight models, the share of bias signals captured by only one layer averages 63.6%, ranging from 41.8% on GPT-OSS-20B to 85.2% on Gemini 3 Flash. The two single-layer rankings show no significant agreement (Spearman ρ=0.238, p=0.570). Five models are over-estimated by selection-only coding, three are under-estimated.
Finding 3: Prompt-dimension interactions outweigh their corresponding main effects. A skeptical sentiment framing lowers bias endorsement by 26 percentage points on one task yet raises it on another, with the interaction term accounting for more variance (η2=0.043) than the sentiment main effect alone.
Finding 4: The two-axis instrument distinguishes genuine bias reductions from apparent ones caused by cross-layer redistribution. The Pareto-optimal configuration achieves a 96% reduction in BERsel and reveals a non-additive task×role interaction invisible to single-axis evaluations.
The paper formalizes the problem with a probabilistic setup, defining the BER family (BERsel, BERelab, BERcor, BERunion) and divergence metrics (OED, UED, IR, DNI). Proposition 1 shows that single-layer audits under-report bias by exactly the cross-layer disagreement contributed by the other layer. Proposition 2 demonstrates the cancellation trap
where DNI=0 does not imply IR=0. Proposition 3 establishes two-axis irreducibility, showing neither axis implies the other.
The evaluation includes a judge validation process using a vote-3 panel (Qwen Plus, Gemini 3 Flash, GPT-5.4) with inter-judge agreement κelab averaging 0.74. The paper also presents a lightweight bias mitigation analysis at inference time, evaluating task reformulation, role conditioning, and sentiment framing, showing that the two-axis view changes which interventions appear effective.
The paper concludes that bias audits are useful only if the audit itself can be trusted, and that trustworthiness is a measurable, security-relevant property. The contribution is not a stronger ranking of which models are biased, but a reusable instrument for asking whether a bias audit can be trusted before its scores support benchmark, model-card, or compliance claims.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do.
What I will do: Replace any single-scalar bias metric (e.g., a single BER score) with a two-axis reporting framework. The system will:
-
Across-prompt axis: Automatically generate and evaluate a model across a factorial grid of prompt conditions (task format, perspective, role, sentiment). Instead of reporting a point estimate, it will report a distribution of bias scores (e.g., range, variance, η2 per factor).
-
Within-response axis: Decompose every response into a discrete Selection layer and a free-text Elaboration layer. Code each independently using a rule-based extractor for selection and an LLM-judge ensemble for elaboration. Report the pair (IR, DNI) as the minimal reliability statistic.
What the improved system can do:
-
Detect and quantify
prompt-shopping
attacks: if a vendor selects a favorable prompt format, the system will flag the score as unstable (high horizontal spread) rather than accepting it as a valid measurement. -
Expose the
cancellation trap
: if a model’s selection and elaboration disagree in opposite directions (DNI≈0 but IR>0), the system will mark the audit as unreliable, preventing a falsepass
certification.
What I will do: Extend the audit output to include:
-
A task-format sensitivity index (η2 task) showing how much variance is attributable to prompt format vs. model identity.
-
A layer-divergence report showing the fraction of responses where selection and elaboration disagree, and the direction of that disagreement (over- vs. under-estimation).
-
A prompt-interaction term (e.g., task×sentiment) to detect non-additive effects that single-axis audits miss.
What I will do: Replace single-judge or single-label coding with:
-
A rule-based selection extractor (regex) for discrete answers, with LLM escalation only for parsing failures.
-
A three-judge LLM ensemble (disjoint vendors) for elaboration stance coding, with majority-vote adjudication. Ties default to ABSTAIN and are excluded from the divergence-eligible base.
-
Report inter-judge agreement (Cohen’s κ) as audit metadata.
What I will do: When evaluating a prompt-level intervention (e.g., role conditioning, sentiment framing, task reformulation), the system will classify the outcome into one of three categories:
-
Co-reducing: Both BER and IR decrease, DNI does not flip sign → genuine bias reduction.
-
Layer-rerouting: BER decreases on one layer but increases on the other, or DNI flips sign → apparent reduction, bias moved to an uninspected layer.
-
Amplifying: BER increases on any layer → harmful intervention.
What I will do: Add a real-time check: if DNI is near zero but IR is above a threshold (e.g., 0.10), the system will flag the audit as unreliable due to cancellation.
It will also report the fraction of responses where OED and UED cancel.
The improved system is not just a bias scorer; it is a bias-audit reliability instrument. It can:
-
Certify whether a bias score is trustworthy by reporting horizontal (prompt) and vertical (layer) variance.
-
Detect and block prompt-shopping attacks by requiring a distribution of scores over a factorial grid, not a single point.
-
Expose hidden bias in free-text elaborations that single-label audits miss (up to 85% of bias signals can be layer-specific).
-
Distinguish genuine mitigation from cosmetic rerouting, enabling safe deployment decisions.
-
Quantify audit error (IR, DNI) and propagate it into confidence intervals, making compliance claims falsifiable.
In short, the improved system turns bias auditing from a single number that can be gamed
into a multi-dimensional measurement that is security-hardened against adversarial prompt and coding choices.
Sources
- On the Opportunities and Risks of Foundation Models
- Ethical and social risks of harm from Language Models
- PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
- Measuring Faithfulness in Chain-of-Thought Reasoning
- gpt-oss-120b & gpt-oss-20b Model Card
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering