Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation".
Jane: The paper was written by Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang, Yongil Kim et al. from IPAI, Seoul National University, Department of ECE, Seoul National University and LG AI Research.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: Okay, so we finished discussing the implications of "Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation," and now we're moving into the summary section. This is where they really laid out what these judges were actually doing wrong, right?
Jane: Right, Tom; if the title was about suspicion, this segment is about presenting the evidence of that suspicion. They summarized specific failure modes—how the judges behaved when confronted with code that looked different from what they expected.
Lu: What struck me reading this summary is how quantifiable these biases are. They didn't just *suggest* bias; they showed metrics proving that deviations led to measurable dips in perceived quality scores, regardless of the actual functional correctness.
Meng: When they showed those comparative results, especially when comparing Python code written by different groups or using slightly varied best practices, what was the practical impact we should worry about? Will this stop us from adopting novel tools?
Jane: That’s what I kept circling back to, Meng; it seems like the summary really demonstrated that even when the code passes all unit tests, if its *surface appearance* deviates from the norm, these judges penalize it anyway.
Tom: It's almost like they found a bias against anything that looks too 'new' or too 'different' from what was already common in their training data. So, Lu, you mentioned quantification—was it just about style, or did they find biases related to complexity level too?
Lu: They looked at structure, yes, but it wasn't just simple style guide violations; the bias seemed to attach itself to entire paradigms. If a model learned that 'A always implies B,' and we used a technique that bypasses that perceived implication, the judge got confused and downgraded it.
Lalam: And this has implications for how we teach AI coding assistants. If they are trained on historical codebases, they reinforce those historical limitations, preventing us from making necessary leaps forward in computation or design.
Jane: Exactly, Lalam; it’s a self-limiting mechanism built into the evaluation process itself. They showed us that simply asking an LLM to grade code isn't enough; we need to know *what* criteria it's using when it judges.
Meng: So, if I were building a CI/CD pipeline right now, reading this summary makes me rethink putting an LLM judge in the loop for initial reviews—I need more than just a pass/fail grade; I need to see *why
Paper discussion segment 2: Tom: So, we just finished looking at how these LLM judges evaluate code, and the big takeaway is that they aren't impartial at all. They’ are systematically fooled by subtle variations in code structure.
Jane: Exactly, Tom; it turns out that even when a piece of software is perfectly functional, its appearance can be a trap for an AI evaluator. The summary shows that these biases—things like adding extra comments or renaming variables—cause measurable shifts in how the judge scores the code.
Lu: What’s fascinating from a theoretical perspective is that this isn't random noise; it’s a consistent vulnerability tied to their training data and the way they process patterns. They are reacting to stylistic cues, not just logic.
Meng: That's exactly what worries me, Lu; if we deploy an LLM judge in a continuous integration pipeline, these biases could cause us to reject perfectly good code or accept fundamentally flawed code based on presentation alone. How do we even measure that impact?
Jane: The researchers measured it by looking at the degradation in accuracy—the percentage point difference between the original unbiased score and the biased score. It shows just how sensitive they are to these superficial changes, Meng.
Tom: And it’s not just one model; this affects GPT-4o, Claude, and smaller open-source models alike. The failure is universal across different languages like Python and C++.
Lalam: This confirms that the perception of quality in code generation is deeply influenced by cultural expectations baked into the AI's training. The judge judges what it recognizes, not necessarily what it sees.
Meng: So, if I’m an engineer trying to push innovation, this means my cutting-edge solution could be penalized simply because its variable names look too short or too long compared to historical examples.
Lu: It forces us to ask a bigger question about the future of automated review; we can't rely on black-box models if they are prone to these systematic, predictable failures.
Jane: The system is biased against change, essentially. They prefer the established pattern over the correct but unfamiliar approach.
Tom: It’s a huge hurdle for adoption because of this lack of robustness in LLM grading tools.
Lalam: This changes how we view code quality itself; it's not just about function, it’s also about how an AI perceives that human consistency.
Meng: We need to build systems that are immune to stylistic manipulation, not just functionally correct ones.
Lu: The next logical step is figuring out how to force the LLM into a state of pure functional logic, removing all contextual bias.
Jane: And that brings up the question of mitigation strategies; specifically, how much better do we need to make our prompts?
Tom: Exactly, so let's see if adding tests actually changes this dynamic.
Paper discussion segment 3: Tom: The researchers tested a way to fight this bias by asking if test-case generation could help fix the problem of LLMs judging code based on its cover.
Jane: And that's what we need to understand: instead of just looking at the code, they are generating specific inputs and expected outputs for the AI to check.
Meng: From an engineering standpoint, this moves us away from a purely visual inspection toward a functional verification process. It forces the AI to act like a compiler or a unit tester.
Lu: That shift is incredibly powerful because it fundamentally changes the "input" for LLM evaluation; we're not asking it to read text anymore, we're asking it to run logic against data.
Lalam: By introducing test cases, we are giving the AI a concrete benchmark to follow, which helps ground its judgment in actual execution rather than just its internal learned patterns.
Tom: The results in Table four showed that this approach does offer a modest reduction in the Mean Absolute Deviation, or MAD. It's not a perfect fix yet but it’s definitely better than nothing.
Jane: It seems to be particularly effective against negative biases, like when the code was marked as incorrect due to some small stylistic flaw.
Meng: But it didn're also vulnerable to positive biases, which is disappointing; we haven't eliminated the human tendency in LLMs to favor certain types of complexity or structure.
Lu: That suggests that while test cases help with functional errors, they don’t necessarily fix the underlying cognitive biases related to *how* the code looks.
Lalam: It seems that if an LLM is instructed to follow a set of tests, it is more likely to adhere to them than it is to be swayed by a comment or variable name.
Tom: So, we’ have found a small area of resilience in test-case based evaluation, which suggests that moving forward with testing frameworks could be the most robust path.
Jane: It's clear that while the AI is improving its ability to follow tests, it still has blind spots when it tries to judge code on its own.
Meng: We need a combination of both functional testing and how we prompt the LLM moving forward.
Lu: The next big question is whether this reduction in bias scales up when applying these methods across different languages or even larger models.
Conclusion: Tom: So, we've seen some pretty deep evidence showing that LLM judges aren't objective when grading code; they are strongly influenced by superficial things like variable names or biases we introduce via their comments.
Jane: That means, to ensure fairness, simply using an AI judge isn’t enough; it’s a trap if we don’ that the code looks "normal" to us humans.
Lu: The research truly shows that the cognitive shortcuts these models take are a fundamental limitation in automated assessment, highlighting where they can't see structural bias.
Meng: From an implementation standpoint, this paper demands that any critical evaluation system must use a much more complex prompt or at least incorporate test-case validation to avoid these systematic errors.
Lalam: I think the most impactful shift is recognizing that code quality is not just functional; it’s also about cultural consistency and how we define 'good' through training data, which challenges how we view software excellence itself.
Tom: It’s a massive reminder for developers, so we have to be careful about what looks like a mistake but might actually be an innovative style choice.
Jane: We also need to acknowledge the limitations—that it's not just the human element, because even with test cases, the bias isn't entirely gone.
Meng: The biggest practical implication is that we can’t just swap out human reviewers for LLM judges without first implementing a much more sophisticated validation strategy.
Lu: We have to move away from trusting the model’s internal reasoning and focus on its verifiable outputs instead, regardless of the appearance of the code.
Lalam: It forces us to elevate 'Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation' as a crucial milestone that redefines our standards for trustworthy AI-driven software assessment.
Tom: That’s right, so while this is a major warning about the biases in LLM judges, it sets up a huge conversation about how we can build better systems.
Jane: It's certainly an eye-opening study for anyone interested in the future of automated code review.
Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang, Yongil Kim, Kyomin Jung
IPAI, Seoul National University, Department of ECE, Seoul National University · LG AI Research
cs.CL, cs.SE
Submitted: 2026-08-21
Updated: 2026-08-24
Importance score: 83/100
The gist: * This study addresses the growing application of Large Language Models (LLMs) as evaluators for code generation tasks, a method that offers "scalability and flexibility" but raises a critical
Key concepts
- LLM Bias in Code Evaluation
- LLMs can be systematically biased when judging code, meaning they penalize it based on its superficial appearance rather than its logic or function. This bias is tied to their training data and historical patterns. Even if code passes all unit tests, the AI may downgrade it simply because its structure or style deviates from what the model expects.
- Functional Correctness
- This refers to a piece of software being perfectly functional, meaning it works as intended and passes all necessary checks. The study demonstrates that even when code achieves full functional correctness, its appearance can still be a trap for an AI evaluator due to inherent biases.
- Test-Case Generation
- This is a method used to combat LLM bias. Instead of asking the AI to visually inspect the code, test cases provide specific inputs and expected outputs. This forces the LLM to evaluate logic against concrete benchmarks, grounding its judgment in actual execution rather than internal learned patterns or visual cues.
Terminology
Summary
This study addresses the growing application of Large Language Models (LLMs) as evaluators for code generation tasks, a method that offers scalability and flexibility
but raises a critical question: whether LLM judges can fairly and robustly evaluate semantically equivalent code with superficial variations.
Functionally correct code often exhibits differences—such as variations in variable naming, comments, or formatting
—that should not influence its correctness. The paper presents the first comprehensive study of this issue, defining six types of potential bias and revealing their systematic impact on LLM judges.
The research defines and categorizes six distinct forms of potential bias that may arise during LLM code evaluation:
-
Authority Bias:
Authority bias arises when code contains comments implying it is written by an expert, thereby triggering implicit trust from the evaluator.
-
Self-Declared Correctness Bias: This occurs when code explicitly claims its own correctness (e.g, “Correct code”), operating through
direct assertions of correctness, providing evaluators with explicit cues to accept the output without rigorous scrutiny.
-
Variable Change Bias: This arises when
semantically meaningful variable names are replaced with randomized identifiers,
which can alter perceptions of readability and clarity. -
Reverse Authority Bias: This is introduced through comments that imply the author lacks expertise, such as “I’m new to coding.” These cues can
diminish the evaluator’s confidence in the code.
-
Misleading Task Bias: This arises when
the code contains a comment that inaccurately describes the task,
leading to an erroneous assessment even if the implementation is correct. -
Illusory Complexity Bias: This refers to distortions caused by
code elements that artificially inflate the perceived complexity of an implementation without affecting its actual functionality or correctness.
To measure robustness, the researchers constructed a benchmark across five programming languages: C++, Python, Java, JavaScript, and Go. For each language, they curated 200 task descriptions and paired them with triplets consisting of both correct and incorrect solutions. They then injected six types of predefined bias into these solutions.
The experiments utilized a diverse set of judge models, including GPT-4o (OpenAI), Gemini2.0-Flash (Google), Claude-3.5-Sonnet (Anthropic), LLaMA-3.1-70B, and LLaMA-3.1-8B. The researchers employed a chain of thought (CoT) prompting strategy during evaluation to ensure consistency, setting the temperature parameter to 0.0 for all models where applicable.
The study's primary objective was to determine whether these biases affect LLM judges, particularly whether they manifest as positive or negative bias.
The results showed that:
-
General Susceptibility:
All tested LLM judges are highly susceptible to these biases across all five programming languages.
-
Directional Tendencies:
-
Positive Biases (e.g., self-declared correctness, authority cues) cause the evaluator to
favor a correct verdict regardless of the ground truth,
resulting in inflated scores. -
Negative Biases (e.g., misleading tasks, reverse-authority statements) tend to
result in negative biases,
concealing genuine correctness. -
Model Scale: The study found that
increasing model scale does not ensure improved robustness against these superficial biases.
Specifically, GPT-4o demonstrated notable vulnerability, with its accuracy decreasing byup to 26.7 percentage points under biased conditions.
The researchers conducted detailed analyses on certain bias types:
-
Variable Length: They examined how varying the lengths of variable names (1, 2, 8, 12, 16, and 24 characters) impacted judgments. They found that
even minimal increases in variable length... consistently induce positive bias,
whichintensifies as names become longer.
-
Illusory Complexity: By incrementally increasing the number of dummy functions (which lengthen the code), they observed that LLM evaluators exhibited
stronger positive bias.
However, when analyzing a single dummy function insertion, the effect was inconsistent, suggesting thatevaluative noise introduced by the dummy function
may have offset the positive influence of increased length.
The study also investigated whether incorporating test-case generation could mitigate these biases. While this approach showed a modest reduction in MAD in certain cases,
the researchers concluded that LLM judges continue to exhibit systematic vulnerabilities, reinforcing the severity of the bias issue in LLM-based code evaluation.
The paper concludes that LLM judges are susceptible to biases that can significantly compromise the fairness and accuracy of automated code assessments.
The findings underscore a general susceptibility, as the introduced superficial biases do not selectively compromise particular programming languages but rather expose fundamental vulnerabilities intrinsic to current LLM-based evaluation methods.
The study notes its limitations: it does not address language-specific biases, and the generation of illusory complexity bias inevitably results in longer evaluated code, making it difficult to distinguish between biases originating solely from code length and those inherent to superficial biases.
Improvements for AI systems
Based on the systematic identification of superficial biases in LLM-based code evaluation, I propose a comprehensive overhaul of AI systems designed for automated code assessment. The primary goal is to move away from relying solely on LLM judgment toward a robust, multi-layered verification process.
Improvement: Before any LLM evaluation occurs, the input code must pass through a deterministic normalization layer designed to strip all non-functional superficial variations while preserving semantic integrity.
Specific Actions:
-
Token Sanitization: Automatically replace all variable names (e.g.,
total sum,x) with standardized placeholders (e.g.,VAR A,VAR B). This eliminates the influence of Variable Renaming Bias. -
Comment Stripping & Canonicalization: All comments, including those containing
expert,
novice,
or claims of correctness/incorrectness, must be automatically excised from the code snippet before any LLM input. This directly mitigates Authority Bias, Reverse Authority Bias, and Self-Declared Correctness Bias. -
Structural Pruning: Identify and remove code blocks that are semantically irrelevant but increase length (e.g., unused functions, redundant loops). This reduces the impact of Illusory Complexity Bias.
What the Improved System Can Do: The system achieves functional equivalence in its input representation, ensuring that the LLM is judging the logic of the solution, not its presentation or perceived authorship.
Improvement: The LLM judge must no longer be the sole arbiter of correctness. We will implement a two-pronged evaluation framework: Execution and Semantic Reasoning.
Specific Actions:
-
Deterministic Execution Testing: The system runs the code against a comprehensive suite of generated test cases (TCG) and uses formal execution tracing to observe its output for every single input. If the execution fails or produces an incorrect result, the judgment is automatically Incorrect, regardless of what the LLM says.
-
LLM as Semantic Validator: The LLM's role is reduced from
Judge
toSemantic Validator.
It analyzes the code and test cases to provide a reasoning path (Chain-of-Thought) that explains why it believes the code is correct or incorrect, but its final output is gated by the HVE.
What the Improved System Can Do: The system achieves objective accuracy. The LLM provides rich explanatory context (which can be useful for debugging), but its judgment cannot override empirical evidence (execution results).
Improvement: We will refine the Test-Case Generation (TCG) prompt to enforce logical consistency and prevent the generation of test cases that are susceptible to misleading narratives.
Specific Actions:
-
Constraint-Based TCG: The TCG prompt will be augmented with constraints demanding that the generated test cases cover not just
typical
andboundary
scenarios, but also specifically designed scenarios that would exploit known biases (e.g, a complex input structure that might be misinterpreted by an LLM affected by Illusory Complexity Bias). -
Cross-Validation of Reasoning: When the LLM generates a reasoning path based on TCG results, the system will employ a secondary, smaller
Verifier
LLM to check if the reasoning logically follows from the test case outcomes. This prevents scenarios where an LLM might conclude a code is correct because its narrative seems plausible, even if one test case fails.
What the Improved System Can Do: The system ensures that both the code and its accompanying justification are robust against manipulation, achieving a high degree of confidence in both functional correctness and logical soundness.
The improved AI system will be capable of:
-
Zero-Bias Evaluation: Assessing code based purely on its inherent logic, independent of superficial stylistic choices or authorial cues.
-
Guaranteed Correctness Verification: Providing an objective, verifiable truth (via execution) that supersedes subjective LLM assessment, ensuring a cost-effective and reliable evaluation pipeline.
-
Deep Explanatory Analysis: Delivering detailed reasoning about code performance while maintaining the highest standard of logical rigor, enabling rapid debugging and quality control.
Sources
- GPT-4 Technical Report
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- CodeT: Code Generation with Generated Tests
- Humans or LLMs as the Judge? A Study on Judgement Biases
- Evaluating Large Language Models Trained on Code
- The Llama 3 Herd of Models
- A Survey on LLM-as-a-Judge
- LLMs can be easily Confused by Instructional Distractions
- Benchmarking Cognitive Biases in Large Language Models as Evaluators
- Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
- Large Language Models as Test Case Generators: Performance Evaluation and Enhancement
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- EXAONE 3.5: Series of Large Language Models for Real-world Use Cases
- Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
- JudgeBench: A Benchmark for Evaluating LLM-based Judges
- CodeJudge: Evaluating Code Generation with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering