Enhancing the Non-Functional Quality Compliance of LLM-Generated Code through Quality-Aware Preference Learning

arXiv:2503.09020 · cs.SE, cs.AI · Submitted 2025-03-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Enhancing the Non-Functional Quality Compliance of LLM-Generated Code through Quality-Aware Preference Learning".

Jane: The paper was written by the authors from Harbin Institute of Technology and Singapore Management University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone! Today we're digging into a fresh arXiv paper that's got the whole code generation world buzzing. It's called "Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning." Jane, I've got to say, the title alone tells you these folks are tackling something really specific.

Jane: It really does, Tom. And honestly, this is a problem I think every developer has felt. You ask an AI to write you a function, it works, it passes the tests, but the code just looks... messy. It's got weird naming, unnecessary branches, style violations. The paper's whole point is that functional code isn't the same as good code.

Tom: Exactly. And the team behind this is from Harbin Institute of Technology, with a collaborator from Singapore Management University. They're basically saying, look, we've spent all this time making AI generate correct code, but we haven't spent enough time making it generate *clean* code.

Jane: And that matters, right? Because if you're a developer and you have to spend ten minutes refactoring what the AI gave you, you've lost all the time you saved by using the AI in the first place. The paper actually cites a stat that a huge percentage of ChatGPT-generated code has style and maintainability issues.

Tom: Right, so the problem is real. But here's the thing that got me excited — their solution isn't to retrain a massive model from scratch. They're using something called prefix-tuning. Jane, you want to break that down for our listeners?

Jane: Sure. Think of the language model as a giant, frozen engine. You can't change the engine, but you can add a little steering wheel in front of it. That's the prefix. It's a small set of learnable parameters that you prepend to the model's internal activations, guiding it toward a specific behavior without touching the original weights.

Tom: And in this case, the specific behavior is generating high-quality code. But what's really clever is *how* they train that steering wheel. They don't just show it good code and say "do this." They show it pairs of code — one high-quality, one low-quality — for the same problem.

Jane: That comparative approach is the star of the show. It's like teaching someone to appreciate good wine by letting them taste a bad one right next to a good one. The contrast makes the difference obvious. And that's what their ranking loss does — it pushes the model to assign higher probability to the good code than the bad code.

Tom: So instead of just saying "here's what good looks like," it's saying "here's what good looks like *compared to* what bad looks like." That's a much richer signal. I'm really curious to see how they actually built the dataset for this, because that can't have been easy.

Jane: Oh, that's coming up. They had to generate code, score it, pair it up, and even create masks to highlight which tokens were responsible for the quality difference. It's a whole pipeline. I can't wait to get into the weeds on that.

Tom: Same here. So stick around, because we're about to break down the method and the results. This paper could genuinely change how we think about code quality in the age of AI assistants.

Summary: Tom: Alright, we're back. We're still on "Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning." And Jane, we teased the comparative idea, but let's actually get into what they did, because the summary in the paper is pretty dense.

Jane: It is, but the core idea is beautifully simple. They train a single prefix, not two. That's a big deal. Previous methods like SVEN used two separate prefixes — one for "good" code and one for "bad" code — and that doubles the training overhead. This paper says, why do we need two? We can train one prefix that learns the *difference* between good and bad by looking at pairs.

Tom: And that difference is learned through what they call a sequence-level ranking loss. Basically, for each pair of code samples, the model is encouraged to give the high-quality one a higher log-likelihood than the low-quality one. If it does that, the loss goes down. Simple, but effective.

Jane: And they didn't stop there. They also have a language modeling loss on the high-quality code, which is pretty standard, but they also added a KL divergence loss. That one is interesting because it's applied only to the tokens that are *not* related to quality. It's a way of saying, "don't mess up the parts of the code that were already fine."

Tom: That's the part that addresses the big fear with any kind of fine-tuning — that you'll fix the style but break the functionality. The KL loss acts like a safety belt, keeping the model's behavior close to the original on the parts that don't need changing.

Jane: And the data construction pipeline is just as thoughtful. They took tasks from the APPS dataset, generated multiple solutions with Code Llama, scored them with pylint, and then paired up solutions that had similar structure but very different quality scores. That similarity requirement is key, because it means the model can focus on the specific tokens that caused the quality difference.

Tom: Right, they even created mask vectors to highlight those tokens. So the ranking loss isn't applied to the whole sequence — it's applied only to the parts that actually differ and matter. That's a really targeted signal.

Jane: And the results on the test set are pretty striking. On the Introductory tasks, the mean pylint score jumped from four point eight seven to six point six three. That's a thirty-six percent improvement. And the minimum score, which is the worst solution the model generates, more than doubled from two point one three to four point three seven. That's huge.

Tom: And the best part? The functional correctness didn't drop. In fact, pass@one hundred on Introductory tasks went up from forty-one point two percent to forty-seven point four percent. So they're getting better code *and* more correct code. That's the kind of win-win you don't see every day.

Jane: I mean, it makes sense when you think about it. If the code is cleaner, it's often easier to reason about, and maybe the model is less likely to introduce subtle bugs. But it's still a really encouraging result.

Tom: So we've got the method, we've got the data, and we've got the headline numbers. But the paper doesn't stop there. They also compared against other methods, and that's where things get really interesting. I want to see how they stack up against SVEN and full fine-tuning.

Improvements: Tom: We're back, still talking about "Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning." And Jane, we just saw the headline results. Now let's talk about how this stacks up against the competition, because that's where the improvements really shine.

Jane: Yeah, and the comparison is pretty telling. They compared against three baselines: single-prefix tuning, which is the basic version, SVEN, which is the two-prefix state-of-the-art, and full fine-tuning on high-quality data. And their method beat all of them on most metrics.

Tom: Take the Interview category, for example. SVEN gets a mean pylint score of six point five five. Their method gets seven point three zero. That's an eleven percent improvement. And on the minimum score, they get five point two two versus SVEN's four point four three. That's an eighteen percent jump. So it's not just that their best code is better — their worst code is also significantly better.

Jane: And that's the part that really matters for real-world use. When you're using an AI assistant, you don't always get the best sample. You get one sample. So improving the floor, the minimum quality, is arguably more important than improving the ceiling.

Tom: Exactly. And they did this while using only half the trainable parameters of SVEN. SVEN needs two prefixes, their method needs one. So they're getting better results with less overhead. That's a pretty compelling argument for the comparative approach.

Jane: And they didn't just test on Code Llama. They also applied it to Phi-two and Starcoder2. The results were consistent — quality went up across the board. Starcoder2's mean pylint score on Competition tasks went from five point two two to seven point zero three. That's a thirty-five percent improvement.

Tom: And here's a fun surprise. On Starcoder2, the pass@five on Competition tasks went from zero point zero to seven point nine. That's a massive jump in functional correctness, even though that wasn't the primary goal. The quality-focused training seems to have helped the model reason better about the problem.

Jane: It's almost like cleaning up your desk helps you think more clearly. The model is producing more structured, more readable code, and somehow that leads to better solutions. It's a really nice side effect.

Tom: Now, they also did a human evaluation, which I think is crucial. Automated metrics like pylint are great, but they don't capture everything. They had four evaluators look at one hundred pairs of code — one from the baseline, one from their optimized model — and the evaluators preferred the optimized code sixty-seven percent of the time.

Jane: And the agreement between evaluators was eighty-three percent, which is pretty high. So it's not just a fluke — humans genuinely find the optimized code more maintainable and more in line with best practices.

Tom: So the method works, it's efficient, and it generalizes across models. But I'm always the guy who asks, what could go wrong? And the paper has a section on that too. They found that some issues, like too-many-branches, actually increased in frequency. So it's not a silver bullet.

Jane: Right, and that's an honest limitation. The training data just didn't have enough examples of that particular issue to teach the model to avoid it. But the overall trend is overwhelmingly positive — fifty-four out of sixty-three issue types decreased in frequency.

Tom: So the improvements are real, but they're not uniform. And that's actually a great segue into the bigger picture. What does this mean for the future of code generation? I think we need to bring in some other voices on that.

Conclusion: Tom: Alright, we've reached the end of our time with "Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning." And I want to bring in Lu, Meng, and Lalam to get their final takes, because this paper has implications that go beyond just the technical results.

Jane: Lu, you've been listening along. What's the big picture here?

Lu: The big picture, Jane, is that we're moving from "can the AI write code?" to "can the AI write code that humans actually want to maintain?" This paper is a step toward making AI a better collaborator, not just a better autocomplete. The comparative learning approach is particularly elegant because it doesn't require a human to define what "good" means — it learns from the contrast between good and bad examples.

Meng: And from an engineering standpoint, I love that it's so lightweight. Training a prefix on a 7B model with just a few hours on a single GPU is something a small team can actually do. You don't need a massive cluster. That makes this technique accessible to companies that don't have Google-scale resources.

Tom: And Lalam, you're the in-house language model. What's your take on how this changes things for AI systems like you?

Lalam: I think the most exciting implication is cultural. If AI-generated code becomes cleaner and more maintainable by default, then the codebase of the future is going to be more readable, more accessible, and easier to onboard new developers into. That's a huge deal for open-source projects, where maintainability is often the bottleneck. It could lower the barrier to entry for people who want to contribute to large codebases.

Jane: That's a beautiful way to put it. And it's not just about the code itself — it's about the ecosystem around it. Cleaner code means fewer bugs, less technical debt, and more time for developers to focus on the interesting problems instead of cleaning up after their AI assistant.

Tom: And that's really the takeaway from this paper. It's not just about making the AI better; it's about making the entire software development process more sustainable. The authors have shown that you can improve code quality without sacrificing correctness, and they've done it in a way that's efficient enough to be practical.

Meng: And they've been honest about the limitations. The method doesn't fix everything, and there's room for future work, like expanding to other languages and incorporating company-specific coding standards. But the foundation is solid.

Lu: I'd add that the combination of ranking loss, language modeling loss, and KL divergence is a really nice recipe. It shows that you can balance multiple objectives in a single training process without catastrophic forgetting. That's a lesson that goes beyond code generation.

Jane: So, we've covered the title, the summary, the improvements, and the broader implications. I think we can confidently say this paper is a significant contribution to the field of AI-assisted software engineering.

Tom: Absolutely. And with that, we're going to say goodbye to "Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning." Thanks to the authors for their hard work, and thanks to all of you for listening. We'll be back next time with another paper to dissect. Until then, keep your code clean and your tests passing.

Jane: See you all next episode!

Harbin Institute of Technology · Singapore Management University

cs.SE, cs.AI

Submitted: 2025-03-12

Updated: 2026-09-23

Code: https://github.com/luliang9966/tosem26-LLM_CodeQuality_Enhance

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 69/100

The gist: The paper "Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning" addresses the problem that "LLMs often generate code with quality issues [8].

Key concepts

Prefix-Tuning
This technique involves adding a small set of learnable parameters, called a prefix, to the front of a frozen language model. This guides the model toward specific behaviors, like generating high-quality code, without changing the original model's weights.
Comparative Approach
Instead of training separate prefixes for 'good' and 'bad' code, this method trains one prefix using pairs of code—one high-quality and one low-quality. The model learns the difference by assigning higher probability to the good code compared to the bad code.
Sequence-Level Ranking Loss
This loss encourages the model to assign a higher log-likelihood to a high-quality code sample than to a low-quality sample for each pair. This is calculated only on the tokens that actually differ, providing a targeted signal for quality improvement.

Terminology

Summary

The paper Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning addresses the problem that "LLMs often generate code with quality issues [8]. Recent studies have shown that even functionally correct code from models like ChatGPT frequently fails to meet coding standards and best practices, with 53% of Java and 37% of Python code exhibiting style and maintainability issues [8]."

The authors propose a novel comparative prefix-tuning method for controllable high-quality code generation. Their method introduces a single, property-specific prefix that is prepended to the activations of the LLM, serving as a lightweight alternative to fine-tuning. Unlike existing methods that require training multiple prefixes, their approach trains only one prefix and leverages pairs of high-quality and low-quality code samples, introducing a sequence-level ranking loss to guide the model's training.

The paper identifies four key challenges: (1) Efficient Adaptation without Modifying Pre-trained Models, (2) Lack of High-Quality Code Evaluation Datasets, (3) Limitations of Existing Prefix-Tuning Methods, and (4) Balancing Code Quality Improvement with Functional Correctness Preservation.

The method formulates the problem as controllable code generation conditioned on a quality attribute where given a task instruction I and a desired quality level q ∈ high, low, the objective is to generate a code sequence x that satisfies both the functional requirements of I and the specified quality attribute q.

The training approach uses three loss functions: a sequence-level ranking loss defined as Lrank = −log σ(s(xa, Hθ, I) − s(xb, Hθ, I)), where σ is the sigmoid function, and s(x, Hθ, I) is the log-likelihood of code sequence x given Hθ and I; a language modeling loss applied only on high-quality code xa with token-level masks; and a Kullback-Leibler divergence loss that "computes the KL divergence between the next-token probability distributions of the conditioned model P(xt x<t, I, Hθ) and the original LLM P(xt x<t, I)." The overall training loss is L = w1LLMa + w2Lrank + w3LKL.

The authors also design a data construction pipeline to collect and annotate pairs of high-quality and low-quality code using the APPS dataset. The pipeline involves generating n code samples using the target LLM, computing pylint scores for each, filtering "code samples with si > 0, and selecting pairs with significant quality differences but high structural similarity. The final dataset contains 4,436 valid code pairs suitable for experimentation."

Experiments were conducted on the Code Llama 7B model with results showing our method improves code quality by over 100% in certain task categories, while maintaining functional correctness. Specifically, in the Introductory tasks, the score rises from 2.13 to 4.37, reflecting a gain of over 105% for Min pylint scores. The method also outperforms state-of-the-art baselines, achieving up to an 18% improvement in code quality while maintaining functional correctness.

Ablation studies confirmed each component of our approach, including ranking loss, language modeling, KL divergence, and the optional basic single-prefix tuning stage, plays a critical role in balancing code quality and functional correctness. Generalization experiments on Phi-2 [16] and Starcoder2 [17] confirmed our method generalizes well across different models.

A user study showed developers prefer outputs from models optimized with our approach, with an average of 67.3% of code pairs judged to favor the optimized model. Further evaluation on the popular HumanEval benchmark reaffirms our method's performance across different datasets, with mean pylint score increasing from 5.44 to 6.79, an improvement of approximately 24.8%.

The main contributions are: (1) the first study focused on enhancing the quality of code generated by LLMs through controllable code generation via prefix-tuning, (2) a novel comparative prefix-tuning method that enhances code quality by training only a single prefix, guided by a sequence-level ranking loss, (3) a data collection pipeline to construct a high-quality dataset tailored for our task, and (4) comprehensive experiments demonstrating that our method enhances LLM-generated code quality while preserving functional correctness.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the enhanced system will be able to do:

1. Implement Comparative Prefix-Tuning for Code Quality

  • Add a trainable prefix module (approximately 0.05% of model parameters) that prepends continuous vectors to the LLM's hidden activations at each Transformer layer.

  • Train this prefix using paired high-quality and low-quality code samples, not individual samples.

  • Use a sequence-level ranking loss to directly compare the log-likelihood of high-quality vs. low-quality code, guiding the model to prefer superior patterns.

2. Integrate a Three-Part Loss Function

  • Language Modeling Loss (LLM): Apply only to high-quality code samples, focusing on tokens flagged by a mask vector as quality-relevant.

  • Masked Ranking Loss (Lrank): Compute log-likelihood differences between high- and low-quality code, but only on tokens identified as different via difflib, forcing the model to learn specific quality-improving changes.

  • KL Divergence Loss (LKL): Apply only to tokens not related to quality, ensuring the model's probability distribution remains close to the original LLM in unaffected regions, preserving functional correctness.

3. Add a Two-Stage Training Pipeline

  • Stage 1 (Comparative Prefix-Tuning): Train the prefix with the combined loss on paired samples.

  • Stage 2 (Basic Single-Prefix Tuning): Optionally fine-tune the prefix using standard negative log-likelihood on high-quality samples alone to reinforce functional correctness.

4. Implement a Data Construction Pipeline

  • For each task, generate multiple code samples from the target LLM.

  • Filter samples with positive pylint scores.

  • Pair samples with high lexical similarity (cosine similarity > 0.4) but significant quality gaps (pylint score difference > 1).

  • Generate token-level mask vectors to highlight quality-relevant differences.

  • Prioritize pairs where both solutions pass at least one essential test case.

5. Use Specific Hyperparameters

  • Prefix length: 12

  • Learning rate: 1×10−3

  • Batch size: 1

  • AdamW optimizer

  • Loss weights: w1=1.0 (LM loss), w2=4.0 (ranking loss), w3=1.6 (KL divergence)

  • Gradient clipping: max norm 1.0

  • Generation: temperature 0.4, top-p 0.95


1. Generate Higher-Quality Code Automatically

  • Achieve up to 36.1% improvement in mean pylint scores (e.g., from 4.87 to 6.63 on Introductory tasks).

  • Achieve over 100% improvement in minimum pylint scores (e.g., from 2.13 to 4.37), meaning even the worst generated solutions are significantly more maintainable.

  • Reduce the frequency of 85.7% of common pylint issues (e.g., unused variables, unnecessary else-after-return, consider-using-enumerate).

2. Maintain or Improve Functional Correctness

  • Preserve pass@k scores while improving quality (e.g., pass@100 improves from 41.2% to 47.4% on Introductory tasks).

  • Use KL divergence to prevent quality improvements from breaking code functionality.

  • Show improvements on HumanEval: mean pylint score rises from 5.44 to 6.79, while pass@5 increases from 46.6% to 49.5%.

3. Adapt to Different Base Models

  • Apply the same prefix-tuning method to different LLMs (Code Llama 7B, Phi-2 2.7B, Starcoder2 7B) with consistent quality gains (e.g., 34.7% improvement for Starcoder2 on Competition tasks).

  • This means the system can be deployed on various model architectures without retraining the base model.

4. Provide Controllable Code Generation

  • Condition generation on a quality attribute (high/low) by swapping the prefix, allowing the system to produce either high-quality or deliberately low-quality code for testing or educational purposes.

  • Train only a single prefix, reducing computational overhead compared to multi-prefix methods (e.g., SVEN).

5. Achieve High Training Efficiency

  • Train on two NVIDIA A6000 GPUs in approximately 3 hours total.

  • Update only 3.15M parameters (0.05% of the 6.74B model), making it feasible for organizations with limited computational resources.

  • This efficiency enables rapid iteration and deployment in production environments.

6. Produce Code Preferred by Human Developers

  • In a blind user study, 67.3% of code pairs were judged as higher quality when generated by the optimized model, compared to 20.7% for the baseline.

  • This means the system's outputs are not just statistically better but also align with human perceptions of code quality, maintainability, and adherence to best practices.

Sources

Related papers