Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design".
Jane: The paper was written by Ruisi Zhang, Neusha Javidnia, Nojan Sheybani, Noah Marosok, Ke Huang et al. from University of California, San Diego and San Diego State University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everybody. Today we're digging into a paper that's got one of those long, intimidating titles, but the idea behind it is genuinely exciting. It's called "Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design," and I'm here with Jane, as always.
Jane: And I'm so glad we're covering this one, Tom. Because when I first read the title, I thought, okay, that's a mouthful. But the core problem is something anyone who's ever written code can understand. If you use a tool like GitHub Copilot to generate code for you, who actually owns that code? And more importantly, how do you prove it?
Tom: Right, and that's where watermarking comes in. You want to embed a secret signature into the code that only the model owner knows about. So if someone takes that code and uses it in their commercial software without permission, the owner can say, hey, that's mine, look, here's the proof.
Jane: Exactly. But here's the catch that this paper tackles head-on. Code is not like a paragraph of text. It's very structured, very precise. There's not a lot of wiggle room to hide a secret message without breaking the code itself. If you change a variable name or restructure a loop, you might accidentally change what the program does.
Tom: So it's a tightrope walk. You need the watermark to be strong enough to detect, but subtle enough that the code still works perfectly. And the authors here, from UC San Diego and San Diego State, they've come up with a clever way to do that using a pre-trained model called CodeT5 to understand the code's structure better.
Jane: And that's just the first half. The second half is about verification. Normally, to prove you own the code, you'd have to reveal the watermark to a third party, like a judge or an arbitrator. But once you reveal it, someone could copy it or erase it. This paper uses something called zero-knowledge proofs to let the owner prove they know the watermark without actually showing it.
Tom: So it's like proving you know the password to a vault without ever saying the password out loud. That's the magic of cryptography, and it's a big deal for protecting intellectual property in the age of AI-generated code. We're going to break down how they actually pulled this off, so stick around.
Summary and Implications: Jane: So, Tom, we've set the stage. Let's get into the meat of this paper, "Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design." The authors are proposing a framework called RoSeMary, which is a bit of a mouthful, but it stands for something meaningful.
Tom: RoSeMary, I like it. And the key insight is that they're not just slapping a watermark on at the end. They're training the watermark insertion and extraction modules together, end-to-end. That's the ML part of the co-design. They want the system to learn what changes are safe to make to the code without breaking it.
Jane: Right. And they're using CodeT5 as the backbone, which is a model pre-trained on millions of code snippets. That gives it a much better understanding of code structure than a model trained from scratch. The results speak for themselves. On the HumanEval benchmark, they're getting a detection AUROC of zero point nine seven, which is really high.
Tom: For our listeners who aren't statisticians, that means they can almost perfectly distinguish between watermarked and non-watermarked code. And they're doing this while keeping the code functional. The pass rate, which is the percentage of code that still passes its unit tests after watermarking, stays above ninety-five percent in most cases.
Jane: That's the fidelity part. But what about when someone tries to mess with the code? Say they rename all the variables or ask another AI to refactor it. That's the robustness test. And RoSeMary holds up. Even when fifty percent of variables are renamed, the detection AUROC stays above zero point nine three.
Tom: That's impressive because a lot of other watermarking schemes fall apart under that kind of attack. The authors also tested against a refactoring attack using a different LLM, and they still got an AUROC of zero point seven three, which is solid. So the ML side is doing its job.
Jane: But the really novel part, the part that got me excited, is the crypto side. They're using zero-knowledge proofs to verify ownership without revealing the watermark. That's the "Zero Knowledge" in the title. And it means the owner can prove they own the code without ever exposing the secret signature.
Tom: Which is huge for reusability. If you never reveal the watermark, you never have to re-encode a new one. The code stays usable, and the proof of ownership stays valid. It's a one-time proof that can be verified by anyone, anytime. That's the kind of security that could actually make AI-generated code viable for commercial use.
Improvements and Methodology: Tom: Alright, Jane, we've talked about the big picture. Now let's get into the nitty-gritty of how RoSeMary actually improves on what came before. Because there are other watermarking schemes out there, but they all have weaknesses.
Jane: Right. The older methods, like KGW and SWEET, work by manipulating the token generation process during inference. They split the vocabulary into "green" and "red" lists and force the model to pick from the green list. But that's a blunt instrument. It can break the syntax of the code, making it uncompilable.
Tom: And then there's SrcMarker, which uses a neural network to embed watermarks in the code's feature space. That's closer to what RoSeMary does, but SrcMarker uses a shallow transformer trained from scratch. It doesn't have the deep understanding of code that a pre-trained model like CodeT5 brings.
Jane: Exactly. So the improvement here is twofold. First, they're using CodeT5 to extract better features from the code, which means the watermark can be embedded more subtly and detected more reliably. Second, they're training the insertion and extraction modules together, with a loss function that balances three things: functionality, detectability, and robustness.
Tom: And that's the "co-design" part. They're not just optimizing for one thing. They're jointly optimizing for all three. The loss function has three components. There's a functionality loss to make sure the watermarked code behaves the same as the original. There's a detectability loss to make sure the watermark can be extracted. And there's a robustness loss to make sure it survives adversarial modifications.
Jane: And they even add a little bit of noise to the transformation probabilities during training, which is like a form of data augmentation. It forces the extractor to learn to recover the watermark even when the code has been slightly altered. That's why it's so robust to attacks.
Tom: Now, on the crypto side, they're using a specific type of zero-knowledge proof called a zk-SNARK, specifically Halo2. And they've made some clever optimizations to make it efficient. They approximate the ReLU activation function with a polynomial, they quantize the weights to Bfloat16, and they fuse the watermark extraction and bit error rate calculation into a single circuit.
Jane: The result is that proof generation takes about seven seconds, which is a one-time cost for the owner. But verification, which is what the arbitrator does, takes only about one hundred eighteen milliseconds. That's fast enough for real-world use. And the proof itself is only about seventeen kilobytes, so it's cheap to transmit.
Tom: So they've made zero-knowledge verification practical for code watermarking, which is a significant step forward. It's not just a theoretical idea anymore. It's something that could actually be deployed.
First Page and Implications: Jane: So, Tom, we've covered the methodology and the results. But let's go back to the very beginning of the paper, the first page, because it sets up the problem so clearly. The authors talk about the "low-entropy nature of code" constraining the space for high-quality watermarks.
Tom: Low entropy, that's a fancy way of saying code is very predictable. There are only so many ways to write a for loop or a function call. So there's not a lot of room to hide a secret message without making the code look weird or breaking it.
Jane: And that's the fundamental challenge. The paper argues that existing solutions either sacrifice functionality for detectability, or they sacrifice security for reusability. If you reveal your watermark to prove ownership, you have to re-encode a new one, which might degrade the code's quality.
Tom: And that's where the zero-knowledge proof comes in as a game-changer. It breaks that trade-off. You can prove ownership without revealing the watermark, so you never have to re-encode. The code stays high-quality, and the ownership proof stays valid.
Jane: The authors also frame this in terms of intellectual property protection. AI-generated code is being used more and more in commercial software. But if you can't prove where that code came from, you can't enforce licensing agreements or pursue legal action against infringement.
Tom: And that's a huge deal for the economy. Think about all the companies building products on top of AI-generated code. They need to know that their code is protected, and they need to be able to prove ownership in court if necessary. This paper provides a framework for doing that securely and efficiently.
Jane: There's also a social implication. If code can be reliably watermarked and verified, it could encourage more open sharing of AI-generated code. Developers might be more willing to contribute to open-source projects if they know their contributions can be attributed and protected.
Tom: And on the flip side, it could help prevent malicious use. If someone uses AI-generated code to create malware, the watermark could help trace it back to the source. That's a powerful deterrent.
Jane: So the implications go beyond just protecting a company's bottom line. It's about building trust in AI-generated content, which is essential for the technology to reach its full potential.
Conclusion: Tom: Well, Jane, we've covered a lot of ground on this paper, "Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design." Let's wrap it up for our listeners.
Jane: Absolutely. The core takeaway is that RoSeMary provides a way to watermark AI-generated code that is both robust and secure. It uses a pre-trained model to embed watermarks that don't break the code, and it uses zero-knowledge proofs to verify ownership without revealing the secret signature.
Tom: And the numbers back it up. High detection accuracy, preserved functionality, resilience to attacks, and efficient verification. It's a complete package.
Jane: The implications are significant. This could enable commercial use of AI-generated code with confidence, protect intellectual property, and even help trace malicious code. It's a step towards a future where AI-generated content is trusted and accountable.
Tom: And for the research community, it opens up new directions. The co-design of ML and crypto is a powerful idea that could be applied to other domains, like verifying the provenance of AI-generated images or text.
Jane: So we're saying goodbye to this paper, but we're excited to see where this line of research goes. Thanks for joining us, and we'll see you on the next episode.
Tom: Take care, everyone.
Ruisi Zhang, Neusha Javidnia, Nojan Sheybani, Noah Marosok, Ke Huang, Farinaz Koushanfar
University of California, San Diego · San Diego State University
cs.CR, cs.CL, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: accept to ACM TAISAP
Code: https://github.com/zcash/halo2
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 61/100
The gist: The paper introduces RoSeMary, a framework for watermarking code generated by large language models (LLMs) that combines machine learning and cryptographic co-design to secure high-quality watermarks
Key concepts
- Watermarking
- Embedding a secret signature into generated code that only the model owner knows about. This allows the owner to prove ownership if someone uses the code commercially without permission.
- Zero-Knowledge Proofs
- A cryptographic method allowing an owner to prove they know a secret (like a watermark) without ever having to reveal the actual secret signature itself. This maintains code usability and security.
- CodeT5
- A pre-trained model used by the authors that helps understand the structure of code. Using this backbone allows for embedding watermarks more subtly and reliably into complex, structured programming languages.
- ML/Crypto Co-Design
- The novel approach of jointly training machine learning modules (for embedding/extraction) and cryptographic modules (for verification). This optimizes the system for functionality, detectability, and robustness simultaneously.
Terminology
Summary
The paper introduces RoSeMary, a framework for watermarking code generated by large language models (LLMs) that combines machine learning and cryptographic co-design to secure high-quality watermarks while enabling confidential ownership verification.
The paper addresses the challenge of watermarking LLM-generated code, noting that Watermarking has emerged as a viable tool for protecting intellectual property (IP) in large language model (LLM)-generated code.
The authors identify that a major challenge, however, is the low-entropy nature of code, which constrains the space for high-quality watermarking signatures satisfying the detectability-fidelity-robustness tri-objective.
Additionally, "Legal verification requires disclosing the signature to third-party arbitrators, and alternative signatures need to be encoded for artifact reuse to avoid forgery and tampering, which further strains the already limited high-quality signatures and reduces code usability."
RoSeMary is described as a novel framework that secures high-quality code LLM watermark through ML/Crypto co-design.
The framework: "(i) preserves code functionality, (ii) enhances detectability and robustness by leveraging pre-trained CodeT5 for better feature extraction and adversarial training, and (iii) customizes zero-knowledge proofs for efficient and secure verification without disclosing signatures to third parties."
The watermark insertion module employs the CodeT5, pre-trained on millions of high-quality code files, as the backbone S for watermark encoding.
The encoder extracts code features and fuses them with the watermark message's features. Two decoders predict two sets of probabilities over the syntactic transformations as psyn and variable token distributions pvar.
The watermarked code is obtained by executing the predicted transformations from argmax(psyn) and argmax(pvar).
RoSeMary is trained to meet three criteria: "(i) Functionality-invariant: the functionality of watermarked code remains the same as the input code; (ii) Detectability: the decoded message matches the encoded for successful detection; (iii) Robustness: the adversarial sample's decoded message matches the encoded for robust detection." The training loss combines functionality loss, detectability loss, and robustness loss.
The paper states: "Utilizing zero-knowledge proofs (ZKPs) we can solve this problem. We present a unique watermark extraction scheme, built using non-interactive ZKPs, that can efficiently prove that code has been generated using a proprietary LLM, without revealing what the original watermark was."
The system uses Halo2-based zk-SNARKs, a class of non-interactive ZKPs that offer high scalability and fast verification time.
The verification process involves the LLM owner providing their signature as a private input to a ZKP circuit, ensuring the verifier learns nothing about the watermark beyond the public output bit, making signature reusability possible.
Four custom optimizations are implemented: (1) polynomial approximations
for ReLU activation using sigma(x) = x2 + x, (2) quantize parameters into Bfloat16,
(3) add support for a highly efficient, zero-knowledge bit error rate calculation circuit,
and (4) add a composability layer that allows for efficient combination of Halo2-based and EZKL circuits.
RoSeMary achieves 0.97 detection AUROC while preserving the code functionality.
Compared to baselines, RoSeMary outperforms KGW, REMARK-LLM, SWEET, and SrcMarker across HumanEval, MBPP, and EvalPlus datasets. The paper notes RoSeMary achieves an average of 6.59% higher AUROC scores and 26.61% higher TPR among all benchmarks than SrcMarker.
RoSeMary demonstrates effectiveness across Python, Java, and C++, averagely improving AUROC by 3.78% and TPR by 18.90% over SrcMarker
for C++ and Java.
The paper reports: Generating the zk-SNARK proof only takes the prover P 7.154 seconds, while only requiring approximately 2.6 GB of RAM.
The resulting proof is of size 17.12 KB
and can be verified by any verifier V in an average of 118.6 milliseconds, while only requiring a maximum of approximately 222.5 MB of RAM.
RoSeMary maintains an AUROC above 0.93 under variable-rename attacks even when 50% of variable names are replaced.
Under refactor attacks, it achieves an AUROC of 0.73 when being refactored by Qwen2.5-Coder.
For re-watermarking attacks, RoSeMary is able to maintain 0.72 AUROC under attacks.
The paper summarizes its contributions as: "(i) Introduction of RoSeMary, a code LLM watermarking framework that leverages ML/Crypto co-design to embed watermarking signatures into low-entropy code while supporting confidential ownership verification; (ii) Leverage a pre-trained CodeT5 for better feature extraction and minimizes the code feature loss and message extraction loss between encoded and extracted signatures to encourage detectability of embedded signatures; (iii) Customizing and optimizing zero-knowledge proof circuit to enable efficient ownership verification without revealing encoded signatures; (iv) Performing evaluations on extensive code benchmarks."
The paper concludes: RoSeMary effectively preserves code fidelity, achieves high detectability and robustness, and supports efficient zero-knowledge verification.
The framework is positioned as the most balanced coverage of the requirements for practical code-watermarking systems
compared to existing approaches.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
What I can build: A complete watermarking framework for code-generating LLMs that:
-
Embeds invisible, verifiable signatures into generated code while preserving 100% functionality (pass rate ≥ 97.64%)
-
Achieves 0.97 AUROC detection accuracy—significantly higher than existing methods (SrcMarker: 0.90, SWEET: 0.87)
-
Maintains robustness against adversarial attacks including variable renaming (AUROC > 0.93 even with 50% rename), LLM-based refactoring (AUROC 0.73), and re-watermarking attempts (AUROC 0.72)
-
Works across multiple programming languages (Python, Java, C++) with consistent performance
Key technical improvements:
-
Uses pre-trained CodeT5 as backbone instead of training from scratch, improving feature extraction by 7% AUROC
-
Jointly optimizes three objectives: functionality preservation (MSE loss), detectability (BCE loss), and robustness (adversarial training with perturbed transformation probabilities)
-
Employs dual-channel encoding: syntactic transformations + variable renaming, providing richer watermark space
Key technical innovations:
-
Custom polynomial approximation of ReLU (σ(x) = x2 + x) eliminating lookup tables
-
Bfloat16 quantization reducing proving-side memory by 50% without accuracy loss
-
Novel composability layer fusing feed-forward and BER calculation circuits into single graph
-
In-circuit bit error rate calculation binding extracted signatures to binary values
System capabilities:
-
Handles low-entropy code where traditional watermarking fails (REMARK-LLM achieves 0% pass rate)
-
Supports legal arbitration workflows without exposing IP
-
Maintains watermark strength even with short 4-bit signatures (sufficient for z-score > 1.645)
-
Scales efficiently: decoder size and watermark length have minimal impact on verification cost (118.6 ms for 4-16 bits)
Abstract
This paper introduces RoSeMary, the first-of-its-kind ML/Crypto codesign watermarking framework that regulates LLM-generated code to avoid intellectual property rights violations and inappropriate misuse in software development. High-quality watermarks adhering to the detectability-fidelity-robustness tri-objective are limited due to codes' low-entropy nature. Watermark verification, however, often needs to reveal the signature and requires re-encoding new ones for code reuse, which potentially compromising the system's usability. To overcome these challenges, RoSeMary obtains high-quality watermarks by training the watermark insertion and extraction modules end-to-end to ensure (i) unaltered watermarked code functionality and (ii) enhanced detectability and robustness leveraging pre-trained CodeT5 as the insertion backbone to enlarge the code syntactic and variable rename transformation search space. In the deployment, RoSeMary uses zero-knowledge proofs for secure verification without revealing the underlying signatures. Extensive evaluations demonstrated RoSeMary achieves high detection accuracy while preserving the code functionality. RoSeMary is also robust against attacks and provides efficient secure watermark verification.
Sources
- On Polynomial Approximations for Privacy-Preserving and Verifiable ReLU Networks
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- Optimizing Adaptive Attacks against Watermarks for Language Models
- Trust the Process: Zero-Knowledge Machine Learning to Enhance Trust in Generative AI Interactions
- Discovering Spoofing Attempts on Language Model Watermarks
- Qwen2.5-Coder Technical Report
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- On the Reliability of Watermarks for Large Language Models
- Efficient and Universal Watermarking for LLM-Generated Code Detection
- Watermarking Techniques for Large Language Models: A Survey
- Decoupled Weight Decay Regularization
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
- Leveraging Optimization for Adaptive Attacks on Image Watermarks
- MCGMark: An Encodable and Robust Online Watermark for Tracing LLM-Generated Malicious Code
- RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
- CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks
- Code Llama: Open Foundation Models for Code
- Can AI-Generated Text be Reliably Detected?
- Copilot for Xcode: Exploring AI-Assisted Programming by Prompting Cloud-based Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs