Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design

summary

Video file (mp4)

The gist

The paper introduces RoSeMary, a framework for watermarking code generated by large language models (LLMs) that combines machine learning and cryptographic co-design to secure high-quality watermarks

In short

The episode discusses 'Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design,' a framework called RoSeMary. The hosts explain how to embed secret, robust watermarks into AI-generated code while maintaining functionality. They detail using zero-knowledge proofs for secure ownership verification, enabling commercial use of AI code.

Key concepts

Watermarking
Embedding a secret signature into generated code that only the model owner knows about. This allows the owner to prove ownership if someone uses the code commercially without permission.
Zero-Knowledge Proofs
A cryptographic method allowing an owner to prove they know a secret (like a watermark) without ever having to reveal the actual secret signature itself. This maintains code usability and security.
CodeT5
A pre-trained model used by the authors that helps understand the structure of code. Using this backbone allows for embedding watermarks more subtly and reliably into complex, structured programming languages.
ML/Crypto Co-Design
The novel approach of jointly training machine learning modules (for embedding/extraction) and cryptographic modules (for verification). This optimizes the system for functionality, detectability, and robustness simultaneously.

Terminology used across episodes

This episode discusses

The paper

Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign · Read on arXiv

Ruisi Zhang, Neusha Javidnia, Nojan Sheybani, Noah Marosok, Ke Huang, Farinaz Koushanfar

University of California, San Diego · San Diego State University

This paper introduces RoSeMary, the first-of-its-kind ML/Crypto codesign watermarking framework that regulates LLM-generated code to avoid intellectual property rights violations and inappropriate misuse in software development. High-quality watermarks adhering to the detectability-fidelity-robustness tri-objective are limited due to codes' low-entropy nature. Watermark verification, however, often needs to reveal the signature and requires re-encoding new ones for code reuse, which potentially compromising the system's usability. To overcome these challenges, RoSeMary obtains high-quality watermarks by training the watermark insertion and extraction modules end-to-end to ensure (i) unaltered watermarked code functionality and (ii) enhanced detectability and robustness leveraging pre-trained CodeT5 as the insertion backbone to enlarge the code syntactic and variable rename transformation search space. In the deployment, RoSeMary uses zero-knowledge proofs for secure verification without revealing the underlying signatures. Extensive evaluations demonstrated RoSeMary achieves high detection accuracy while preserving the code functionality. RoSeMary is also robust against attacks and provides efficient secure watermark verification.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design".

Jane: The paper was written by Ruisi Zhang, Neusha Javidnia, Nojan Sheybani, Noah Marosok, Ke Huang et al. from University of California, San Diego and San Diego State University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everybody. Today we're digging into a paper that's got one of those long, intimidating titles, but the idea behind it is genuinely exciting. It's called "Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design," and I'm here with Jane, as always.

Jane: And I'm so glad we're covering this one, Tom. Because when I first read the title, I thought, okay, that's a mouthful. But the core problem is something anyone who's ever written code can understand. If you use a tool like GitHub Copilot to generate code for you, who actually owns that code? And more importantly, how do you prove it?

Tom: Right, and that's where watermarking comes in. You want to embed a secret signature into the code that only the model owner knows about. So if someone takes that code and uses it in their commercial software without permission, the owner can say, hey, that's mine, look, here's the proof.

Jane: Exactly. But here's the catch that this paper tackles head-on. Code is not like a paragraph of text. It's very structured, very precise. There's not a lot of wiggle room to hide a secret message without breaking the code itself. If you change a variable name or restructure a loop, you might accidentally change what the program does.

Tom: So it's a tightrope walk. You need the watermark to be strong enough to detect, but subtle enough that the code still works perfectly. And the authors here, from UC San Diego and San Diego State, they've come up with a clever way to do that using a pre-trained model called CodeT5 to understand the code's structure better.

Jane: And that's just the first half. The second half is about verification. Normally, to prove you own the code, you'd have to reveal the watermark to a third party, like a judge or an arbitrator. But once you reveal it, someone could copy it or erase it. This paper uses something called zero-knowledge proofs to let the owner prove they know the watermark without actually showing it.

Tom: So it's like proving you know the password to a vault without ever saying the password out loud. That's the magic of cryptography, and it's a big deal for protecting intellectual property in the age of AI-generated code. We're going to break down how they actually pulled this off, so stick around.

Summary and Implications: Jane: So, Tom, we've set the stage. Let's get into the meat of this paper, "Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design." The authors are proposing a framework called RoSeMary, which is a bit of a mouthful, but it stands for something meaningful.

Tom: RoSeMary, I like it. And the key insight is that they're not just slapping a watermark on at the end. They're training the watermark insertion and extraction modules together, end-to-end. That's the ML part of the co-design. They want the system to learn what changes are safe to make to the code without breaking it.

Jane: Right. And they're using CodeT5 as the backbone, which is a model pre-trained on millions of code snippets. That gives it a much better understanding of code structure than a model trained from scratch. The results speak for themselves. On the HumanEval benchmark, they're getting a detection AUROC of zero point nine seven, which is really high.

Tom: For our listeners who aren't statisticians, that means they can almost perfectly distinguish between watermarked and non-watermarked code. And they're doing this while keeping the code functional. The pass rate, which is the percentage of code that still passes its unit tests after watermarking, stays above ninety-five percent in most cases.

Jane: That's the fidelity part. But what about when someone tries to mess with the code? Say they rename all the variables or ask another AI to refactor it. That's the robustness test. And RoSeMary holds up. Even when fifty percent of variables are renamed, the detection AUROC stays above zero point nine three.

Tom: That's impressive because a lot of other watermarking schemes fall apart under that kind of attack. The authors also tested against a refactoring attack using a different LLM, and they still got an AUROC of zero point seven three, which is solid. So the ML side is doing its job.

Jane: But the really novel part, the part that got me excited, is the crypto side. They're using zero-knowledge proofs to verify ownership without revealing the watermark. That's the "Zero Knowledge" in the title. And it means the owner can prove they own the code without ever exposing the secret signature.

Tom: Which is huge for reusability. If you never reveal the watermark, you never have to re-encode a new one. The code stays usable, and the proof of ownership stays valid. It's a one-time proof that can be verified by anyone, anytime. That's the kind of security that could actually make AI-generated code viable for commercial use.

Improvements and Methodology: Tom: Alright, Jane, we've talked about the big picture. Now let's get into the nitty-gritty of how RoSeMary actually improves on what came before. Because there are other watermarking schemes out there, but they all have weaknesses.

Jane: Right. The older methods, like KGW and SWEET, work by manipulating the token generation process during inference. They split the vocabulary into "green" and "red" lists and force the model to pick from the green list. But that's a blunt instrument. It can break the syntax of the code, making it uncompilable.

Tom: And then there's SrcMarker, which uses a neural network to embed watermarks in the code's feature space. That's closer to what RoSeMary does, but SrcMarker uses a shallow transformer trained from scratch. It doesn't have the deep understanding of code that a pre-trained model like CodeT5 brings.

Jane: Exactly. So the improvement here is twofold. First, they're using CodeT5 to extract better features from the code, which means the watermark can be embedded more subtly and detected more reliably. Second, they're training the insertion and extraction modules together, with a loss function that balances three things: functionality, detectability, and robustness.

Tom: And that's the "co-design" part. They're not just optimizing for one thing. They're jointly optimizing for all three. The loss function has three components. There's a functionality loss to make sure the watermarked code behaves the same as the original. There's a detectability loss to make sure the watermark can be extracted. And there's a robustness loss to make sure it survives adversarial modifications.

Jane: And they even add a little bit of noise to the transformation probabilities during training, which is like a form of data augmentation. It forces the extractor to learn to recover the watermark even when the code has been slightly altered. That's why it's so robust to attacks.

Tom: Now, on the crypto side, they're using a specific type of zero-knowledge proof called a zk-SNARK, specifically Halo2. And they've made some clever optimizations to make it efficient. They approximate the ReLU activation function with a polynomial, they quantize the weights to Bfloat16, and they fuse the watermark extraction and bit error rate calculation into a single circuit.

Jane: The result is that proof generation takes about seven seconds, which is a one-time cost for the owner. But verification, which is what the arbitrator does, takes only about one hundred eighteen milliseconds. That's fast enough for real-world use. And the proof itself is only about seventeen kilobytes, so it's cheap to transmit.

Tom: So they've made zero-knowledge verification practical for code watermarking, which is a significant step forward. It's not just a theoretical idea anymore. It's something that could actually be deployed.

First Page and Implications: Jane: So, Tom, we've covered the methodology and the results. But let's go back to the very beginning of the paper, the first page, because it sets up the problem so clearly. The authors talk about the "low-entropy nature of code" constraining the space for high-quality watermarks.

Tom: Low entropy, that's a fancy way of saying code is very predictable. There are only so many ways to write a for loop or a function call. So there's not a lot of room to hide a secret message without making the code look weird or breaking it.

Jane: And that's the fundamental challenge. The paper argues that existing solutions either sacrifice functionality for detectability, or they sacrifice security for reusability. If you reveal your watermark to prove ownership, you have to re-encode a new one, which might degrade the code's quality.

Tom: And that's where the zero-knowledge proof comes in as a game-changer. It breaks that trade-off. You can prove ownership without revealing the watermark, so you never have to re-encode. The code stays high-quality, and the ownership proof stays valid.

Jane: The authors also frame this in terms of intellectual property protection. AI-generated code is being used more and more in commercial software. But if you can't prove where that code came from, you can't enforce licensing agreements or pursue legal action against infringement.

Tom: And that's a huge deal for the economy. Think about all the companies building products on top of AI-generated code. They need to know that their code is protected, and they need to be able to prove ownership in court if necessary. This paper provides a framework for doing that securely and efficiently.

Jane: There's also a social implication. If code can be reliably watermarked and verified, it could encourage more open sharing of AI-generated code. Developers might be more willing to contribute to open-source projects if they know their contributions can be attributed and protected.

Tom: And on the flip side, it could help prevent malicious use. If someone uses AI-generated code to create malware, the watermark could help trace it back to the source. That's a powerful deterrent.

Jane: So the implications go beyond just protecting a company's bottom line. It's about building trust in AI-generated content, which is essential for the technology to reach its full potential.

Conclusion: Tom: Well, Jane, we've covered a lot of ground on this paper, "Robust Zero Knowledge Verifiable Watermarking of Code LLMs with ML/Crypto Co-Design." Let's wrap it up for our listeners.

Jane: Absolutely. The core takeaway is that RoSeMary provides a way to watermark AI-generated code that is both robust and secure. It uses a pre-trained model to embed watermarks that don't break the code, and it uses zero-knowledge proofs to verify ownership without revealing the secret signature.

Tom: And the numbers back it up. High detection accuracy, preserved functionality, resilience to attacks, and efficient verification. It's a complete package.

Jane: The implications are significant. This could enable commercial use of AI-generated code with confidence, protect intellectual property, and even help trace malicious code. It's a step towards a future where AI-generated content is trusted and accountable.

Tom: And for the research community, it opens up new directions. The co-design of ML and crypto is a powerful idea that could be applied to other domains, like verifying the provenance of AI-generated images or text.

Jane: So we're saying goodbye to this paper, but we're excited to see where this line of research goes. Thanks for joining us, and we'll see you on the next episode.

Tom: Take care, everyone.

More episodes

← Home