ProxyEraseAgent: Blind Watermark Removal in the Wild

summary

Video file (mp4)

The gist

The gist The key challenge in single-image blind watermark removal is not merely how to transform the image, but how to obtain a useful direction for removal without access to the hidden decoder.

In short

ProxyEraseAgent recovers blind watermarks by using a knowledge base of public watermarking systems to find 'proxy decoders.' It ranks these decoders based on their response to an input image and then uses their feedback to plan an adaptive, multi-step attack path. This method successfully removes watermarks with a 94.8% success rate across 11 different systems, showing that external feedback can guide removal when the target decoder is unknown.

Key concepts

Knowledge Base Construction
This involves gathering information from publicly available image-watermarking systems. It creates a collection of encoders and decoders (Em, Dm) along with metadata (πm) like resolution and configuration. This base allows the agent to test different potential decoders without needing access to the actual target system.
Response-Guided Proxy Retrieval
This stage evaluates candidate decoders from the knowledge base by measuring their confidence in decoding an input image. A higher response strength (Rm(x)) means the decoder is more confident. The calibrated retrieval score (Sm(x_w)) ranks these decoders, ensuring the agent selects those most likely to provide useful feedback.
Pareto Beam Search
This is the planning mechanism that determines the removal steps. It explores various attack operations like geometric distortion or compression by tracking two metrics: proxy watermark disruption and perceptual distortion. It keeps only non-dominated candidates, ensuring the agent finds an effective removal path without exceeding a set perceptual quality limit.
Blind Watermark Removal
The core challenge is removing a hidden watermark from an image when you do not know the secret decoder used to embed it. ProxyEraseAgent solves this by using external, known watermarking systems as proxies. These proxies provide necessary feedback to guide the removal process toward the correct solution.

Terminology used across episodes

This episode discusses

The paper

ProxyEraseAgent: Blind Watermark Removal in the Wild · Read on arXiv

Jun Yao, Chao Wang, Yupeng Qiu, Zehua Ma, Weiming Zhang, †Bin Liu‡ Han Fang

University of Science and Technology of China · National University of Singapore

Invisible image watermark removal has received growing attention. Despite substantial progress, existing attacks face a tension between practicality and specificity. Attacks exploiting detector outputs, decoder responses, or paired images can be tailored to the watermark decision boundary, but require information rarely available in realistic scenarios. Conversely, attacks based on compression, geometric distortion, or reconstruction are easily deployed from a single watermarked image, but remain largely open-loop: they apply generic transformations without knowing if the image is moving toward watermark failure. Thus, the key challenge in single-image blind watermark removal is not merely how to transform the image, but how to obtain a useful removal direction without accessing the hidden decoder. To bridge this gap, we propose ProxyEraseAgent, an agent-driven framework recovering attack specificity through proxy decoder responses. Publicly available watermarking schemes provide a natural knowledge base of candidate decoders, where some are informative for a given unknown image. Our insight is that a decoder producing a strong calibrated response to the query image likely shares a nearby decoding boundary with the hidden target mechanism. ProxyEraseAgent ranks these decoders by calibrated response strength and uses the top ones as proxy boundary estimators. Their responses then guide a progressive search over heterogeneous removal operations (e.g., geometric distortion, JPEG compression, image reconstruction, and gradient perturbation) under perceptual-quality constraints. Experiments across 11 watermarking systems show ProxyEraseAgent achieves a 94.8% attack success rate, demonstrating the effectiveness of response-guided proxy retrieval and feedback-driven sequential planning for blind watermark removal.

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "ProxyEraseAgent: Blind Watermark Removal in the Wild".

Elias: The gist The key challenge in single-image blind watermark removal is not merely how to transform the image,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we've seen that this paper proposes ProxyEraseAgent as a way to solve that blind watermark removal problem where you can't access the target decoder >

Elias: The authors are pointing out a gap in current research: existing attacks are either dependent on prior knowledge you don't have, or they’re just random changes that don't adapt to whether they’re making progress >

Nadia: They argue that the key challenge is getting a useful direction for removal without knowing the hidden decoder, and they suggest a proxy-in-the-loop system to bridge that gap >

Priya: Essentially, they are using external models as proxies to estimate whether a transformation is actually disrupting the watermark, which gives them an adaptive path forward >

Elias: They propose retrieving informative proxy decoders by calibrating their responses against a knowledge base of known systems >

Nadia: And then they use those retrieved responses to guide sequential removal planning through geometric distortion, compression, and other operations >

Priya: The importance here is that it moves the problem from a generic transformation task to an attack-specific one guided by feedback from external models >

Elias: The results they show across eleven watermarking systems achieved a ninety-four point eight percent attack success rate, which they claim proves their method is effective without needing access to the target decoder > <ref:2610.11290#pg3,a 94.8% attack success rate>

Nadia: That ninety-four point eight percent success rate is what really stands out when you look at the comparison with other baseline methods because it shows consistent effectiveness across different watermarking types > <ref:2610.11290#pg3>

Priya: It’s not just that it works on one system; it’s that they found a way to make the removal process adaptive based on external, public information >

Elias: That's what matters for the cryptographic side too, because it shows how you can derive useful attack guidance from models you don't even own or control >

Conclusion: Nadia: So looking at "ProxyEraseAgent: Blind Watermark Removal in the Wild," we see that they’ve tackled a really hard problem of blind watermark removal by bringing external knowledge into the attack process >

Elias: The authors, including Yao and Wang, are essentially showing how you can build a closed-loop system where you don't need to query the hidden decoder for every single step >

Nadia: It moves the focus from just finding any transformation that looks good to finding transformations that have been specifically chosen because the external models suggest they’re on the right track >

Priya: For someone listening, this means that in a real-world scenario where you can't talk to the watermark encoder or decoder, you can still build a strategy for removal using publicly known systems as guides >

Elias: It implies that the robustness of these watermarks isn't just about the math of the embedding, but how well an attacker can use surrounding information to navigate that embedding space >

Nadia: The implication is that future research in this area should focus on making these proxy retrievals even more robust so they work better when the target decoder is completely hidden >

Priya: And we need to see if this approach scales beyond eleven systems, because if it can consistently guide the attack across many different watermarking mechanisms, that’s where it gets really useful >

Elias: It suggests that understanding how external models react to image changes gives us a new way to evaluate watermark security in practical settings >

More episodes

← Home