ProxyEraseAgent: Blind Watermark Removal in the Wild

arXiv:2610.11290 · cs.CR · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "ProxyEraseAgent: Blind Watermark Removal in the Wild".

Elias: The gist The key challenge in single-image blind watermark removal is not merely how to transform the image,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we've seen that this paper proposes ProxyEraseAgent as a way to solve that blind watermark removal problem where you can't access the target decoder >

Elias: The authors are pointing out a gap in current research: existing attacks are either dependent on prior knowledge you don't have, or they’re just random changes that don't adapt to whether they’re making progress >

Nadia: They argue that the key challenge is getting a useful direction for removal without knowing the hidden decoder, and they suggest a proxy-in-the-loop system to bridge that gap >

Priya: Essentially, they are using external models as proxies to estimate whether a transformation is actually disrupting the watermark, which gives them an adaptive path forward >

Elias: They propose retrieving informative proxy decoders by calibrating their responses against a knowledge base of known systems >

Nadia: And then they use those retrieved responses to guide sequential removal planning through geometric distortion, compression, and other operations >

Priya: The importance here is that it moves the problem from a generic transformation task to an attack-specific one guided by feedback from external models >

Elias: The results they show across eleven watermarking systems achieved a ninety-four point eight percent attack success rate, which they claim proves their method is effective without needing access to the target decoder > <ref:2610.11290#pg3,a 94.8% attack success rate>

Nadia: That ninety-four point eight percent success rate is what really stands out when you look at the comparison with other baseline methods because it shows consistent effectiveness across different watermarking types > <ref:2610.11290#pg3>

Priya: It’s not just that it works on one system; it’s that they found a way to make the removal process adaptive based on external, public information >

Elias: That's what matters for the cryptographic side too, because it shows how you can derive useful attack guidance from models you don't even own or control >

Conclusion: Nadia: So looking at "ProxyEraseAgent: Blind Watermark Removal in the Wild," we see that they’ve tackled a really hard problem of blind watermark removal by bringing external knowledge into the attack process >

Elias: The authors, including Yao and Wang, are essentially showing how you can build a closed-loop system where you don't need to query the hidden decoder for every single step >

Nadia: It moves the focus from just finding any transformation that looks good to finding transformations that have been specifically chosen because the external models suggest they’re on the right track >

Priya: For someone listening, this means that in a real-world scenario where you can't talk to the watermark encoder or decoder, you can still build a strategy for removal using publicly known systems as guides >

Elias: It implies that the robustness of these watermarks isn't just about the math of the embedding, but how well an attacker can use surrounding information to navigate that embedding space >

Nadia: The implication is that future research in this area should focus on making these proxy retrievals even more robust so they work better when the target decoder is completely hidden >

Priya: And we need to see if this approach scales beyond eleven systems, because if it can consistently guide the attack across many different watermarking mechanisms, that’s where it gets really useful >

Elias: It suggests that understanding how external models react to image changes gives us a new way to evaluate watermark security in practical settings >

Jun Yao, Chao Wang, Yupeng Qiu, Zehua Ma, Weiming Zhang, †Bin Liu‡ Han Fang

University of Science and Technology of China · National University of Singapore

cs.CR

Submitted: 2026-10-08

Updated: 2026-10-08

License: http://creativecommons.org/licenses/by/4.0/

The gist: The gist The key challenge in single-image blind watermark removal is not merely how to transform the image, but how to obtain a useful direction for removal without access to the hidden decoder.

Key concepts

Knowledge Base Construction
This involves gathering information from publicly available image-watermarking systems. It creates a collection of encoders and decoders (Em, Dm) along with metadata (πm) like resolution and configuration. This base allows the agent to test different potential decoders without needing access to the actual target system.
Response-Guided Proxy Retrieval
This stage evaluates candidate decoders from the knowledge base by measuring their confidence in decoding an input image. A higher response strength (Rm(x)) means the decoder is more confident. The calibrated retrieval score (Sm(x_w)) ranks these decoders, ensuring the agent selects those most likely to provide useful feedback.
Pareto Beam Search
This is the planning mechanism that determines the removal steps. It explores various attack operations like geometric distortion or compression by tracking two metrics: proxy watermark disruption and perceptual distortion. It keeps only non-dominated candidates, ensuring the agent finds an effective removal path without exceeding a set perceptual quality limit.
Blind Watermark Removal
The core challenge is removing a hidden watermark from an image when you do not know the secret decoder used to embed it. ProxyEraseAgent solves this by using external, known watermarking systems as proxies. These proxies provide necessary feedback to guide the removal process toward the correct solution.

Terminology

Summary

The gist The key challenge in single-image blind watermark removal is not merely how to transform the image, but how to obtain a useful direction for removal without access to the hidden decoder.

ProxyEraseAgent Framework

ProxyEraseAgent is an agent-driven framework that recovers attack specificity through proxy decoder responses The framework turns a feedback-free removal problem into a proxy-guided decision process through two coupled stages. First, it performs response-guided proxy decoder retrieval where it evaluates candidate decoders from the knowledge base and ranks them according to their calibrated responses to the input. Second, it uses the responses of these proxy decoders to guide removal path planning.

Knowledge Base Construction

To obtain external watermark-related feedback without accessing the target decoder, a heterogeneous knowledge base is constructed from publicly available image-watermarking systems This knowledge base is denoted as K =

M m=1, where Em and Dm denote the encoder and decoder of the m-th watermarking system, respectively, and πm contains the system-specific metadata required for inference, including input preprocessing, image resolution, message length, and decoding configuration. The decoder Dm provides watermark-related responses for proxy retrieval and subsequent attack guidance.

Response-Guided Proxy Retrieval

The Response-guided Proxy Retrieval Agent evaluates candidate decoders in the knowledge base and ranks them according to their calibrated responses to the input For an input image x, the decoding uncertainty is quantified using the mean binary entropy Hm(x) The decoder response strength is defined as Rm(x) = 1 − Hm(x), where a larger Rm(x) indicates more confident bit predictions. The calibrated retrieval score is defined as Sm(x w) = Rm(x w)/Rm(Em(xw)) + ε, where the numerator measures the response of decoder Dm to the unknown watermarked image, and the denominator measures its response to a reference image watermarked by its own encoder, and ε is a small constant for numerical stability. The top-ranked decoders are retained as proxy decoders.

Proxy-Guided Attack Planning Agent

The Proxy-Guided Attack Planning Agent uses the responses of these proxy decoders to guide removal path planning Instead of committing to a fixed transformation, the agent progressively explores a heterogeneous attack space, including geometric distortion, JPEG compression, image reconstruction, and gradient-based perturbation. The attack action space is constructed from four complementary families: geometric distortion, compression distortion, regeneration, and adversarial perturbation. The search proceeds by recursively generating an attack sequence (a1... at) starting from x0 = x w. Each candidate state is evaluated along two dimensions: proxy watermark disruption and perceptual distortion.

Pareto Beam Search

The Proxy-Guided Attack Planning Agent uses Pareto beam search to construct the attack sequence Starting from the original watermarked image, each retained state is expanded at every search step using all admissible attack operations and their parameter settings. The resulting candidates are evaluated according to the proxy disruption score B(x) and perceptual distortion L(x). Candidates whose perceptual distortion exceeds the predefined search budget are discarded A candidate xi is dominated by another candidate xj if B(xj) ≥ B(xi), L(xj) ≤ L(xi), with at least one strict inequality. Only non-dominated candidates are retained, and a Pareto beam of width three is maintained at each search depth. The search proceeds for at most T steps, and the final attack result is selected under the perceptual-quality constraint of LPIPS ≤ 0.05.

Experimental Results

Experiments across 11 image watermarking systems show that ProxyEraseAgent achieves an attack success rate of 94.8% This demonstrates the effectiveness of response-guided proxy retrieval and feedback-driven sequential planning for blind watermark removal. The main comparison with watermark-removal baselines shows that ProxyEraseAgent achieves an overall cASR of 94.8% across 11 watermarking systems, substantially outperforming all baseline methods. This indicates that response-guided proxy retrieval and sequential attack planning can provide effective attack-specific guidance without querying the hidden target decoder. The results show that ProxyEraseAgent achieves the best performance on 9 out of 11 watermarking systems, demonstrating more consistent effectiveness across heterogeneous watermarking mechanisms.

Ablation Studies

Under the blind setting where the target decoder is excluded from the knowledge base, randomly selecting three proxy decoders achieves a cASR of 89.00% Selecting proxy decoders according to calibrated entropy responses achieves a cASR of 94.82% with K = 3. Increasing the number of proxy decoders to K = 3 improves the cASR to 94.82%. Using only the highest-ranked proxy decoder (K = 1) achieves a cASR of 90.36%. Increasing K to 5 further improves the cASR to 95.36%, suggesting that additional proxy signals can still provide marginal benefits. Under the condition of complete knowledge base coverage, the cASR increases to 99.73% with K = 3. Increasing the search depth from one to four achieves a cASR of 83.91% to 94.82%, respectively. The average selected sequence length increases from 1.00 at depth one to 2.96 at depth four, indicating that the planner actively utilizes multi-step attack paths when beneficial.

Conclusion

ProxyEraseAgent retrieves informative proxy decoders from publicly available watermarking models using calibrated responses, and leverages their feedback to adaptively construct effective attack paths Experiments across 11 watermarking systems show that ProxyEraseAgent achieves an overall cASR of 94.8%, outperforming the strongest fixed attack by 15.3%. More importantly, our results demonstrate that the key challenge of blind watermark removal can be addressed by recovering informative feedback from external watermarking models and using it to guide adaptive, image-specific attack planning. This provides a new perspective for evaluating watermark robustness in realistic settings where the underlying watermarking mechanism is unknown and inaccessible.

References

Bang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal, Yuancheng Xu, Chenghao Deng, Sicheng Zhu, Abdirisak Mohamed, Yuxin Wen, Tom Goldstein et al.

Improvements for AI systems

  1. Bold header: ProxyEraseAgent Framework

This agent-driven framework recovers attack specificity through proxy decoder responses by ranking candidate decoders based on their calibrated response strength, allowing for a targeted removal path rather than generic transformations.

  1. Bold header: Response-Guided Proxy Retrieval

The system quantifies decoding uncertainty using the mean binary entropy, defining the decoder response strength as Rm(x) = 1 − Hm(x), which is then normalized via a self-calibration score, making heterogeneous decoder responses comparable.

  1. Bold header: Adaptive Attack Path Planning

ProxyEraseAgent uses the feedback from proxy decoders to guide removal path planning, progressively exploring a heterogeneous attack space (geometric distortion, JPEG compression, image reconstruction, and gradient-based perturbation) to maximize watermark disruption while satisfying perceptual quality constraints.

  1. Bold header: Robustness Across BER Thresholds

The framework demonstrates robustness across different evaluation criteria because it consistently maintains a clear performance advantage over baseline methods across different BER thresholds, proving the improvement is not limited to the default threshold of BER ≥ 0.1.

  1. Bold header: Ablation Study on Proxy Decoder Selection

The research confirms that selecting multiple proxy decoders, such as Entropy 3 Included 99.73%, significantly improves success rates compared to using only a single decoder, indicating that aggregating feedback from multiple proxy decoders provides more robust guidance.

  1. Bold header: Sequential Search Depth Optimization

The system dynamically adjusts its search strategy by varying the maximum search depth (up to four) and utilizing Pareto beam search to select the candidate with the highest proxy disruption score under perceptual constraints, resulting in an optimal balance between effectiveness and image fidelity.

Abstract

Invisible image watermark removal has received growing attention. Despite substantial progress, existing attacks face a tension between practicality and specificity. Attacks exploiting detector outputs, decoder responses, or paired images can be tailored to the watermark decision boundary, but require information rarely available in realistic scenarios. Conversely, attacks based on compression, geometric distortion, or reconstruction are easily deployed from a single watermarked image, but remain largely open-loop: they apply generic transformations without knowing if the image is moving toward watermark failure. Thus, the key challenge in single-image blind watermark removal is not merely how to transform the image, but how to obtain a useful removal direction without accessing the hidden decoder. To bridge this gap, we propose ProxyEraseAgent, an agent-driven framework recovering attack specificity through proxy decoder responses. Publicly available watermarking schemes provide a natural knowledge base of candidate decoders, where some are informative for a given unknown image. Our insight is that a decoder producing a strong calibrated response to the query image likely shares a nearby decoding boundary with the hidden target mechanism. ProxyEraseAgent ranks these decoders by calibrated response strength and uses the top ones as proxy boundary estimators. Their responses then guide a progressive search over heterogeneous removal operations (e.g., geometric distortion, JPEG compression, image reconstruction, and gradient perturbation) under perceptual-quality constraints. Experiments across 11 watermarking systems show ProxyEraseAgent achieves a 94.8% attack success rate, demonstrating the effectiveness of response-guided proxy retrieval and feedback-driven sequential planning for blind watermark removal.

Sources

Related papers