DWBench: Holistic Evaluation of Watermark for Dataset Copyright Auditing
cs.CR
Submitted: 2026-02-14
Updated: 2026-09-14
Comments: 26 pages, 4 figures. Accepted to ACM CCS 2026
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: The surging demand for large-scale datasets in deep learning has heightened the need for effective copyright protection, given the risks of unauthorized use to data owners.
Terminology
Abstract
The surging demand for large-scale datasets in deep learning has heightened the need for effective copyright protection, given the risks of unauthorized use to data owners. Although the dataset watermark technique holds promise for auditing and verifying usage, existing methods are hindered by inconsistent evaluations, which impede fair comparisons and assessments of real-world viability. To address this gap, we organize existing methods according to two key dimensions, implementation and verification, to support a consistent analysis and evaluation pipeline across tasks. Based on this framework, we develop DWBench, a unified benchmark and open-source toolkit for systematically evaluating image dataset watermark techniques in classification and generation tasks. Using DWBench, we assess 25 representative methods under standardized conditions, perturbation-based robustness tests, multi-watermark coexistence, and multi-user interference. To enable accurate and reproducible benchmarking, we use TPR@5%FPR for unified sample-level comparison and introduce the verification success rate (VSR) for dataset-level auditing. Key findings reveal that standard single-watermark evaluations tend to overestimate practical auditability. Methods that verify reliably in isolation often suffer from performance degradation at low watermarked-sample ratios, while yielding ambiguous ownership evidence in complex multi-user and multi-watermark settings. We hope that DWBench can facilitate advances in watermark reliability and practicality, thus strengthening copyright safeguards in the face of widespread AI-driven data exploitation.
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs