WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency
Z Sun, Q Jiang, S Sheng, L Xiang
cs.CR, cs.AI
Submitted: 2026-07-14
Comments: 21 page
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Large language models (LLMs) have achieved remarkable success but raise growing concerns about content provenance and misuse, motivating the need for reliable watermarking techniques.
Terminology
Abstract
Large language models (LLMs) have achieved remarkable success but raise growing concerns about content provenance and misuse, motivating the need for reliable watermarking techniques. However, these techniques have rarely been adopted in practice mainly for two reasons: i) severely degraded model performance, and ii) additional inference overhead. To confirm the problem, we construct a comprehensive benchmark spanning different generation tasks to systematically evaluate 9 representative watermarking methods. We found almost all existing methods are designed for text fluency, but not for restricted and complicated tasks, and their overhead prevents them from deployment in latency-critical systems. To address i) and ii), we propose an LLM watermarking scheme WaterMoE for the growingly popular Mixture-of-Experts (MoE) LLMs. WaterMoE embeds watermarking signals through controlled perturbation into the expert selection at each router, which accumulates to token selection shift at the final output. In contrast to watermarking as a post-processing token-sampling approach, WaterMoE embeds watermark within the inference loop incurring negligible quality degradation and computational overhead. Extensive experiments demonstrate that our method achieves a fidelity performance close to the unwatermarked and consistently outperforms state-of-the-art watermarking methods on the benchmark, with up to 4 times speedup, incurring merely 1% additional inference latency compared to native generation. The results demonstrate the capability of WaterMoE to be deployed in real-world tasks.
Sources
- Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
- Understanding the planning of LLM agents: A survey
- A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
- LLM Agents for Education: Advances and Applications
- RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience
- Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
- DeepSeek-V3 Technical Report
- Qwen3 Technical Report
- Mixtral of Experts
- Measuring Coding Challenge Competence With APPS
- Training Verifiers to Solve Math Word Problems
- Measuring Massive Multitask Language Understanding
- Instruction-Following Evaluation for Large Language Models
- WritingBench: A Comprehensive Benchmark for Generative Writing
- Robust Distortion-free Watermarks for Language Models
- Unbiased Watermark for Large Language Models
- An Unforgeable Publicly Verifiable Watermark for Large Language Models
- A Semantic Invariant Robust Watermark for Large Language Models
- Black-Box Detection of Language Model Watermarks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs