Feedback Coding Enables Inference-Time Covert Agentic Communication
cs.IT, cs.CR, math.IT
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: As large language models (LLMs) are increasingly used to automate digital interactions, users can leverage LLM-generated text as cover for covert communication within seemingly benign conversations.
Terminology
Abstract
As large language models (LLMs) are increasingly used to automate digital interactions, users can leverage LLM-generated text as cover for covert communication within seemingly benign conversations. Existing LLM steganography, however, is predominantly white-box, requiring the sender and receiver to share the cover statistics, typically through access to the model weights and prompt. Black-box schemes remove this requirement by allowing the receiver to operate solely on the generated text, but current approaches rely on fixed-length, open-loop watermarking techniques that suffer from high decoding error rates under variable-length token generation. We recast black-box LLM steganography as a sequential communication problem with causal, noiseless feedback: every generated token is observed by both parties and can guide subsequent embedding. Based on this perspective, we introduce Burnashev Adaptive Posterior Matching (BAM), a feedback-coding scheme that combines posterior matching with a decode-and-confirm phase. The design is inspired by classical information-theoretic feedback-coding principles, while its security is established through a cryptographic reduction proof. Across three open-weight language models, we demonstrate that BAM attains 0-0.1% empirical message error on an 8-bit payload in around 50 tokens, across 1000 trials, versus 10-17% for the strongest black-box baseline at comparable length. Building on the proposed steganography algorithm, we demonstrate the feasibility of an end-to-end communication protocol that achieves high communication rates across multiple conversational settings.
Sources
- MC$^2$Mark: Distortion-Free Multi-Bit Watermarking for Long Messages
- Perfectly Secure Steganography Using Minimum Entropy Coupling
- BiMark: Unbiased Multilayer Watermarking for Large Language Models
- ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
- MirrorMark: Generalizable Mirrored Sampling for Multi-bit LLM Watermarking
- StealthInk: A Multi-bit and Stealthy Watermark for Large Language Models
- Sequentiality and Adaptivity Gains in Active Hypothesis Testing
- Provable Secure Steganography Based on Adaptive Dynamic Sampling
- Undetectable Conversations Between AI Agents via Pseudorandom Noise-Resilient Key Exchange
- Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
Related papers
- Clipped Affine Policy: Low-Complexity Near-Optimal Online Power Control for Energy Harvesting Communications over Fading Channels
- Discrepancy for Random Linear Codes
- A New Approach to Code Smoothing Bounds
- Contextual Memory-Enhanced Source Coding for Low-SNR Communications
- Symmetry-Enforced Quadratic Approximate-Degradability Bounds for Noisy Landau-Streater Channels
- Anonymous Shamir's Secret Sharing via Reed-Solomon Codes Against Permutations, Insertions, and Deletions