CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking".
Nadia: As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts regarding "CORE-BREW." The goal is to synthesize these descriptions into a single, comprehensive, long,
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So, we're talking about CORE-BREW today, which is this new framework for watermarking Large Language Models using Logarithmic Likelihood Ratios for soft decoding instead of just hard decisions. Elias, what catches your eye about the title and the authors?
Elias: The title itself points to a shift in how we handle robustness; moving from simple hard decisions to principled soft decision decoding is the big move here, which is exactly what I'm interested in. The authors are Joeun Kim and HoEun Kim, researchers from DGIST Daegu.
Priya: From my side, I’m curious about what this means for the actual data we’re measuring; does this framework fundamentally change the fidelity of the embedded information when we run adversarial tests?
Nadia: Exactly, Priya; it's not just a tweak, it's a complete re-thinking of how we verify those watermarks against edits. This paper proposes CORE-BREW as a way to make LLM provenance reliable under heavy editing attempts.
Elias: It’s important to understand that the authors are tackling the problem where existing ECC-based watermarks often discard crucial token-level reliability information during their hard decisions, which is a major weakness in their approach.
Priya: And what does this LLR analysis they introduce actually tell us about the underlying probability distribution of an LLM's output? Is that information accessible in a meaningful way?
Nadia: It gives them closed-form per-token Logarithmic Likelihood Ratios, or LLRs, which are the soft evidence needed for principled soft decision decoding. This lets them leverage the nuance of the model's output probabilities instead of just a yes or no answer.
Elias: And that calibration is key because they use a Constant Hit-Rate extension to block-wise BREW, targeting a fixed hit-rate p, which establishes this position-homogeneous reliability scale for the channel. That sounds like they’re trying to fix the context dependency issue that plagued previous soft decoding attempts.
Priya: Fixing that context dependency seems vital; if the channel model is position-homogeneous, then we can actually trust those LLRs more when we measure robustness under paraphrasing attacks.
Nadia: Right, and this leads us into the core of what they propose: two distinct detection modes—the Strict-Safe Decoder and the FPRCalibrated Decoder. These are designed to give users control over their verification needs.
Elias: The Strict-Safe mode focuses on preserving fidelity by strictly adhering to the designated codeword acceptance region, which maintains a high degree of fidelity regarding known constraints.
Priya: So, one mode prioritizes absolute safety and adherence to boundaries, while the other must be balancing detection power against how often we get false positives during testing.
Nadia: Precisely; the FPRCalibrated mode uses likelihood-based scoring and lightweight list decoding to precisely map out that trade-off between False Positive Rate and True Positive Rate.
Title and authors: Elias: It sounds like they’ve built a system where you can tune the sensitivity of the watermarking mechanism based on whether you need high certainty or better detection power when things get messy.
Priya: That tuning capability is what makes it practical for real-world deployment; we don't want an oversensitive system that flags every slightly paraphrased document as compromised.
Nadia: It’s about managing that trade-off effectively, and they’ve also added several safeguards to keep things stable during the process.
Elias: I noticed they introduced entropy-aware erasure safeguards to handle low-entropy contexts, treating skipped or low-confidence positions as erasures with zero LLR evidence. That prevents extreme biasing of the decoding process when the base model output is almost deterministic.
Priya: That stability is important because I worry that if we push these systems too hard, we could accidentally introduce semantic degradation in those low-entropy areas, and this safeguard seems to prevent that quality loss.
Nadia: It does, and they also included window shifting capabilities which further helps ensure alignment robustness during the detection phase. These little adjustments seem designed to keep the process smooth when tokens are slightly misaligned.
Elias: The experiments they ran on open-source LLMs like OPT-1 point 3B and Mistral-7B show that CORE-BREW significantly improves low false positive discrimination and maintains comparable semantic quality, even against token-level edits.
Priya: That's the most critical part for us; if it robustly handles those token-level attacks without making the watermarked text look garbage to a human reader, then this moves from a theoretical curiosity to something useful for auditing.
Nadia: It does, and the authors provide complete proofs in Appendix C covering things like the Constant Hit-Rate property and FPR-Calibrated score-based tail bounds. That level of mathematical backing is reassuring when we're dealing with security claims.
Elias: Mathematically, they’ve shown how to prove the bounds on the detection performance in both modes, which solidifies the theoretical foundation for using these LLRs effectively.
Priya: So, when we look at real applications like verifiable provenance or auditing policy compliance tags, this framework seems to offer a much more reliable signal than what we saw in earlier ECC-based methods.
Nadia: It definitely offers a different kind of evidence; instead of just confirming presence or absence, it provides a quantified measure of how much the text has been perturbed while still recovering the watermark.
Elias: The implications here are that we can move toward more sophisticated integrity checks in dynamic environments where minor edits are expected, like verifying generated code blocks or policy drafts
Nadia's application point five: .
Priya: I think the real impact is in content streams, where CORE-BREW could distinguish genuine watermarked outputs from sophisticated paraphrased adversarial text designed to bypass simpler detectors
Nadia's application point four: .
Nadia: Exactly; it gives us a tool that can operate at a higher level of semantic scrutiny than what we’ve seen with prior methods like DetectGPT, which is what we need for high-stakes verification.
Title and authors: Elias: If we look at the underlying cryptography, the assumption is that the LLRs are well-defined via the constant hit-rate calibration, meaning it relies on a consistent probability structure across all time steps. That consistency is what makes it robust against context variability.
Priya: It sounds like a significant step forward in measurement research because it moves the performance evaluation beyond just binary detection metrics and into the realm of continuous likelihood scores.
Nadia: So, to wrap up this discussion on CORE-BREW: it’s a principled framework that takes LLM watermarking from heuristic embedding to a mathematically grounded soft decision process.
Elias: It’s about creating a channel model that respects the underlying probability distribution, allowing for better control over the FPR versus TPR trade-off through those two distinct detection modes.
Priya: For us, it means we can measure watermarking robustness with much higher precision and confidence in a variety of challenging editing scenarios
Nadia's application point nine: .
Nadia: And for the applications we discussed, whether it’s verifying provenance or auditing policy tags, this framework provides a solid foundation for ensuring accountability in AI-generated text
Priya's application point two: .
Elias: The main thing to keep in mind is that the method itself relies on the calibration being correct; if that constant hit-rate assumption doesn't hold perfectly, the LLRs might become inaccurate.
Priya: And one limitation we see in their work is that they focus heavily on token-level edits; we need to see how this performs when the attacks are more complex, like full paraphrasing, which is where prior methods struggled
CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking: .
Nadia: That’s a fair point; they did show results against paraphrase attempts, but we need more data on truly adversarial content streams to confirm its limits
CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking: .
Elias: Well, it’s certainly an improvement over the hard decision methods like MPAC or Qu et al. because they explicitly incorporate error correction codes into their design.
Priya: It seems the authors are committed to addressing these gaps through their work on CORE-BREW, which is a lot to take in terms of how much they've refined the core signal handling
CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking: .
Nadia: It’s certainly a deep dive into the mechanics of robust watermarking, and I think it sets a new standard for what we expect from these provenance systems.
Elias: Indeed, the shift to LLR-based decoding suggests that principled soft evidence is now the path forward for reliable multi-bit embedding in LLMs
CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking: .
Priya: So, we've covered the title, how it works, and what it means for our measurements. That gives us a good overview of the CORE-BREW paper today.
The paper's summary: Nadia: So, we just finished digging into the mechanics of CORE-BREW, and now Priya needs to unpack what that actually means for us in terms of real-world security and measurement.
Elias: Exactly. The core concept is moving away from those blunt hard decisions when decoding watermarks and instead using a Logarithmic Likelihood Ratio analysis, which gives us much richer evidence about the token's authenticity.
Priya: From my side, what I’m seeing in the summary is that they’ve essentially built a calibrated channel model that accounts for the underlying probability distribution of the LLM output rather than just treating it as a binary yes or no.
Nadia: It sounds like they’ve managed to get explicit control over the trade-off between detection power and how often we get false positives using those two distinct decoding modes.
Elias: That calibration through the constant hit-rate extension is what makes their LLR computation principled, which is a huge step forward from previous heuristic methods we’ve seen in this area.
Priya: And I see them introducing entropy-aware erasure safeguards, which prevents the system from becoming unstable when the base model starts producing really predictable text.
Nadia: That stability is what matters because if the mechanism starts making wild guesses in low-entropy contexts, all our verification efforts fall apart and we lose trust in the output.
Elias: Furthermore, they’ve detailed how those safeguards interact with their window shifting and list decoding techniques to manage the FPR–TPR curve precisely.
Priya: So, what this shows us is that we can now theoretically tune a watermarking system to be either extremely conservative for safety checks or aggressively sensitive for detection power in specific scenarios.
Nadia: That tuning capability is huge because it means we can deploy these tools where the risk profile changes constantly, like in real-time content moderation streams.
Elias: And when you look at the empirical validation they did, they’ve shown that CORE-BREW maintains strong discrimination against token-level edits while keeping semantic quality high, which is a tough balance to strike.
Priya: That result is telling because it means we aren't sacrificing the readability of the watermarked text just to make it more robust against simple paraphrasing attacks.
Nadia: It gives us a solid foundation for things like verifiable provenance in journalism or policy documents, where ensuring accountability is paramount.
Elias: I wonder, though, if the assumption they make about a position-homogeneous reliability scale holds up perfectly when we move from small models to the larger LLMs we're actually deploying in production.
Priya: That’s a fair point; scalability and generalization across different model architectures are always where we need to be cautious when evaluating these kinds of theoretical proofs.
Nadia: Well, that leads us right into the exploitation question, Elias—who could possibly exploit this system cheaply if they knew the exact calibration parameters?
Elias: That’s what I want to know next; understanding the cost of exploiting a principled soft decision process is where we figure out its real-world security value.
The paper's improvements: Tom: We’ve got the summary of CORE-BREW, and now we need to talk about what they propose as improvements to make this framework even better than what we have today for multi-bit watermarking.
Nadia: So, I’m looking at how they suggest refining the detection modes; they aren't just offering two options, but a calibrated approach that lets you really tune the sensitivity based on your specific security needs.
Elias: That calibration process is interesting because it moves beyond a fixed threshold and lets us exploit soft evidence beyond just the nominal radius of error, which should theoretically increase detection power significantly.
Priya: What’s really striking to me are these entropy-aware erasure safeguards; they’re designed to keep the system stable when the base model output is almost completely deterministic, preventing any quality degradation there.
Nadia: That stability is what we need for long-term deployment; if the mechanism starts making wild guesses in those low-entropy areas, all our verification efforts fall apart and we lose trust in the output.
Elias: I agree with Priya on that; it addresses a key weakness where hard decision methods often fail because they treat all tokens equally regardless of their context or confidence level.
Priya: And those window shifting capabilities they mentioned seem like a clever way to ensure alignment robustness during the detection phase, helping to smooth out token misalignment issues.
Nadia: It sounds like the authors are building a system that’s not just robust against simple edits but also handles the nuances of context and low-confidence predictions effectively.
Elias: Indeed, and I’m looking at how they handle the FPR–TPR trade-off through score-based decoding; it gives us a mathematical way to quantify exactly how much detection power we gain for every bit of false positives we accept.
Priya: That quantification is vital because it allows us to make informed decisions about deployment, telling us exactly what level of false positives we can tolerate versus the level of attack sophistication we can detect.
Nadia: If this works as described, it means we could have much more reliable auditing tools for policy compliance tags in massive legal documents because the signal remains strong even under minor paraphrasing.
Elias: The underlying assumption they rely on is that this constant hit-rate calibration remains accurate across the entire sequence, and if that assumption breaks down due to context variability, the LLRs might become unreliable.
Priya: That’s a valid concern; we need to see if their proofs hold up when we test against more complex, long-form adversarial content streams where contextual shifts are severe.
Nadia: So, while the proposed improvements sound very promising for real-world applications like verifying generated code integrity, it hinges entirely on the mathematical consistency of that initial calibration step.
Elias: Exactly; if the calibration isn't perfect, you’re just getting a fancy soft decoder with a potentially inaccurate probability landscape.
Conclusion: Tom: We’ve reached the conclusion of our discussion on CORE-BREW, where we’ve covered the technical details and what this framework means for AI security and measurement research.
Nadia: So, to recap, CORE-BREW takes LLM watermarking away from simple hard decisions by using LLR analysis to create a principled soft decoding process that offers tunable control over detection power and false positive rates.
Elias: That’s the high-level summary; it’s essentially a mathematically grounded method for embedding multi-bit signals into AI outputs using probabilistic evidence rather than just fixed boundaries.
Priya: What I think is the biggest implication is that we can finally move toward measuring watermarking robustness with much higher precision, which opens up new avenues for auditing policy compliance tags in large documents.
Nadia: It’s exciting because it suggests that verifiable provenance isn't just a theoretical concept anymore; we have a framework that can provide quantified evidence of integrity against editing attacks.
Elias: I'm still focused on the cryptographic assumptions, though; the entire system relies heavily on the accuracy of that constant hit-rate calibration to define those per-token LLRs reliably.
Priya: And from a measurement standpoint, seeing how it handles low-entropy contexts with those erasure safeguards shows a real commitment to maintaining quality in all operational regimes.
Nadia: It’s definitely a lot to take in about how much this refines the mechanics of robust watermarking compared to earlier heuristic methods.
Elias: I think the next big challenge for anyone working on this is rigorously testing those theoretical bounds when we introduce more complex, context-dependent adversarial scenarios than what they focused on initially.
Priya: That’s exactly where my focus will be; we need to push them to show how this handles genuine paraphrasing attacks beyond just simple token substitution.
Nadia: Well, it’s clear that CORE-BREW sets a new standard for what we expect from systems that claim to verify the integrity of AI-generated text.
Elias: Indeed; the move to LLR-based decoding is definitely the path forward for reliable multi-bit embedding in LLMs.
Joeun Kim, HoEun Kim, *Young-Sik Kim
Department of AI, DGIST
cs.CR, cs.CL
Submitted: 2026-06-23
Updated: 2026-09-29
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
The gist: As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts regarding "CORE-BREW." The goal is to synthesize these descriptions into a single, comprehensive, long,
Key concepts
- Channel Shaping
- This is the process of calibrating the watermark channel using a Constant Hit-Rate (CHR) extension. It sets a fixed reliability scale ($ ext{p}^ ext{star}$) to ensure that the watermark signal's strength is consistently represented across all tokens, which is necessary for accurate LLR computation.
- Logarithmic Likelihood Ratios (LLRs)
- LLRs are the core evidence used in CORE-BREW. They represent the log-likelihood ratio of observing a specific token compared to an alternative. These values serve as principled soft evidence, allowing the system to make nuanced decisions beyond simple hard boundaries.
- Dual Detection Modes
- The framework offers two distinct decoding modes: a Strict-Safe Decoder for maximum fidelity and a FPRCalibrated Decoder for better detection power. This allows users to choose between prioritizing guaranteed preservation or maximizing the ability to detect subtle watermarks.
Terminology
Summary
As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts regarding CORE-BREW.
The goal is to synthesize these descriptions into a single, comprehensive, long, and detailed summary that accurately reflects the technical depth of the work while ensuring no critical contribution is overlooked.
Here is the combined, detailed summary:
The paper introduces CORE-BREW, a novel and robust framework designed for multi-bit watermarking of Large Language Models (LLMs). The central innovation of CORE-BREW lies in transforming the traditional hard decision process into a principled, softdecision decoding mechanism by strategically embedding watermark signals directly into the model's output logits during the inference phase. This approach is specifically engineered to enhance robustness against a wide spectrum of adversarial attacks, including token-level edits, paraphrasing attempts, and general substitution/deletion/insertion attacks.
The foundational technical contribution of CORE-BREW is its sophisticated method for channel shaping the watermark channel. Instead of relying on arbitrary embedding methods, CORE-BREW employs a Constant Hit-Rate (CHR) extension to the existing block-wise BREW technique. This calibration targets a fixed hit-rate parameter (p), which establishes a position-homogeneous reliability scale for the watermark channel. This calibration is crucial because it enables the computation of closed-form, per-token Logarithmic Likelihood Ratios (LLRs). These LLRs serve as the principled soft evidence necessary for principled softdecision decoding, moving beyond simple hard decision boundaries to leverage the nuanced probability distributions inherent in LLM outputs.
CORE-BREW distinguishes itself by offering two distinct detection modes, effectively separating the framework's guarantees from its detection capabilities:
-
Strict-Safe Decoder: This mode focuses on preservation and guarantee. It operates by strictly adhering to the designated codeword acceptance region, ensuring that the bounded-distance verification is maintained, thereby preserving a high degree of fidelity concerning known constraints.
-
FPRCalibrated Decoder: This mode prioritizes detection power while explicitly managing the trade-off between False Positive Rate (FPR) and True Positive Rate (TPR). It utilizes likelihood-based scoring combined with lightweight list decoding techniques to precisely characterize this critical FPR–TPR trade-off, allowing users to tune the system based on their specific operational requirements.
To ensure practical applicability and mitigate quality degradation in challenging scenarios, CORE-BREW incorporates several advanced safeguards:
-
Entropy-Aware Erasure Safeguards: These mechanisms are introduced to handle low-entropy contexts (i.e., when the base model output is near-deterministic). In such cases, skipped or low-confidence positions are treated as erasures—contributing zero LLR evidence—effectively preventing extreme biasing of the decoding process and maintaining stability.
-
Window Shifting: The framework supports window shifting capabilities, which further aids in achieving alignment robustness during the detection phase.
The authors provide extensive empirical validation demonstrating that CORE-BREW significantly outperforms prior multi-bit watermarking baselines across diverse testing conditions. The experiments are comprehensive, evaluating the framework on:
-
Various open-source LLMs (e.g., OPT-1.3B, Mistral-7B).
-
Multiple datasets and model configurations.
-
A wide array of attacks, specifically token-level edits and complex paraphrasing scenarios.
The results consistently show improved low-FPR discrimination and superior robustness while crucially maintaining comparable semantic quality of the target text, confirming that the increased robustness does not come at a significant cost to the generated content's utility. The theoretical underpinnings are rigorously supported by complete proofs in Appendix C, covering key properties such as the Constant Hit-Rate property, Strict-Safe false-positive bounds, and FPR-Calibrated score-based tail bounds.
The research is committed to reproducibility. The authors provide open access to data and code, including anonymized code releases for both CORE-BREW-Strict and CORE-BREW-Cal, configuration files, environment requirements, and executable scripts, allowing researchers to fully reproduce the main experimental results. Furthermore, the paper thoroughly addresses the broader societal implications of watermarking LLMs. It discusses potential positive uses (e.g., provenance verification) alongside significant risks (misuse, false attribution). Explicit safeguards regarding data release and model checkpoint distribution are detailed to mitigate redistribution risks, adhering toNeurIPS Code of Ethics standards.
**In essence, CORE-BREW is a principled framework that elevates LLM watermarking from heuristic embedding to a mathematically grounded softdecision process.
Improvements for AI systems
As a fastidious researcher, I have analyzed CORE-BREW in detail. This framework offers significant advancements over prior ECC-based watermarking by replacing hard-decision decoding with principled soft-decision LLR analysis, effectively calibrating the watermark channel and providing explicit trade-off control between detection power (TPR) and false positive rate (FPR).
Here are the specific improvements this system enables in AI applications:
)Specific Improvements Enabled by CORE-BREW
-
CORE-BREW-Cal Enables high TPR (up to 97%) while maintaining near-zero FPR (0.002), providing the best overall performance across clean settings and under paraphrasing attacks, significantly outperforming baseline methods like MPAC and Qu et al.
-
Strict-Safe Mode Guarantees strict adherence to the original bounded-distance acceptance region (inherited combinatorial FPR control), making it ideal for high-stakes verification where false positives are absolutely unacceptable, even if it sacrifices some detection power compared to the calibrated mode.
-
Entropy-Aware Erasure Safeguards Prevents semantic degradation and quality loss in low-entropy contexts (e.g., near deterministic token generation), ensuring that the watermark mechanism remains stable across all operational regimes without requiring excessively large, potentially damaging logit shifts.
-
Robustness to Token-Level Attacks Maintains strong discrimination against synonym substitution and insertion/deletion attacks compared to baselines, crucial for verifying content integrity in dynamic environments like code generation or policy drafting where minor edits are common.
-
FPR-Calibrated Mode (Score-Based Decoding) Allows the system to exploit soft evidence beyond the nominal Hamming radius, enabling recovery of watermarked content even when local token alignment is slightly perturbed, thus increasing detection power without compromising false positive control through score-based rejection and candidate budget management.
)What the Improved AI System Can Do (Specific Applications)
-
Verifiable Provenance for LLM Outputs The system can be deployed to provide cryptographic proof that a generated text snippet originated from a specific, authorized model instance using a secret key, ensuring accountability in journalism or public policy documents.
-
Audit of Policy Compliance Tags It can reliably verify the presence of embedded policy tags (multi-bit payloads) within large legal or technical documents, ensuring that compliance signals have not been maliciously altered by minor editing or paraphrasing attempts.
-
Detection in Adversarial Content Streams In real-time content moderation pipelines, CORE-BREW can distinguish between genuine watermarked outputs and sophisticated paraphrased adversarial text designed to evade simpler detectors like DetectGPT, offering a high-confidence signal against semantic rewriting.
-
High-Stakes Code Integrity Verification For software development (e.g., verifying generated code blocks), the system can ensure that critical security tags or user identifiers embedded in the code have not been subtly altered by insertion or deletion attacks, providing a strong guarantee of integrity for high-security outputs.
Sources
- Watermarking Language Models with Error Correcting Codes
- Mistral 7B
- Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking
- Necessary and Sufficient Watermark for Large Language Models
- ParaNMT-50M: Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine Translations
- Attacking Neural Text Detectors
- OPT: Open Pre-trained Transformer Language Models
- BERTScore: Evaluating Text Generation with BERT
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs