ResearchPod Summary
Existing multi-bit watermarking schemes for LLMs often rely on hard-decision decoding, where each token is mapped to a single bit. This approach discards the model's inherent token-level probability information, making the watermark vulnerable to token-level edits, insertions, and paraphrasing. The authors ask whether incorporating soft-decision decoding—which uses log-likelihood ratios (LLRs) to represent token reliability—can improve watermark robustness and detection power without sacrificing false-positive control.
To enable soft-decision decoding, the authors introduce CORE-BREW, which calibrates the watermark channel by forcing a constant hit-rate (p-star) at every generation step. By adaptively adjusting logit offsets, the system ensures that the probability mass assigned to the target token list remains fixed, regardless of the prompt or context. This calibration allows for the calculation of principled, closed-form LLRs.
To maintain text quality, the authors implement entropy-aware safeguards that treat low-entropy or highly predictable positions as "erasures" (zero LLR contribution) rather than forcing large, quality-degrading logit shifts. The system supports two detection modes:
Experiments on open-source LLMs demonstrate that CORE-BREW significantly improves discrimination at low FPR levels compared to existing ECC-based baselines. By leveraging soft information and erasure-aware decoding, the system maintains higher robustness against token-level attacks and paraphrasing while preserving semantic quality comparable to unwatermarked text. The two-mode decoding architecture provides flexibility, allowing users to choose between strict safety guarantees and higher detection sensitivity.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.