ResearchPod Summary
Logit-based text watermarking embeds secret, pseudo-random green lists into the generation process of large language models to enable reliable provenance tracking. However, its effectiveness is governed by a fundamental trade-off between detectability and semantic distortion. Because existing analytical tools provide limited guidance for principled parameter selection, practical deployments typically rely on ad hoc heuristic tuning. This paper addresses this limitation by developing a rigorous statistical framework that explicitly quantifies the relationships between watermark hyperparameters, statistical power, and semantic distortion.
The authors formulate watermark detection as a sequence-level hypothesis test. By modeling the green-list mechanism under a non-informative Dirichlet prior for next-token probability vectors, the paper derives explicit closed-form expressions for the marginal green-token probability under watermarking. Furthermore, by characterizing the autoregressive dependence of generated tokens through an information decay assumption, the authors establish the asymptotic distribution of the aggregate test statistic under the alternative hypothesis. This yields a tractable, closed-form approximation of the statistical detection power as a function of the watermark parameters.
To manage semantic fidelity, the authors use the expected token-wise Kullback-Leibler divergence between the watermarked and original distributions as the measure of distortion. They demonstrate that for a fixed green-list fraction, the KL divergence is strictly increasing in the logit shift parameter, establishing a one-to-one correspondence between the logit bias and the induced distortion. This enables researchers to reduce the two-dimensional hyperparameter search over the green-list fraction and logit bias down to a single-dimensional numerical optimization problem under explicit constraints.
The framework provides concrete procedures for initializing parameter searches and identifying Pareto-optimal configurations. By parameterizing watermark strength directly in terms of a distortion budget or a target detection power, practitioners can compute optimal settings prior to generation rather than relying on trial-and-error. Empirical validation across multiple language models and datasets demonstrates that this framework consistently identifies configurations that outperform heuristic strategies.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.