ResearchPod Summary
As large language models (LLMs) become ubiquitous, distinguishing their output from human-authored text is critical for maintaining trust and attribution. Existing watermarking schemes often struggle with robustness; they are frequently defeated by paraphrasing or translation, which alter the specific token sequences or local n-grams that surface-level watermarks rely on. This paper introduces Dual-Embedding Watermarking (DEW) to address these limitations by leveraging semantic representations to create a more resilient, signal-based watermark.
DEW operates by computing watermark biases based on the cosine similarity between projected context embeddings and candidate token embeddings. By using two separate embedding models—one for the preceding context and one for the candidate tokens—the method ensures that semantically related tokens receive similar watermark signals. These signals are obfuscated using pseudo-random matrices seeded with a secret key, which are then applied as zero-centered biases to the model's logits during generation. During detection, the algorithm reconstructs these biases to compute a document-level score, which is then evaluated against a statistical null distribution to determine the presence of the watermark.
Experimental results across multiple LLMs demonstrate that DEW significantly outperforms existing semantic and surface-level watermarking baselines in robustness. Specifically, DEW maintains high detection rates after paraphrasing and remains notably more detectable than prior methods after translation into languages such as German and French. Furthermore, DEW achieves this robustness with lower computational overhead than other semantic watermarking schemes, while maintaining text quality comparable to unwatermarked generations as measured by perplexity and oracle-based preference scores.
DEW provides a practical, efficient, and robust solution for the attribution of LLM-generated content. By moving beyond simple token-hash dependencies toward a semantic, signal-processing approach, it offers a viable path for safeguarding AI-generated text against common adversarial transformations like translation and paraphrasing, which are currently major hurdles for responsible AI deployment.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.