Pith. sign in

REVIEW 2 cited by

Watermarking Language Models for Many Adaptive Users

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11109 v2 pith:6GUGWACA submitted 2024-05-17 cs.CR cs.AIcs.CL

classification cs.CRcs.AIcs.CL
keywords watermarkingtextzero-bitlanguageschemesusersadaptiveeven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study watermarking schemes for language models with provable guarantees. As we show, prior works offer no robustness guarantees against adaptive prompting: when a user queries a language model more than once, as even benign users do. And with just a single exception (Christ and Gunn, 2024), prior works are restricted to zero-bit watermarking: machine-generated text can be detected as such, but no additional information can be extracted from the watermark. Unfortunately, merely detecting AI-generated text may not prevent future abuses. We introduce multi-user watermarks, which allow tracing model-generated text to individual users or to groups of colluding users, even in the face of adaptive prompting. We construct multi-user watermarking schemes from undetectable, adaptively robust, zero-bit watermarking schemes (and prove that the undetectable zero-bit scheme of Christ, Gunn, and Zamir (2024) is adaptively robust). Importantly, our scheme provides both zero-bit and multi-user assurances at the same time. It detects shorter snippets just as well as the original scheme, and traces longer excerpts to individuals. The main technical component is a construction of message-embedding watermarks from zero-bit watermarks. Ours is the first generic reduction between watermarking schemes for language models. A challenge for such reductions is the lack of a unified abstraction for robustness -- that marked text is detectable even after edits. We introduce a new unifying abstraction called AEB-robustness. AEB-robustness provides that the watermark is detectable whenever the edited text "approximates enough blocks" of model-generated output.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Administrative Law's Fourth Settlement: AI and the Scrutable State

    cs.CY 2026-02 conditional novelty 6.0 of 10

    The recent Supreme Court retrenchment in administrative law is best read as a scrutable-size reaction to government's opacity, and AI plus new audit-based doctrines could restore capability and accountability at once.

  2. From Forensics to Ecosystems: Rethinking Watermarks for Generative AI Oversight

    cs.CY 2026-08 conditional novelty 5.0 of 10

    Watermarks should be repurposed from forensic identification of individual AI outputs to ecosystem-level measurement of aggregate synthetic content saturation.

Pith tools