REVIEW 7 cited by
SoK: Watermarking for AI-Generated Content
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As the outputs of generative AI (GenAI) techniques improve in quality, it becomes increasingly challenging to distinguish them from human-created content. Watermarking schemes are a promising approach to address the problem of distinguishing between AI and human-generated content. These schemes embed hidden signals within AI-generated content to enable reliable detection. While watermarking is not a silver bullet for addressing all risks associated with GenAI, it can play a crucial role in enhancing AI safety and trustworthiness by combating misinformation and deception. This paper presents a comprehensive overview of watermarking techniques for GenAI, beginning with the need for watermarking from historical and regulatory perspectives. We formalize the definitions and desired properties of watermarking schemes and examine the key objectives and threat models for existing approaches. Practical evaluation strategies are also explored, providing insights into the development of robust watermarking techniques capable of resisting various attacks. Additionally, we review recent representative works, highlight open challenges, and discuss potential directions for this emerging field. By offering a thorough understanding of watermarking in GenAI, this work aims to guide researchers in advancing watermarking methods and applications, and support policymakers in addressing the broader implications of GenAI.
Forward citations
Cited by 7 Pith papers
-
Adaptively Robust LLM Monitoring via Activation Watermarking
Activation Watermarking embeds a secret keyed direction in an LLM's hidden states so policy-violating responses can be detected by a cosine test, cutting adaptive-jailbreak evasion relative to guard models.
-
Pay for The Second-Best Service: A Game-Theoretic Approach Against Dishonest LLM Providers
A delegation mechanism makes near-truthful behavior approximately dominant for LLM API providers, and a matching impossibility result caps user utility at the second-best honest service.
-
Optimizing Token Choice for Code Watermarking: An RL Approach
An RL-trained policy adaptively biases token choices to watermark LLM-generated code while preserving executable behavior.
-
Watermark in the Classroom: A Conformal Framework for Adaptive AI Usage Detection
Standard, hierarchical, and weighted conformal prediction applied to LLM watermark scores can control false-positive rates when detecting guideline-violating AI edits in simulated classroom essays.
-
Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice
Obfuscation reduces detection of N-gram watermarked code to random guessing, and the authors prove this is unavoidable under a distribution consistency assumption.
-
A Crack in the Bark: Leveraging Public Knowledge to Remove Tree-Ring Watermarks
VAE-recovered latent surrogates make Tree-Ring watermarks removable: ROC-AUC drops from 0.993 to 0.153 with little image quality loss.
-
First-Place Solution to NeurIPS 2024 Invisible Watermark Removal Challenge
A competition-winning pipeline removes 95.7% of StegaStamp and TreeRing watermarks on the NeurIPS 2024 benchmark by combining VAE fine-tuning, diffusion purification, and translation tricks.
Discussion (0). Continue with ORCID to comment.