REVIEW 12 cited by
Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Current image watermarking methods are vulnerable to advanced image editing techniques enabled by large-scale text-to-image models. These models can distort embedded watermarks during editing, posing significant challenges to copyright protection. In this work, we introduce W-Bench, the first comprehensive benchmark designed to evaluate the robustness of watermarking methods against a wide range of image editing techniques, including image regeneration, global editing, local editing, and image-to-video generation. Through extensive evaluations of eleven representative watermarking methods against prevalent editing techniques, we demonstrate that most methods fail to detect watermarks after such edits. To address this limitation, we propose VINE, a watermarking method that significantly enhances robustness against various image editing techniques while maintaining high image quality. Our approach involves two key innovations: (1) we analyze the frequency characteristics of image editing and identify that blurring distortions exhibit similar frequency properties, which allows us to use them as surrogate attacks during training to bolster watermark robustness; (2) we leverage a large-scale pretrained diffusion model SDXL-Turbo, adapting it for the watermarking task to achieve more imperceptible and robust watermark embedding. Experimental results show that our method achieves outstanding watermarking performance under various image editing techniques, outperforming existing methods in both image quality and robustness. Code is available at https://github.com/Shilin-LU/VINE.
Forward citations
Cited by 12 Pith papers
-
Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking
Major watermarking benchmarks omit cross-lingual, cultural, and demographic reporting, creating a pluralistic evaluation gap that current governance mandates ignore.
-
LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts
A latent-cascaded video generation framework with dual frequency-split experts reports state-of-the-art 2K/4K video generation on VBench, FIDpatch, and human preference.
-
LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
I2VWM uses video-like training distortions and optical-flow frame alignment to keep image watermarks decodable in AI-generated videos made from that image.
-
Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
SSTrack trains a Vision Transformer tracker without frame-wise box labels by combining forward global search, backward local association, and instance contrastive learning, and reports state-of-the-art self-supervised...
-
Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
IP-FVR restores degraded face videos with consistent identity by conditioning a video diffusion model on a reference photo of the same person.
-
Signals of Provenance: Practices & Challenges of Navigating Indicators in AI-Generated Media for Sighted and Blind Individuals
Both sighted and blind/low-vision users frequently overlook platform AI labels and rely on titles, comments, and other content cues, with blind users further hindered by inaccessible label design.
-
CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration
CineVision integrates scriptwriting with real-time visual pre-visualization, and a small lab study suggests it reduces workload and improves director-cinematographer mutual understanding compared with an AI image tool...
-
CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation
A text-guided SAM2 variant with cross-modal attention, semantic prompt generation, and a similarity-sorted memory bank achieves top Dice and surface scores on seven public multi-organ CT datasets.
-
Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
FG-PAN improves zero-shot brain tumor subtype classification by aligning refined visual patch features with LLM-generated fine-grained text prototypes.
-
FADE: Adversarial Concept Erasure in Flow Models
FADE combines adversarial training with trajectory preservation to erase concepts from diffusion models, reporting state-of-the-art erasure on Stable Diffusion benchmarks, but the evidence is incomplete and the theore...
-
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
A YOLO-family detector named Butter claims state-of-the-art efficiency on KITTI, BDD100K, and Cityscapes, but the paper's loss equations, parameter counts, and baseline comparisons contain contradictions that undermin...
-
TRACE: Trajectory-Constrained Concept Erasure in Diffusion Models
TRACE combines a closed-form cross-attention nullification with a late-timestep fine-tuning loss to erase concepts from diffusion models, claiming better erasure and fidelity than published baselines.
Discussion (0). Sign in to comment.