{"id":"5f49fc8e-925e-40ba-a093-9c8c580f4fd3","arxiv_id":"2412.03283","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Using a proxy diffusion model and a single watermarked reference image, an attacker can imprint or erase Tree-Ring and Gaussian Shading watermarks on arbitrary images.","lead":"This paper shows that two popular 'semantic' watermarks used to mark AI-generated images can be forged or removed by an attacker who has only one watermarked example image and an unrelated public image-generation model. The result matters because watermarking is a leading proposal for detecting and attributing AI content, and the attacks work even across different model architectures.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing assumption is cross-model latent transfer (Eq. 4); it is empirically supported for tested pairs but remains unexplained and is weak for Tree-Ring on FLUX.1 (23% imprint, 35% reprompt+), so the 'fundamental/unrelated' claim is stronger than the evidence.","rationale":"The paper is honest, well-executed, and releases code, and the main empirical results are substantial: both attacks work across multiple target models with only one reference image, and the baseline comparison correctly identifies that prior Gaussian Shading attacks relied on key/nonce reuse. The reader's conditional verdict is appropriate. My stress-test pass converges on the same weakest assumption: the cross-model transfer of latent-space optimization is the load-bearing step, and it is demonstrated empirically rather than explained. The FLUX.1/Tree-Ring numbers are the clearest evidence that the property is not universal: 23% imprint detection and 35% enhanced reprompting are far from 'almost perfect', and the paper's own Sec. F admits that autoencoder similarity does not fully explain transfer. This does not invalidate the central claim for the tested configurations, but it does mean the 'fundamental vulnerability of semantic watermarks' framing is stronger than what the worst-case evidence supports. A revision should add confidence intervals for detection rates, discuss the transfer mechanism more carefully, and scope the claims to the models and watermark variants actually tested. No rejection-level concern was identified; the conditional accept with these caveats is the right verdict.","tokens_in":60053,"tokens_out":9714,"duration_ms":103720,"concrete_test":"Run the Imprint-Forgery and Reprompting attacks with a deliberately maximally dissimilar proxy/target pair, e.g., SD2.1 (4-channel VAE, UNet) against a target LDM with a 16-channel VAE and a DiT backbone that shares no training ancestry with SD2.1, and measure Tree-Ring TPR@1%FPR at 150 steps. If the detection rate does not exceed chance (or does not increase monotonically with the proxy-space L2 reduction), then Eq. (4) transfer is not a general property and the 'fundamental vulnerability' claim must be scoped to Gaussian-Shading-like signals or to proxy-target pairs with sufficiently aligned VAE geometry. Additionally, report confidence intervals for the FLUX.1 Tree-Ring numbers to determine whether the 23% figure is statistically distinguishable from an ineffective attack.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a single watermarked reference image and an arbitrary proxy model suffice to forge or remove semantic watermarks 'even with different latent spaces and architectures'. The mechanism that carries the attack across models is the implicit assumption behind Eq. (4): reducing Euclidean distance between latents in the proxy model's space also reduces the watermark-relevant distance in the target model's space. This is not a mathematical guarantee; it is an empirical transfer property. The paper's own transferability analysis (Sec. 4.5, Sec. F, Fig. 20) shows that autoencoder functional similarity correlates with transfer but does not explain it, and the Mitsua/Common Canvas case is explicitly flagged as a counterexample to a purely autoencoder-based explanation. Moreover, the strongest cross-architecture case, Tree-Ring against FLUX.1, obtains only 0.23 detection at 150 imprinting steps and 0.35 with enhanced reprompting (Tabs. 1 and 3), while Gaussian Shading on the same target reaches 0.95/0.88. Because Tree-Ring and Gaussian Shading differ in signal strength and embedding structure (zero-bit pattern vs. multi-bit signed latents), the claim that 'semantic watermarks' as a class are fundamentally vulnerable depends on a transfer property that has only been demonstrated for a handful of model pairs and fails to be strong in the most dissimilar tested pair. If cross-model transfer is not a general property of LDMs, the attack success rates drop sharply for deployment configurations outside the tested zoo, and the abstract's 'even with different latent spaces and architectures' overstates the evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two black-box forgery attacks against semantic watermarks (Tree-Ring and Gaussian Shading) in latent diffusion models, plus a removal variant. Using a proxy model (by default SD2.1), the attacker first inverts a single watermarked reference image to obtain its estimated initial latent noise. The Imprint-Forgery attack then optimizes Eq. (4) to bring the inverted latent of a clean cover image closer to that watermarked latent, while the Imprint-Removal attack targets the negated latent via Eq. (5). The Reprompting attack instead regenerates a new image from the extracted latent with a different prompt, optionally probing multiple prompts and Gaussian Shading bin resamplings. The evaluation covers four target models (SD2.1-Anime, SDXL, PixArt-Σ, FLUX.1), a 7×7 transferability matrix over seven proxy/target models, ablations over samplers and inversion steps, a threshold-defense analysis, and comparisons with averaging, regeneration, adversarial embedding, and surrogate baselines.","tokens_in":60248,"tokens_out":4730,"duration_ms":48474,"significance":"The paper makes a timely and practically relevant contribution to the security evaluation of semantic watermarks. Its main experiments are carefully set up: watermark detection is judged by the original verifiers, Gaussian Shading is deployed with fresh keys and nonces per image (avoiding a known deployment pitfall), the code is released, and the authors are transparent about cases where the attack is weaker, notably Tree-Ring against FLUX.1. The baseline comparison clarifies that previously proposed attacks often fail when Gaussian Shading is implemented correctly. If the underlying cross-model latent transfer is a general property, the single-reference, black-box attack would be an important negative result for current semantic watermarking. However, the evidence for generality is the paper's main vulnerability: the transfer mechanism is only empirically demonstrated on a limited set of model pairs and is not fully explained, so the abstract's claim that unrelated models with different architectures suffice is stronger than the data shown.","major_comments":[{"comment":"","section":"Abstract; Sec. 4.2, Table 1; Sec. 4.4, Table 3"},{"comment":"","section":"Sec. 4.5; Sec. F, Fig. 20"},{"comment":"","section":"Sec. 4.6, Fig. 9 and Sec. G, Fig. 21"}],"minor_comments":[{"comment":"","section":"Sec. 4.1 / Table 1 footnote"},{"comment":"","section":"Sec. 3.1, Eq. (4)"},{"comment":"","section":"Sec. 4.4, Reprompt+"},{"comment":"","section":"Sec. F, Fig. 20"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and well-executed, but the gap between the strong 'fundamental vulnerability' framing and the concrete Tree-Ring/FLUX.1 numbers is the main substantive issue. A major revision that either narrows the claims or strengthens the cross-architecture evidence would make the contribution solid. The unexplained transfer mechanism is the key risk for the paper's long-term impact."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Single-reference black-box forgery and removal works against both Tree-Ring and Gaussian Shading across several target models, including cross-architecture transfer from a UNet proxy to DiT targets. This is a real result, not an artifact: the evaluation is careful, the baselines are handled fairly, and the code is public on GitHub.\\n\\nWhat is genuinely new: prior forgery attacks needed many watermarked images, white-noise image requests, or access to the target model. Here a single watermarked image plus an unrelated public proxy (SD2.1) suffices. The paper also corrects an important confound in prior work: the averaging and surrogate attacks against Gaussian Shading only worked because keys and nonces were reused; under correct deployment those attacks fail, while the new attack still works.\\n\\nThe paper does its homework. Four target models, a 7x7 transferability matrix, ablations over samplers and inversion steps, and a full baseline comparison. The authors are transparent about the Gaussian Shading deployment ambiguity, even citing their own related paper on it, which is legitimate. Attack success is judged by the original schemes' external verifiers, not by anything fitted in this paper.\\n\\nThe soft spots are real but not fatal. The load-bearing assumption is that reducing Euclidean distance between latents in the proxy space also transfers to the target's verification. This is empirically supported for the tested pairs but not explained; the paper's own autoencoder-similarity analysis shows the correlation is incomplete (Mitsua vs Common Canvas). And the strongest cross-architecture case is the weakest: Tree-Ring against FLUX.1 forges only 23% at 150 imprinting steps and 35% with reprompt+. So 'fundamental vulnerability' and 'even with different latent spaces and architectures' is stronger than the worst-case numbers support. The paper acknowledges this, which helps, but the abstract framing is a notch beyond the evidence.\\n\\nOne minor presentation issue: the body tables don't show variance, though the supplementary does. That's easy to fix.\\n\\nThis paper is for anyone building or auditing watermarking schemes for diffusion models. It deserves a serious referee. My recommendation: accept with revisions that soften the universality claim and, if possible, probe the transfer mechanism further. The central result holds.","headline":"Single-reference black-box forgery and removal against Tree-Ring and Gaussian Shading is a real, carefully evaluated result that deserves serious peer review; only the 'fundamental vulnerability' framing slightly overstates the weak Tree-Ring/FLUX.1 case.","tokens_in":60881,"tokens_out":2538,"would_cite":true,"duration_ms":25863,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single watermarked image and an unrelated diffusion model are enough to forge semantic watermarks.","keywords":["semantic watermarking","diffusion models","watermark forgery","watermark removal","Tree-Ring","Gaussian Shading","black-box attacks","DDIM inversion"],"falsifier":"Run the Imprint-Forgery attack over a grid of proxy–target pairs chosen to minimize latent compatibility, for instance a 16-channel DiT-based autoencoder as target with a 4-channel UNet-based proxy, or pairs selected for near-zero functional cosine similarity between their latents. If detection rates on Gaussian Shading at FPR $10^{-6}$ stay at chance for all such pairs even beyond 150 optimization steps, the claim that unrelated models with different latent spaces suffice for forgery would be falsified; if even one such pair transfers, the claim survives at its weakest point, which the 23% Tree-Ring result against FLUX.1 already marks.","tokens_in":59745,"feed_emoji":"🖼️","tokens_out":20213,"duration_ms":162655,"temperature":0.7,"pith_summary":"Tree-Ring and Gaussian Shading protect AI-generated images by embedding a watermark into the initial noise of a diffusion model's generation, recoverable only by inverting that model, with the security of the scheme resting on the target model staying secret. The paper claims this secrecy is unnecessary to break. An attacker armed with one watermarked reference image and any diffusion model of their own — even one with a different architecture and an independently trained autoencoder — can imprint the watermark onto arbitrary real images, strip it from watermarked images, or generate new images that verify as watermarked and are attributed to the reference image's user. If the claim holds, watermark-based detection and attribution of AI content can be spoofed without any access to the protected model, and a single publicly posted generated image becomes enough to compromise the system.","feed_headline":"Forge or erase diffusion watermarks with one image and a proxy model","feed_subtitle":"Both forgery and removal work against Tree-Ring and Gaussian Shading with no access to the target model.","key_machinery":"The central object is the inverse DDIM sampler $\\mathcal{I}_{0\\to T}(z_0; u)$, which walks a latent back along the denoising trajectory of a model $u$ to recover the initial noise in which the watermark lives. All attacks run inside an attacker-chosen proxy model $\\Theta_A = (E_A, u_A, D_A)$; the target model is never queried. The Imprint-Forgery attack is carried by the loss $$\\mathcal{L}_{\\text{forgery}}(\\delta) = \\left| \\mathcal{I}_{0\\to T}(\\tilde{z}$_0^{{(c)}}$ + \\delta; u_A) - \\tilde{z}$_T^{{(w)}}$ \\right|^2,$$ minimized over a perturbation $\\delta$ of the cover image's latent with gradient checkpointing to backpropagate through the whole inversion, while a mask preserves sensitive regions such as faces and text. The mechanism that makes the attack work is transferability: a distance reduction in the proxy's latent space appears as a distance reduction in the target's latent space as well. The paper measures this across seven models and finds a correlation with the functional similarity of the models' autoencoders, but notes the correlation does not fully explain the transfer, so transferability itself, rather than any single shared component, is the load-bearing phenomenon. For Reprompting, the recovered noise $\\tilde{z}_T^{(w)}$ is the only artifact reused.","core_discovery":"On the paper's own terms, the discovery is that watermark forgery and removal reduce to an optimization problem in the attacker's own latent space, not the target's. Given a watermarked image, the attacker inverts it with a proxy model to recover the watermarked initial noise $\\tilde{z}_T^{(w)}$; the Imprint-Forgery attack then minimizes, by gradient descent through the proxy's inverse DDIM sampler, the Euclidean distance between $\\tilde{z}_T^{(w)}$ and the inverted noise of a clean cover image, pulling the cover image's latent toward the watermark. The reduced distance transfers to the target model's latent space, so the target verifier detects the watermark in the modified cover image, and the same construction with the target noise negated removes it instead. The Reprompting attack feeds the recovered $\\tilde{z}_T^{(w)}$ back into the proxy generator with a new prompt, resampling values inside the recovered sign bins for Gaussian Shading. The authors report near-perfect detection and attribution for Gaussian Shading on all four target models and, for Tree-Ring, above 84% on three of four; the hardest case is FLUX.1, with its different latent geometry, where Tree-Ring detection reaches only 23% at 150 steps, and the authors themselves note a corrected data-processing error in earlier PSNR figures that leaves the detection-rate claims unaffected.","pith_inferences":["Because the attacks never touch the target model, the paper's results imply that inversion-based watermark schemes whose security rests on the model being secret have already lost that basis; a scheme that wants to survive this attack family likely must bind verification to a target-specific transformation, such as a decoder-dependent re-encoding, that a proxy's latent geometry cannot replicate.","The imperfect correlation between attack success and autoencoder functional similarity suggests a practical pre-deployment audit: compute latent cosine similarity between a candidate watermark host and publicly available models of several families; high similarity would predict vulnerability, while the paper's anomalous pairs, where dissimilar autoencoders still transfer, mark where a second predi","Operationally, the one-image requirement means any watermarked image that becomes public, such as a social-media post or shared screenshot, is a sufficient oracle for forgery, so an operator should assume the watermark is compromised at the moment of publication rather than when misuse appears."],"forward_implications":["Detection stops separating real from generated: with Imprint-Forgery, an arbitrary clean image becomes verified as watermarked, so a provider's \"is this AI-generated?\" flag can be raised against images the model never produced.","Attribution stops being trustworthy: because the forged latent recovers the reference image's message bits, the service attributes the forged or reprompted content to the innocent user who posted the single reference image.","Removal uses the same machinery: negating the recovered watermark noise erases both detection and attribution of genuinely generated content, with Gaussian Shading detection falling to zero after 100 steps on all tested targets.","Threshold tightening is not a defense: forged images and legitimately watermarked images under common perturbations such as JPEG, salt-and-pepper noise, and rotation produce overlapping p-values and bit accuracies, so no stricter threshold separates them.","The single-reference requirement changes the threat model: earlier baselines needed thousands of watermarked images or knowledge of the target model, while these attacks need one public image and an off-the-shelf proxy model."],"supporting_citations":[{"why":"Defines Tree-Ring, the zero-bit semantic watermark whose circular ring pattern in the initial latent's frequency domain is the forgery target in the detection scenario.","marker":"[41]"},{"why":"Defines Gaussian Shading, the multi-bit watermark whose encrypted message selects the sign bins of the initial latent; it supplies the detection and attribution targets the attacks must satisfy.","marker":"[45]"},{"why":"Supplies the inverse DDIM sampler used to recover the initial latent noise of both the watermarked reference image and the cover images in the proxy model.","marker":"[27]"},{"why":"Provides the averaging-attack baseline that the paper's attacks are compared against; it requires thousands of reference images and fails on Gaussian Shading, which highlights the single-image contribution.","marker":"[44]"},{"why":"Provides the surrogate-attack baseline, which needs training on many watermarked images and fails against a secure Gaussian Shading instantiation; the paper's training-free attacks are defined against these baselines.","marker":"[32]"},{"why":"Supplies Stable Diffusion 2.1, the default proxy model used in all main experiments and the reference point for the transferability grid.","marker":"[30]"}],"fun_headline_variants":["One image, proxy model: forge or erase semantic watermarks","Diffusion watermarks broken with a single reference image","Semantic watermarks fall to black-box forgery attacks","Tree-Rings and Gaussian Shading cracked by proxy models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The attacks assume that pulling a forged image's latent noise closer to the watermark's noise inside the attacker's own model also pulls it closer inside the target model's latent space, and the paper's own evidence shows this transfer is real but incomplete, since against FLUX.1, a model with an unfamiliar latent geometry, Tree-Ring forgery succeeds in only 23% of cases at 150 steps.","fun_headline_variants_meta":{"raw":{"variants":["One image, proxy model: forge or erase semantic watermarks","Diffusion watermarks broken with a single reference image","Semantic watermarks fall to black-box forgery attacks","Tree-Rings and Gaussian Shading cracked by proxy models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000338,"raw_usage":{"total_tokens":1911,"prompt_tokens":1033,"completion_tokens":878,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":811}},"tokens_in":649,"tokens_out":878,"duration_ms":5994,"temperature":1.0,"reasoning_tokens":811,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:33:24.504656+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Imprint-Forgery attack over a grid of proxy–target pairs chosen to minimize latent compatibility, for instance a 16-channel DiT-based autoencoder as target with a 4-channel UNet-based proxy, or pairs selected for near-zero functional cosine similarity between their latents. If detection rates on Gaussian Shading at FPR $10^{-6}$ stay at chance for all such pairs even beyond 150 optimization steps, the claim that unrelated models with different latent spaces suffice for forgery would be falsified; if even one such pair transfers, the claim survives at its weakest point, which the 23% Tree-Ring result against FLUX.1 already marks.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines Tree-Ring, the zero-bit semantic watermark whose circular ring pattern in the initial latent's frequency domain is the forgery target in the detection scenario."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines Gaussian Shading, the multi-bit watermark whose encrypted message selects the sign bins of the initial latent; it supplies the detection and attribution targets the attacks must satisfy."},{"cited_title":"Mokady, A","cited_arxiv_id":null,"evidence_quote":"Supplies the inverse DDIM sampler used to recover the initial latent noise of both the watermarked reference image and the cover images in the proxy model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the averaging-attack baseline that the paper's attacks are compared against; it requires thousands of reference images and fails on Gaussian Shading, which highlights the single-image contribution."},{"cited_title":"Saberi, V","cited_arxiv_id":null,"evidence_quote":"Provides the surrogate-attack baseline, which needs training on many watermarked images and fails against a secure Gaussian Shading instantiation; the paper's training-free attacks are defined against these baselines."},{"cited_title":"Rombach, A","cited_arxiv_id":null,"evidence_quote":"Supplies Stable Diffusion 2.1, the default proxy model used in all main experiments and the reference point for the transferability grid."}],"review_version":1}