{"id":"4c3f3ed7-1815-4aaf-bd7d-afa820e3f3d1","arxiv_id":"2608.10091","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A decoder-aware training objective makes signpost watermarks coexist with four image and three video watermarking systems, and video watermark coexistence is demonstrated for the first time.","lead":"This paper trains a lightweight 'signpost' watermark that is optimized to survive alongside other watermarks in images and video, so a media asset can carry a routing signal pointing to its provenance watermark. It shows that video watermarks can coexist for the first time, and that training with a decoder-aware objective improves coexistence and preserves visual quality.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline coexistence gain is not attributed to the proposed loss: the TrustMark-P result is measured with the signpost as WM2, while Eq. (1) and Table 4 only optimize/evaluate the signpost-first direction, leaving architecture and JND as confounds.","rationale":"The reader's stated weakest assumption (access to frozen encoders) is a real deployment limitation, but it does not threaten the paper's internal claim that the objective works when such access exists. The concern above is more load-bearing because it targets the paper's strongest evidence for that internal claim. In good faith, the method may well work; the paper's Table 4 demonstrates that the coexistence loss materially improves signpost-as-WM1 robustness (e.g., TrustMark-P overlay: 0.841 to 0.960) and solo signpost accuracy (0.957 to 0.977). The issue is specifically that the headline cross-system result, the one cited in the abstract-level claim, is not covered by that ablation. A noCoexist signpost with the same architecture/JND has solo PSNR 53.0 dB, even slightly higher than the full model's 52.3 dB, so it is plausible that it already interferes less than ZOETROPE. The proposed test directly separates the coexistence loss contribution from the architecture/JND contribution in the WM2 direction. Pending that test, the verdict should remain conditional: the central claim is plausible and partially supported, but its strongest quantitative evidence needs an additional control. I do not elevate to REJECT because the signpost-first robustness improvements are direct and the WM2-direction gain may well survive the control. I differ from the reader's identified weakest assumption: the more immediate risk is not access to proprietary encoders but whether the headline effect is attributable to the proposed objective at all.","tokens_in":12213,"tokens_out":10955,"duration_ms":104367,"concrete_test":"Take the noCoexist checkpoint from Table 4 (same architecture and JND, coexistence loss disabled), apply it as WM2 over TrustMark-P as WM1, and measure TrustMark-P's noise-average bit accuracy under the same 17-augmentation protocol. If this accuracy is close to Ours's 0.929, the proposed loss is not the cause of the headline gain; if it falls near ZOETROPE's 0.687, the loss is causal. This single comparison resolves whether the confound lands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest specific result — TrustMark-P noise-average bit accuracy rising from 0.687 under ZOETROPE to 0.929 under the proposed signpost (Table 3, TM-P row, Ours column) — is measured with TrustMark-P as WM1 and the signpost as WM2. The coexistence loss in Eq. (1), however, trains only the opposite direction: after the signpost is applied first, a frozen secondary encoder is overlaid, and the signpost decoder f_theta is supervised to recover s from the doubly-watermarked image (Section 3, 'Frozen secondary encoders'). No loss term involves any secondary decoder, and no ablation reports secondary-decoder accuracy with the signpost as WM2. Table 4's noCoexist versus Ours comparison measures signpost bit accuracy in the signpost-first direction only. The ZOETROPE baseline also differs in architecture, training data regime, and JND guidance, with a much lower solo PSNR (45.8 dB vs 52.4 dB), so the larger interference it causes could be due to a denser, less perceptually tuned residual rather than to the absence of the coexistence objective. Consequently, the headline improvement cannot currently be attributed to 'decoder-aware training;' it may be an emergent property of a higher-quality sparse residual. This is a gap in the internal evidence for the central claim, distinct from the access-to-encoders deployment limitation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 'signpost watermarking': a lightweight watermark trained to coexist with a set of independently developed image and video watermarking systems, acting as a routing signal to a registry of such systems. The method uses a UNet encoder and ResNet-50 decoder, and its training objective (Eq. 1) adds a coexistence term that applies frozen secondary encoders on top of the signpost stego image and requires the signpost decoder to still recover the payload. The paper reports that (i) emergent coexistence extends to video watermarking, and (ii) decoder-aware training substantially improves coexistence, with the largest gains for TrustMark-based systems. Experiments cover four image watermarking systems and three video systems, with PSNR/VMAF quality metrics and bit-accuracy under clean and augmented conditions, plus single-decoder and leave-one-out ablations.","tokens_in":12524,"tokens_out":5969,"duration_ms":57094,"significance":"If the causal claim holds, this is a genuinely useful contribution: it provides the first systematic study of video watermark coexistence and offers a concrete architecture and training objective for an interoperability layer on top of proprietary watermarking systems. The ablation design is thoughtful: the noCoexist-versus-Ours comparison in Table 4 cleanly isolates the effect of the coexistence loss in the signpost-first direction, and the paper correctly identifies the TrustMark family as the most interference-sensitive watermarkers. The residual visualizations and PSNR analysis are also informative. However, the paper's headline improvement in secondary-watermark decoding accuracy is not yet attributable to the proposed objective, because the key comparison (Table 3, TrustMark-P row: 0.687 under ZOETROPE vs 0.929 under Ours) is measured with the signpost as the second watermark, whereas the training objective and the ablation table only evaluate the signpost-first direction.","major_comments":[{"comment":"The central claim that decoder-aware training improves coexistence is not yet supported by the evidence as presented. The strongest specific result, TrustMark-P noise-average bit accuracy rising from 0.687 under ZOETROPE to 0.929 under the proposed signpost (Table 3, TM-P row, Ours column), is measured with TrustMark-P as WM1 and the signpost as WM2. However, the coexistence loss in Eq. (1) is evaluated only in the opposite direction: the signpost is applied first, a frozen secondary encoder is overlaid, and the signpost decoder is trained to recover s from the doubly-watermarked image. No loss term involves any secondary decoder, and the ablation in Table 4 reports signpost bit accuracy only in the signpost-first direction. Because the noCoexist baseline is never evaluated in the WM2 direction, the improvement over ZOETROPE confounds the proposed objective with architectural and quality differences (sparse residual, JND guidance, solo PSNR 52.4 dB vs 45.8 dB). To support the attribution, please add an experiment that trains a noCoexist signpost with the same architecture, data, JND guidance, and comparable PSNR, and measures secondary-watermark bit accuracy when that signpost is overlaid as WM2. Without such a control, the headline claim in the abstract that watermarks 'can be trained with a decoder-aware objective to improve coexistence' should be narrowed to the signpost-first direction demonstrated in Table 4.","section":"§4.3, Table 3; §3, Eq. (1)"},{"comment":"No error bars, confidence intervals, or significance tests are reported anywhere in the experimental section. The image evaluation uses N=1000 held-out images and the video evaluation uses only 155 test videos, and many of the ablation differences in Table 4 are small (e.g., video rows differing by 0.002-0.004). Without measures of variance it is difficult to assess whether the ranking of ablated variants, or the claimed improvements such as the Table 4 image 'Ours' solo accuracy of 0.977 vs 0.957 for noCoexist, are reliable. Please report standard errors or bootstrap confidence intervals for the key comparisons, especially the headline 0.929 vs 0.687 in Table 3 and the Ours-vs-noCoexist deltas in Table 4.","section":"§4.1, Tables 2-4"},{"comment":"The method requires access to the frozen encoder weights of every watermarking system the signpost must coexist with, as stated in Section 3. The introduction, however, emphasizes that watermarking systems are 'often kept secret to reduce attacks.' The paper does not address this tension: for proprietary or secret encoders, the proposed optimization cannot be applied, and the 'practical path' of a registry of systems implicitly requires every registered system to expose its encoder. Please state this limitation explicitly and discuss under what trust assumptions the signpost remains practical, or temper the interoperability claim accordingly.","section":"§3, 'Frozen secondary encoders'; §1"}],"minor_comments":[{"comment":"There are several typographical artifacts: 'T able 1', 'T able 2', 'T able 3', and 'T able 4' have an errant space; 'asignpostwatermarkcanbeextendedthrough' is missing spaces in the introduction; 'The question follows:can' is missing a space after the colon; and 'for example a integer' should read 'for example an integer.'","section":"§1, §2, general"},{"comment":"The sentence 'These failures are not a concern, as two competing signposts sharing a single image is not a design requirement and two signposts are never deployed together' dismisses a case that may occur if multiple independent signpost providers operate; please either justify this design assumption or acknowledge it as a limitation.","section":"§4.2"},{"comment":"The evaluation protocol states that 'the first watermark (hereafter, WM1) which is applied to a clean cover image' should be 'the first watermark (hereafter, WM1), which is applied'; please fix the missing comma and consider explaining why the signpost-first direction is the primary training scenario while the WM2 direction is the one highlighted in the abstract.","section":"§4.1"},{"comment":"The JND-guided loss in Eq. (2) is motivated by perceptual masking, but no comparison is provided to a version without JND guidance in the ablation table; since JND is a confounder in the headline WM2-direction comparison, adding such an ablation would strengthen the isolation of the coexistence objective.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the community, but the central claim needs a cleaner causal test. The key issue for the editor is whether the authors can supply the missing WM2-direction control experiment; if they can, the paper would be significantly stronger. I would also gently encourage the authors to release code or data for reproducibility, since the training set is proprietary and the method is not otherwise accessible to independent verification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper asks a timely question: can a lightweight signpost watermark be trained to coexist cleanly with independently deployed watermarks? It does two genuinely new things. It provides the first empirical study of coexistence among video watermarking systems, and it introduces a decoder-aware training objective that explicitly optimizes the signpost to survive the application of frozen secondary watermarks. The writing is clear, the ablations in the signpost-first direction are internally consistent, and the video result is a real gap in the literature. The paper is worth engaging, but the central attribution claim has a load-bearing gap.\n\nThe coexistence loss in Eq. (1) trains the signpost as WM1: the signpost is applied first, a frozen secondary encoder is overlaid, and the signpost decoder is supervised to recover its payload from the doubly-watermarked image. Table 4 ablate this direction and it works. But the headline result, TrustMark-P rising from 0.687 under ZOETROPE to 0.929 under the proposed signpost (Table 3), is in the opposite direction: there the signpost is WM2, overlaid on top of TrustMark-P. The loss contains no term that measures the secondary watermark's decoder accuracy, and no ablation reports reverse-direction accuracy for noCoexist versus the full model. So the improvement over ZOETROPE is confounded with architecture, JND guidance, training data, and residual sparsity (solo PSNR 45.8 versus 52.4 dB). The paper's own conclusion makes the same claim, but the evidence is not there yet.\n\nThere are also smaller, more ordinary soft spots: no error bars anywhere, a proprietary training set, no released code, and a deployment assumption that the signpost trainer has access to frozen encoders of every system it must coexist with. The paper acknowledges that watermarking systems are often kept secret but does not address the tension beyond a registry suggestion.\n\nNone of this makes the paper a bad one. The SP-first ablation is a solid demonstration that the objective improves signpost robustness under secondary overlays. The video coexistence observations are new and useful. What is missing is a symmetric evaluation: training with the signpost as WM2, or at least reporting noCoexist performance in the reverse direction so the loss can be separated from the quality/architecture confounds. A reviewer should ask for exactly that experiment.\n\nFor a reader working on provenance or watermark interoperability, this paper is worth a serious referee slot. I would send it to review, but I would expect the authors to close the reverse-direction attribution gap before publication.","headline":"The coexistence loss and the headline result face opposite directions, so the paper's flagship gain over ZOETROPE is not yet attributed to the method.","tokens_in":13004,"tokens_out":3349,"would_cite":true,"duration_ms":32755,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that watermark coexistence is a trainable objective: a lightweight signpost watermark can be optimized so that independently built watermarking systems layered on the same image or video keep decoding, rather than relying…","keywords":["signpost watermarking","watermark coexistence","decoder-aware training","image watermarking","video watermarking","content provenance","layered provenance signaling","imperceptible watermarking"],"falsifier":"Run the same signpost training with the coexistence loss disabled ($\\lambda_c = 0$) on identical data, and compare TrustMark-P's noise-averaged bit accuracy when the signpost is applied first and TrustMark-P second; the paper reports roughly $0.84$ for the ablated model and $0.96$ for the full model, so a rerun that fails to reproduce a gap near that size would falsify the claim that the coexistence term is the cause. A complementary decisive test is a watermarking system held out of training: if its overlay accuracy shows no improvement over the no-coexistence signpost, the benefit is memorization of the training set rather than general coexistence.","tokens_in":12014,"feed_emoji":"🖼️","tokens_out":12984,"duration_ms":110404,"temperature":0.7,"pith_summary":"This paper tries to turn the coexistence of visual watermarks from an accidental property into an explicit design goal. It argues that a lightweight 'signpost' watermark can be trained so that when other independently built watermarks are layered on top of the same image or video, both the signpost and the other watermarks keep decoding accurately. The paper first shows that the emergent coexistence previously observed for image watermarks also holds for video watermarks, then introduces a decoder-aware training objective that places frozen copies of the other watermark encoders inside the signpost's loss. The central result is that TrustMark-P, the most interference-sensitive system tested, keeps $0.929$ noise-averaged bit accuracy when the signpost is layered underneath it, versus $0.687$ with the previous signpost baseline ZOETROPE, while the signpost itself reaches $52.4$ dB solo PSNR. If right, this makes layered provenance signaling practical: one standardized signpost can route a decoder to the right provenance system without scanning all of them.","feed_headline":"Trained coexistence lifts TrustMark accuracy to 0.93","feed_subtitle":"A lightweight signpost watermark now lets provenance systems layer on top without sacrificing their payload decode.","key_machinery":"The load-bearing mechanism is the coexistence loss term $\\lambda_c \\sum_k L_{\\mathrm{BCE}}(f_\\theta(\\hat{x}_k), s)$, evaluated after each frozen secondary watermark encoder $k$ is applied on top of the signpost's stego image $\\hat{x}$. This term forces the signpost decoder $f_\\theta$ to read the 32-bit secret $s$ from the doubly-watermarked image, which in turn forces the signpost encoder to place its residual in signal space complementary to each secondary system. It is used with a frozen set of image encoders (PixelSeal, InvisMark, TrustMark-P, MaskWM) and, for video, VideoSeal, InvisMark, TrustMark-Q, and FlowMark, plus a Just Noticeable Difference (JND)-modulated perceptual loss that lets the encoder concentrate energy where distortion is least visible.","core_discovery":"The central claim is that coexistence is trainable. The paper trains a signpost encoder-decoder with a UNet encoder (32-bit secret lifted and upsampled into the cover image) and a ResNet-50 decoder, using a loss that evaluates signpost decoding after each of several frozen secondary watermark encoders has been overlaid on the signpost's output. The coexistence term $\\lambda_c \\sum_k L_{\\mathrm{BCE}}(f_\\theta(\\hat{x}_k), s)$ pushes the signpost's residual into spatial and frequency regions left free by the secondary systems. As a result TrustMark-P's noise-averaged bit accuracy under overlay rises from $0.687$ with ZOETROPE to $0.929$ with the proposed signpost, and the signpost's own solo PSNR is $52.4$ dB; video experiments show the same pattern, with TrustMark-Q improving from at most $0.880$ under other overlays to $0.945$ under the signpost. Ablations with the coexistence loss disabled, with single frozen decoders, and with leave-one-out subsets indicate all frozen decoders contribute and that removing any one lowers overlay accuracy.","pith_inferences":["Because the method needs frozen encoder access at training time, its real-world reach is limited to watermarking systems whose encoders are available; systems kept secret to resist attacks cannot be optimized for, so the registry-based interoperability path works only for open or licensable systems.","The signpost's residual is sparse and spatially localized at $52.4$ dB PSNR; this likely makes it more fragile under adversarial editing or strong compression than dense global-pattern watermarks, a risk not covered by the paper's augmentation suite and worth testing directly.","A held-out generalization test—training on a few open watermarkers then measuring coexistence with an entirely different, never-seen watermarker—would show whether the objective learns general complementarity or memorizes the training systems; the paper's FM-32 result hints at transfer while its ZOETROPE result shows a held-out signpost can still disrupt a sensitive system.","The registration model also has a bootstrap problem: early deployment needs a critical mass of open watermarking systems to train against, so the signpost's value grows with the size of the registry rather than being available from day one."],"forward_implications":["A standardized signpost watermark can act as a routing layer: its payload indexes a public list of provenance watermarking systems, so a decoder checks one signpost instead of testing all $N$ watermark decoders.","Video watermarking systems, previously unexamined for coexistence, display the same emergent compatibility as image systems, and decoder-aware training improves it, so video provenance can also use signpost-based routing.","The objective gives the largest gains for the watermark family that suffers most under naive overlay: TrustMark-P image accuracy rises from $0.687$ to $0.929$ and TrustMark-Q video from at most $0.880$ to $0.945$, with solo quality at $52.4$ dB PSNR.","Removing any one frozen decoder from training lowers overlay accuracy, meaning full-coverage training data, not a single critical pairing, is what produces the best coexistence.","The signpost is not a new watermarking standard but an interoperability layer, so independent systems can join a shared provenance ecosystem without agreeing on one algorithm."],"supporting_citations":[{"why":"Shows independently trained image watermarking systems can already coexist and be ensembled; this is the premise the paper turns into an explicit training objective.","marker":"[23]"},{"why":"Prior image-only signpost watermark trained without the coexistence objective; it is the baseline whose TrustMark-P accuracy of 0.687 the proposed method improves to 0.929.","marker":"[10]"},{"why":"TrustMark supplies the TrustMark-P and TrustMark-Q frozen secondary encoders used in training, the curriculum the signpost training follows, and a quality and robustness baseline.","marker":"[7]"},{"why":"VideoSeal is one of the frozen video watermark encoders used to supervise video signpost training and evaluate coexistence.","marker":"[17]"},{"why":"FlowMark (and its no-coexistence variant FM-32) serves as a frozen video secondary encoder and as a partly held-out signpost in the video coexistence analysis.","marker":"[2]"},{"why":"PixelSeal is a frozen image watermark encoder used in coexistence training and evaluation.","marker":"[26]"},{"why":"InvisMark is a frozen image and video watermark encoder used across both signpost training sets.","marker":"[32]"},{"why":"MaskWM is a frozen image watermark encoder whose structured mask residual studies the effect of a visually dominant secondary system.","marker":"[20]"}],"fun_headline_variants":["Coexistence by design: overlay accuracy 0.93","From serendipity to trainable coexistence: 0.93","Layered provenance: trainable coexistence to 0.93","Trainable coexistence lifts overlay bit accuracy to 0.93","Signpost watermark enables layered provenance at 0.93"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes the signpost trainer has access to frozen copies of every watermarking system it must coexist with at training time; for systems kept secret to resist attacks, that access does not exist, so the central interoperability path fails for those systems.","fun_headline_variants_meta":{"raw":{"variants":["Coexistence by design: overlay accuracy 0.93","From serendipity to trainable coexistence: 0.93","Layered provenance: trainable coexistence to 0.93","Trainable coexistence lifts overlay bit accuracy to 0.93","Signpost watermark enables layered provenance at 0.93"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001451,"raw_usage":{"total_tokens":5821,"prompt_tokens":902,"completion_tokens":4919,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":4832}},"tokens_in":518,"tokens_out":4919,"duration_ms":32898,"temperature":1.0,"reasoning_tokens":4832,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:14:29.271587+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same signpost training with the coexistence loss disabled ($\\lambda_c = 0$) on identical data, and compare TrustMark-P's noise-averaged bit accuracy when the signpost is applied first and TrustMark-P second; the paper reports roughly $0.84$ for the ablated model and $0.96$ for the full model, so a rerun that fails to reproduce a gap near that size would falsify the claim that the coexistence term is the cause. A complementary decisive test is a watermarking system held out of training: if its overlay accuracy shows no improvement over the no-coexistence signpost, the benefit is memorization of the training set rather than general coexistence.","supporting_citations":[{"cited_title":"In: Intl","cited_arxiv_id":null,"evidence_quote":"Shows independently trained image watermarking systems can already coexist and be ensembled; this is the premise the paper turns into an explicit training objective."},{"cited_title":"IEEE Computer Graphics and Applications (IEEE CG&A), in press","cited_arxiv_id":null,"evidence_quote":"Prior image-only signpost watermark trained without the coexistence objective; it is the baseline whose TrustMark-P accuracy of 0.687 the proposed method improves to 0.929."},{"cited_title":"FlowMark: Mask-Guided Video Watermarking","cited_arxiv_id":"2607.05261","evidence_quote":"FlowMark (and its no-coexistence variant FM-32) serves as a frozen video secondary encoder and as a partly held-out signpost in the video coexistence analysis."},{"cited_title":"Advances in Neural Information Processing Systems 38, 146313–146346 (2026)","cited_arxiv_id":null,"evidence_quote":"MaskWM is a frozen image watermark encoder whose structured mask residual studies the effect of a visually dominant secondary system."}],"review_version":1}