REVIEW 3 major objections 5 minor 28 references
Screen-Conditioned Watermarking Against Multi-Screen Collusion Attacks
T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read CoMSMark claims screen-conditioned watermark residuals keep copyright marks readable above 90% under multi-screen collusion while keeping forged marks near chance.
desk verdict CoMSMark names a real new attack and a plausible defense, but 'collusion-resistant' is only shown against the exact mean-residual estimator the loss is trained on. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a screen-ID-conditioned residual encoder: a style-modulation branch changes channel responses and a spatial-fusion branch shapes spatial distributions, turning a screen embedding into per-screen residual patterns derived from the shared watermark bits. A dual-head decoder separates copyright recovery from screen identification. A collusion-suppression loss L_coll suppresses the mean residual across a batch and maximizes prediction entropy on forged samples. The image-agnostic encoding paradigm lets one residual serve all images on a given screen.
What would settle it
Run the same collusion protocol but average only captures from the same screen ID, subtract a clean image of that screen, and check whether the recovered residual drives forged-image watermark accuracy well above 50%; if it does, the collusion-resistance claim is false.
Extended reading notes
Core claim
The central claim is that multi-screen collusion can be defeated by reducing the shared component of watermark residuals across screens, not by making the residuals stronger. CoMSMark conditions residual generation on a discrete screen ID through channel modulation and spatial fusion, so each screen shows a different watermark pattern that still carries the same copyright bit string. A collusion-suppression loss directly minimizes the norm of the mean residual over a batch of screens and pushes decoder predictions on forged samples toward 0.5, so averaging captured copies no longer yields a usable pattern. Because the residual is generated from the watermark and screen ID alone, independent
Load-bearing premise
The defense is exercised against one specific attack family — estimating the shared residual by subtracting unpaired clean images and averaging over screens, then removing or adding it with a scalar strength — and any adversary using a different estimator is outside the tested claim.
Editorial extensions
If this is right
- If CoMSMark works as claimed, copyright verification survives collusive removal attacks even when the attacker uses up to 100 screens.
- Forged watermarked images decode at chance level, so an adversary cannot create plausible counterfeit watermarked copies from estimated residuals.
- Screen-level source attribution is feasible without enlarging the watermark payload, because the screen identity is carried as a separate conditioning channel.
- Large-scale distribution becomes efficient: residuals are computed once per (watermark, screen) pair, not per image.
- The defense extends to practical multi-screen display settings, with reported accuracy above 98% across capture distances, angles, and device combinations.
Reading between the lines
- An untested but natural next attack, going beyond the paper, is averaging only copies from the same screen ID before subtracting clean images; that would not cancel the screen-specific residual and could expose the shared component more directly.
- Because screen IDs are discrete (128 identities), an adversary could enumerate IDs and probe which one matches a captured screen; a larger or continuous screen-attribution space might be more robust.
- The paper's attack model assumes the adversary has access to unwatermarked clean images; if those are unavailable, the estimated residual is weaker, so the suppression loss may be doing less work than the experiments suggest.
- The reported near-50% forged accuracy holds against the trained mean-residual attack; a surrogate decoder trained to invert screen-specific residuals is an untested but plausible bypass.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper identifies a multi-screen collusion threat for screen-shooting watermarking: when the same copyright watermark is distributed across different screens, the shared residual component can be estimated by subtracting unpaired clean images and averaging, then used for watermark removal or forgery. The authors propose CoMSMark, an image-agnostic encoder that conditions the watermark residual on a discrete screen ID via channel modulation and spatial fusion, a dual-head decoder that recovers both the watermark and the screen ID, and a collusion suppression loss that penalizes the mean residual and encourages uncertain predictions on forged samples. Experiments with physical screen-shooting captures report watermark decoding accuracy above 90% under the tested removal attack, forged-watermark accuracy near 50%, reliable screen ID attribution, and competitive robustness across capture distances, angles, and device combinations.
Significance. The multi-screen collusion attack is a realistic and under-explored threat in screen-shooting watermarking, and the paper proposes a novel defense direction that jointly addresses copyright verification, screen-level source attribution, and image-agnostic encoding. The design is technically coherent: screen-conditioned residuals are a natural way to diversify the watermark pattern across screens, and the progressive training schedule is a sensible way to balance competing objectives. The physical capture experiments are a strength, and the reported robustness under varied capture conditions is useful. However, the central 'collusion-resistant' claim is currently supported only against the specific mean-residual attack family that the training losses are explicitly designed to defeat. If the scheme is shown to resist a broader class of adaptive attacks, it would be a meaningful contribution to the field; as it stands, the evidence is too narrow to support the general claim.
major comments (3)
- [Abstract and §Experimentation (Collusion Resistance Evaluation)] The headline claims—'average watermark accuracy above 90% under removal attacks' and 'forged-watermark accuracy near 50%'—are established only against the residual estimator defined in Eq. (2) and the removal/forgery operation in Eq. (7). This is the same estimator family that the training objective directly optimizes against: L_msl (Eq. 15) minimizes the batch mean residual \bar R that the attack estimates, and L_adl (Eq. 17) pushes decoder output on forged samples, built from that same \bar R in Eq. (16), toward 0.5. The reported values therefore demonstrate resistance to the trained mean-residual attack, not to multi-screen collusion in general. The paper needs experiments or analysis against at least one out-of-family adversary, for example same-screen residual estimation, screen clustering, or a surrogate decoder trained to invert residuals.
- [§Proposed Method, Eq. (1) and image-agnostic encoder] Because the residual is generated independently of image content, R_s in Eq. (1) is a fixed function of (W, S) for each screen. An adversary who collects several captures from the same screen ID can average them over different image contents to estimate R_s directly; content and noise will tend to cancel, while the screen-specific residual is preserved. L_msl suppresses only the cross-screen batch mean, so it does not prevent this same-screen estimator. This is a concrete, untested attack that would break the 'multi-screen' framing of the defense. Please add an experiment in which the adversary estimates the residual from same-screen captures and applies Eq. (7), and discuss whether the threat model in Eq. (2) is the only one considered.
- [§Experimentation, Figs. 5-6 and Table 3] All headline accuracy numbers are single-point estimates with no error bars, repeated captures, or multiple random test subsets. This is particularly important for the forgery claim: CoMSMark's forged-watermark ACC ranges from 53% to 56% in content-diverse collusion and from 62.02% to 53.53% in content-aligned collusion. A value of 62% is materially above the 50% chance level, and without confidence intervals the claim that the method 'remains close to random guessing' is not substantiated. Please report means and standard deviations over repeated physical captures or multiple random test selections, and state the number of independent trials.
minor comments (5)
- [Abstract] Typo: 'this multi-screen collusion attacks has not been' should be 'this multi-screen collusion attack has not been'.
- [Table 1] The PSNR column for FPSMark and Ours is formatted as '32.9935.93' and '35.8535.85'; add a separator or align the table entries.
- [§Experimentation, Table 3] The PSNR values under collusive removal and forgery are nearly identical across all methods and appear to be dominated by the attack strength and residual scaling rather than by the watermarking method itself. The discussion of these PSNR values should be qualified accordingly.
- [§Experimentation, Table 4] The table caption and text contain spacing errors ('Angleresults areaveraged' / 'averagedover'). Also clarify whether the angle results are averaged over left/right directions only or also over multiple captures.
- [§Proposed Method, Eq. (11)] The YUV transformation γ is used in L_mse but its definition is only given in words. Please provide the explicit mapping or a reference.
Circularity Check
Collusion-resistance results are trained against the exact attack estimator used in evaluation; removal/forgery ACCs reduce partly to the Lcoll objective.
-
fitted input called prediction
[Proposed Method / Training Strategy (Eqs. 14–18); Collusion Resistance Evaluation (Figs. 5–6 and accompanying text)]
"The mean residual is computed as \bar R = 1/B \sum_{i=1}^B R_i, and suppressed by Lmsl = \|\bar R\|_2^2. ... To simulate collusion forgery, we add \bar R to clean images: \tilde I_o = I_o + \bar R. ... Ladl = 1/BL \|p_col −0.5\|_1. ... The estimated residual is then exploited for watermark removal and forgery: Irem = I_w − α \hat R_K, I_for = I_o + α \hat R_K. ... The resulting residual estimates are averaged to approximate the shared watermark residual, which is then subtracted from the captured watermarked images."
The attack estimator in Eq. 2 averages residuals across screens to approximate \hat R_K ≈ (1/K) ∑ R_s, and Eq. 7 uses that average for removal/forgery. The training loss computes the same batch quantity \bar R (Eq. 14), explicitly minimizes its norm (Eq. 15), and constructs forged samples by adding \bar R to clean images while driving decoder probabilities toward 0.5 (Eqs. 16–17). The evaluation then reports removal and forgery ACC using exactly this mean-residual estimator. Thus the headline results—removal ACC above 90% and forgery ACC near 50%—are measurements of the optimized objectives, not independent out-of-sample predictions. The paper does not test screen-specific residual estimation, surrogate decoders, or adaptive estimators, so the general claim of resisting multi-screen collus
full rationale
The central circularity is that the collusion-suppression training objective and the collusion-resistance evaluation use the same mean-residual estimator. Lmsl suppresses the norm of the batch mean residual (Eq. 14–15), and Ladl explicitly trains the decoder to output 0.5 on forged images built with that mean residual (Eq. 16–17). The reported ACC values under collusive removal and forgery are therefore partly by construction: the defense is optimized to defeat the exact attack it is then tested on. This is a fitted-input-called-prediction pattern rather than a definitional identity, because removal ACC is not literally a loss term, but the attack strength is directly minimized by Lmsl and the forgery target is directly set by Ladl. The paper gives no evidence against alternative estimators, such as averaging only same-screen captures, which would recover screen-specific residuals that Lmsl does not suppress. However, the paper has substantial non-circular content: visual quality (PSNR/SSIM), screen-ID attribution accuracy, and robustness under distances, angles, and device combinations are evaluated against external baselines and do not reduce to the training objective. There is no load-bearing self-citation chain; the image-agnostic encoder is credited to prior work but is not used to justify the security claim. Overall the central collusion-resistance claim is partially circular, but the framework has independent contributions, so a score of 6 is appropriate.
Assumptions & free parameters
free parameters (3)
- Screen ID vocabulary size =
128 identities (0–127)
- Loss weights =
λ_w=10, λ_e=5, λ_s=2, λ_coll=0.5; λ3=λ4=1
- Mini-batch size for collusion suppression =
16
assumptions (4)
- domain assumption Screen-shooting distortion is modeled additively: I_w^s = I_o^s + R^s + N^s (Eq. 1).
- domain assumption The attacker can collect independently sampled clean images and subtract them from watermarked captures (Eq. 2).
- domain assumption Averaging over content-diverse collusion makes the content-related residual tend to zero (Eq. 4).
- domain assumption The differentiable noise layer faithfully represents the real screen-shooting channel.
Cite this review
Pith. "Pith review of Screen-Conditioned Watermarking Against Multi-Screen Collusion Attacks." pith.science (2026). https://pith.science/paper/3PICIFGC
@misc{pith2026260723553,
author = {Pith},
title = {Pith review of: Screen-Conditioned Watermarking Against Multi-Screen Collusion Attacks},
year = {2026},
howpublished = {\url{https://pith.science/paper/3PICIFGC}},
note = {Machine review of arXiv:2607.23553}
}
read the original abstract
Screen-shooting poses a significant threat to confidential information protection. While existing screen-shooting watermarking methods enable copyright verification, the copyrighted images carrying the same copyright watermark across different screens often exhibit highly similar and estimable watermark patterns. These shared patterns can be exploited for watermark removal and forgery, a threat we term the multi-screen collusion attack. To mitigate this threat, we propose CoMSMark, a collusion-resistant image-agnostic watermarking framework for multi-screen shooting, which reduces shared residual components across screens to resist multi-screen collusion attacks. Specifically, we incorporate screen ID through a style modulation mechanism, enabling the encoder to generate screen-specific watermark residuals for reliable source attribution. We further introduce a collusion suppression loss that reduces shared residual components and encourages high-entropy predictions for forged samples, improving resistance to collusion attacks. Finally, to enable efficient large-scale distribution, CoMSMark employs an image-agnostic encoding paradigm that generates watermark residuals independently of image content. Extensive experiments demonstrate that CoMSMark effectively resists both collusion-based watermark removal and forgery. It maintains an average watermark accuracy above 90% under removal attacks while keeping forged-watermark accuracy near 50%. Moreover, CoMSMark achieves competitive robustness under diverse screen-shooting conditions, including varying capture distances and angles.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
IEEE Transactions on Information Forensics and Security , volume=
Screen-shooting resilient watermarking , author=. IEEE Transactions on Information Forensics and Security , volume=. 2018 , publisher=
2018
-
[2]
IEEE Transactions on Circuits and Systems for Video Technology , volume=
Robust high-capacity watermarking over online social network shared images , author=. IEEE Transactions on Circuits and Systems for Video Technology , volume=. 2020 , publisher=
2020
-
[3]
Proceedings of the 29th ACM International Conference on Multimedia , pages=
Mbrs: Enhancing robustness of dnn-based watermarking by mini-batch of real and simulated jpeg compression , author=. Proceedings of the 29th ACM International Conference on Multimedia , pages=
-
[4]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Stegastamp: Invisible hyperlinks in physical photographs , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[5]
Proceedings of the 30th ACM International Conference on Multimedia , pages=
Pimog: An effective screen-shooting noise-layer simulation for deep-learning-based watermarking network , author=. Proceedings of the 30th ACM International Conference on Multimedia , pages=
-
[6]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Light field messaging with deep photographic steganography , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[7]
IEEE Transactions on Circuits and Systems for Video Technology , volume=
Flexible partial screen-shooting watermarking with provable robustness , author=. IEEE Transactions on Circuits and Systems for Video Technology , volume=. 2025 , publisher=
2025
-
[8]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Sim-to-Real: An Unsupervised Noise Layer for Screen-Camera Watermarking Robustness , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Show all 28 references
-
[9]
arXiv preprint arXiv:2406.09026 , year=
Steganalysis on digital watermarking: Is your defense truly impervious? , author=. arXiv preprint arXiv:2406.09026 , year=
-
[10]
IEEE Transactions on Pattern Analysis and Machine Intelligence , year=
Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models , author=. IEEE Transactions on Pattern Analysis and Machine Intelligence , year=
-
[11]
Pattern Recognition , pages=
Secure distribution: Anti-collusion watermarking via Spectral Weight Modulation in latent diffusion models , author=. Pattern Recognition , pages=. 2026 , publisher=
2026
-
[12]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
RoPaSS: Robust Watermarking for Partial Screen-Shooting Scenarios , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[13]
IEEE Transactions on Circuits and Systems for Video Technology , year=
Robust image watermarking with synchronization using template enhanced-extracted network , author=. IEEE Transactions on Circuits and Systems for Video Technology , year=
-
[14]
Proceedings of the AAAI conference on artificial intelligence , volume=
Film: Visual reasoning with a general conditioning layer , author=. Proceedings of the AAAI conference on artificial intelligence , volume=
-
[15]
IEEE Transactions on Information Forensics and Security , volume=
Screen-shooting robust watermark based on style transfer and structural re-parameterization , author=. IEEE Transactions on Information Forensics and Security , volume=. 2025 , publisher=
2025
-
[16]
Proceedings of the European Conference on Computer Vision (ECCV) , pages=
Zhu, Jiren and Kaplan, Russell and Johnson, Justin and Fei-Fei, Li , title =. Proceedings of the European Conference on Computer Vision (ECCV) , pages=
-
[17]
Advances in Neural Information Processing Systems , volume=
Udh: Universal deep hiding for steganography, watermarking, and light field messaging , author=. Advances in Neural Information Processing Systems , volume=
-
[18]
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=
The unreasonable effectiveness of deep features as a perceptual metric , author=. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=
-
[19]
Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 , pages=
Microsoft coco: Common objects in context , author=. Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 , pages=. 2014 , organization=
2014
-
[20]
arXiv preprint arXiv:1711.05101 , year=
Decoupled weight decay regularization , author=. arXiv preprint arXiv:1711.05101 , year=
-
[21]
IEEE Transactions on Circuits and Systems for Video Technology , volume=
Robust histogram shape-based method for image watermarking , author=. IEEE Transactions on Circuits and Systems for Video Technology , volume=. 2014 , publisher=
2014
-
[22]
IEEE Transactions on Multimedia , volume=
DCT-based image watermarking using subsampling , author=. IEEE Transactions on Multimedia , volume=. 2003 , publisher=
2003
-
[23]
2020 International Conference on Communication and Signal Processing , pages=
Image security enhancement using DCT & DWT watermarking technique , author=. 2020 International Conference on Communication and Signal Processing , pages=. 2020 , organization=
2020
-
[24]
International Conference on Learning Representations , volume=
Watermark anything with localized messages , author=. International Conference on Learning Representations , volume=
-
[25]
Information Sciences , pages=
Versatile and harmless deepfake proactive forensics via conditional watermarking , author=. Information Sciences , pages=. 2025 , publisher=
2025
-
[26]
IEEE Transactions on Multimedia , volume=
Screen-Shooting Resistant Watermarking with Grayscale Deviation Simulation , author=. IEEE Transactions on Multimedia , volume=. 2024 , publisher=
2024
-
[27]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Learning invisible markers for hidden codes in offline-to-online photography , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[28]
IEEE Transactions on Cybernetics , volume=
RIHOOP: Robust invisible hyperlinks in offline and online photographs , author=. IEEE Transactions on Cybernetics , volume=. 2020 , publisher=
2020
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.