{"id":"b209c8c1-e016-44f2-9a7e-c210557587f7","arxiv_id":"2412.00341","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that argues for combining physics-driven optimization with cross-modal adversarial learning, but offers no concrete framework, experiments, or evidence.","lead":"This preprint is a review of research that combines physics-informed optimization with cross-modal adversarial learning, plus a sketched proposal for a unified framework. It contains no experiments, derivations, or concrete implementation, so the central promise remains an unsupported outline.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that physics-constrained perturbations improve cross-modal robustness rests entirely on the unverified premise in §3.3 that attacks stay effective and physically plausible after transfer across modalities; no derivation, experiment, or exact formulation is provided.","rationale":"I read the paper as a survey and research proposal rather than a completed research contribution. The strongest claim in §3.4 goes beyond what the review establishes, and the premise identified by the reader is indeed the one that must be true for that claim to hold. My concern is more specific than the reader's: physical constraints and cross-modal transfer are in direct tension, because stronger physical plausibility typically means a smaller, more restricted perturbation set, while transferability relies on the source perturbation aligning with the target modality's decision geometry. The paper provides no formal statement of this trade-off and no evidence that it can be resolved. I also note that the abstract says the study examines experimental outcomes, but no experiments are reported; the proposed framework is mentioned only as a future direction, with no equations or algorithmic description. The reader's UNVERDICTED verdict is therefore appropriate. The reference misassignments and the empty citation in §2.2 independently support the conclusion that the survey's evidentiary foundation is unreliable, but they are secondary to the missing verification of the central synergy. The concrete test I propose would settle whether the premise has empirical support; until that test is run, the paper should not be treated as establishing the claimed robustness benefit.","tokens_in":5875,"tokens_out":4466,"duration_ms":45478,"concrete_test":"Run a minimal RGB-to-infrared transfer attack with and without a physical constraint. Use SYSU-MM01 or RegDB and a white-box source model. Generate (a) unconstrained PGD, (b) perturbations constrained by the atmospheric-scattering model cited in §2.3, and (c) random perturbations matched in norm. Measure source attack success, target model rank-1 retrieval drop, and a physical-plausibility score. The §3.4 claim requires (b) to outperform (c) on transfer while matching it on plausibility. If the authors do not supply code, the test still requires them to provide the exact constraint set and optimization objective; without that, the claim is unfalsifiable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's abstract and §3.4 assert that combining cross-modal adversarial learning with physics-driven optimization enables the creation of more robust, adaptable, and secure image retrieval systems. The load-bearing premise appears in §3.3: adversarial attacks in the context of RGB images may be adapted to infrared images, maintaining their effectiveness while ensuring physical plausibility. This premise is doing all the work; if physical constraints do not preserve attack efficacy or cross-modal transfer, the proposed fusion has no demonstrated benefit. The paper offers no optimization objective, no constraint set, and no evaluation for this premise. The tension is real: physical constraints restrict the admissible perturbation set, which tends to reduce attack success, while cross-modal transfer requires gradient alignment across modalities, which is not guaranteed by physical plausibility alone. The Discussion (§4) acknowledges trade-offs for augmentation and adversarial training but never acknowledges that the core fusion premise faces the same trade-off. The relevant references do not supply the missing support: [4] concerns dehazing and PDE solvers, not cross-modal attacks, and [3]/[19] are cross-modal attacks without physical constraints. Citation errors, including FGSM attributed to the dropout paper [17], PGD to backpropagation [16], and an empty reference for evolutionary optimization for perturbations, further weaken the survey's evidentiary basis. This is not an internal contradiction, so the correct response is not rejection; it is that the central claim is currently unverified and unfalsifiable as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a short review-style manuscript that argues for combining cross-modal adversarial learning with physics-driven optimization to improve the robustness, transferability, and security of image retrieval systems. The paper opens with a general introduction to image retrieval, then surveys adversarial learning, cross-modal attack methods, and physics-informed optimization in Section 2. Section 3 presents a 'Methodology' that describes cross-modal data augmentation, adversarial training, and physics-driven adversarial optimization, but this section contains no equations, algorithms, or experimental results. The central claim, stated in Sections 3.3 and 3.4, is that physics-constrained adversarial perturbations can be transferred across modalities (e.g., RGB to infrared) while remaining both effective and physically plausible. Section 4 discusses limitations of data augmentation and adversarial training but does not address the core feasibility question of the proposed fusion. The paper ends with a brief conclusion and a reference list that contains multiple citation mismatches.","tokens_in":6088,"tokens_out":2942,"duration_ms":28361,"significance":"If the core claim were substantiated—that physics-constrained perturbations preserve adversarial effectiveness across modalities and that jointly integrating cross-modal adversarial learning with physics-driven optimization improves robustness—this would be a meaningful contribution to the security and multimodal retrieval literature. The paper also provides a structurally organized survey that identifies relevant topics and points to a plausible research direction. However, the manuscript does not deliver a derivation, a formal problem statement, a single experiment, or a comparison against baselines. As written, the central thesis is an assertion rather than a demonstrated result, and the citation errors in a review paper further undermine its evidentiary value. The strengths are the clear taxonomy and the identification of a genuine gap; the weakness is the absence of any verifiable technical content supporting the proposed framework.","major_comments":[{"comment":"The load-bearing claim is the bullet 'Cross-Domain Transferability': 'adversarial attacks in the context of RGB images may be adapted to infrared images, maintaining their effectiveness while ensuring physical plausibility.' This premise is necessary for the paper's central argument that physics-driven adversarial optimization improves cross-modal robustness, but no optimization objective, constraint set, algorithmic procedure, or evaluation is given. The manuscript never specifies what 'physical plausibility' means formally, nor how effectiveness is maintained after transferring a perturbation across modalities. Without a concrete formulation or experimental evidence, the reader cannot assess whether the premise holds.","section":"Section 3.3"},{"comment":"The manuscript presents a 'Methodology' section that claims a novel approach but provides no equations, pseudocode, or formal definitions. Phrases such as 'optimization under physical constraints' and 'the integration of physics-driven methods ensures that adversarial perturbations are consistent with real-world scenarios' are never made precise. In a paper that claims to examine 'theoretical foundations and experimental outcomes' (Abstract), the absence of any formal or empirical support for the proposed framework is a load-bearing gap: the claimed synergy between physics-driven constraints and cross-modal adversarial learning is not demonstrated.","section":"Sections 3.3 and 3.4"},{"comment":"A review paper must have accurate citations, but Section 2.1 contains systematic reference mismatches: FGSM is attributed to [17], which is the Dropout paper by Srivastava et al.; PGD is attributed to [16], which is the backpropagation paper by Rumelhart et al.; Universal Adversarial Perturbations are attributed to [3], which is actually a cross-modal attack paper; and Szegedy et al.'s adversarial examples work is cited as [15], which is LeCun et al.'s 'Deep learning' survey. Additionally, Section 2.2 contains an empty reference for 'evolutionary optimization for perturbations []'. These errors materially impair the survey's reliability and must be corrected.","section":"Section 2.1 and References"},{"comment":"The Discussion acknowledges trade-offs for data augmentation and adversarial training, but it never acknowledges the central trade-off of the proposed fusion: physical constraints restrict the set of admissible perturbations, which tends to reduce attack success, while cross-modal transfer requires gradient alignment across modalities, which is not guaranteed by physical plausibility alone. The paper therefore omits the key technical risk of its own core proposal. A concrete test would be to measure attack success rate and physical realism for RGB-generated perturbations applied to infrared inputs, with and without physics constraints; the manuscript provides no such analysis or evidence.","section":"Section 4"}],"minor_comments":[{"comment":"The title contains a spacing artifact, 'Adversa rial Learning,' which should be corrected to 'Adversarial Learning.'","section":"Title and Abstract"},{"comment":"The abstract describes the paper as a review that examines 'theoretical foundations and experimental outcomes,' but Section 3 presents a new 'Methodology' without any theoretical derivation or experimental outcome. The genre inconsistency should be resolved, for instance by clearly labeling Section 3 as a position/proposal rather than an established method.","section":"Abstract and Section 3"},{"comment":"The empty reference for 'evolutionary optimization for perturbations' should be filled, or the sentence should be removed. In a review, an unresolved placeholder is a quality issue.","section":"Section 2.2"},{"comment":"Several references, especially those in the [3]-[12] range, appear to be self-citations to arXiv preprints by the same research group. While this is not itself an error, the review would benefit from citing a broader set of established works in adversarial robustness and cross-modal learning to support the survey's claims.","section":"References"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to be an early draft: it has citation errors, an empty reference, a formatting artifact in the title, and a methodology section that makes unsupported assertions. The concentration of citations on a single research group is notable and may reflect narrow coverage rather than a comprehensive review. The central technical claim is not supported by any derivation or experiment, and the paper's current scope would need substantial expansion—adding a formal framework and evaluation—to be publishable in a serious venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a review preprint, not a research paper. The one genuinely new thing is the concluding sketch of a unified physics-driven cross-modal adversarial framework, but that sketch has no equations, no algorithm, and no experiment. What the paper does well is assemble three literatures—data augmentation, adversarial learning, and physics-informed optimization—and point at a real gap: few existing methods combine physical constraints with cross-modal perturbation transfer. The early sections read like a competent textbook summary, especially the data augmentation taxonomy.\n\nThat is where the credit ends. The central claim in Sections 3.3–3.4, that physics-constrained perturbations improve cross-modal robustness and transferability, is asserted in prose and never supported by a derivation, an optimization objective, or an evaluation. The paper acknowledges trade-offs for augmentation and adversarial training in Section 4 but never confronts the same tension in its own fusion premise: restricting perturbations to physically plausible sets can reduce attack efficacy, and cross-modal transfer is not guaranteed by physical realism alone. The stress-test note gets this right.\n\nThe citation problems are more than cosmetic. FGSM is attributed to the dropout paper, PGD to the backpropagation paper, and there is an empty reference for evolutionary perturbation methods. The survey leans heavily on a concentrated cluster of arXiv preprints by one group (refs [3–12], [19]); self-citation is not itself a flaw, but the mismatched attributions suggest the literature was not read carefully enough to support a review.\n\nI would not send this to peer review in its current form. It is a review with no new result and with accuracy problems that undermine its usefulness as a survey. If the authors substantially revise—fix the citations, clearly separate the proposed framework from the reviewed material, and add at least a formal problem statement—it might become a passable survey for readers new to the area. As is, I would desk-reject. It is not a serious enough piece to warrant referee time.","headline":"A shallow survey-style preprint whose only novel element is a sketched future framework; the central fusion claim is asserted, not tested, and the citation errors are real.","tokens_in":6655,"tokens_out":2074,"would_cite":false,"duration_ms":19696,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that cross-modal adversarial learning and physics-driven optimization can be fused so that adversarial perturbations remain physically realistic while transferring across modalities, yielding more robust and secure…","keywords":["cross-modal adversarial learning","physics-driven optimization","image retrieval","adversarial perturbations","infrared and RGB matching","data augmentation","model robustness","neural PDE solvers"],"falsifier":"Take a retrieval model trained on paired RGB and infrared images, generate a physics-constrained perturbation and a standard perturbation of the same magnitude, and measure each attack's success when transferred across modalities; if the physics-constrained attack falls to chance while the standard one stays effective, the central claim fails.","tokens_in":5629,"feed_emoji":"🛡️","tokens_out":7891,"duration_ms":70073,"temperature":0.7,"pith_summary":"This review argues that cross-modal adversarial learning and physics-driven optimization belong in a single framework, and that fusing them yields adversarial perturbations that are both physically realistic and transferable across modalities such as RGB and infrared. The proposed methodology has three parts: cross-modal data augmentation, adversarial training, and physics-guided perturbation generation. The paper claims that these parts together address modality discrepancy, limited data, and insufficient model robustness, producing image retrieval systems that are more secure, adaptable, and interpretable. A sympathetic reader should care because if the synthesis holds, it turns adversarial examples from a security nuisance into a tool for building physically grounded, multi-sensor models.","feed_headline":"Physically constrained attacks claimed to transfer across RGB and IR","feed_subtitle":"One framework pairs physical realism with adversarial training to harden cross-modal retrieval.","key_machinery":"The load-bearing mechanism is a physics-constrained adversarial optimization loop coupled to cross-modal data augmentation. The attack side uses gradient-based perturbation generators such as the fast gradient sign method and projected gradient descent inside an adversarial training schedule; the physics side adds constraints drawn from domain models, illustrated by atmospheric scattering for image dehazing and by physics-informed objectives from neural PDE solvers. The mechanism is intended to restrict the search for adversarial examples to a physically plausible manifold, so that every successful perturbation is also a realistic variation a sensor might actually record, and so that perturbations learned in one modality remain valid in another.","core_discovery":"On its own terms, the paper's central claim is that constraining adversarial perturbations with physical models—rather than letting them be arbitrary gradient noise—improves both attack and defense in cross-modal settings. The paper states that physics-guided adversarial optimization keeps perturbations consistent with real-world capture conditions, and that such perturbations can be transferred from one modality to another, for example from RGB to infrared images, while preserving both effectiveness and physical plausibility. It further claims that the same principle applies beyond vision, in scientific computing tasks such as neural PDE solvers with sparse data, where adversarial learning under physical constraints increases robustness. The conclusion is that a unified framework joining physical principles to adversarial optimization is a viable pathway for multi-modal learning systems.","pith_inferences":["A direct way to stress-test the framework beyond the paper is to pit the physics-constrained attack against an unconstrained attack of equal perturbation budget on a visible-infrared retrieval benchmark; the core premise predicts no drop in transferability, while the alternative predicts a drop.","The paper implies that physical realism and attack efficacy are complementary, but they may trade off: the tighter the physical constraint set, the smaller the space of effective perturbations, so the constraint set itself may need to be learned per domain.","An untested extension suggested by the paper is to use physics simulators as a data generator for rare capture conditions, so that physically consistent adversarial examples double as training data for the target modality."],"forward_implications":["If the framework works as described, retrieval models trained with it should withstand both deliberate adversarial attacks and natural variations in illumination, camera hardware, and sensor noise.","Attacks generated on RGB images should transfer to infrared images without losing success rate, while remaining visually and physically plausible.","The same physics-guided optimization should improve robustness of neural PDE solvers trained on sparse data, since the perturbation set is aligned with physical constraints.","The proposed unified framework would give researchers a single set of principles and evaluation criteria for comparing augmentation, attack, and defense methods across modalities."],"supporting_citations":[{"why":"Supplies the cross-modality perturbation synergy attack that anchors the paper's cross-modal transferability claim.","marker":"[3]"},{"why":"Presents the cross-modality perturbation synergy method for person re-identification, the attack framework the proposed approach extends.","marker":"[19]"},{"why":"Provides the physics-guided adversarial learning result for neural PDE solvers, the paper's main evidence for combining physics with adversarial optimization.","marker":"[4]"},{"why":"Defines the extreme-capture-environment robustness problem that motivates the data augmentation and defense components.","marker":"[5]"},{"why":"Underlies the multi-modal data learning strategy that the cross-modal augmentation component builds on.","marker":"[6]"},{"why":"Provides the local feature masking robustness method cited as a precursor defense for the unified framework.","marker":"[12]"}],"fun_headline_variants":["Physics-constrained attacks cross RGB and IR","Adversarial learning gets a physics backbone","From RGB to IR: physics-guided attacks","Unified physics-adversarial framework for multi-modal","Physical constraints make attacks transferable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument depends on one premise: that a perturbation constrained to obey physical laws can still fool a model trained on a different modality, so realism and attack effectiveness survive the addition of physical constraints.","fun_headline_variants_meta":{"raw":{"variants":["Physics-constrained attacks cross RGB and IR","Adversarial learning gets a physics backbone","From RGB to IR: physics-guided attacks","Unified physics-adversarial framework for multi-modal","Physical constraints make attacks transferable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1539,"prompt_tokens":867,"completion_tokens":672,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":483,"completion_tokens_details":{"reasoning_tokens":606}},"tokens_in":483,"tokens_out":672,"duration_ms":6432,"temperature":1.0,"reasoning_tokens":606,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:28:19.524286+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a retrieval model trained on paired RGB and infrared images, generate a physics-constrained perturbation and a standard perturbation of the same magnitude, and measure each attack's success when transferred across modalities; if the physics-constrained attack falls to chance while the standard one stays effective, the central claim fails.","supporting_citations":[{"cited_title":"Cross-modality perturbation syn- ergy attack for person re-identiﬁcation","cited_arxiv_id":null,"evidence_quote":"Presents the cross-modality perturbation synergy method for person re-identification, the attack framework the proposed approach extends."},{"cited_title":"Beyond Augmentation: Empowering Model Robustness under Extreme Capture Environments","cited_arxiv_id":"2407.13640","evidence_quote":"Defines the extreme-capture-environment robustness problem that motivates the data augmentation and defense components."},{"cited_title":"Beyond Dropout: Robust Convolutional Neural Networks Based on Local Feature Masking","cited_arxiv_id":"2407.13646","evidence_quote":"Provides the local feature masking robustness method cited as a precursor defense for the unified framework."}],"review_version":1}