{"id":"a3a8f374-945b-4472-8acd-f17c34f2e594","arxiv_id":"2412.13709","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Binary tape patterns designed with a black-box genetic search and rendered on 3D human models can hide people from a NIR-based YOLOv5 detector, with 87.9% average physical attack success at 3-5 m.","lead":"This paper demonstrates a fully passive physical attack on nighttime surveillance cameras that use near-infrared illumination: binary patterns cut from reflective and insulating tapes, pasted on clothing, make a YOLO-based human detector miss the person. The result matters because it shows a low-cost reliability risk in NIR AI systems that are widely deployed for security.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Physical evidence targets official COCO-pretrained YOLOv5, not an NIR-trained detector, so the 87.9% ASR alone does not support the paper's central generalization to NIR-based surveillance AI.","rationale":"The reader flagged generalization from one detector and a small NIR finetuning set. I agree, and I want to sharpen it: the physical experiment, which is the paper's main evidence for real-world impact, does not use an NIR-trained detector at all; it uses official COCO YOLOv5. This is exactly the transfer step that must hold for the Sec. 7 conclusion. The digital result against the finetuned model supports the mechanism but is weak evidence by itself because the training set is small. A re-run on a large NIR-trained detector is the natural check. The paper has real strengths: the tape-based physical intensity manipulation is plausible and demonstrated, and the digital search is described in enough detail to reproduce. My concern is not that the attack is fake; it is that the empirical scope is narrower than the conclusion. Thus the conditional verdict stands without change.","tokens_in":13393,"tokens_out":7805,"duration_ms":73022,"concrete_test":"Repeat the physical protocol of Table 2 at 3/4/5 m with the same tape patterns, evaluating a detector set that includes (i) a YOLOv5s finetuned on a large NIR corpus with many identities (e.g., RegDB/SYSU-MM01 plus the authors' data) and (ii) the authors' NIR-finetuned YOLOv5s. Report per-distance ASR with per-subject confidence intervals. If the large-NIR-trained detector's ASR is materially below 87.9% (or the finetuned model's physical ASR drops), the evidence does not establish the paper's broad NIR-AI conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. 5.2.2 states that the physical experiments 'use them to attack the official YOLO.' Official YOLOv5 is pretrained on COCO RGB data, not on NIR imagery. The only NIR-trained detector is the YOLOv5s finetuned on 13 identities / 13,000 frames (Sec. 5.1), and it appears only in the digital evaluation (Tab. 1). Therefore the headline physical result, 87.93% average ASR in Tab. 2, is demonstrated against a model whose training distribution is RGB, not against the NIR surveillance models the conclusion (Sec. 7) warns about. The finetuned NIR model's high digital ASR (94.14%) is suggestive, but the finetuning set is small, and no physical evaluation is run on it. For the central claim that NIR-based AI raises significant reliability concerns to hold, one must assume the physical attack transfers to detectors trained on substantially larger and more varied NIR data. That transfer is the load-bearing, untested step. Additionally, Tab. 2 reports no error bars or number of frames/subjects, so the 87.93% figure could be dominated by one subject or by correlated frames.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that near-infrared (NIR) surveillance imagery is inherently vulnerable to physically realizable adversarial attacks because of color and texture loss in the NIR band and because the co-located placement of NIR illuminants and cameras allows simple intensity manipulation via retro-reflective and insulating tapes. The authors design binary clothing patterns using a black-box genetic algorithm over SMPL-based 3D human renderings, combine these patterns with real NIR backgrounds, and evaluate the resulting attack in both digital and physical settings against YOLOv5-based detectors. The paper reports strong digital attack success rates (up to 94.14% ASR on a finetuned NIR detector, Table 1) and a physical attack success rate of 87.93% on average against the official COCO-pretrained YOLOv5 (Table 2). The central claim is that these results reveal significant reliability concerns for nighttime surveillance systems powered by NIR AI algorithms.","tokens_in":13737,"tokens_out":2545,"duration_ms":24715,"significance":"If the results hold, the paper would be one of the first to demonstrate a fully passive, low-cost physical attack specifically targeting NIR-based human detection, an important and underexplored security domain. The work has several concrete strengths: it identifies a physical mechanism (co-located illumination plus retro-reflective materials) that is well grounded in imaging optics; it uses a black-box, 3D-aware search that avoids gradient access; it includes physical validation with qualitative RGB/NIR stealthiness results; and it releases code. The digital evaluation includes error bars and compares against reasonable baselines (all black, all white, random). However, the physical evaluation is limited in scope and, as detailed in the major comments, the experimental bridge between the physical demonstration and the paper's broad conclusion about NIR-based surveillance AI is incomplete.","major_comments":[{"comment":"The physical attack experiments are conducted only against the official COCO-pretrained YOLOv5, not against the NIR-finetuned detector described in Sec. 5.1. The 87.93% average ASR in Table 2 therefore demonstrates an attack on an RGB-trained model, not on the NIR-based surveillance models that the Sec. 7 conclusion warns about. The digital results on the finetuned NIR model (94.14% ASR, Table 1) are suggestive, but no physical evaluation is reported on that model. To support the central claim, the authors should either run physical attacks against the NIR-finetuned detector (or a larger NIR-trained detector) or provide a concrete argument and supporting evidence for why the physical attack transfers from the official YOLO to NIR-trained detectors.","section":"Sec. 5.2.2, Table 2"},{"comment":"Table 2 reports no error bars, no number of subjects, no number of frames, and no per-subject breakdown for the physical experiments. The aggregate 87.93% ASR could be dominated by a single subject or by many correlated frames from one pose sequence. The paper should report per-subject and per-frame statistics, confidence intervals, and a description of the capture protocol (number of subjects, actions, angles, distances) so that the physical result is properly interpretable.","section":"Table 2, Sec. 5.2.2"},{"comment":"The NIR-finetuned detector is trained on only 13 identities and 13,000 video frames. This is a very small training set, and the digital ASR of 94.14% on this model may partly reflect the model's low capacity and limited variation rather than a fundamental vulnerability of NIR-based AI in general. The conclusion in Sec. 7 extrapolates from this single small model to 'NIR-based AI' at large. The authors should temper the claim or support it with experiments on a larger, more diverse NIR training set or on additional architectures.","section":"Sec. 5.1"}],"minor_comments":[{"comment":"There is a typo: 'disussed' should be 'discussed'.","section":"Sec. 4.3.1"},{"comment":"The term 'homeomorphism' is used loosely; consider 'shared topology' or 'consistent mesh topology' for clarity.","section":"Sec. 4.3.1, Fig. 8"},{"comment":"The notation 'Ours (5black)' and 'Ours (1black)' is not defined in the main text; please explain that these refer to whether the head/hands/feet parts (5 parts) or only the head (1 part) are fixed to black.","section":"Sec. 5.1, Table 1"},{"comment":"The ASR definition could be clarified for the case where the detector produces no true-positive labels in the no-attack condition; currently N_0 might be zero for some inputs.","section":"Eq. (7)"},{"comment":"The sentence 'we can see that both AC and ASR become worse as the distance of the person becomes closer to the camera' appears to be an error: the numbers in Table 2 improve (AC decreases and ASR increases) at shorter distances, except for the 3m row where ASR is lower. Please rephrase to describe the actual trend (degradation at very close range due to visible head/hands/feet and non-black insulating tape).","section":"Sec. 5.2.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a computer vision or security venue. The main technical gap is the mismatch between the physical experiment (official RGB-trained YOLOv5) and the paper's broad conclusion about NIR-based surveillance AI. This is fixable with additional experiments or a carefully scoped rewrite, but it is load-bearing for the central claim. The small finetuning dataset further limits the generality. I would encourage the authors to address the physical-evaluation statistics and to add an experiment, even a limited one, on the NIR-finetuned detector."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first black-box physical attack on NIR-based human body detection that I know of, and the core physics story—retro-reflective tape under co-located 850nm LED illumination creating high-intensity patches on clothing—is solid and clearly explained. The digital search pipeline (SMPL render + genetic algorithm) is standard but applied in a new modality, and the digital evaluation against both a finetuned NIR YOLOv5 and the official checkpoint shows the searched patterns beat simple baselines like all-black and all-white. There is also a held-out set of 2,140 rendered maps, which is a reasonable generalization check. Credit where due: the paper is honest about its limitations (head/hands/feet texture, near-distance failures) and the code is promised on GitHub.\n\nThe soft spots are real but not disqualifying. The stress-test note lands: the physical experiments in Sec. 5.2.2 use the official COCO-pretrained YOLOv5, not the NIR-finetuned model. So the headline 87.93% physical ASR demonstrates that the pattern fools a generic YOLO on NIR-captured frames, but it does not by itself establish the paper's central generalization that \"NIR-based AI\" in surveillance systems is vulnerable. The finetuned NIR detector only appears in the digital table, and it was trained on 13 identities / 13,000 frames, which is tiny relative to real surveillance training sets. The digital simulation also treats tapes as exact 0/255 with no modeling of retro-reflector angular falloff or distance-dependent intensity, which may explain why physical ASR drops at 3m. Table 2 reports no error bars or per-subject/frame counts, so the average could be inflated by correlated frames.\n\nNone of that undermines the basic result: it is a plausible, low-cost, fully passive attack on a YOLO detector operating on NIR imagery. What is missing is evidence that it transfers to detectors actually trained on large-scale NIR data. That is an addressable gap—test on a larger NIR dataset or on a second NIR-trained architecture—rather than a fatal flaw.\n\nWho is this for? Researchers in adversarial ML and surveillance security will want to read it; it opens a useful new axis (NIR modality, passive materials) that RGB patch work mostly ignored. The paper deserves a serious referee, but I would ask for a physical run against the NIR-finetuned model and a less sweeping conclusion in Sec. 7.\n\nRecommendation: engage with it, but treat the physical ASR as an existence proof for generic YOLO, not a proven risk to deployed NIR surveillance AI.","headline":"First physical NIR human-detector attack; the core physics is solid, but the headline physical result runs against a generic YOLO, not a NIR-trained detector, so the broad conclusion needs tempering.","tokens_in":14207,"tokens_out":2020,"would_cite":true,"duration_ms":18485,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Binary patterns of retro-reflective and insulating tape, searched in digital space on 3D human models and pasted onto clothing, hide a person from a YOLOv5 near-infrared human detector in physical tests, with an average attack success…","keywords":["near-infrared imaging","physical adversarial attack","surveillance cameras","human detection","retro-reflective tape","YOLO detector","black-box genetic search","NIR texture loss"],"falsifier":"Capture the same physical setup with the 850-nm LED moved off-axis from the camera and measure the attack success rate of the unchanged tape pattern; if the detector still fails around 87.93% of frames, the claim that retro-reflective brightening under co-located geometry drives the attack is falsified.","tokens_in":13252,"feed_emoji":"📷","tokens_out":10777,"duration_ms":96691,"temperature":0.7,"pith_summary":"This paper tries to establish that near-infrared (NIR) surveillance AI is fundamentally more brittle than RGB-based AI because NIR images lose color and texture information, and because the camera's auxiliary LED sits almost on the lens axis. The authors show that cheap retro-reflective tape and black insulating tape, arranged into binary patterns on clothing, can manipulate the local brightness of NIR images and fool a YOLOv5 human detector. The patterns are searched in a simulated 3D world using a black-box genetic algorithm, then physically realized; the physical attack hides a person from the detector in 87.93% of tested frames at 3–5 meters. If the claim holds, nighttime surveillance systems powered by NIR AI are exposed to a fully passive, low-cost evasion threat.","feed_headline":"Tape patterns hide people from NIR surveillance AI 87.93% of the time","feed_subtitle":"Retro-reflective and insulating tape rearrange near-infrared brightness, hiding a person from YOLOv5 at 3 to 5 meters.","key_machinery":"The load-bearing object is retro-reflective tape: because it reflects light back along the incoming direction, it appears bright in a co-located camera/illuminant geometry no matter how the tape is oriented, letting the wearer write $0$ or $255$ brightness values onto the NIR image. The search machinery is a black-box genetic algorithm over 31 semantic body-part binary patterns, rendered on SMPL-based 3D human models from random viewpoints and composited with real NIR backgrounds, with the detector's average confidence as the fitness function.","core_discovery":"The central claim is that the design of NIR nighttime surveillance itself creates an attack surface: the camera's color channels respond almost identically at NIR wavelengths, the spectral reflectance of dyed fabric flattens so clothing textures disappear, and the nearly co-located 850-nm LEDs and camera make it easy to change image brightness with retro-reflective material. The authors demonstrate that covering parts of clothing with retro-reflective tape, which returns light toward its source and therefore appears bright regardless of orientation, and black insulating tape, which appears dark, lets an attacker write binary brightness patterns directly onto the NIR image. They search those body-part patterns with a black-box genetic algorithm over 3D human models without needing model weights or gradients, then paste the winning pattern on a person. On a YOLOv5 detector the pattern achieves an average physical attack success rate of 87.93% at 3–5 meters, and the NIR-finetuned detector is attacked even more easily in digital tests, supporting the claim that NIR-based AI image understanding is fragile.","pith_inferences":["Not settled by the paper: a detector trained on a much larger, more diverse NIR corpus, or a different architecture, might resist the same pattern; the evidence covers one YOLOv5 checkpoint and a 13-identity finetuning set.","A natural next experiment is to move the 850-nm LED off the lens axis; if the physical attack success rate collapses, the paper's co-location mechanism is confirmed as the operative cause.","Because retro-reflective tape brightens NIR regardless of its visible color, the same attack could be rendered as normal-looking clothing, a design direction the paper demonstrates qualitatively but does not quantify with a user study."],"forward_implications":["A person wearing the tape pattern can evade a YOLOv5 human detector in physical NIR surveillance footage at 3–5 meters, with an average attack success rate of 87.93%.","The attack is fully passive and inexpensive, requiring no knowledge of model parameters, no gradient access, and no active light-emitting hardware.","Because the NIR-finetuned YOLOv5 was easier to attack than the official checkpoint, specialized NIR training does not by itself remove the vulnerability.","The pattern transfers across camera angles and distances because the search is performed on 3D human models rendered from arbitrary viewpoints.","Altering the co-located camera/LED geometry is a fundamental defense, since the tape's brightening effect depends on light returning along the viewing axis."],"supporting_citations":[{"why":"Supplies the spectral-reflectance analysis of dyed fabrics used to ground the claim that clothing texture fades in NIR.","marker":"[18]"},{"why":"Supplies evidence on NIR surveillance frames losing visible color and texture, supporting the vulnerability analysis.","marker":"[36]"},{"why":"Provides the public RGB-NIR person dataset used to illustrate how clothes appear as uniform, textureless regions in NIR.","marker":"[35]"},{"why":"Provides the SMPL 3D body model used to render segmented human shapes for the digital pattern search.","marker":"[19]"},{"why":"Defines the YOLO object-detection architecture family that the physical attack targets.","marker":"[26]"}],"fun_headline_variants":["Tape patterns hide people from NIR night-surveillance AI","Retro-reflective tape defeats night-time AI human detection","87.93% attack rate: tape fools NIR surveillance YOLO","NIR surveillance AI fooled by tape patterns on clothing","Tape on clothes blinds near-infrared human detectors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that one detector model and a small near-infrared training set represent what real surveillance systems use; if actual systems are trained on far more varied data, the demonstrated success rate may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Tape patterns hide people from NIR night-surveillance AI","Retro-reflective tape defeats night-time AI human detection","87.93% attack rate: tape fools NIR surveillance YOLO","NIR surveillance AI fooled by tape patterns on clothing","Tape on clothes blinds near-infrared human detectors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000406,"raw_usage":{"total_tokens":2135,"prompt_tokens":996,"completion_tokens":1139,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":1054}},"tokens_in":612,"tokens_out":1139,"duration_ms":8823,"temperature":1.0,"reasoning_tokens":1054,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:52:06.024988+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Capture the same physical setup with the 850-nm LED moved off-axis from the camera and measure the attack success rate of the unchanged tape pattern; if the detector still fails around 87.93% of frames, the claim that retro-reflective brightening under co-located geometry drives the attack is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the spectral-reflectance analysis of dyed fabrics used to ground the claim that clothing texture fades in NIR."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies evidence on NIR surveillance frames losing visible color and texture, supporting the vulnerability analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the public RGB-NIR person dataset used to illustrate how clothes appear as uniform, textureless regions in NIR."}],"review_version":1}