{"id":"503ca45a-dcbd-4b89-a3d8-67d4c5837f82","arxiv_id":"2606.30493","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Behavioral mapping of ICS attacks reveals dataset-specific physical patterns and shows binary evaluation metrics substantially overestimate detection performance.","lead":"The paper maps ICS sensor traces to five physical behavior types (drift, spike, oscillation, repetition, switching) across SWaT, WADI and HAI datasets and finds that binary normal/attack labels hide large performance drops when models must predict the specific behavior. A smart generalist should read it to see how standard benchmark scores in critical-infrastructure security can overstate real detection capability.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Sufficiency of the five physical primitives to represent attack behavioral diversity is unvalidated","rationale":"The reader's weakest_assumption directly identifies the load-bearing point for the strongest_claim. The full-text placeholder does not alter this; the argument's dependence on the primitives' adequacy remains the clearest internal risk. No other assumption (e.g., Random Forest choice) is more central to the binary-vs-proxy comparison.","tokens_in":1762,"tokens_out":339,"duration_ms":19511,"concrete_test":"Sample 200 attack windows across SWaT/WADI/HAI; have two ICS domain experts independently assign behavioral labels (allowing additional categories or 'none'); compute Cohen's kappa with the paper's automated mapping and count windows requiring new primitives. If kappa < 0.6 or >15% require new categories, the five-primitive set is insufficient for the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that aggregate binary metrics limit visibility into performance across behavioral proxies, evidenced by the SWaT macro-F1 drop from 85.44% (binary) to 37.84% (behavior-proxy multiclass)—depends on the five primitives (drift, spike, oscillation, repetition, switching) being sufficient to capture relevant diversity. The framework maps traces to these primitives to create the multiclass labels and to show dataset-specific distributions. Without shown coverage (e.g., fraction of attack windows assigned to a primitive, inter-rater agreement, or comparison against a larger candidate set of behaviors), the multiclass degradation could reflect incomplete or arbitrary labeling rather than a genuine blind spot in binary evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes a behavioral characterization framework for ICS intrusion detection that maps multivariate process traces to five physical primitives: drift, spike, oscillation, repetition, and switching. It applies this to the SWaT, WADI, and HAI datasets to reveal dataset-specific behavioral distributions and shows that an indicative Random Forest baseline exhibits significant performance degradation when evaluated under behavior-proxy multiclass prediction compared to standard binary labeling (e.g., SWaT macro F1 from 85.44% to 37.84%). The authors argue for complementing binary benchmarking with behavior-stratified evaluation.","tokens_in":1904,"tokens_out":403,"duration_ms":24937,"significance":"If the five primitives are shown to be sufficient, the work would provide a valuable demonstration of how aggregate binary metrics can obscure performance variations across different attack behaviors in ICS security, encouraging more nuanced evaluation practices. The empirical application to three public benchmarks and the concrete metric comparisons offer a practical illustration of the proposed limitation.","major_comments":[{"comment":"The central claim that binary metrics limit visibility into performance across behavioral proxies (Abstract) rests on the assumption that the five chosen primitives (drift, spike, oscillation, repetition, switching) sufficiently capture the relevant behavioral diversity of cyber-physical attacks. However, the manuscript provides no evidence of coverage (e.g., fraction of attack windows assigned to each primitive), inter-rater agreement for labeling, or comparison against a broader set of candidate behaviors, which could mean the observed multiclass degradation reflects incomplete labeling rather than a genuine blind spot.","section":"Behavioral characterization framework (as described in the abstract and methods)"}],"minor_comments":[{"comment":"The baseline is described only as 'indicative'; more detail on the Random Forest setup (features, hyperparameters, train/test split) would strengthen the evaluation claims.","section":"Evaluation section"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the thoughtful review and the opportunity to clarify the behavioral characterization framework. We address the major comment point by point below.","responses":[{"response":"We agree that the original manuscript lacks explicit quantitative coverage statistics and does not discuss inter-rater agreement or comparisons to alternative behavior sets. The five primitives were selected from ICS process dynamics literature to represent core physical effects of attacks on sensor/actuator signals. In the revision we will add a table reporting the fraction of attack windows assigned to each primitive per dataset (e.g., WADI repetition dominance) to demonstrate coverage. The mapping procedure uses deterministic, rule-based thresholds on signal statistics rather than manual annotation, rendering traditional inter-rater agreement inapplicable; we will expand the methods section to detail these rules and their rationale. A systematic comparison against a larger candidate behavior taxonomy is a worthwhile direction for future work but lies beyond the scope of the present study, whose primary contribution is to illustrate how behavior-stratified evaluation exposes limitations of binary metrics. The observed multiclass F1 degradation remains informative even under the current primitives, as it directly shows performance variation across the behaviors that are present.","revision_made":"partial","referee_comment":"[Behavioral characterization framework (as described in the abstract and methods)] The central claim that binary metrics limit visibility into performance across behavioral proxies (Abstract) rests on the assumption that the five chosen primitives (drift, spike, oscillation, repetition, switching) sufficiently capture the relevant behavioral diversity of cyber-physical attacks. However, the manuscript provides no evidence of coverage (e.g., fraction of attack windows assigned to each primitive), inter-rater agreement for labeling, or comparison against a broader set of candidate behaviors, which could mean the observed multiclass degradation reflects incomplete labeling rather than a genuine blind spot."}],"tokens_in":1381,"tokens_out":384,"duration_ms":24245,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is that binary normal-versus-attack labels on public ICS datasets can mask real differences in how attacks manifest physically, and the authors demonstrate this with a drop in macro F1 from 85% binary to 38% when predicting the five behavior classes on SWaT.\n\nThey define five primitives (drift, spike, oscillation, repetition, switching), map the traces from SWaT, WADI, and HAI, and report that the datasets sit in different parts of that space. WADI leans on repetition, HAI on drift and oscillation, SWaT on stealthier patterns. A random forest baseline then shows the multiclass scores degrade relative to binary ones. That concrete comparison on existing benchmarks is the useful piece.\n\nThe work is straightforward and stays within the ICS evaluation niche. It does not overclaim broader impact.\n\nThe soft spot is the primitives themselves. Nothing in the abstract shows they cover the attacks adequately or that the mapping from raw traces is reproducible; the stress-test concern about unvalidated sufficiency holds on the available text. The baseline is labeled only indicative, so the performance gap is suggestive rather than tightly controlled.\n\nThis is for ICS security people who run or review detection papers and want to think about evaluation blind spots. A reader already working on those three datasets will get the most out of it.\n\nI would send it to peer review. The central observation is worth referee time even if the primitives section needs more detail and validation.","headline":"The paper shows binary ICS metrics can hide behavioral differences across attacks but the five primitives need justification for the claim to land cleanly.","tokens_in":2392,"tokens_out":372,"would_cite":false,"duration_ms":25757,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Binary normal-versus-attack labels in ICS intrusion detection hide large performance gaps across distinct attack behaviors.","keywords":["ICS intrusion detection","behavioral characterization","binary labeling","cyber-physical attacks","SWaT","WADI","HAI","multiclass evaluation"],"falsifier":"Apply the same five-primitive mapping to a new ICS dataset containing documented attacks and check whether the resulting behavioral distribution matches any of the three studied datasets or instead requires additional primitives.","tokens_in":2669,"feed_emoji":"📉","tokens_out":654,"duration_ms":20636,"temperature":0.7,"pith_summary":"The paper maps raw sensor traces from three public ICS testbeds into five physical primitives and finds that attacks occupy different regions of this space than normal operation. Datasets themselves differ sharply: one is dominated by repetition, another by drift and oscillation, and the third by frozen telemetry. When a baseline detector is scored with behavior-proxy multiclass labels instead of binary labels, aggregate F1 falls sharply, for example from 85 percent to 38 percent on one dataset. The authors therefore argue that conventional binary benchmarks conceal blind spots that behavior-stratified evaluation would expose.","feed_headline":"Binary ICS scores drop from 85% to 38% when attack behaviors are distinguished","feed_subtitle":"Five physical primitives reveal that datasets occupy separate regions and that aggregate metrics conceal performance gaps across attack type","key_machinery":"The behavioral characterization framework that converts raw multivariate traces into five interpretable physical primitives.","core_discovery":"A behavioral characterization framework maps multivariate process traces into the primitives drift, spike, oscillation, repetition, and switching; when applied to SWaT, WADI, and HAI, it shows that attack windows exhibit clear shifts relative to normal operation, that the three datasets occupy largely distinct regions of behavioral space, and that binary aggregate metrics therefore limit visibility into detector performance across behavioral proxies.","pith_inferences":["Detectors tuned only on binary labels may systematically under-perform on repetition-heavy or drift-heavy attack classes.","New public ICS benchmarks could be released with the five-primitive labels already attached to speed adoption of stratified evaluation.","Targeted incident response could route alerts according to the dominant primitive observed rather than a single attack flag."],"forward_implications":["Attack windows produce measurable shifts away from normal-operation distributions in the five-primitive space.","The three datasets occupy largely non-overlapping regions, with WADI dominated by repetition, HAI by sustained drift and oscillation, and SWaT by stealthier frozen behavior.","A Random Forest baseline shows macro F1 dropping from 85.44 percent under binary evaluation to 37.84 percent under behavior-proxy multiclass prediction on SWaT, with comparable drops on the other two datasets.","Behavior-stratified evaluation is needed to expose performance blind spots that aggregate binary scores conceal."],"fun_headline_variants":["ICS binary metrics limit visibility into attack behavior performance","Attack behaviors reveal dataset biases in SWaT WADI and HAI","Five physical primitives map ICS attacks beyond binary labels","Binary evaluation hides behavioral diversity in public ICS benchmarks"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The five chosen physical primitives capture enough of the behavioral diversity present in real cyber-physical attacks.","fun_headline_variants_meta":{"raw":{"variants":["ICS binary metrics limit visibility into attack behavior performance","Attack behaviors reveal dataset biases in SWaT WADI and HAI","Five physical primitives map ICS attacks beyond binary labels","Binary evaluation hides behavioral diversity in public ICS benchmarks"]},"model":"grok-4.3","cost_usd":0.005122,"raw_usage":{"total_tokens":2513,"prompt_tokens":713,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":51224500,"prompt_tokens_details":{"text_tokens":713,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1738,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":713,"tokens_out":62,"duration_ms":19705,"temperature":1.0,"reasoning_tokens":1738,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T05:13:18.976309+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Apply the same five-primitive mapping to a new ICS dataset containing documented attacks and check whether the resulting behavioral distribution matches any of the three studied datasets or instead requires additional primitives.","supporting_citations":[],"review_version":1}