{"id":"b76db91c-b416-4353-b900-d0e88e3b2062","arxiv_id":"2506.07882","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"On a real-world TalTech SOC NIDS alert dataset, DeepLIFT explanations for an LSTM alert classifier scored highest on faithfulness, robustness, complexity, and analyst-based reliability among LIME, SHAP, Integrated Gradients, and DeepLIFT.","lead":"This paper compares four explainable AI methods for an LSTM model that prioritizes network intrusion alerts from a real security operations center. It reports that DeepLIFT produces the most faithful, robust, simple, and analyst-aligned explanations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DeepLIFT's 'consistent outperformance' claim is internally contradicted by Table 3, where Integrated Gradients is significantly better on Low Complexity.","rationale":"The stress-test pass read the paper in good faith: the goal is clear, and the benchmark design is reasonable. The central claim, however, is that DeepLIFT consistently outperforms LIME, SHAP, and Integrated Gradients across all four quality criteria. The paper's own Table 3 refutes this. On the Low Complexity metric, the pairwise comparisons label IG as the better performer over DeepLIFT with p=1.07e-42, and Table 2's means align with IG having lower entropy. Since the paper defines lower entropy as preferable, DeepLIFT loses one of the four axes. This is a direct, internal contradiction that does not depend on external consensus or on the underspecified LSTM sequence construction. The reader's weakest_assumption (missing LSTM input shaping) is a serious reproducibility flaw, and the RMA standard deviation of 25.28 for a proportion-like metric also undermines the reliability comparison, but the complexity result is more decisive because it contradicts the specific wording 'consistently outperformed' with the authors' own significance tests. A simple independent recomputation would settle it. Therefore, the reader's REJECT verdict is appropriate; my specific concern differs, so I mark disagreement with the reader's identified weakest assumption.","tokens_in":13972,"tokens_out":6175,"duration_ms":68297,"concrete_test":"Independently recompute the Low Complexity entropy (Eq. 14) for Integrated Gradients and DeepLIFT on the same 2000-point test subset, using the authors' published explanation scores or a reimplementation of Captum's IG and DeepLIFT on the trained LSTM. If the IG mean entropy is lower than DeepLIFT's and a paired Wilcoxon signed-rank test still favors IG with p<0.05, then the claim that DeepLIFT consistently outperforms is false as written and the abstract/conclusion must be revised to acknowledge IG's superiority on complexity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and conclusion assert that DeepLIFT consistently outperformed the other three XAI methods across the evaluation framework (faithfulness, complexity, robustness, reliability). This is contradicted by the paper's own results for the complexity criterion. In Section 3.4.4, lower entropy (Eq. 14) is defined as the better (lower-complexity) explanation. Table 2 reports a mean Low Complexity of 2.1745 for Integrated Gradients versus 2.2635 for DeepLIFT, so IG is better by the paper's own definition. Table 3 confirms this is statistically significant: under Low Complexity, the pairwise comparison 'IG vs Deep Lift' is marked 'I' (IG is the better performer) with p = 1.07e-42. Thus DeepLIFT does not outperform on at least one of the four headline criteria. Because the central claim is explicitly 'consistently outperformed,' it is internally inconsistent with the reported evidence. The conclusion's statement that DeepLIFT shows 'superior performance ... across these evaluation metrics' cannot be reconciled with Table 3. While the LSTM sequence-construction issue and the extreme RMA standard deviations also raise concerns, this direct self-contradiction is sufficient to invalidate the primary conclusion.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates four post-hoc feature-attribution XAI methods (LIME, SHAP, Integrated Gradients, DeepLIFT) for explaining an LSTM model trained to prioritize NIDS alerts from a real SOC dataset. It introduces a four-criterion evaluation framework (faithfulness, complexity, robustness, reliability), applies pairwise Wilcoxon tests, and claims that DeepLIFT consistently outperforms the other methods across all criteria and that XAI feature rankings agree with features identified by SOC analysts.","tokens_in":14171,"tokens_out":2708,"duration_ms":32602,"significance":"If the central claim were valid, the paper would provide a practical, data-grounded recommendation for choosing an explainer for LSTM-based NIDS alert prioritization and a reusable evaluation framework. The use of a real SOC dataset, collaboration with analysts, and statistical pairwise comparisons are genuine strengths. However, the main claim is internally contradicted by the reported results, and the reliability evaluation is circular by construction; these issues prevent the paper from supporting its stated significance.","major_comments":[{"comment":"The abstract and conclusion state that DeepLIFT 'consistently outperformed' the other XAI methods across the evaluation framework. This is directly contradicted by the paper's own results on the Low Complexity metric. Section 3.4.4 defines lower entropy (Eq. 14) as lower (better) complexity; Table 2 reports a mean Low Complexity of 2.1745 for Integrated Gradients versus 2.2635 for DeepLIFT, and Table 3 marks the IG-vs-DeepLIFT pairwise comparison as 'I' (IG better) with p = 1.07e-42. Thus DeepLIFT does not outperform on at least one of the four headline criteria, and the stated central conclusion is not supported by the reported evidence.","section":"Abstract, §4 Table 3, §3.4.4"},{"comment":"The LSTM model is trained on alert group records, but the paper never specifies how the static feature vectors are converted into the time-ordered sequences an LSTM requires. No input shape, sequence length, feature ordering, or time-step construction is described; Eqs. (1)–(6) give only generic LSTM update equations. Since the explanations are explanations of this LSTM, the model being explained is underspecified, and the reported attributions could reflect an arbitrary or unjustified temporal structuring of the data.","section":"§3.2"},{"comment":"The Reliability metrics (RMA and RRA) use a ground truth mask constructed from the five SOC analyst features listed in Table 1, and Section 4 then presents agreement with those same analyst features as evidence validating the XAI methods. This is circular: any explainer that ranks the five analyst-selected features highly will score well on reliability, so the later 'validation' is not independent. The paper should either use a genuinely independent ground truth or explicitly frame reliability as measuring concordance with the analyst features rather than as validation of explanation correctness.","section":"§3.4.1, §4, Table 1"},{"comment":"The reported RMA standard deviations are implausibly large relative to the means: DeepLIFT has 0.7812 ± 25.2805 and LIME has 0.6234 ± 9.7008. This suggests numerical instability, extreme outliers, or an error in computation or reporting. Because RMA is one of the two reliability metrics supporting the DeepLIFT recommendation, this instability must be explained before any reliability conclusion can be accepted.","section":"Table 2"}],"minor_comments":[{"comment":"The text says 'Table 5 shows the results' when the displayed table is labeled Table 2; the reference should be corrected.","section":"§4"},{"comment":"The faithfulness correlation definition contains malformed notation, including 'B∈( |d| |B|)' and a trailing parenthesis in 'xB = xi|i ∈ B}'; the subset-sampling procedure and baseline value should be stated clearly.","section":"§3.4.2, Eq. (12)"},{"comment":"The max sensitivity definition contains a garbled condition 'M (x) =M (x)(z)'; the neighborhood definition and distance metric should be written precisely.","section":"§3.4.3, Eq. (13)"},{"comment":"The text uses both 'Relevancy' and 'Relevance' for the same metrics; the terminology should be made consistent.","section":"§3.4.1"},{"comment":"The paper mentions '10 features' in the example explanations, but the dataset description includes a large number of AttrSimilarity fields; the total feature dimensionality used in the LSTM and in the explainers should be stated explicitly.","section":"§3.2 and §4"}],"recommendation":"reject","confidential_remarks":"The internal contradiction in the central claim (Tables 2 and 3) and the circular reliability validation are substantial and would require reworking the paper's main conclusion and evaluation design. The missing LSTM sequence-construction details also make the explained model underspecified for readers. These issues go beyond local corrections, so I cannot recommend acceptance or minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read the TalTech XAI-for-NIDS alert paper. The short version: there is a real benchmark here worth looking at, but the advertised conclusion—DeepLIFT consistently wins—does not survive contact with the paper's own numbers.\n\nThe application is legitimate: real SOC alert data (Suricata at TalTech), an LSTM prioritization model, four standard explainers, and a four-criterion evaluation with Wilcoxon tests. That's a concrete, relevant comparison. The SOC analyst involvement to define a ground-truth feature set is a sensible touch. The paper is honest enough to report all pairwise statistics in Table 3.\n\nBut the central claim is internally contradicted. Under Low Complexity, lower entropy is better, and Table 2 shows IG at 2.17 vs DeepLIFT at 2.26. Table 3 marks IG as the significantly better explainer for that metric (p=1e-42). So DeepLIFT does not 'consistently outperform'—it loses on one of the four headline criteria. The conclusion simply ignores its own result. That alone invalidates the abstract claim. Second, the LSTM is underspecified. The input features are static alert-group attributes; the paper never explains how they are turned into a sequence. Without that, the XAI explanations are explaining an ill-defined model. Third, the reliability ground truth uses the same SOC analyst features that are later presented as independent validation—that's circular. Fourth, the RMA standard deviation for DeepLIFT is 25.28, which makes the 0.78 mean meaningless. The 2000-point subset selection is unexplained, and no code/data are released.\n\nThese are fixable in principle. The multi-metric comparison and the analyst-ground-truth process are worth keeping. But as it stands, the paper's primary conclusion is not supported.\n\nThis is for referees and SOC researchers who want a cautionary example of XAI evaluation pitfalls. It deserves a serious referee rather than desk rejection, but it needs major revision before publication.\n\nMy recommendation: send it to peer review but with an explicit request that the authors resolve the complexity contradiction, specify the LSTM input construction, break the circularity in reliability, and release code/data. If they can't, it stays a cautionary note.","headline":"Useful benchmark idea, but the headline claim is contradicted by the paper's own Table 3, and the LSTM is underspecified; needs major revision before it can support a DeepLIFT recommendation.","tokens_in":14740,"tokens_out":2932,"would_cite":false,"duration_ms":33608,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that DeepLIFT is the most trustworthy post-hoc explainer for LSTM-based prioritization of network intrusion alerts, as measured by faithfulness, complexity, robustness, and reliability.","keywords":["explainable AI","network intrusion detection","alert prioritization","LSTM","DeepLIFT","feature attribution","security operations center","XAI evaluation"],"falsifier":"Train the same LSTM on versions of the same alert data with randomly permuted time-step order, then recompute classification accuracy and the four explanation metrics; if accuracy and DeepLIFT's advantage remain essentially unchanged, the claimed temporal structure is irrelevant and the comparison describes a model that did not need to be an LSTM. A second, stronger check is to run the same evaluation on a different SOC's alert dataset and see whether DeepLIFT still wins on reliability.","tokens_in":13746,"feed_emoji":"🛡️","tokens_out":3751,"duration_ms":43248,"temperature":0.7,"pith_summary":"The paper tries to establish that post-hoc explainability for deep-learning-based NIDS alert prioritization can be evaluated on four criteria, and that one method, DeepLIFT, dominates the others. Using a real 60-day Suricata alert log from a university security operations center, the authors train an LSTM to classify alert groups as important or irrelevant, then compare LIME, SHAP, Integrated Gradients, and DeepLIFT. They claim DeepLIFT consistently wins on every evaluation metric, and that its top features overlap with the five features SOC analysts identify as most important. If the claim holds, security operations centers can adopt DeepLIFT as a default explainer for LSTM-based alert triage without sacrificing fidelity to the model's decision process.","feed_headline":"DeepLIFT beats LIME, SHAP, and IG at explaining NIDS alerts","feed_subtitle":"Four-metric test on real SOC data favors DeepLIFT's faithful, stable, simple, analyst-aligned explanations.","key_machinery":"The central mechanism is the combined evaluation framework: an LSTM classifier for alert-group priority, four feature-attribution explainers (LIME, SHAP, Integrated Gradients, DeepLIFT), and a four-criterion scorecard made of named metrics—high faithfulness correlation, monotonicity, max sensitivity, low-complexity entropy, relevance mass accuracy, and relevance rank accuracy. The reliability metrics use a ground-truth mask built from the five features that TalTech SOC analysts identified as essential for alert significance, turning each explainer's attribution vector into scalar scores that can be compared across 2000 test points.","core_discovery":"The authors claim that DeepLIFT, which propagates activation differences from a reference input, produces explanations for the LSTM's alert-priority decisions that are more faithful, more robust, less complex, and more reliable than LIME, SHAP, and Integrated Gradients. Concretely, DeepLIFT achieves the highest high-faithfulness correlation (0.7559), the lowest max sensitivity (0.0008), low complexity (2.2635), and the best reliability scores against the analyst-defined ground truth (relevance mass accuracy 0.7812, relevance rank accuracy 0.6754). Pairwise Wilcoxon signed-ranks tests show significant differences favoring DeepLIFT across all metrics. The paper further claims that global SHAP analysis ranks the SOC-analyst-identified features among the most impactful, validating the practical usefulness of the explanations.","pith_inferences":["Editorial inference: if the LSTM's input sequences were constructed arbitrarily from static alert-group records, the four explainers may be describing a model that does not actually exploit temporal structure; a natural check is to retrain on randomly shuffled sequences and see whether the classification performance and DeepLIFT's advantage persist.","Editorial inference: the conclusion may transfer to other sequence models such as GRUs or transformers only if reference-based propagation methods behave similarly on their attributions; this is a testable extension beyond the paper.","Editorial inference: the reliability scores depend on the five analyst-selected features; at a different SOC with a different ground-truth feature set, the ranking of explainers could shift.","Editorial inference: DeepLIFT's low max sensitivity suggests it may produce stable explanations for streaming alert data, but this needs direct verification on online, non-static inputs."],"forward_implications":["For LSTM-based NIDS alert triage, DeepLIFT is the recommended post-hoc explainer among the four tested in this study.","The four-metric framework can be reused by SOC engineers to benchmark explainability tools before deployment.","Analyst-identified features such as SignatureMatchesPerDay, Similarity, SCAS, SignatureID, and SignatureIDSimilarity can serve as a practical ground truth for checking explanation reliability.","Explanation quality can be assessed independently of model accuracy, so SOCs may not need to trade detection performance for transparency.","The statistical comparison procedure offers a template for future studies selecting among explainable-AI methods."],"supporting_citations":[{"why":"Supplies the DeepLIFT method that the paper identifies as the best-performing explainer.","marker":"[24]"},{"why":"Defines Integrated Gradients, one of the four compared explainers.","marker":"[23]"},{"why":"Defines LIME, one of the four compared explainers.","marker":"[25]"},{"why":"Defines SHAP, one of the four compared explainers.","marker":"[26]"},{"why":"Provides the faithfulness, robustness, and complexity metrics used in the evaluation framework.","marker":"[28]"},{"why":"Provides the relevance rank accuracy and relevance mass accuracy metrics used for reliability evaluation.","marker":"[29]"},{"why":"Establishes the practice of using Wilcoxon signed-ranks tests for comparing explainers.","marker":"[41]"},{"why":"Describes the stream-clustering-guided dataset construction that produced the real-world NIDS alert data used for training and evaluation.","marker":"[22]"}],"fun_headline_variants":["DeepLIFT tops explainability for NIDS alerts on real SOC data","DeepLIFT beats LIME, SHAP, IG in NIDS alert explanation quality","For NIDS alerts, DeepLIFT explains better than LIME, SHAP, IG","Real SOC data: DeepLIFT wins XAI comparison for alert priorities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the LSTM processes a meaningful time-ordered sequence, but the paper never states how the static alert-group feature vectors are converted into sequences, so if that ordering is arbitrary the entire model and its explanations rest on an unjustified inductive bias.","fun_headline_variants_meta":{"raw":{"variants":["DeepLIFT tops explainability for NIDS alerts on real SOC data","DeepLIFT beats LIME, SHAP, IG in NIDS alert explanation quality","For NIDS alerts, DeepLIFT explains better than LIME, SHAP, IG","Real SOC data: DeepLIFT wins XAI comparison for alert priorities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000237,"raw_usage":{"total_tokens":1534,"prompt_tokens":998,"completion_tokens":536,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":450}},"tokens_in":614,"tokens_out":536,"duration_ms":6092,"temperature":1.0,"reasoning_tokens":450,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:23:24.532454+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same LSTM on versions of the same alert data with randomly permuted time-step order, then recompute classification accuracy and the four explanation metrics; if accuracy and DeepLIFT's advantage remain essentially unchanged, the claimed temporal structure is irrelevant and the comparison describes a model that did not need to be an LSTM. A second, stronger check is to run the same evaluation on a different SOC's alert dataset and see whether DeepLIFT still wins on reliability.","supporting_citations":[{"cited_title":"Learning important features through propagating activation differences,","cited_arxiv_id":null,"evidence_quote":"Supplies the DeepLIFT method that the paper identifies as the best-performing explainer."},{"cited_title":"Axiomatic attribution for deep networks,","cited_arxiv_id":null,"evidence_quote":"Defines Integrated Gradients, one of the four compared explainers."},{"cited_title":"Why should I trust you? Explaining the predictions of any classifier,","cited_arxiv_id":null,"evidence_quote":"Defines LIME, one of the four compared explainers."},{"cited_title":"A unified approach to interpreting model predictions,","cited_arxiv_id":null,"evidence_quote":"Defines SHAP, one of the four compared explainers."},{"cited_title":"CLEVR-XAI: A benchmark dataset for the ground truth evaluation of neural network explanations,","cited_arxiv_id":null,"evidence_quote":"Provides the relevance rank accuracy and relevance mass accuracy metrics used for reliability evaluation."},{"cited_title":"How can I choose an explainer? An application-grounded evaluation of post-hoc explanations,","cited_arxiv_id":null,"evidence_quote":"Establishes the practice of using Wilcoxon signed-ranks tests for comparing explainers."},{"cited_title":"Stream clustering guided supervised learning for classifying NIDS alerts,","cited_arxiv_id":null,"evidence_quote":"Describes the stream-clustering-guided dataset construction that produced the real-world NIDS alert data used for training and evaluation."}],"review_version":1}