{"id":"6ae3f9a4-5a8d-47eb-bb5d-cfbb3089c73b","arxiv_id":"2507.03993","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MalVol-25 is a new dataset of 30 clean and infected Windows memory dumps from 15 malware variants, intended for ML and agentic AI detection research.","lead":"This paper describes MalVol-25, a volatile memory dataset of 30 clean and infected snapshots covering 15 malware families on Windows 7, 8.1, 10 and 11. The authors position it as a resource for machine learning and agentic AI detection, but provide little concrete evidence of its contents, validation, or practical value.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of a diverse, ML/RL-ready dataset is undercut by the reported scale (30 dumps, one infected capture per malware/OS pair) and by the total absence of per-snapshot infection evidence; the paper's own Sections III-C, IV, and VII expose these gaps.","rationale":"The reader's weakest assumption was that infected snapshots contain active, correctly labelled malware behaviour, and the present stress test agrees that this is the central soft spot. I mark agreement as partial because the issue is sharper and broader than 'no per-sample evidence': even if every label were correct, the reported dataset of 30 dumps with a single infected capture per malware/OS pair cannot support the paper's claims of diversity across operating systems, of 'detailed behavioural and environmental features' sufficient for RL state/transition modelling, or of the multimodal time-synchronized data described in Section VII. These are internal inconsistencies with the stated methodology, not merely missing empirical proof. The paper does have independent support in the form of a public DOI and a clearly described virtualized acquisition environment, so the artifact may be exactly what the authors intend; that is why I do not move the verdict to REJECT. A concrete download-and-verify step can settle whether the missing validation evidence exists in the dataset package. If it does, the conditional acceptance is justified; if it does not, the central claim should be treated as unverified. The reader's CONDITIONAL verdict remains the correct call, so the recommendation is UNCHANGED.","tokens_in":8424,"tokens_out":3129,"duration_ms":37758,"concrete_test":"Download MalVol-25 from DOI 10.21227/kg5b-nf37, compute SHA-256 hashes for every dump, and, for each of the 15 infected dumps, run Volatility 3 plugins (windows.pslist, windows.pstree, windows.malfind, windows.dlllist, windows.netscan) plus a diff against the paired clean dump. Require documented, malware-specific artifacts in at least 14 of 15 infected dumps and no artifacts in the clean dumps. Independently check whether the package actually contains the claimed time-series, multimodal data, and label schema described in Section VII; if the artifacts or labels are absent, the Section IV and Section VII claims are unsupported and the abstract's 'well-validated' assertion should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is not simply that infected snapshots exist, but that they contain active, correctly labelled malware behaviour in sufficient quantity and detail to support the stated ML and agentic-AI use cases. Section III-C reports only 30 total dumps (15 clean plus 15 infected), with one infected capture per malware/OS row in Table I. There are no repeated trials, no timestamps, and no temporal series; the claimed 'modelling system states and transitions' and RL readiness in Section VII cannot be supported by the described acquisition protocol. Section VII further claims time-series memory snapshots synchronised with multimodal data such as network traffic, system call traces, and logs, but none of those modalities appear in Section III's methodology. The phrase 'multiple operating systems' in the abstract and Section III-B is materially weaker than stated: Table I covers only Windows 7, 8.1, 10, and 11, all variants of one OS family. Section IV asserts that the Volatility Framework and manual inspection confirmed artifacts and showed clear differences between clean and infected states, but no per-sample outputs are provided or referenced: no hashes, no process names, no Volatility plugin results, no network evidence, and no label files. Section VI concedes virtualization constraints and narrow coverage, yet the abstract and Section V claim a 'reliable resource that supports consistent, repeatable investigations'. A 15-infected-snapshot dataset with one capture per sample and no published validation evidence cannot establish that claim from the manuscript alone. The concern is therefore not a disagreement with the community consensus; it is an internal evidentiary gap between the reported procedure and the claims of diversity, validation, and RL-readiness.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces MalVol-25, a volatile-memory dataset constructed by running 15 malware samples (mostly ransomware, with a worm, a trojan, and injection/crypto malware) in isolated Windows 7/8.1/10/11 virtual machines, capturing a clean baseline memory dump and an infected memory dump for each pair. The authors claim the resulting 30 dumps are validated via Volatility Framework artifact extraction and manual inspection, documented with cryptographic checksums, and available at IEEE Dataport. They further claim the dataset supports machine learning, agentic AI, reinforcement-learning state-transition modelling, and multimodal analysis. The paper contains no per-snapshot artifact evidence, no label schema, no checksum manifest, and no experimental evaluation; several broad claims in the abstract and Sections IV-VII exceed what the methodology and reported scale can support.","tokens_in":8790,"tokens_out":6514,"duration_ms":65683,"significance":"If the dataset actually contains correctly labelled, verified infected snapshots, it would be a modest but useful addition to the small set of public memory-dump datasets, particularly because it provides paired clean/infected captures and reports ethical-approval and containment procedures. The strengths are a controlled experimental setup, use of a public repository (IEEE Dataport), a clear pairing design, and a stated intent to combine automated and manual validation. There are no equations or fitted parameters, so the usual circularity concerns do not apply; the heavy reliance on the authors' own reference [3] in the motivation does not by itself invalidate the dataset. However, the manuscript currently does not provide the evidence needed to assess those strengths, and the stated scale and missing modalities fall short of the abstract's and Section VII's promises.","major_comments":[{"comment":"The described dataset contains only 30 memory dumps (one clean and one infected per malware/OS pair), but the abstract and Section VII claim the dataset enables modelling system states and transitions and RL-based malware detection. With no repeated trials, no timestamps, and no temporal sequence, one snapshot per state cannot support state-transition or reinforcement-learning claims. Either add the missing captures or substantially weaken these claims.","section":"Section III-C and Table I"},{"comment":"The validation claim is unsupported. Section IV asserts that the Volatility Framework and manual inspection revealed anomalous processes and network patterns, but it provides no plugin names, no process names, no network connection lists, no injected-code indicators, no artifact counts, and no per-sample results. Since the dataset's central value depends on correct labels and verified infection, include a validation appendix or companion manifest with per-snapshot evidence and a concrete label schema.","section":"Section IV"},{"comment":"The paper states that cryptographic checksums and standardised naming were used to preserve data integrity, but no checksums, file manifest, naming convention, or directory structure are shown. Without a manifest listing each file and its hash, the integrity and replicability claims cannot be checked by readers. Add this information either in the paper or as a linked file in the dataset repository.","section":"Section V and Data Availability"},{"comment":"Section VII claims that the dataset includes 'time-series memory snapshots synchronised with multimodal data such as network traffic, system call traces, and logs.' Section III describes only RAM snapshot acquisition; no such synchronised multimodal collection is described in the methodology, and no corresponding data files are identified. This claim should be removed or the described data must be added to the dataset.","section":"Section VII"},{"comment":"The abstract and Section III-B emphasise 'multiple operating systems' and 'diverse platforms,' but Table I covers only Windows 7, 8.1, 10, and 11—four releases of the same OS family. The OS-diversity claim is overstated and should be reworded to 'multiple Windows versions' unless non-Windows systems are added to the dataset.","section":"Abstract and Section III-B"}],"minor_comments":[{"comment":"Figure 1 appears as a caption with no visible diagram in the manuscript; please include the actual figure or remove the reference.","section":"Figure 1"},{"comment":"There are typographical artifacts such as 'V olatility' and 'MALV ADA' with irregular spacing, and missing spaces in 'SectionII' and 'SectionIII'; a careful proofreading pass is needed.","section":"General formatting"},{"comment":"Section VI concedes virtualization constraints and narrow coverage, while Section VIII calls the dataset 'comprehensive' and 'well-validated'; these statements should be reconciled to avoid overclaiming.","section":"Section VI and Section VIII"},{"comment":"The Data Availability statement gives a DOI but no inventory of file names, sizes, or formats; a brief table of contents for the dataset would help readers understand what is being released.","section":"Data Availability"},{"comment":"Reference [5] contains a mixed DOI prefix that does not match the journal article, and reference [3] is the authors' own prior work cited repeatedly in the motivation; the former should be corrected and the latter should not be used as the sole basis for the RL-readiness claims.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central claims depend entirely on the actual contents of the IEEE Dataport deposit. I would recommend that an editor verify that the DOI resolves and that the deposit contains at least the claimed 30 dumps, a manifest with hashes, and per-sample label files. The paper also leans heavily on reference [3] (the authors' own work) in the motivation and future-work sections; this is not improper, but it should not substitute for independent related-work support. If the dataset is as small and unimodal as the text suggests, the journal should consider whether a 30-dump single-OS-family dataset meets its significance bar even after the missing evidence is supplied."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a small memory-dump dataset (30 dumps: 15 clean, 15 infected, across 15 named Windows malware samples) with a DOI, and that artifact is the real contribution. The paper around it, though, makes claims about diversity, validation, and RL-readiness that the manuscript itself does not support.\n\nCredit where due: the dataset is new. Table I lists a spread of ransomware, worms, trojans, and injection/crypto malware, matched to Windows 7/8.1/10/11. The acquisition pipeline — VM isolation, clean baseline snapshot, infection, post-infection snapshot, Volatility analysis — is standard, but executing it on a new set of samples and publishing the snapshots is a legitimate resource. Related work is broad and correctly places this against FabIoT, CMD_2024, MALVADA, Sadek et al., and others. No quantitative comparison is attempted, but the paper doesn't claim to beat those datasets; it claims to be a new, complementary resource.\n\nThe soft spots are real and load-bearing. Section III-C says the dataset \"comprises 30 clean and infected live memory dumps\" — that is one infected capture per sample, no repeated trials, no timestamps, no temporal series. Section IV asserts Volatility and manual inspection \"confirmed artifacts\" and showed \"clear differences,\" but gives no actual outputs: no hashes, no process lists, no plugin results, no network evidence, no label schema. Section V claims checksums but doesn't show them. Section VII describes time-series snapshots synchronized with network traffic, system call traces, and logs — none of which appear in the methodology. The abstract's \"multiple operating systems\" is actually four versions of one OS family. The RL and agentic-AI framing is speculative relative to the data.\n\nThese gaps are not just cosmetics. If the infected snapshots don't actually contain active, correctly labelled malware behaviour, the dataset is just 15 memory dumps of unknown quality. The paper provides no way to tell from the text alone.\n\nI'd send this to peer review rather than desk-reject, because the artifact is concrete, has a DOI, and a referee can check it. But the review would need to require: a dataset card with hashes and acquisition scripts, per-sample infection evidence (Volatility outputs, process names, malware hashes), a realistic label schema, and a rewrite that brings the RL/multimodal claims in line with what was actually collected. With that, it becomes a modest but usable testbed. As it stands, it's a conditional: the data might be fine, but the manuscript doesn't demonstrate it.","headline":"A small new memory-dump dataset with a DOI, but the manuscript's validation and RL-readiness claims are unsupported by the evidence presented.","tokens_in":9283,"tokens_out":2918,"would_cite":false,"duration_ms":29859,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MalVol-25 pairs clean and infected RAM dumps to give malware-detection models a system-state view.","keywords":["malware dataset","volatile memory","memory forensics","machine learning","agentic AI","reinforcement learning","incident response","ransomware"],"falsifier":"For each of the 15 named malware families, inspect the corresponding infected memory dump for that family's distinctive artifacts (for example, WannaCry's mutex or ransom-note remnants, Cerber's process names, or GandCrab's encryption traces); if a substantial share of claimed-infected dumps contains no trace of the named malware while the clean baseline does, the central claim of correctly labelled, diverse infection states collapses.","tokens_in":8188,"feed_emoji":"💾","tokens_out":5714,"duration_ms":56700,"temperature":0.7,"pith_summary":"This paper claims that a systematically generated dataset of 30 clean and infected volatile memory snapshots, spanning 15 malware families across four Windows versions, can fill a gap in training data for machine learning and agentic AI malware detection and response. It argues that combining automated malware execution in isolated virtual machines with memory forensics and manual validation yields snapshots that capture behavioural and environmental features at the system-state level. If correct, the dataset lets models learn state transitions from clean to infected memory, supporting reinforcement-learning and adaptive incident-response research. The paper frames its contribution as a reproducible, ethically governed alternative to prior datasets that it says lack diversity, detailed labelling, or memory-level depth.","feed_headline":"30 RAM dumps pair clean and infected states for malware AI","feed_subtitle":"Labelled memory snapshots across 15 malware families and four Windows versions give detection models a state-change view.","key_machinery":"The central object is a paired memory snapshot: a clean RAM dump and an infected RAM dump of the same virtual machine, taken under standardised timing after controlled malware execution. The dataset is built by automated malware execution in an isolated, regularly reset virtualised environment, with snapshots later examined using established memory-forensics tooling and manual cross-checking. This pairing is what enables state-transition modelling: the clean snapshot defines the baseline state, the infected snapshot defines the post-infection state, and their difference supplies the behavioural and environmental features for machine-learning and reinforcement-learning models.","core_discovery":"The central claim is that MalVol-25 provides a diverse, labelled, and detailed volatile-memory resource for training and testing malware detection and response systems, where current datasets are limited in diversity, labelling, or detail. The authors construct it by pairing each malware sample with a clean baseline memory snapshot and a post-infection snapshot taken after the malware has had time to execute and manifest behaviour, across Windows 7, 8.1, 10, and 11. They argue that the resulting 30 dumps encode the state change caused by infection, and that automated forensic analysis and manual inspection confirm visible differences in processes, network patterns, and injected code between clean and infected states. The paper's stated significance is that system-state and transition modelling of this kind is what reinforcement-learning and agentic AI approaches need to make detection and response decisions.","pith_inferences":["The dataset's practical value will depend on per-sample verification artifacts (hashes, process lists, timestamps, infection logs) being published alongside the dumps; the paper currently asserts validation without showing them.","At 30 dumps and 15 families, the resource is more likely a testbed or benchmark seed than a training-scale corpus, and pairing it with existing larger datasets may be needed for deep-learning experiments.","If the infection-timing protocol is standardised and released as part of the toolkit, the dataset could be extended to continuous time-series snapshots, which would materially strengthen the claimed reinforcement-learning use case.","An independent audit using different memory-forensics tooling than the authors used would be a cheap and direct way to test label reliability and strengthen trust in the clean/infected pairing."],"forward_implications":["Researchers can train and benchmark memory-based malware detectors on a common, labelled set of clean and infected dumps instead of ad hoc collections.","The paired clean/infected structure supports reinforcement-learning formulations where system states and transitions are explicit, enabling detection and response strategies beyond static classification.","Multi-OS and multi-family coverage allows evaluation of whether detection generalises across Windows versions and malware categories such as ransomware, trojans, and worms.","The documented, integrity-checked release supports reproducible comparison of forensic and machine-learning pipelines, a step toward standard benchmarks in memory-level malware analysis."],"supporting_citations":[{"why":"Justifies the dataset's stated purpose of supporting reinforcement-learning and agentic AI for malware investigation and incident response.","marker":"[3]"},{"why":"Provides the prior single-device IoT malware behaviour dataset against which the paper measures its multi-OS diversity.","marker":"[5]"},{"why":"Provides the prior cloud malware dataset with static and dynamic features that the paper contrasts with memory-level snapshots.","marker":"[6]"},{"why":"Provides the prior execution-trace dataset with detailed labels whose missing memory-level detail the paper aims to supply.","marker":"[7]"},{"why":"Provides the prior sandbox-based behavioural dataset and feature-sharing platform that motivates the paper's reproducibility emphasis.","marker":"[11]"},{"why":"Provides the closest prior memory-snapshot dataset of compromised hosts, the baseline the paper's family and OS diversity extends.","marker":"[19]"}],"fun_headline_variants":["Memory dumps capture infection state change for malware AI","Clean and infected RAM snapshots power malware detection models","MalVol-25 dataset offers 30 labeled memory dumps for detection","Dataset pairs clean and infected memory for adaptive defenses","Volatile memory dataset enables state-aware malware detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's central premise is that each infected snapshot actually contains active, correctly labelled malware behaviour; the paper says malware was given time to execute and that validation was performed, but it provides no per-sample evidence such as malware hashes, process lists, or infection logs to verify that the right malware ran in each dump.","fun_headline_variants_meta":{"raw":{"variants":["Memory dumps capture infection state change for malware AI","Clean and infected RAM snapshots power malware detection models","MalVol-25 dataset offers 30 labeled memory dumps for detection","Dataset pairs clean and infected memory for adaptive defenses","Volatile memory dataset enables state-aware malware detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000137,"raw_usage":{"total_tokens":1116,"prompt_tokens":879,"completion_tokens":237,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":159}},"tokens_in":495,"tokens_out":237,"duration_ms":4796,"temperature":1.0,"reasoning_tokens":159,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:56:58.842426+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For each of the 15 named malware families, inspect the corresponding infected memory dump for that family's distinctive artifacts (for example, WannaCry's mutex or ransom-note remnants, Cerber's process names, or GandCrab's encryption traces); if a substantial share of claimed-infected dumps contains no trace of the named malware while the clean baseline does, the central claim of correctly labelled, diverse infection states collapses.","supporting_citations":[{"cited_title":"Creation of a Dataset Modeling the Behavior of Malware in IoT Devices","cited_arxiv_id":null,"evidence_quote":"Provides the prior single-device IoT malware behaviour dataset against which the paper measures its multi-OS diversity."},{"cited_title":"and Long, H.V","cited_arxiv_id":null,"evidence_quote":"Provides the prior cloud malware dataset with static and dynamic features that the paper contrasts with memory-level snapshots."},{"cited_title":"and Álvarez, P","cited_arxiv_id":null,"evidence_quote":"Provides the prior execution-trace dataset with detailed labels whose missing memory-level detail the paper aims to supply."},{"cited_title":"and Venter, H., 2023","cited_arxiv_id":null,"evidence_quote":"Provides the prior sandbox-based behavioural dataset and feature-sharing platform that motivates the paper's reproducibility emphasis."},{"cited_title":"and Binder, A.,","cited_arxiv_id":null,"evidence_quote":"Provides the closest prior memory-snapshot dataset of compromised hosts, the baseline the paper's family and OS diversity extends."}],"review_version":1}