{"id":"e6e2ad96-7443-4bbd-b7ab-42c1c3e38bbf","arxiv_id":"2508.19552","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new open-source simulator and a characterized 200TB, 25M-frame synthetic radio dataset with 100 modulation classes, intended to train AI models for spectrum sensing.","lead":"This paper introduces CSRD2025, a claimed 25 million frame synthetic radio dataset for training AI spectrum sensors, plus an open-source simulator that generates it. The resource is about 10,000 times larger than the common RadioML 2018 dataset and includes 100 modulation types with detailed labels. The dataset itself is not hosted for download; users must regenerate it with the provided code.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sim2Real bridging claim rests on uncalibrated simulation; no OTA/HIL transfer test is provided.","rationale":"I read the full manuscript and checked the internal quantitative claims. The scale arithmetic is plausible: 10M scenarios with 1–4 receivers yields ~25M frames, and ~200TB/25M ≈ 8MB per frame is consistent with 1M complex samples per frame on average. The framework description is detailed and the modular design is a genuine contribution. However, the central scientific value of CSRD2025 is that it is 'specifically engineered to bridge the Sim2Real gap.' That claim requires evidence that models trained on the synthetic data transfer to real or hardware-emulated RF. The paper offers no such evidence: Section III.B describes many realistic components, but none are calibrated against measurements; Section IV.A quantifies scale and diversity, not fidelity. Because the authors themselves identify the Sim2Real gap as the main limitation of synthetic data in Section II.B.1, the absence of any transfer evaluation is load-bearing. This is not an internal inconsistency or fabrication concern; it is a missing-validation concern. A small OTA/HIL cross-evaluation would settle it. I also considered the fact that the 200TB dataset is not directly downloadable, but the paper is explicit about this and provides a generator; reproducibility is important but secondary to whether the generated data actually serves real spectrum sensing. The reader's conditional verdict is appropriate and I would not change it.","tokens_in":16661,"tokens_out":6633,"duration_ms":69304,"concrete_test":"Generate a small labeled OTA or HIL benchmark (e.g., 5 modulations from the 100 classes: QPSK, 16QAM, OFDM, GMSK, FM; 5–10 SNRs; USRP captures or Colosseum channel emulator with wired ground truth). Train a standard detector/classifier only on CSRD2025 and test on this benchmark; compare against the same model trained on a matched small real/HIL set. If CSRD-trained performance is near chance or falls by more than ~20% absolute mAP/accuracy relative to real-trained performance, the Sim2Real bridging claim fails. The test can be run with the open-source generator and a few hours of RF capture.","verdict_should_be":"UNCHANGED","load_bearing_attack":"CSRD2025's stated purpose is to bridge the Sim2Real gap and provide high-fidelity data for spectrum-sensing LAMs. For that central claim to hold, the simulated channel and impairment distributions in Section III.B must produce signals whose discriminative features match real over-the-air captures. Nothing in the paper tests this: no OTA or HIL recordings, no comparison against real-world subsets of datasets like RadioML or SPREAD, no calibration of the PA/phase-noise/IQ-imbalance ranges in Table II to specific hardware, and no path-loss validation of the OSM ray-tracing module. The tenfold larger scale and 100 classes would not help if models learn MATLAB-specific artifacts rather than physical signal structure. This is a missing-evidence rather than internal-inconsistency problem: the internal arithmetic is consistent (10M scenarios × ~2.5 receivers ≈ 25M frames; ~200TB/25M ≈ 8MB/frame), but fidelity is asserted, not demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ChangShuoRadioData (CSRD), a modular MATLAB-based simulation framework for generating synthetic RF data, and presents CSRD2025, a benchmark instance claimed to contain over 25,000,000 recorded frames (~200 TB) of passband IQ data from about 10,000,000 simulation scenarios. The dataset is claimed to cover 100 modulation classes, statistical and OSM-based ray-tracing channel models, diverse transmitter/receiver RF impairments, COCO-format spectrogram annotations, and standardized 8:1:1 train/validation/test splits. The authors position CSRD2025 as a Stage-1 synthetic resource engineered to 'bridge the Sim2Real gap' for spectrum-sensing large AI models. The paper contains no model training or transfer experiments; its central evidence is the described generation pipeline and internal distributional statistics.","tokens_in":16854,"tokens_out":6377,"duration_ms":72250,"significance":"If the claims are substantiated, this could be a valuable community resource: the stated scale (≈200 TB, 25M frames) is orders of magnitude larger than RadioML 2018.01A, and the combination of 100 modulation classes, passband signal representation, impairment modeling, OSM ray tracing, and COCO annotations could support a range of spectrum-sensing tasks. The framework's config-driven design, use of SigMF-like metadata, and fixed frame-index splits are strengths for reproducibility. However, the significance currently rests on two unvalidated pillars: (1) the 'high fidelity' and 'Sim2Real gap' claims are asserted on the basis of simulation components rather than demonstrated against over-the-air or hardware-in-the-loop data, and (2) the dataset itself is not hosted, so the central artifact cannot be independently inspected. The paper therefore reads as a promising resource description rather than a demonstrated benchmark.","major_comments":[{"comment":"The abstract and Section III.B claim that CSRD2025 is 'specifically engineered to bridge the Sim2Real gap' and that the pipeline provides 'high-fidelity' data. No experiment in the paper tests transfer to real radio environments: there are no OTA recordings, no hardware-in-the-loop measurements, no calibration of the impairment ranges in Table II to specific hardware, and no evaluation of models trained on CSRD2025 against real-world datasets such as RadioML 2018.01A or SPREAD. The internal arithmetic is consistent, but the load-bearing premise that synthetic features match physical signal structure is unvalidated. This is the central claim of the paper, so it should be supported, or the claims should be scaled back to 'intended to bridge the Sim2Real gap'.","section":"§III.B and §IV.A (Abstract)"},{"comment":"The paper describes CSRD2025 as a large-scale dataset benchmark, but footnote 1 states that the 200 TB dataset is not hosted for direct download and can only be fully reproduced using the framework, configurations, and fixed random seeds. No commit hash, version tag, checksum manifest, or exact generator environment is given, and no representative subset is provided. This makes the central artifact inaccessible for independent verification or immediate use. A dataset paper should either host a usable subset (e.g., 100 GB–1 TB) or provide a precise, versioned, executable recipe plus a manifest of emitted files, so readers can reproduce the corpus and confirm the claimed 25M-frame/200TB statistics.","section":"Footnote 1, §IV.F"},{"comment":"There is an unresolved inconsistency between the generation parameters and the reported signal-duration distribution. Table II states a symbol rate of 30–50 kHz and 500–2000 symbols per segment, which gives a maximum per-segment duration of about 0.067 s. Figure 10, however, reports signal durations extending to roughly 0.8 s and a long tail beyond that. If 'signal duration' includes multiple segments with idle intervals or is measured differently from 'segment duration', that needs to be stated precisely; otherwise the reader cannot tell which parameter bounds actually generated the data. This matters because the variable frame length and burst structure are advertised as key features.","section":"Table II and Fig. 10 (§IV.E)"},{"comment":"Several dataset statistics, especially in Fig. 10 and the accompanying text, are said to be derived from a 'representative sample' or 'sample data' without specifying the sample size, the sampling procedure, or the fraction of the 25M frames covered. Since the full dataset is not publicly accessible, these summary statistics are the only evidence of the corpus's internal distribution. For a benchmark paper, the authors should either report full-corpus statistics or provide a rigorous sampling plan with confidence intervals. I also note the absence of any proof-of-concept evaluation (e.g., AMC accuracy, object-detection mAP on the COCO annotations) that would demonstrate label correctness and practical usability of the dataset.","section":"§IV.E"}],"minor_comments":[{"comment":"The text says OSM map files for '25 distinct geographic locations' are used, then mentions 'downloading 10 representative 2km x 2km examples for each category' with nine named categories. This is inconsistent: 9 categories × 10 examples = 90 files, not 25. Clarify the number of environments and examples actually included.","section":"§III.B.4 (Ray tracing)"},{"comment":"The example metadata shows a MasterClockRate of 1.11e6 Hz while the dataset includes signal bandwidths up to ~800 kHz. Without down-conversion or band-pass filtering details, this seems to violate Nyquist. The authors should comment on the relationship between master clock, carrier frequency, and reported bandwidth in the metadata example.","section":"§IV.C, Listing 1"},{"comment":"The phrase 'open-source framework' should be qualified: the implementation uses MATLAB and the Communications Toolbox, so reproducibility depends on commercially licensed software. State the exact dependencies, toolbox versions, and a license for the framework code. The author footnote describing an equal contribution as 'mainly for supplying the MATLAB license' is not appropriate for a scientific paper and should be removed.","section":"Data and Code Access"},{"comment":"The COCO annotation pipeline is not fully specified: the STFT window, FFT size, hop length, and any scaling/gain normalization need to be stated so that bounding-box coordinates and spectrogram image sizes can be reproduced exactly by third parties.","section":"§IV.D"},{"comment":"There are minor presentation issues: 'Barcelon `es' in the Fig. 5 caption, the Fig. 7 axis labels are difficult to read, and there is no complete machine-readable list of the 100 modulation classes apart from the figure. Please include a table of all classes in an appendix and fix the typographical issues.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the paper is a resource paper whose stated value depends on both scale and realism. The scale is internally plausible, but the realism/Sim2Real claim is currently unsupported by any external validation, and the dataset itself is not downloadable. These are not unfixable, but they require substantial additional work (transfer experiments, versioned code, and at least partial data access) before the resource can credibly serve as a benchmark. If the authors are unable to provide transfer evidence, the paper should be reframed as a generator description rather than a Sim2Real-bridging dataset. I see no circularity or internal-fabrication issue; the main risk is overclaiming from unvalidated simulation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a resource paper, not a results paper. The authors release an open-source MATLAB-based generator and describe CSRD2025, a synthetic dataset benchmark with ~25 million frames (~200 TB), 100 modulation classes, variable frame lengths, ray-tracing from OSM data, and COCO-format spectrogram annotations. If the scale is real, it is the largest public synthetic RF dataset by a wide margin—roughly 10,000x RadioML 2018.01A in volume, and much broader in class count and scenario diversity. The framework is largely assembled from existing MATLAB Communications Toolbox components, but the configuration-driven modular design, SigMF-like metadata, and standardized 8:1:1 splits are useful and the paper's internal arithmetic checks out.\n\nWhat the paper does well: it lays out a reproducible pipeline (modulation library, impairments, channel models, interference scheduling, annotation conversion) and gives honest statistics about class imbalance, SNR distribution, and signal durations. The comparison table with SPREAD, TorchSig, RadioML, etc. is helpful for positioning.\n\nNow the soft spots. The dataset itself is not downloadable; you get the generator and a promise of fixed random seeds. No commit hash is pinned, so the exact configuration used to produce CSRD2025 is not independently verifiable. More importantly, the headline claim—that the data \"bridges the Sim2Real gap\"—rests entirely on unvalidated domain assumptions. There is no OTA or HIL transfer experiment, no comparison against real-world captures, no calibration of impairment parameter ranges to specific hardware, and no path-loss validation of the ray-tracing module. The authors are explicit about the three-stage strategy (synthetic -> HIL -> real) and position CSRD2025 as Stage 1, which is honest, but they then repeatedly call the data \"high-fidelity\" and say it is \"engineered to bridge the Sim2Real gap.\" That is overreach. A reader who trains on CSRD2025 could be learning MATLAB artifacts rather than physical signal structure. Also, the \"10,000 times larger than RadioML\" framing is apples-to-oranges (passband vs baseband, variable vs fixed frame length), though Table I mostly saves them on that.\n\nThe bottom line: this is a useful contribution for anyone who wants a configurable synthetic RF data generator and a target specification for large-scale spectrum-sensing benchmarks. It is not a validated path to real-world performance. The paper should go to peer review, but the referee should require a pinned code release, a small sample of generated data or checksums, and either at least one transfer experiment with real or HIL data or a toned-down claim. My own verdict would be conditional, not reject.\n\nI'd bring it to a reading group if the discussion is about dataset engineering and Sim2Real evaluation. I'd cite it if I needed a reference for a large synthetic spectrum-sensing resource, while flagging the unvalidated fidelity.","headline":"A genuinely large and well-specified synthetic RF dataset plus an open-source generator, but the paper's central Sim2Real bridging claim is asserted rather than tested—worth refereeing, not yet worth believing.","tokens_in":17369,"tokens_out":1603,"would_cite":true,"duration_ms":20810,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces an open-source simulation framework and a 25-million-frame, ~200 TB synthetic radio dataset with 100 modulation classes, designed to supply the scale that large AI models need for spectrum sensing.","keywords":["synthetic radio dataset","spectrum sensing","large AI models","modulation classification","ray tracing channel model","RF impairments","COCO object detection","signal metadata"],"falsifier":"Take a fixed test set of real over-the-air recordings with known modulation types and SNRs, then train a detector and classifier on a CSRD2025 subset. If accuracy does not improve with dataset size, or is no better than the same model trained on a far smaller existing synthetic benchmark, the scale-driven Sim2Real claim fails. A minimal check would compare per-class precision on real captures against performance on held-out simulated frames.","tokens_in":1516,"feed_emoji":"📡","tokens_out":1573,"duration_ms":68308,"temperature":0.7,"pith_summary":"This paper tries to remove the data bottleneck for training large AI models in wireless spectrum sensing by offering a way to generate synthetic radio data at unprecedented scale. It introduces an open-source, modular simulator that models the full transmit-receive chain: 100 modulation types, statistical and ray-traced channels from real map data, transmitter and receiver impairments, and multi-antenna links. Using it, the authors characterize CSRD2025, roughly 25 million frames and 200 TB of passband IQ data, about 10,000 times larger than the 2018 benchmark, with per-signal ground-truth metadata and COCO-format spectrogram annotations for object detection. If the simulation realism holds, this would give spectrum-sensing researchers a training resource comparable in scale to the image and text corpora behind large models.","feed_headline":"25 million synthetic radio frames aim to feed spectrum-sensing AI","feed_subtitle":"A 200 TB passband dataset with 100 modulation classes targets the scale gap behind large wireless AI models.","key_machinery":"The load-bearing mechanism is the CSRD simulator's modular frame-by-frame pipeline: an engine samples scenario parameters from JSON configuration files, instantiates modulators, event scheduling, transmitter and receiver impairment models, and either statistical or ray-traced channels, then archives IQ plus metadata. The two features that do the real work for scale are the stochastic event controller that places one to three signal segments per transmitter with controlled spectral overlap, and the perfect ground-truth generator that derives COCO annotations from simulation parameters, making supervised time-frequency detection possible without hand labeling.","core_discovery":"The central discovery is a dataset, not a theorem: synthetic RF data can be manufactured at the scale and diversity required by scaling laws. CSRD2025 contains 100 modulation classes—analog, single-carrier, multi-carrier OFDM/SCFDMA, and OTFS—with frame lengths varying from 2,000 to 4,000,000 samples and one to four transmitters and receivers; 10% of scenarios are ray-traced through map-derived 3D environments. Every frame carries exhaustive JSON ground truth (modulation, timing, bandwidth, channel, impairments, per-signal SNR), and a conversion pipeline turns IQ data into spectrograms with COCO bounding boxes. The claimed payoff is that this is the first benchmark large enough and diverse e","pith_inferences":["Editorial inference: If the simulation fidelity holds, CSRD2025 could support pre-training a spectrum foundation model whose representations transfer to narrow real-world sensing tasks with only a small amount of labeled over-the-air data—the endgame of the authors' three-stage strategy, which they leave implicit.","Editorial inference: The tiered class distribution, with abundant common modulations and sparse high-order and OTFS variants, may create a long-tail learning problem; models may need class-balanced sampling or augmentation to avoid bias, a property worth testing directly.","Editorial inference: The COCO annotations could also be used to benchmark time-frequency segmentation or few-shot novel-class detection, tasks the paper does not explore but that the same labels would support.","Editorial inference: A direct testable extension is to measure accuracy on a small real-world radio capture benchmark after training on CSRD2025 subsets; the Sim2Real claim would be supported if performance degrades gracefully as channel and impairment fidelity are reduced in ablation."],"forward_implications":["Training and evaluation of spectrum-sensing models can move from tens of gigabytes to hundreds of terabytes, approaching the dataset sizes that large-model scaling laws call for.","Object-detection models can be applied directly to spectrograms for joint time-frequency localization and modulation classification using the provided COCO annotations.","The standardized 8:1:1 frame-index splits make results across different studies directly comparable.","Users can generate custom datasets of arbitrary scale and parameter distributions by editing the configuration files, so the framework extends beyond the specific CSRD2025 instance.","The 25 ray-traced environment types provide site-specific propagation diversity that statistical fading alone cannot capture."],"supporting_citations":[{"why":"Supplies the scaling-law argument that dataset size must grow with model size, motivating the massive scale of the proposed dataset.","marker":"[4]"},{"why":"The 18 GB, 24-class benchmark whose scale and baseband format are the main comparison point for CSRD2025.","marker":"[30]"},{"why":"A passband spectrum-monitoring dataset used in the comparison table to position CSRD2025's scale and frame count.","marker":"[32]"},{"why":"A wideband passband benchmark used in the comparison table for scale and frame-length context.","marker":"[36]"},{"why":"Defines the signal metadata standard that the CSRD JSON annotation structure mirrors.","marker":"[18]"},{"why":"Provides the COCO annotation format that the spectrogram bounding-box labels follow.","marker":"[19]"},{"why":"An early synthetic radio benchmark cited as the foundation for automatic modulation classification and as a scale baseline for later datasets.","marker":"[22]"},{"why":"A hardware-in-the-loop platform used to motivate the three-stage data-construction strategy and the Sim2Real gap.","marker":"[29]"},{"why":"A virtual-signal large-model work that CSRD extends toward large-scale spectrum-sensing data generation.","marker":"[26]"},{"why":"A later improved dataset in the same lineage, used to compare scale and realism across benchmarks.","marker":"[38]"}],"fun_headline_variants":["Synthetic radio dataset hits 25M frames for AI sensing","200TB radio data: 100 modulations, one dataset","Massive synthetic RF dataset targets spectrum-sensing AI","CSRD2025: 25M synthetic frames to train wireless AI","Radio AI gets 10,000x bigger synthetic training set"],"cache_read_input_tokens":19200,"weakest_assumption_plain":"The whole Sim2Real value rests on the assumption that the simulated channels, RF impairments, and SNR calculations faithfully reproduce real over-the-air signal behavior; the paper provides no measurement collected from real hardware or over the air to test that transfer.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic radio dataset hits 25M frames for AI sensing","200TB radio data: 100 modulations, one dataset","Massive synthetic RF dataset targets spectrum-sensing AI","CSRD2025: 25M synthetic frames to train wireless AI","Radio AI gets 10,000x bigger synthetic training set"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000646,"raw_usage":{"total_tokens":2850,"prompt_tokens":835,"completion_tokens":2015,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":1929}},"tokens_in":579,"tokens_out":2015,"duration_ms":14572,"temperature":1.0,"reasoning_tokens":1929,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:40:54.395182+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fixed test set of real over-the-air recordings with known modulation types and SNRs, then train a detector and classifier on a CSRD2025 subset. If accuracy does not improve with dataset size, or is no better than the same model trained on a far smaller existing synthetic benchmark, the scale-driven Sim2Real claim fails. A minimal check would compare per-class precision on real captures against performance on held-out simulated frames.","supporting_citations":[{"cited_title":"Over-the-air deep learning based radio signal classification,","cited_arxiv_id":null,"evidence_quote":"The 18 GB, 24-class benchmark whose scale and baseband format are the main comparison point for CSRD2025."},{"cited_title":"Spectrogram data set for deep-learning-based rf frame detection,","cited_arxiv_id":null,"evidence_quote":"A passband spectrum-monitoring dataset used in the comparison table to position CSRD2025's scale and frame count."},{"cited_title":"Large Scale Radio Frequency Wideband Signal Detection & Recognition","cited_arxiv_id":"2211.10335","evidence_quote":"A wideband passband benchmark used in the comparison table for scale and frame-length context."},{"cited_title":"Sigmf: The signal metadata format,","cited_arxiv_id":null,"evidence_quote":"Defines the signal metadata standard that the CSRD JSON annotation structure mirrors."},{"cited_title":"Microsoft coco: Common objects in context,","cited_arxiv_id":null,"evidence_quote":"Provides the COCO annotation format that the spectrogram bounding-box labels follow."},{"cited_title":"Learning robust general radio signal detection using computer vision methods,","cited_arxiv_id":null,"evidence_quote":"An early synthetic radio benchmark cited as the foundation for automatic modulation classification and as a scale baseline for later datasets."},{"cited_title":"Colosseum: Large-scale wireless exper- imentation through hardware-in-the-loop network emulation,","cited_arxiv_id":null,"evidence_quote":"A hardware-in-the-loop platform used to motivate the three-stage data-construction strategy and the Sim2Real gap."},{"cited_title":"Vslm: Virtual signal large model for few-shot wideband signal detection and recognition,","cited_arxiv_id":null,"evidence_quote":"A virtual-signal large-model work that CSRD extends toward large-scale spectrum-sensing data generation."},{"cited_title":"Rml22: Realistic dataset generation for wireless modulation classification,","cited_arxiv_id":null,"evidence_quote":"A later improved dataset in the same lineage, used to compare scale and realism across benchmarks."}],"review_version":1}