{"id":"62430434-8763-418f-a668-cc4acbb1d201","arxiv_id":"2506.07799","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A physics-embedded neural network, fed with matched-filter images from a cooperative ISAC network, detects off-grid drones at a simulated 97.55% detection rate.","lead":"This paper proposes using a network of cellular base stations as a passive imager that reconstructs a 3D map of drones from scattered wireless signals. A neural network cleans up a coarse physical projection of the signals, and in simulation it detects 97.55% of drones while producing few false alarms.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 97.55% detection rate is not reproducible because the paper never specifies how the continuous DNN output is thresholded or otherwise converted into binary detections; the DR/FAR/OSPA values in Tables III-IV and Eq. (25) are therefore ambiguous.","rationale":"The paper is a simulation-based algorithmic study with a plausible architecture (matched-filter initialization plus residual CNN and OHEM losses) and extensive ablations, including a Sionna ray-tracing test. The reader's conditional verdict is well founded. My concern is more specific than the reader's: before worrying about real-world multipath or extended targets, the reported detection statistics are not reproducible from the manuscript because the binary decision rule is absent. This directly undermines the abstract's key number, independent of model realism. The OSPA inconsistency (Eq. (25) vs. Table III's OSPA=50 for an empty estimate) suggests that the evaluation code implements something different from the published metric definition, reinforcing the need for code release. I also note a possible forward-model mismatch between the continuous integral in Eq. (7) and its discretization in Eqs. (10)-(11), where the path-loss exponent appears to change; this deserves a separate check, but the detection-rule gap is more load-bearing because it makes the headline claim undefined. If the authors specify the decision rule and confirm the numbers are robust to reasonable thresholds, the conditional acceptance can be upgraded.","tokens_in":23139,"tokens_out":9373,"duration_ms":111573,"concrete_test":"Obtain the released code (github.com/kiwi1944/LAEImager) and locate the evaluation routine that computes DR, FAR, and OSPA for Tables III-IV. Enumerate the exact rule used to declare a detection (e.g., threshold on the output image, or selection of the top-M voxels). Then re-run Net-10 on the 10,000 test images under two alternative rules: (i) a single global threshold set to 0.5 times the maximum output value, and (ii) selection of the M-hat highest-magnitude voxels with M-hat taken from the OSPA definition. If the resulting DR/FAR differ materially from 97.55% / 3.22%, the headline claim is threshold-dependent. Also recompute OSPA for an all-zero output using Eq. (25) with c_3=1; if it is not 1, the formula in the paper differs from the code used to produce the tables.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a detection rate, but the manuscript never defines the decision rule that converts the DNN's continuous output into binary detections. Section V-A defines DR as 'the proportion of correctly identified targets in the reconstructed image' and FAR as 'the proportion of falsely detected targets that do not exist in the ground truth,' without specifying a threshold or a top-M selection. The DNN in Sec. IV-B is a regression network; its output is not sparse by construction (except when the L1 term in Eq. (22) drives it to zero, as in Net-4/Net-6). Table III reports DNN-y with DR=0 and OSPA=50, and Table IV reports Net-10 with DR=97.55% and FAR=3.22%; these numbers depend entirely on how a detection is declared. The OSPA formula in Eq. (25) with c_3=1 predicts an OSPA of at most 1 for an all-zero output, yet the tables report OSPA=50, indicating that the implemented metric or its parameters differ from the text. Without the evaluation code or an explicit decision rule, the headline improvement over SP (46.52% DR) cannot be verified, and the comparison across methods may be computed under different detection conventions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper addresses low-altitude UAV surveillance by formulating it as a compressed-sensing-based imaging problem in a cooperative ISAC network. The authors derive a linear forward model relating CSI measurements to scattering coefficients of voxels in a region of interest, analyze sensing capability via the point spread function, and propose a two-stage physics-embedded learning approach: a matched-filter initial estimate A^H y followed by a CNN residual refinement trained with OHEM-based loss functions. The paper reports extensive 2D and 3D simulation results, including a Sionna ray-tracing urban canyon scenario, and claims a 97.55% detection rate, outperforming conventional CS (SP) and black-box DNN baselines under off-grid conditions.","tokens_in":23375,"tokens_out":6303,"duration_ms":65043,"significance":"If the quantitative claims are reliable, the paper makes a useful contribution by connecting ISAC-based cellular sensing to sparse imaging and by showing that a hybrid model-based/learning approach can mitigate off-grid errors. The PSF analysis provides practical configuration guidelines, and the Sionna experiment is a step toward realistic validation. The OHEM loss design is a sensible adaptation to the extreme sparsity of aerial images. However, the evaluation protocol is not sufficiently specified to support the headline numbers, and several internal inconsistencies undermine confidence in the reported metrics.","major_comments":[{"comment":"The manuscript defines DR and FAR only verbally and never specifies the decision rule that maps the continuous DNN output to binary detections. Since the network is a regression model, the reported DR=97.55% and FAR=3.22% depend entirely on an unspecified thresholding or selection procedure. Moreover, Eq. (25) with c3=1 gives OSPA=1 for an all-zero output, but Table III (DNN-y) and Table IV (Net-4, Net-6) report OSPA=50, indicating that the implemented metric deviates from the text. This makes the headline quantitative claims ambiguous and not reproducible.","section":"Section V-A and Eq. (25), Tables III and IV"},{"comment":"The configuration behind the 97.55% DR claim is Net-10 (OHEM-2, α=1), but the negative sample ratio η used for this network is not reported. Section V-C3 demonstrates that η strongly affects DR (e.g., for OHEM-2, DR saturates only for η≥25), so the headline result is configuration-specific and cannot be reproduced from the information given.","section":"Section V-C2, Table IV"},{"comment":"The text states that 'Perfect image reconstruction is achieved at d0=3m' (Fig. 8(f)), but Table II reports for Mode A at d0=3m a DR of 96.47% and a FAR of 12.60%. These numbers contradict the notion of perfect reconstruction; the statement should be corrected or the discrepancy explained.","section":"Section V-B2, Table II and Fig. 8"},{"comment":"The abstract's claim of 'significantly outperforming traditional CS-based methods' is not uniformly supported by the metrics. Net-10, the configuration with the 97.55% DR, has an OSPA of 29.94, which is worse than the SP algorithm's OSPA of 27.75. The paper should either clarify that the superiority claim refers only to DR/FAR, or report results for the configuration that also improves OSPA (e.g., Net-9).","section":"Section V-C1, Tables III and IV"},{"comment":"The forward model assumes point scatterers, a single-bounce LOS path, perfect BS synchronization, and calibration-based removal of static background, and the Sionna experiment only injects residual interference without a baseline comparison. The claims of 'all-weather, around-the-clock sensing' and practical applicability are stronger than what the simulation validation can support. The authors should temper these claims or provide additional validation (e.g., comparison against a conventional localization baseline in the Sionna scenario).","section":"Sections II-C and V-C5"}],"minor_comments":[{"comment":"The DNN architecture description (six residual blocks, channel list [64,128,128,128,64,32]) omits kernel sizes, activation functions, normalization layers, and batch size, which are needed for reproducibility.","section":"Section V-C2"},{"comment":"The abstract and several places use 'UA Vs' with a space; the standard 'UAVs' is used elsewhere in the text (e.g., Section I).","section":"Abstract and throughout"},{"comment":"Reference [4] misspells 'Available' as 'Avilable', and references [5] and [6] misspell 'communication' as 'comunication'.","section":"References"},{"comment":"The meaning of the '/' entries for the noise power of Datasets 5 and 6 is unclear; please state the parameter settings explicitly.","section":"Section V-C4 and Table V"},{"comment":"The statement that 'Part of the source code ... will be soon accessed' should be updated to reflect the actual availability status, as the code repository is not accessible at the time of review.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper comes from a group with a strong track record and the core idea is interesting. However, the evaluation metrics need to be pinned down. I recommend asking the authors for either the evaluation code or a detailed description of the detection rule, and for clarification of the OSPA implementation, before the paper can be accepted. The OSPA inconsistency in particular suggests that the reported numbers may have been produced by a different metric than the one stated in Eq. (25)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real methodological combination, not a repackaging. Feeding the fixed matched-filter projection A^H y into a residual CNN trained with an OHEM-based loss, for cooperative multi-BS ISAC imaging, is new relative to the literature they cite. The PSF analysis of antenna, subcarrier, and voxel configurations is a practical addition, and the experimental campaign is extensive: 100k training samples, multiple baselines, and a Sionna ray-tracing 3D case that goes beyond pure line-of-sight simulation. I think the central claim is plausible: the proposed physics-embedded imager does substantially outperform SP and black-box DNNs under off-grid conditions. The code is promised but not yet available, which matters.\n\nThe soft spots, in decreasing order. First, the headline 97.55% DR is not reproducible as reported. The paper never states the decision rule that converts the DNN's continuous output into binary detections. Section V-A defines DR and FAR verbally, with no threshold and no top-M selection rule. The OHEM-2 hyperparameters for Net-10, notably eta, are not reported even though Section V-C3 shows eta strongly controls the DR/FAR tradeoff. Second, the OSPA numbers are internally inconsistent with Eq. (25): with c3=1, an all-zero output should give OSPA = 1, yet DNN-y is reported at 50. Either the OSPA cutoff or normalization in the text differs from the implementation, and that undermines the cross-method comparisons. Third, the paper claims \"perfect image reconstruction\" at d0=3m while Table II lists FAR=12.60% for that same row. That is an overstatement. Fourth, the work is simulation-only and rests on point-scatterer, single-bounce LOS, synchronized BSs, and calibrated background removal; the Sionna experiment relaxes LOS partially but has no baseline comparison and only varies residual interference.\n\nThe citation pattern looks standard for the area, and the papers are placed appropriately. This is clearly serious work by people who understand the problem. It deserves a serious referee: it is valuable for the ISAC/CS imaging community, but the authors should be required to release code, specify the detection rule and all hyperparameters, fix the OSPA inconsistency, and soften the \"perfect\" language before I would rely on the numeric claims.","headline":"A genuinely new physics-embedded pipeline for cooperative ISAC off-grid imaging, backed by extensive simulation, but the headline 97.55% DR is not yet reproducible because the detection rule and key hyperparameters are unspecified.","tokens_in":663,"tokens_out":2357,"would_cite":true,"duration_ms":59669,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that low-altitude drones can be imaged by a cooperative cellular network that treats the airspace as a sparse 3D image, and that a matched-filter plus residual-CNN pipeline, trained with a loss focused on the hardest…","keywords":["low-altitude economy","integrated sensing and communication","compressed sensing imaging","off-grid error","UAV surveillance","physics-embedded deep learning","online hard example mining","point spread function"],"falsifier":"Run the same trained network on measured CSI from two or more synchronized base stations in a built-up area, with a cooperative drone carrying a GPS receiver to provide ground truth, and form the matched-filter input from the measured channel; if the detection rate on off-grid positions falls toward the subspace-pursuit baseline or the false alarm rate rises sharply as multipath and extended-target effects appear, the central claim fails.","tokens_in":22906,"feed_emoji":"📡","tokens_out":9402,"duration_ms":103532,"temperature":0.7,"pith_summary":"Commercial drone flights below 300 meters need surveillance, and this paper proposes getting it from the cellular network itself: several base stations cooperate, transmit sensing signals, and reconstruct a 3D image of the airspace from raw channel measurements. The central claim is that the off-grid mismatch, which occurs because drones rarely sit exactly on the predefined image grid, can be repaired by a two-stage physics-embedded network: first project the measurements onto the grid with the matched filter $\\mathbf{A}^H\\mathbf{y}$, then let a residual convolutional network refine that coarse image. The paper also claims that a loss based on online hard example mining, which makes training focus on the few target voxels and the worst-predicted empty voxels, is what prevents the network from collapsing to an all-zero image; with it, the simulated system reaches a 97.55% detection rate at a 3.22% false alarm rate. If true, this gives an all-weather, around-the-clock surveillance option that needs no extra radar hardware and avoids the error propagation and data-association problems of two-step localization.","feed_headline":"Cooperative cell towers detect off-grid drones 97.55% of the time","feed_subtitle":"A matched-filter backprojection plus a hard-example-trained network turns ordinary cellular signals into 3D airspace images.","key_machinery":"The load-bearing object is the sensing matrix $\\mathbf{A}$, whose columns are the steering vectors (discretized bistatic channel responses) of all voxels; the matched-filter initialization $\\hat{\\boldsymbol{\\sigma}}_{\\mathrm{pri}}=\\mathbf{A}^H\\mathbf{y}$ in Eq. (20) carries the physics into the network, giving the CNN a geometrically meaningful input rather than raw CSI. The refinement is a residual CNN made of convolutional blocks with skip connections. The third piece is the OHEM-based loss in Eq. (21): positive samples are the few nonzero target voxels, negative samples are empty voxels, and only the $\\eta M$ negative voxels with largest prediction error are kept; the OHEM-2 variant normalizes the two groups separately so that increasing $\\eta$ does not drown out the targets. The sparsity penalty $\\alpha\\|\\hat{\\boldsymbol{\\sigma}}\\|_1$ in Eq. (22) keeps the output sparse. The paper also uses the point spread function as a target-independent design metric: lower maximum PSF sidelobes predict better voxel discrimination, and the metric is used to justify sparse antenna arrays, larger bandwidths, and intermediate voxel sizes.","core_discovery":"The discovery the authors are trying to establish is that off-grid drones can be imaged directly from raw channel state information without first localizing or associating targets. The paper models each base-station pair's received signal as a linear combination $\\mathbf{y}=\\mathbf{A}\\boldsymbol{\\sigma}+\\mathbf{z}$ of voxel scattering coefficients, derives the point spread function $\\mathrm{PSF}(n_1,n_2)=|\\langle \\mathbf{A}(:,n_1),\\mathbf{A}(:,n_2)\\rangle|/(\\|\\mathbf{A}(:,n_1)\\|_2\\|\\mathbf{A}(:,n_2)\\|_2)$ to guide antenna, bandwidth, and voxel choices, and then shows that the off-grid error breaking the on-grid model can be absorbed by a learned refinement. The proposed imager first computes $\\hat{\\boldsymbol{\\sigma}}_{\\mathrm{pri}}=\\mathbf{A}^H\\mathbf{y}$, a projection that costs little and keeps all candidate information, and then trains a residual CNN to map that projection to the true image. With the OHEM-2 loss, which normalizes positive and negative sample losses separately and adds a sparsity penalty $\\alpha\\|\\hat{\\boldsymbol{\\sigma}}\\|_1$, the network reaches 97.55% detection rate and 3.22% false alarm rate in the simulated off-grid test set, clearly exceeding subspace pursuit, the raw projection, a black-box DNN fed with $\\mathbf{y}$, and the same DNN fed with subspace-pursuit output.","pith_inferences":["Editorial extension: the paper's geometry is static; a natural next step, which the authors list as future work, is to feed successive frames into a tracker and replace detection with track-before-detect, which should raise detection rate for small-RCS drones.","Editorial extension: a field experiment with a cooperative GPS-equipped drone would be the decisive test; the paper's adaptability experiments only vary noise power and UAV count in simulation, not rich multipath or synchronization error.","Editorial extension: because the matched-filter input $\\mathbf{A}^H\\mathbf{y}$ is cheap and the network only refines it, one could retrain the CNN quickly for a new base-station geometry by regenerating $\\mathbf{A}$ and keeping the same architecture, making the method a candidate for site-specific calibration.","Editorial extension: OHEM-2's separate normalization should generalize to any extremely sparse regression task with rare positives, such as radar point clouds and sparse channel estimation, though the paper only demonstrates it on imaging."],"forward_implications":["If the central claim is right, a cooperative ISAC network can monitor airspace without extra radar hardware, using ordinary cellular transmissions and the existing backhaul to fuse all base-station measurements.","The joint monostatic-plus-multistatic mode, where every base station hears every transmitting base station, gives the best image quality and is worth standardizing for sensing.","System designers can use the PSF sidelobe metric to choose antenna spacing, subcarrier count, bandwidth, and voxel size before deployment, trading resolution against reconstruction accuracy.","The OHEM-2 training recipe, with a tuned negative-sample ratio $\\eta$ and sparsity weight $\\alpha$, should be the default for sparse-image networks where positive voxels are rare.","The same two-stage recipe (matched-filter projection followed by learned refinement) is claimed to transfer to other compressed-sensing off-grid problems such as channel estimation."],"supporting_citations":[{"why":"Supplies the subspace-pursuit algorithm used as the on-grid CS baseline and as the input alternative that the proposed method outperforms.","marker":"[24]"},{"why":"Provides the cooperative multi-BS sensing model and least-squares channel estimation from which the paper builds its measurement equation.","marker":"[11]"},{"why":"Supplies the ISAC beamforming and coverage design that separates communication from sensing and orients beams at the ROI.","marker":"[12]"},{"why":"Establishes the on-grid CS imaging formulation with a sparse scattering coefficient vector and defines the discrete sensing matrix.","marker":"[25]"},{"why":"Introduces online hard example mining, the computer-vision technique the paper adapts into its OHEM-1 and OHEM-2 losses.","marker":"[43]"},{"why":"Supports the integral bistatic channel model in Eq. (7) and the use of point-spread-function-style analysis for imaging.","marker":"[49]"},{"why":"Exemplifies the learning-after-physics-processing pipeline that the paper's matched-filter-plus-DNN structure follows.","marker":"[38]"},{"why":"Provides the taxonomy of physics-embedded learning approaches that frames the proposed two-stage design.","marker":"[36]"}],"fun_headline_variants":["AI and cell towers spot off-grid drones 97.55% of the time","Cellular signals turn into 3D drone airspace images","Off-grid drone detection hits 97.55% with cooperative network","Physics-embedded AI images drones from raw cell data","Rare UAVs found 97.55% via hard-example trained network"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole method depends on each drone being a single point reflector seen through one clean line-of-sight path, with perfectly synchronized base stations and all static echoes and background scattering removed by calibration; if any of those fail, the linear equation the network learns to invert is no longer the right model.","fun_headline_variants_meta":{"raw":{"variants":["AI and cell towers spot off-grid drones 97.55% of the time","Cellular signals turn into 3D drone airspace images","Off-grid drone detection hits 97.55% with cooperative network","Physics-embedded AI images drones from raw cell data","Rare UAVs found 97.55% via hard-example trained network"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00027,"raw_usage":{"total_tokens":1703,"prompt_tokens":1104,"completion_tokens":599,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":506}},"tokens_in":720,"tokens_out":599,"duration_ms":6999,"temperature":1.0,"reasoning_tokens":506,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:25:41.841605+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same trained network on measured CSI from two or more synchronized base stations in a built-up area, with a cooperative drone carrying a GPS receiver to provide ground truth, and form the matched-filter input from the measured channel; if the detection rate on off-grid positions falls toward the subspace-pursuit baseline or the false alarm rate rises sharply as multipath and extended-target effects appear, the central claim fails.","supporting_citations":[{"cited_title":"Cooperative sensing for 6G mobile cellular networks: Feasibility, performance and field trial,","cited_arxiv_id":null,"evidence_quote":"Provides the cooperative multi-BS sensing model and least-squares channel estimation from which the paper builds its measurement equation."},{"cited_title":"Towards seamless sensing coverage for cellular multi-static integrated sensing and communication,","cited_arxiv_id":null,"evidence_quote":"Supplies the ISAC beamforming and coverage design that separates communication from sensing and orients beams at the ROI."},{"cited_title":"Joint multi- user communication and sensing exploiting both signal and environment sparsity,","cited_arxiv_id":null,"evidence_quote":"Establishes the on-grid CS imaging formulation with a sparse scattering coefficient vector and defines the discrete sensing matrix."},{"cited_title":"Training region-based object detectors with online hard example mining,","cited_arxiv_id":null,"evidence_quote":"Introduces online hard example mining, the computer-vision technique the paper adapts into its OHEM-1 and OHEM-2 losses."},{"cited_title":"Fourier transform-based wavenumber domain 3D imaging in RIS-aided communication systems,","cited_arxiv_id":null,"evidence_quote":"Supports the integral bistatic channel model in Eq. (7) and the use of point-spread-function-style analysis for imaging."},{"cited_title":"Deep convolution network for direction of arrival estimation with sparse prior,","cited_arxiv_id":null,"evidence_quote":"Exemplifies the learning-after-physics-processing pipeline that the paper's matched-filter-plus-DNN structure follows."},{"cited_title":"Physics-embedded machine learning for electromagnetic data imaging: Examining three types of data-driven imaging methods,","cited_arxiv_id":null,"evidence_quote":"Provides the taxonomy of physics-embedded learning approaches that frames the proposed two-stage design."}],"review_version":1}