{"id":"43209d86-9d5d-42eb-be5d-847d80673c2d","arxiv_id":"2509.08528","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":12,"one_line_summary":"A simulation-trained deep network combining spectral and angular denoising reduces noise in real multispectral CT projections from the ESRF BM18 beamline.","lead":"This paper trains a neural network to denoise multispectral CT projections from the ESRF BM18 beamline using only simulated data, then shows it reduces noise on real phantom scans. The result matters because multispectral CT is photon-starved and currently sacrifices scan time or image quality; a simulation-trained denoiser could make such scans practical without requiring large real labeled datasets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulator geometry for energy-to-line mapping is internally inconsistent (III.C uses alpha=60 deg, 50 prisms; I.A describes 70.52 deg tip and 100 prisms), so simulation-trained spectral correlations may not match real BM18 bins; this needs an explicit sensitivity check before the sim-to-real claim i","rationale":"The reader's weakest assumption was that the simulated noise and signal model faithfully represents the real BM18 detector chain, including the energy-to-detector-line mapping. My review identifies a concrete, testable internal inconsistency within that mapping: the simulation equations in Section III.C use a prism tip angle of 60 degrees and 50 prisms, while the experimental setup in Section I.A is described as 70.52 degrees and 100 prisms. Since the spectral subnetwork is trained to exploit correlations between adjacent energy bins, a systematic error in the mapping could cause the network to learn spectral structure that does not correspond to the real data. The paper's own statement that 'notable discrepancies' exist between simulation and experiment confirms this is not a hypothetical concern, yet no sensitivity analysis or calibration-transfer step is reported. The experimental validation is also limited: Table 4 reports only four energy bins, uses a 50x averaged noisy reconstruction as reference, and lacks error bars, so it cannot by itself resolve whether the observed transfer is robust to the geometry mismatch. I am not claiming the result is false; rather, the evidence as presented is conditional on an unverified simulator configuration. The reader's CONDITIONAL verdict remains appropriate, so no change is needed. The proposed retraining test would settle whether the concern actually lands: if the network is insensitive to the prism geometry, the claim is robust; if performance degrades with the corrected geometry, the paper's central claim needs revision.","tokens_in":15630,"tokens_out":6660,"duration_ms":75993,"concrete_test":"Regenerate the training set using the geometry described in I.A (alpha=70.52 deg, 100 prisms) while keeping the noise model and architecture fixed; retrain from scratch and recompute Table 4 and Figure 12. If the experimental NRMSE/SNR changes materially (e.g., >10% relative on energy bins 50/99) or the proposed network no longer consistently beats DnCNN, the current simulation configuration is load-bearing; if results are unchanged, the network is robust to the mapping and the concern is resolved. Releasing the simulator configuration would also allow independent verification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central sim-to-real claim requires the simulated data to reproduce the BM18 energy-to-detector-line mapping, because the spectral subnetwork is trained on correlations among adjacent energy bins. Section III.C, Eqs. (2)-(6), defines the refraction model with alpha=60 deg and n_prism=50; Section I.A states the actual prisms have a 70.52 deg tip angle and the array contains 100 prisms. The row assignment in Eq. (6) is directly proportional to the total refraction angle, so a 17% change in tan(alpha/2) and a factor-2 change in n_prism shift every simulated energy-bin boundary substantially. The K-edge calibration described in I.A (26-91 keV) calibrates the real setup, but the paper does not state that this calibration is folded into the training simulator, and no ablation checks sensitivity to these parameters. The paper's own admission of 'notable discrepancies' between simulation and experiment (Section I) makes this omission material. If the training spectral mapping is wrong, the experimental NRMSE gains (Table 4) could be driven mostly by the spatial/angular subnetwork, while the spectral component learns a mismatched structure; this would not invalidate denoising, but it would weaken the specific claim of simulation-trained multispectral transfer across the stated 22.9-400 keV range.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a deep-learning denoiser for multispectral CT projections acquired at the ESRF BM18 beamline, where a prism array disperses the polychromatic beam onto detector rows corresponding to energy bins. The architecture combines a spectral-spatial subnetwork (based on AODN) and an angular-spatial subnetwork (based on PaCNet) via stacked generalization and attention modules. Training is done exclusively on simulated projections produced by a physics-based model incorporating Beer-Lambert absorption, prism refraction, scintillator conversion, double Poisson noise, PRNU, dark signal, and readout noise. The authors validate on a synthetic test set from the same simulator and on real BM18 scans of two phantoms, reporting improved PSNR/SSIM/LPIPS over NLM/TV/DnCNN on synthetic data and lower NRMSE relative to a 50x averaged reference plus improved ROI SNR on real data.","tokens_in":16033,"tokens_out":6624,"duration_ms":77099,"significance":"The application is timely: multispectral CT at synchrotron facilities is photon-starved, and a denoiser that transfers from simulation to real data would enable faster acquisitions. The paper's strengths are the explicit physics-based noise model, the use of external phantom datasets, and a head-to-head comparison with DnCNN. However, the central sim-to-real claim is only as strong as the simulator's fidelity, and the geometry inconsistency identified below is currently unaddressed. If properly resolved, the method is a useful engineering contribution, although the architecture itself is largely assembled from existing building blocks and the synthetic results are in-distribution.","major_comments":[{"comment":"The simulator in Section III.C uses n_prism=50 and alpha=60 deg in Eqs. (2)-(6), while Section I.A describes the actual BM18 prism array as 100 prisms with a 70.52 deg tip angle. Since Eq. (6) maps the total refraction angle to detector row, doubling n_prism and increasing tan(alpha/2) by ~22% materially changes every simulated energy-bin boundary. The paper does not state that the K-edge calibration of Section I.A was folded into the training simulator, nor does it provide a sensitivity analysis. This directly affects the spectral subnetwork, which is trained on correlations among k=64 adjacent bins; a mismatched energy axis could cause the network to learn incorrect spectral structure. Please clarify which geometry was used in training and either retrain with the correct parameters or demonstrate that the real-data results are robust to this discrepancy.","section":"Section III.C vs. Section I.A, Eqs. (2)-(6)"},{"comment":"The synthetic test set is generated by the same simulator used for training (different objects, same noise model and geometry). The high PSNR/SSIM/LPIPS values therefore partly measure the network's ability to invert the training simulator and should not be used to support sim-to-real generalization. Please relabel Table 3 as an in-distribution benchmark and add an out-of-distribution test (e.g., different noise levels, different prism parameters, or real-data-only evaluation) to substantiate the generalization claim.","section":"Section IV.A, Table 3"},{"comment":"The real-data evaluation uses a 50x averaged reconstruction as reference, which is itself noisy and not ground truth, and the 'mean normalized NRMSE' is not defined. Only four energy bins (0, 5, 50, 99) are reported numerically, and the SNR analysis covers only the low-Z phantom. To support the claim of performance over the full 22.9-400 keV range, please define the metric precisely, report bin-wise curves or error bars for both phantoms, and discuss how the 50x averaging affects the comparison.","section":"Section IV.B, Table 4 and Fig. 12"}],"minor_comments":[{"comment":"Duplicate phrase: 'First, spatial features of First, spatial features of one band and projection angle are extracted.'","section":"Section III.A"},{"comment":"'The denoising block from [6] was modified' likely refers to reference [26] (AODN), not [6]. Please correct the citation.","section":"Section III.A"},{"comment":"Typos and inconsistent notation: 'LeRU' should be 'Leaky ReLU'; Table 4 caption 'Desnoing' should be 'Denoising'; 'unfilted', 'horizontaly', 'refrac' should be corrected; and the dark current/readout noise notation 'sigma=darkcurrent!' and 'sigma=0.8!' needs units and explanation.","section":"Throughout"},{"comment":"The normalization of detector lines uses the maximal possible gray value from simulation. Please state explicitly how the same normalization is applied to real experimental data, since a mismatch here could affect the network's input distribution.","section":"Section III.D"},{"comment":"The paper's self-description as literal quotes from two conference papers is unusual for a journal submission. Please rewrite into a single coherent document with consistent nomenclature and avoid duplicated or truncated sentences.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The geometry inconsistency between Section III.C (50 prisms, 60 degrees) and Section I.A (100 prisms, 70.52 degrees) is the main technical risk and must be resolved before acceptance. The paper currently reads as a mechanical merger of two conference papers; the authors should be asked to produce a coherent journal version and to address the simulator-fidelity question directly. The experimental validation on real phantoms is valuable, but without a sensitivity analysis or corrected retraining, the central sim-to-real claim remains insufficiently supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a transparently assembled merge of two conference papers, and the method is a plausible engineering contribution, but the paper contains an internal inconsistency that weakens the sim-to-real claim until it is addressed. The simulation (Section III.C) uses a 60° prism tip angle and 50 prisms, while the experimental setup (Section I.A) is described with a 70.52° tip angle and 100 prisms. The distance also disagrees: 10 m in the introduction versus 28 m in the setup section. Because the energy-to-detector-row mapping in Eq. (6) is proportional to the prism refraction angle, these discrepancies shift every simulated energy boundary. The paper admits 'notable discrepancies' between simulation and experiment but never checks sensitivity to these specific parameters. This is not a minor detail; it's a load-bearing inconsistency.\n\nWhat the paper does well: the architecture is sensible — an angular-spatial branch inspired by PaCNet and a spectral branch inspired by AODN, combined with stacked generalization. The simulation pipeline is detailed, including double Poisson noise, PRNU, dark signal and readout noise, which gives the training data a degree of realism. Validating on two custom phantoms with DnCNN as a baseline is a reasonable first step, and the authors are honest that the manuscript reuses their own conference papers.\n\nThe soft spots are real but mostly addressable. The synthetic test set is generated by the same simulator used for training, so the excellent PSNR/SSIM values partly reflect the network inverting its own training distribution. The real-data quantitative evaluation is thin: two phantoms, four energy bins, and no error bars. No code or trained model is provided, so the results aren't reproducible. And the geometry mismatch above is the main barrier: before the sim-to-real claim can stand, the authors need to either retrain with corrected geometry or show that the network's output is insensitive to a factor-of-two change in prism count and a 17% change in tip angle.\n\nWho this is for: researchers working on spectral CT denoising or synchrotron imaging will find the architecture and noise model useful, even though the novelty is modest. The paper deserves a serious referee, but not in its current form. The right call is to send it to review and ask for major revision: fix the geometry inconsistency or add a sensitivity analysis, provide error bars on the real-data metrics, and ideally release the trained model.","headline":"Transparent conference-paper merge with a useful architecture, but an unflagged simulation/experiment geometry mismatch weakens the sim-to-real claim until sensitivity is shown.","tokens_in":16540,"tokens_out":4007,"would_cite":false,"duration_ms":37255,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural network trained exclusively on simulated projections denoises real multispectral CT scans at the BM18 beamline across 22.9–400 keV.","keywords":["computed tomography","multispectral imaging","deep learning","denoising","simulation-to-real transfer","synchrotron","energy-resolved imaging","convolutional neural networks"],"falsifier":"Scan a third phantom with known material inserts, denoise its projections with the trained network, and compare the lowest-energy-bin reconstructions against a heavily averaged reference; the transfer claim fails if NRMSE does not improve over raw data or if new streak artifacts appear where the simulator predicted none. Alternatively, record blocked-beam flat frames at BM18 and compare measured dark-signal and readout-noise statistics with the simulated Gaussian model.","tokens_in":15566,"feed_emoji":"🩻","tokens_out":11571,"duration_ms":120149,"temperature":0.7,"pith_summary":"Multispectral CT splits the X-ray beam into narrow energy bins, so each bin receives far fewer photons than an energy-integrating detector would, making scans either very noisy or impractically long. This paper claims that a neural network trained exclusively on simulations of the BM18 beamline can denoise real experimental projections across the full 22.9–400 keV range, outperforming a single-frame CNN baseline and classical denoisers while preserving structural detail. The method exploits two redundancies in the data—similarity between adjacent energy bins and similarity between neighboring projection angles—through two sub-networks combined by stacked generalization. If the claim holds, multispectral CT can be run at practical exposure times and its low-flux energy bins become usable for material characterization.","feed_headline":"Simulation-trained network cleans real multispectral CT scans","feed_subtitle":"Uses spectral and angular redundancies to cut noise across 22.9–400 keV energy bins while preserving structure.","key_machinery":"Two-path neural network plus a physics-based simulator. The spectral path, adapted from a hyperspectral-image denoiser, extracts spatial and spectral features across 64 adjacent energy bins with attention-guided convolutions. The angular path, adapted from a video-denoising patch-craft method, builds patch frames from neighboring projections of the same energy bin and applies an angular-post filter to keep projection-to-projection continuity. Stacked generalization—training a third convolution layer on the two frozen subnetworks' feature maps—fuses the paths. The simulator generates training data by modeling Beer-Lambert absorption, silicon prism refraction, scintillator conversion, two Pois","core_discovery":"The central claim is that a neural network trained only on simulated BM18 projections denoises real experimental multispectral CT data. The network pairs a spectral-spatial subnetwork (denoising each energy bin with 64 adjacent bins) with an angular-spatial subnetwork (denoising via neighboring projections of the same bin), fusing both by stacked generalization. On two custom phantoms, it gave lower normalized RMSE than raw data and a DnCNN baseline across sampled energy bins, and higher SNR in homogeneous regions over the first 100 bins, with no added reconstruction artifacts down to roughly 0.076% of detector full scale.","pith_inferences":["The paper validates on two phantoms; I would not extrapolate the sim-to-real transfer to arbitrary samples until a third, unseen object is tested.","The same two-path design could likely be retrained for other spectral CT geometries, but the prism-dispersion and detector-noise models would need to be re-derived for each setup.","The reported low-signal limit (roughly 50 expected gray values) suggests that emphasizing low-flux bins in training, or modeling the noise floor more explicitly, could extend the usable spectral range beyond the first 100 detector lines.","Because the experimental comparison uses a 50-times-averaged reconstruction as reference, the NRMSE numbers measure agreement with an average; a noise-removal test on raw projection residuals would more directly confirm that fine structure is preserved."],"forward_implications":["Multispectral scans at BM18 can use short exposure times and recover image quality by denoising, avoiding scan times that would otherwise stretch from hours to days.","Denoising in projection space produces cleaned data that can be fed directly into standard filtered backprojection without changing the reconstruction pipeline.","Low-flux energy bins—previously too noisy to use—become analyzable, extending the usable spectral range down to about 22.9 keV.","The spectral and angular sub-networks together outperform either alone and beat the single-frame DnCNN baseline, suggesting both redundancies matter."],"supporting_citations":[{"why":"Earlier paper by the same authors supplying the simulation-based network and synthetic-data evaluation that this work merges with experimental results.","marker":"[1]"},{"why":"Earlier paper supplying the experimental BM18 phantom results and the NRMSE/SNR comparisons reported here.","marker":"[2]"},{"why":"DnCNN, the single-frame CNN baseline trained on the same synthetic data for comparison.","marker":"[16]"},{"why":"PaCNet video-denoising method whose patch-craft angular path is adapted for projection-to-projection denoising.","marker":"[23]"},{"why":"AODN hyperspectral-denoising architecture whose spectral-spatial path is adapted for energy-bin denoising.","marker":"[26]"},{"why":"Stacked generalization, the ensemble strategy used to fuse the two sub-networks.","marker":"[27]"},{"why":"CBAM attention module used to guide channel and spatial feature refinement.","marker":"[34]"},{"why":"Public 3D surface dataset used to generate the training volumes for the simulator.","marker":"[41]"},{"why":"Detector datasheet supplying quantum efficiency, dark-current, and readout-noise parameters used in the noise model.","marker":"[43]"}],"fun_headline_variants":["Sim-only training denoises real multispectral CT","Multispectral CT noise tamed by simulation-trained net","Denoising real CT scans with simulated training data","Spectral+angular denoising from simulated CT training","Sim-trained network suppresses noise in real spectral CT"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the simulated noise and signal model faithfully represents the real BM18 detector chain; if the simulator misses a noise source or miscalibrates the energy-to-detector-line mapping, the demonstrated transfer from simulation to experiment on two phantoms could fail on other samples or energy bins.","fun_headline_variants_meta":{"raw":{"variants":["Sim-only training denoises real multispectral CT","Multispectral CT noise tamed by simulation-trained net","Denoising real CT scans with simulated training data","Spectral+angular denoising from simulated CT training","Sim-trained network suppresses noise in real spectral CT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000186,"raw_usage":{"total_tokens":1176,"prompt_tokens":770,"completion_tokens":406,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":340}},"tokens_in":514,"tokens_out":406,"duration_ms":5349,"temperature":1.0,"reasoning_tokens":340,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T20:28:19.684754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Scan a third phantom with known material inserts, denoise its projections with the trained network, and compare the lowest-energy-bin reconstructions against a heavily averaged reference; the transfer claim fails if NRMSE does not improve over raw data or if new streak artifacts appear where the simulator predicted none. Alternatively, record blocked-beam flat frames at BM18 and compare measured dark-signal and readout-noise statistics with the simulated Gaussian model.","supporting_citations":[],"review_version":1}