{"id":"a9e4f58c-fbe5-4a0d-8efd-16edf18ecca9","arxiv_id":"2412.15670","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A conditional latent diffusion model with a custom VQGAN compressor, offset noise, and temporal adaptive thresholding achieves state-of-the-art bone suppression in chest X-rays.","lead":"BS-LDM is a latent diffusion model that removes bones from chest X-rays to produce soft tissue images, beating prior methods on perceptual quality and improving lung disease detection. It uses a custom image compressor, offset noise, and a time-dependent contrast rule, and introduces a private 818-pair hospital dataset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"JSRT is described as containing CXR/DES soft-tissue pairs, but the public JSRT database contains only radiographs; the external half of the SOTA claim rests on unverifiable ground-truth provenance.","rationale":"The reader's weakest assumption concerned universal ground-truth validity across scanners. My concern is more specific and more load-bearing: the JSRT evaluation, which is half of the headline comparison, is based on 'CXR and DES images from the JSRT dataset' that the cited public dataset does not contain. Unlike an ordinary generalization caveat, this is a concrete provenance gap. The claim that BS-LDM outperforms state-of-the-art methods on JSRT depends on the correctness of the paired DES ground truth; if the pairs are absent or were synthesized, the numerical advantages on JSRT are not evidence of superiority. The paper has independent support: the SZCH-X-Rays experiments, the clinical reader study, the downstream classification on the Shenzhen dataset, and the released code. Those support the narrower claim that BS-LDM works on data from the same GE Discovery XR656 scanner and transfers to Shenzhen CXRs. However, they do not shore up the JSRT comparison. A conditional acceptance remains possible only if the authors can resolve provenance; otherwise the JSRT claim should be withdrawn or reclassified. I therefore recommend UNVERDICTED rather than REJECT: the evidence is insufficient at present, but not nonsensical. The reader and I partially agree: we both target the ground-truth premise, but I locate the failure as a factual dataset-description issue rather than only a scanner-generalization issue. Other concerns, such as hyperparameters being selected on evaluation data and the temporal-thresholding description, are secondary once the JSRT ground truth cannot be verified.","tokens_in":17927,"tokens_out":7666,"duration_ms":66437,"concrete_test":"Ask the authors to release the 241 JSRT pairs or to document their exact origin. A single decisive check: inspect the public JSRT release cited as reference [10] for any DES soft-tissue files. If none exist, obtain the authors' pair-generation procedure or the actual images, then recompute Table I JSRT rows using only ground-truth images whose provenance is independently verifiable. If the pairs cannot be produced, remove the JSRT comparison from the SOTA claim and report the method only on SZCH-X-Rays.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"Section III-A states: 'we processed 241 pairs of CXR and DES images from the JSRT dataset, the largest open-source collection available.' Reference [10], the cited JSRT database, contains 247 chest radiographs with and without nodules; it does not contain dual-energy subtraction (DES) soft-tissue images. There is therefore no publicly documented source for the 241 'DES' soft-tissue pairs used to compute the JSRT rows in Table I. If those soft-tissue images were generated by an algorithm or borrowed from another institution, the paper does not say so. Inversion and contrast adjustment, the only operations described for JSRT, cannot convert a CXR into a soft-tissue ground truth; they would at most change display polarity. This makes the JSRT portion of the central claim—BS-LDM outperforms SOTA on JSRT—non-reproducible and potentially circular if the paired 'ground truth' was itself produced by a learned bone-suppression method. The concern is not primarily about cross-scanner generalization; it is an internal inconsistency with the cited dataset that undermines half of the quantitative comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents BS-LDM, a conditional latent diffusion model for bone suppression in high-resolution chest X-rays. The method compresses images with a custom VQGAN (ML-VQGAN), trains the diffusion model conditioned on CXR via channel concatenation, adds offset noise to the forward process, and applies a time-dependent clipping threshold during sampling. The authors introduce a new paired CXR/DES dataset (SZCH-X-Rays, 818 pairs), compare against seven prior methods on SZCH-X-Rays and JSRT, and report radiologist and automated downstream evaluations on Shenzhen chest X-rays. The central claim is that BS-LDM outperforms state-of-the-art bone suppression methods on both SZCH-X-Rays and JSRT while improving downstream lung-disease detection.","tokens_in":18175,"tokens_out":6225,"duration_ms":50647,"significance":"If the results are reproducible, this would be a strong practical contribution and likely the first latent-diffusion approach for this task. The paper includes broad comparisons, a clinical reader study, and an external downstream evaluation, and the authors state that code is available. However, the JSRT half of the comparison rests on a data provenance statement that is contradicted by the cited public database, and the reported numbers for hyperparameter sweeps are selected on the same test sets used for the headline comparison. Both points must be resolved before the performance claims can be accepted.","major_comments":[{"comment":"The paper states that \"we processed 241 pairs of CXR and DES images from the JSRT dataset, the largest open-source collection available\" and uses these pairs for the JSRT rows of Table I. The public JSRT database cited as [10] contains 247 chest radiographs, with and without nodules, and does not contain dual-energy subtraction soft-tissue images; the operations described (inversion and contrast adjustment) cannot produce a soft-tissue ground truth. The BSR metric in Eq. (12) also requires a bone image B, which is not part of JSRT. The JSRT results in Table I, Table II, Fig. 5, and Fig. 8(b) are therefore based on unverifiable or non-existent ground truth. The authors must either provide a verifiable source for these 241 DES pairs, with documentation and release, or remove the JSRT-based comparisons and revise the abstract, introduction, and conclusions accordingly.","section":"Section III-A, Table I"},{"comment":"The hyperparameters lambda (offset noise weight), omega and b (temporal threshold slope/intercept), and the ML-VQGAN loss weights in Tables III-IV were selected by sweeps on SZCH-X-Rays and JSRT, and the same datasets are then used to report the headline results in Table I. This is a selection-on-test-set procedure: the reported improvements may reflect tuning to the evaluation data rather than a model property. The authors should use the 8:1:1 validation split for hyperparameter selection and report results for the chosen configuration on the untouched test split.","section":"Section III-F, Fig. 8, Table I"},{"comment":"Thirteen image pairs with severe motion artifacts, pleural effusions, and pneumothorax were excluded from SZCH-X-Rays, and the diagnostic utility assessment uses only 79 lesion-containing pairs. These exclusions remove exactly the clinically challenging cases where bone suppression would matter most, yet the abstract and conclusion state that the results \"underscore its clinical value\" without this qualification. The authors should report performance on the excluded cases or explicitly restrict the clinical claim to the studied population.","section":"Section III-A, Section III-G"},{"comment":"Across many comparisons in Table I, the reported standard deviations overlap substantially (e.g., BSR 0.976 +/- 0.018 vs. 0.961 +/- 0.022; PSNR 33.224 +/- 3.577 vs. 32.181 +/- 3.296 for BS-Diff), but no significance tests or confidence intervals are provided. The claim of consistent superiority over all baselines would be strengthened by paired significance testing or effect sizes on the matched image pairs.","section":"Table I, Section III-D"}],"minor_comments":[{"comment":"There is a typo: \"Nervertheless\" should be \"Nevertheless\".","section":"Section II-D"},{"comment":"The axis labeling in Fig. 8(b) is confusing: the text reports b = 1.4, while the axis appears to be labeled with a \"b (x 10^-3)\" tick pattern; the units and the roles of omega and b should be clarified.","section":"Fig. 8(b)"},{"comment":"The SZCH-X-Rays dataset is not made publicly available despite the abstract emphasizing its compilation; code availability alone is insufficient for reproducing Table I on that dataset.","section":"Section III-B, code availability"},{"comment":"The clinical evaluation does not state whether the radiologists were blinded to the CXR versus generated soft-tissue condition, and the selection process for the 79 abnormal cases is not described; these details should be added.","section":"Section III-G"}],"recommendation":"major_revision","confidential_remarks":"The JSRT provenance issue is serious enough that I would ask the authors to supply the actual 241 soft-tissue images or a chain-of-custody description before any acceptance. If the images cannot be produced, the JSRT claims should be deleted and the paper re-scoped to the SZCH-X-Rays dataset. The selection-on-test hyperparameter analysis also needs to be redone or justified with a validation split. I do not see an equation-level circularity in the method itself; the concerns are about data provenance and evaluation protocol."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the SZCH-X-Rays experiment looks like a real advance: first latent-diffusion bone suppression, trained on a decent-size paired hospital dataset, with sensible training tricks (offset noise, temporal adaptive thresholding) and careful downstream and reader evaluations. Second, the JSRT half of the comparison is built on something that does not exist as described. That needs to be settled before trusting the claim that BS-LDM is SOTA on public data.\n\nThe architecture is reasonable. The ML-VQGAN compression with multi-level loss is a standard but solid choice, the 22-hour training on one A100 is practical, and the downstream Shenzhen results (sensitivity gains across three classifiers) line up with the intended clinical use. I also credit the authors for including a reader study with multiple radiologist experience levels rather than stopping at image metrics.\n\nNow the soft spots, in proportion. The JSRT issue is load-bearing. Reference [10] is the JSRT database, which contains 247 plain chest radiographs, not dual-energy subtraction (DES) soft-tissue images. The paper says it processed 241 CXR/DES pairs from JSRT with inversion and contrast adjustment. Those operations cannot produce a soft-tissue ground truth; they only flip and stretch intensities. So the JSRT rows in Table I, and the claim of outperforming SOTA on JSRT, rest on unverifiable provenance. If the soft-tissue images were generated by another algorithm, especially a learned bone suppressor, the comparison becomes circular. This is not a minor error; it removes the external validation from half the paper.\n\nBeyond that: hyperparameters were selected with the evaluation datasets in view (offset noise lambda, threshold slope and intercept, VQGAN loss weights), so the reported numbers are probably optimistic. The ablation tables lack error bars and significance tests. The test set excludes pleural effusions and pneumothorax, so the clinical claims do not cover those common conditions. And while code is public, the main dataset and model weights are not, so the SZCH numbers cannot be reproduced by others.\n\nNet: the core idea deserves a serious referee, and the architecture is worth engaging with, but I would require a major revision. The authors need to disclose exactly where the JSRT soft-tissue images came from, provide a downloadable processed dataset or model outputs, and redo the hyperparameter selection on a proper validation split. If the JSRT provenance cannot be fixed, the JSRT claims should be removed or replaced with an honestly described dataset.","headline":"The SZCH-X-Rays result looks like a genuine step forward for bone suppression, but the JSRT half of the comparison rests on soft-tissue ground truth that the cited JSRT database does not contain.","tokens_in":18698,"tokens_out":3275,"would_cite":false,"duration_ms":28509,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BS-LDM, a conditional latent diffusion model, generates soft-tissue chest X-rays with a bone-suppression ratio of 0.976 on the new SZCH-X-Rays dataset, surpassing seven prior methods.","keywords":["bone suppression","chest X-ray","latent diffusion models","dual-energy subtraction","soft tissue imaging","VQGAN","offset noise","temporal adaptive thresholding"],"falsifier":"Run BS-LDM on paired CXR/DES images collected from a different scanner, hospital, or exposure setting and compare BSR and LPIPS against the same-scanner test results; a large drop would show the learned mapping is tied to one acquisition setup rather than a general bone-suppression solution.","tokens_in":1786,"feed_emoji":"🩻","tokens_out":1988,"duration_ms":51250,"temperature":0.7,"pith_summary":"The paper claims that a conditional latent diffusion model, BS-LDM, can turn a single high-resolution chest X-ray into a soft-tissue image with bones largely removed, approaching the output of dual-energy subtraction without requiring specialized two-exposure equipment. On the authors' new 818-pair hospital dataset, SZCH-X-Rays, and on the public JSRT dataset, BS-LDM outperforms seven existing bone-suppression methods on all four reported metrics. Clinical and automated downstream evaluations indicate that the generated soft-tissue images improve detection of lung lesions compared with the original X-rays. If the claim holds, bone suppression becomes practical for routine single-exposure chest X-rays at a lower computational cost than earlier diffusion-based approaches.","feed_headline":"Bone suppression hits 0.976 with conditional latent diffusion","feed_subtitle":"The authors report gains over seven prior methods and improved lesion detection for radiologists and classifiers.","key_machinery":"The central object is ML-VQGAN, a vector-quantized generative adversarial network constrained by a multi-level hybrid loss that combines L1, perceptual, adversarial, and quantization terms to build a perceptually faithful latent space. It carries the argument by compressing high-resolution chest X-rays into a low-dimensional latent manifold where conditional latent diffusion becomes computationally feasible while preserving texture detail. The second mechanism is offset noise, which augments Gaussian noise with zero-frequency bias to compensate for the greater resistance of low-frequency image content to standard noise injection. The third is temporal adaptive thresholding, which clips the latent variable during each reverse sampling step using a threshold $s=\\omega t+b$ that expands over time, preventing pixel saturation while allowing contrast to match real soft-tissue images.","core_discovery":"The central claim is that BS-LDM, an end-to-end framework built on a conditional latent diffusion model, achieves state-of-the-art bone suppression in high-resolution chest X-rays. The model compresses each 1024x1024 image into a 4x128x128 latent space using a vector-quantized GAN trained with a multi-level hybrid reconstruction loss, then runs a diffusion denoising process conditioned on the CXR latent by channel concatenation. Two additions target the low-frequency errors typical of diffusion models: offset noise in the forward process injects zero-frequency bias to correct luminance drift, and a temporal adaptive thresholding strategy clips latent pixels with a threshold that grows linearly with the sampling timestep. On SZCH-X-Rays the method reports a bone suppression ratio of 0.976, MSE of 0.00060, PSNR of 33.224 dB, and LPIPS of 0.051; on JSRT it reports 0.922, 0.00071, 34.312 dB, and 0.049, all better than the compared baselines. Ablation studies show that removing either offset noise or temporal adaptive thresholding substantially degrades low-frequency fidelity and pixel-intensity alignment.","pith_inferences":["The paper does not test cross-scanner generalization for the generation task, so a natural extension is to evaluate BS-LDM on paired CXR/DES data from other vendors, exposure settings, or post-processing pipelines; the current ground truth comes from a single GE Discovery XR656 unit.","The offset-noise and temporal-adaptive-thresholding fixes are specific to low-frequency drift, which suggests they could transfer to other medical image translation tasks with similar brightness or contrast instability, though the paper only demonstrates them for bone suppression.","The construction of the SZCH-X-Rays dataset and the reprocessed JSRT negative-image pairs may serve as a benchmark for future bone-suppression work, but the JSRT processing steps are not independently validated against clinically acquired dual-energy images.","The claim that BS-LDM preserves lesions relies on radiologist review and automated classifiers rather than direct registration of lesions between CXR and generated tissue images; a pixel-level lesion-preservation analysis would strengthen the clinical conclusion."],"forward_implications":["If BS-LDM performs as reported, single-exposure chest X-rays can receive bone suppression close to dual-energy subtraction quality without needing specialized DES hardware or extra radiation.","Radiologists reading the generated soft-tissue images would be expected to detect more lung lesions: the paper reports junior radiologist F1 rising from 0.51 to 0.63 and senior F1 from 0.60 to 0.75.","Automated classifiers trained on chest X-rays would improve when fed BS-LDM soft-tissue images, with sensitivity gains of about 13.5%, 3.24%, and 3.18% for AlexNet, DenseNet, and ResNet on the Shenzhen dataset.","Because BS-LDM's inference time is about 77.7% of BS-Diff and DDPM, high-resolution diffusion-based bone suppression is more practical in clinical settings than prior diffusion baselines."],"supporting_citations":[{"why":"Rombach et al. latent diffusion models provide the underlying LDM architecture that BS-LDM adapts for bone suppression.","marker":"[25]"},{"why":"Chen et al. BS-Diff is the pixel-space diffusion baseline whose resolution and low-frequency limitations this paper directly addresses.","marker":"[24]"},{"why":"Ho et al. DDPM supplies the diffusion forward and reverse process formalism, including the static thresholding approach the paper builds on.","marker":"[21]"},{"why":"Guttenberg's offset-noise blog post is the source of the offset-noise idea used in the forward process.","marker":"[26]"},{"why":"Ho and Salimans classifier-free guidance introduces dynamic thresholding, which the temporal adaptive thresholding strategy extends.","marker":"[33]"},{"why":"Shiraishi et al. JSRT dataset is one of the two evaluation datasets, with its 241 pairs processed into negative images for testing.","marker":"[10]"},{"why":"Jaeger et al. Shenzhen chest X-ray dataset is the external dataset used for the automated downstream evaluation.","marker":"[44]"},{"why":"Hogeweg et al. defines the Bone Suppression Ratio metric, the primary quantitative measure of bone removal.","marker":"[35]"},{"why":"Zhang et al. defines LPIPS, the perceptual similarity metric used to evaluate detail preservation.","marker":"[37]"}],"fun_headline_variants":["Latent diffusion model strips bone from chest X-rays, hits 0.976","BS-LDM: conditional latent diffusion for high-res bone suppression","Diffusion-based bone suppression scores 0.976 on CXR dataset","Offset noise + adaptive threshold boost X-ray bone removal to 0.976","Chest X-ray bone suppression reaches 0.976 via latent diffusion"],"cache_read_input_tokens":20864,"weakest_assumption_plain":"The load-bearing premise is that dual-energy-subtraction soft-tissue images from one GE scanner, together with inverted and contrast-adjusted JSRT images, define what correct bone-suppressed output looks like for all chest X-rays.","fun_headline_variants_meta":{"raw":{"variants":["Latent diffusion model strips bone from chest X-rays, hits 0.976","BS-LDM: conditional latent diffusion for high-res bone suppression","Diffusion-based bone suppression scores 0.976 on CXR dataset","Offset noise + adaptive threshold boost X-ray bone removal to 0.976","Chest X-ray bone suppression reaches 0.976 via latent diffusion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000897,"raw_usage":{"total_tokens":3912,"prompt_tokens":1044,"completion_tokens":2868,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":2771}},"tokens_in":660,"tokens_out":2868,"duration_ms":17887,"temperature":1.0,"reasoning_tokens":2771,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:11:55.026775+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run BS-LDM on paired CXR/DES images collected from a different scanner, hospital, or exposure setting and compare BSR and LPIPS against the same-scanner test results; a large drop would show the learned mapping is tied to one acquisition setup rather than a general bone-suppression solution.","supporting_citations":[{"cited_title":"Bs-diff: Effective bone suppression using conditional diffusion models from chest x-ray images,","cited_arxiv_id":null,"evidence_quote":"Chen et al. BS-Diff is the pixel-space diffusion baseline whose resolution and low-frequency limitations this paper directly addresses."},{"cited_title":"Diffusion with offset noise,","cited_arxiv_id":null,"evidence_quote":"Guttenberg's offset-noise blog post is the source of the offset-noise idea used in the forward process."},{"cited_title":"Shiraishi, S","cited_arxiv_id":null,"evidence_quote":"Shiraishi et al. JSRT dataset is one of the two evaluation datasets, with its 241 pairs processed into negative images for testing."},{"cited_title":"Suppression of translucent elongated structures: applications in chest radiography,","cited_arxiv_id":null,"evidence_quote":"Hogeweg et al. defines the Bone Suppression Ratio metric, the primary quantitative measure of bone removal."},{"cited_title":"The unreasonable effectiveness of deep features as a perceptual metric,","cited_arxiv_id":null,"evidence_quote":"Zhang et al. defines LPIPS, the perceptual similarity metric used to evaluate detail preservation."}],"review_version":1}