{"id":"a52fd7b8-f299-4b07-a585-b82e9644b9c1","arxiv_id":"2506.13443","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"PRO is a text-conditioned latent diffusion model that generates CT sinograms, and its synthetic data improves some CT reconstruction tasks but degrades others.","lead":"A CT imaging model, PRO, generates synthetic X-ray projection data (sinograms) from text prompts and then converts them into CT images. The authors say this improves downstream CT reconstruction, but their own tables show gains in only a few settings and losses in others.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Tables V and VI contradict the paper's central claim that PRO synthetic data significantly improves low-dose and sparse-view reconstruction.","rationale":"The reader's weakest_assumption focused on whether simulated fan-beam projections and the undocumented head dataset are valid stand-ins for real clinical CT data. Those are legitimate external-validity concerns, but the more fundamental problem is internal: even under the paper's own simulated AAPM benchmark, the central claim of significant downstream improvement fails in a majority of reported settings. Tables V and VI show that replacing AAPM training data with PRO-generated data mostly degrades or leaves unchanged PSNR and SSIM, with the largest degradations occurring in the most challenging settings (180-view sparse-view and 1e4 low-dose). Because the abstract, introduction, and conclusions assert a significant improvement, and because Section IV-G itself only claims 'comparable performance,' the empirical support is internally inconsistent. No error bars or significance tests are provided, so the few favorable point estimates cannot be distinguished from noise. I did not find a need to challenge the novelty of text-conditioned projection-domain synthesis, which is plausible, but the central application claim is not established by the evidence presented. Since the reader already reached REJECT, my stress-test does not move the verdict; it identifies the same empirical weakness from a different angle, hence UNCHANGED.","tokens_in":17181,"tokens_out":4072,"duration_ms":41783,"concrete_test":"Using the publicly released code and data, re-run the Section IV-G protocol for both downstream methods: train GMSD and OSDM on the released 2,000 synthetic PRO images and on the 2,000 AAPM images under identical seeds, hyperparameters, and test sets. Report per-setting PSNR/SSIM differences with at least three independent training seeds, plus paired statistical tests. If the reproduced tables match the paper's point estimates and a paired sign test shows no significant majority of settings favoring PRO, then the 'significantly improves' claim should be revised to 'comparable' or 'setting-dependent.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the Abstract, Section I, and Section III-A, is that incorporating PRO-generated projection data significantly improves low-dose and sparse-view CT reconstruction. Section IV-G is the only direct test of this claim, and the paper's own tables contradict it. In Table V (sparse-view, GMSD), PRO improves PSNR in only 1 of 4 view settings (90 views: 37.321 vs 37.150 dB) and degrades PSNR at 60, 120, and 180 views (34.417 vs 34.905; 38.792 vs 39.897; 40.586 vs 42.091). SSIM improves only at 60 and 90 views, and drops at 120 views (0.9481 vs 0.9586) and 180 views (0.9625 vs 0.9706). In Table VI (low-dose, OSDM), PRO improves PSNR in only 1 of 3 noise settings (5e4: 41.74 vs 41.20 dB) and is worse at 1e5 (42.49 vs 42.62 dB) and 1e4 (34.56 vs 37.43 dB); SSIM is worse in all three settings (0.9777 vs 0.9899; 0.9670 vs 0.9857; 0.9242 vs 0.9683). Out of eight primary quality comparisons, PRO wins only three. The accompanying text in Section IV-G itself says 'comparable performance,' not 'significantly improves.' Because these tables are the paper's evidence for the central claim, the claim is unsupported as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces PRO, a two-stage framework for generating synthetic CT projection data (sinograms) using a latent diffusion model conditioned on anatomical text prompts, followed by FBP reconstruction and a CNN-based SharpNet refinement. The authors claim that PRO is the first projection-domain CT synthesis method that operates entirely without prior measurements, and that incorporating PRO-generated projection data significantly improves downstream sparse-view and low-dose CT reconstruction. Experiments use simulated fan-beam projections of AAPM abdominal slices, compare generation quality via FID/IS/KID against image-domain and projection-domain baselines, and test downstream performance by training GMSD and OSDM on PRO-generated versus original AAPM data.","tokens_in":17492,"tokens_out":7522,"duration_ms":73201,"significance":"If the downstream gains were real, projection-domain synthesis with prompt control would be a valuable contribution to CT data augmentation and to physics-aware generative modeling. The paper provides public code, detailed geometry settings for the abdomen simulator, and a broad set of generation-quality metrics, which are strengths. However, the headline claim of significant downstream improvement is not supported by the paper's own tables, and several claims about physical simulation and the head-prompt experiments are not backed by the described implementation and data. The generation-quality results are also mixed, with PRO not being best on FID in Table I.","major_comments":[{"comment":"The central claim that incorporating PRO-synthesized data 'significantly improves' downstream low-dose and sparse-view reconstruction is contradicted by the quantitative results. In Table V, PRO data yields higher PSNR than AAPM data in only 1 of 4 view settings (90 views, 37.321 vs 37.150 dB) and lower PSNR at 60, 120, and 180 views (e.g., 40.586 vs 42.091 dB at 180 views); SSIM also drops at 120 and 180 views (0.9481 vs 0.9586 and 0.9625 vs 0.9706). In Table VI, PRO improves PSNR in only 1 of 3 noise levels (5e4, 41.74 vs 41.20 dB) and is worse at 1e5 and 1e4 (e.g., 34.56 vs 37.43 dB at 1e4); SSIM is worse in all three noise settings. Section IV-G's own text describes the results as 'comparable performance,' but the Abstract, Introduction, and Conclusions claim significant improvement. This load-bearing claim must either be supported by new experiments (e.g., multiple seeds, confidence intervals, or a synthetic-plus-real augmentation setting) or be removed and reframed throughout the paper.","section":"Section IV-G, Tables V and VI; Abstract and Section I"},{"comment":"The 'head' prompt experiments are not reproducible because the 500 head CT images used to train the model are never described in the Data Specification. Section IV-A specifies only the AAPM abdominal dataset, including slice counts and geometry; no source, preprocessing, or train/test split is given for the head images. Since prompt-conditioned generation across anatomies is one of the paper's key contributions, the dataset description must be added and the head-prompt evaluation (validated on '100 held out images') must be defined precisely.","section":"Section V (Discussion) and Section IV-A"},{"comment":"The manuscript repeatedly claims that projection-domain synthesis explicitly incorporates or simulates beam hardening, scattering, and material attenuation. The actual forward model in Section IV-A is Siddon's ray-driven algorithm applied to AAPM CT slices, which computes ray line integrals without modeling beam hardening or scatter. Either the physics claims must be restricted to what the simulator implements (geometry and linear attenuation), or the simulator must be extended accordingly; as written, the claims overstate the method's physical fidelity.","section":"Section III-A and Section IV-A"},{"comment":"The downstream experiment protocol is incomplete. The manuscript does not state the test set used for GMSD and OSDM evaluation, whether the reported numbers are means over multiple runs, or any uncertainty or statistical significance measure. Given that the favorable PSNR differences are 0.17–0.54 dB, significance cannot be assessed. This information is necessary to evaluate the 'significantly improves' claim even for the entries where PRO appears better.","section":"Section IV-G"}],"minor_comments":[{"comment":"The text states that PRO achieves the second lowest FID score 'indicating higher visual fidelity' and then emphasizes outperformance on IS and KID; however, StyleGAN has a lower FID (0.4018) than PRO (0.4241). The discussion should acknowledge that PRO does not win on FID in this comparison.","section":"Section IV-D, Table I"},{"comment":"The caption of Table II says 'GENERATING 100 CT IMAGES' but the table and surrounding text also report results for 1000 images; the caption and text should be aligned.","section":"Section IV-E and Table II"},{"comment":"The number of DDIM sampling steps is given as 250 in Section IV-B, but Section IV-B later states 'the first stage model conducts 72 steps of DDIM sampling,' and the Appendix also uses 72 steps; please clarify which protocol applies to each result.","section":"Section IV-B and Appendix"},{"comment":"There are several typographical issues, including 'neglects the easurement process' in Section I, 'Downsapling' in Table IV, and inconsistent terminology for the 4x latent configuration; these should be corrected.","section":"Throughout"}],"recommendation":"reject","confidential_remarks":"The decision is reject because the paper's own downstream experiments contradict its central claim of significant improvement; this is not a presentation issue but a fundamental mismatch between the stated contribution and the evidence. Substantial new experiments, or a major reframing of the paper as a generation-quality study with only comparable downstream performance, would be needed before resubmission. I also note that the head-prompt experiments, which support a key contribution, lack a data description and cannot be verified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is unconditional, text-conditioned synthesis of whole sinograms. Prior projection-domain diffusion work was limited to inpainting partial sinograms, so moving to full generation from noise with anatomical prompts is a real step, and the code is public. The two-stage DMPD-plus-SharpNet pipeline, with FBP in between, is a sensible design and the authors do compare against GAN and diffusion baselines in the projection domain.\n\nThat said, the paper's headline claim does not survive contact with its own numbers. Tables V and VI are the only direct test of the data-augmentation story, and out of eight PSNR/SSIM comparisons, PRO-trained models win three. At 120 and 180 views in sparse-view GMSD, and at the highest noise level in low-dose OSDM, the synthetic data clearly hurts. The text in Section IV-G says \"comparable performance,\" which is accurate, but the abstract and introduction say \"significantly improves.\" That mismatch is not a minor wording issue; it is the central claim of the paper.\n\nThere are other soft spots, in proportion. The Discussion's head-prompt experiment relies on 500 head CT images that never appear in the Data Specification, so the reader cannot tell what geometry or preprocessing was used. There are no error bars on the downstream reconstruction numbers, which matters when the gaps are a few tenths of a dB. The KID text misquotes Table I (calling 0.3038 \"0.0028,\" which is actually the standard deviation). And the paper claims projection-domain synthesis can model beam hardening and scatter, but the actual simulator is Siddon's ray-driven fan-beam projection, which does neither; those are aspirations, not capabilities demonstrated here.\n\nFor whom is this paper? CT researchers working on projection-domain generative models or synthetic data augmentation. The core idea is worth discussing, and the code availability makes it easy to probe. But as submitted, the empirical support for the main claim is inconsistent with the stated conclusions.\n\nMy recommendation: this deserves a serious referee, not a desk reject. The novelty is legitimate and the architecture is coherent. But it needs major revision: reframe the contribution honestly, describe the head data, add uncertainty quantification, and fix the KID error. If those are addressed, the synthesis part is citable; the downstream claim needs to be either demonstrated consistently or withdrawn.","headline":"Real novelty in text-conditioned sinogram synthesis, but the paper's own tables undercut the main data-augmentation claim.","tokens_in":18053,"tokens_out":2136,"would_cite":false,"duration_ms":24015,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports a projection-domain latent diffusion model, PRO, that generates CT sinograms from scratch under anatomical text prompts; the synthetic data match or improve low-dose and sparse-view reconstruction.","keywords":["CT synthesis","projection domain","sinogram generation","latent diffusion model","text-conditioned generation","low-dose CT reconstruction","sparse-view CT reconstruction","data augmentation"],"falsifier":"Generate sinograms with PRO, inject the same physical effects a real scanner has (measured detector noise, beam hardening, scatter), reconstruct and train the downstream reconstruction networks on them; if performance falls to parity with or below training on real low-dose projections, the claim that projection-domain synthesis preserves acquisition physics is refuted.","tokens_in":16966,"feed_emoji":"🩻","tokens_out":9774,"duration_ms":86365,"temperature":0.7,"pith_summary":"PRO is an attempt to move CT data synthesis out of the reconstructed-image domain and into the raw projection (sinogram) domain. The paper reports that a latent diffusion model trained on simulated fan-beam projections of abdominal CT slices can generate new sinograms from scratch, guided by text prompts such as \"head\" or \"body\" through task-specific latent spaces. When the generated sinograms are reconstructed with filtered back projection and lightly refined by a CNN, the resulting CT images are statistically close to real scans. The paper then shows that training downstream low-dose and sparse-view reconstruction networks on 2,000 of these synthetic images instead of the original low-dose CT challenge data maintains or improves PSNR/SSIM in several settings. The point is that projection-domain synthesis can serve as a foundation-model-style source of training data, because it preserves measurement-domain physics that image-domain generators discard.","feed_headline":"Synthetic CT sinograms rival real data for reconstruction training","feed_subtitle":"A text-prompted latent diffusion model generates raw CT projections, and the synthetic data match or beat real scans.","key_machinery":"The load-bearing object is DMPD, a latent diffusion model whose forward and reverse processes run on a compressed latent code of the sinogram rather than on pixels of the reconstructed image. A text encoder turns prompts such as \"head\" or \"body\" into embeddings that select one of several task-specific latent spaces, so each anatomy class gets its own generative trajectory; the denoising U-Net is conditioned on those embeddings through cross-attention. After sampling, the decoder maps the latent back to a sinogram, filtered back projection converts it to an image, and SharpNet, a small U-Net-like CNN trained on synthetic noise pairs, removes residual noise and sharpens detail. The two-stage split is what lets the model keep physical consistency from the projection domain while compensating for the lossy compression of latent diffusion in the image domain.","core_discovery":"The central claim is that CT synthesis belongs in the Radon domain: instead of generating reconstructed images, PRO generates raw sinograms with a latent diffusion model (DMPD), then applies filtered back projection followed by a lightweight CNN refiner (SharpNet). The paper argues this is the first framework that synthesizes CT projection data without any prior measurements and with text-prompt control. The generated projections are claimed to capture cross-detector correlations, material attenuation, and view-dependent anatomical structure that image-domain methods lose, and the downstream experiments, replacing real training data with 2,000 PRO-generated samples in sparse-view and low-dose reconstruction, are offered as evidence that the synthetic data are faithful enough to purpose.","pith_inferences":["The paper leaves the 500 head CT images used for the \"head\" prompt out of its Data Specification, so the cross-anatomy generalization claim is not yet reproducible from the manuscript alone.","A test the authors do not run is mixed-data training: blending PRO-generated sinograms with real projections in controlled proportions, which would directly probe the method's value as an augmentation tool under realistic dose constraints.","The realism ceiling is set by Siddon's ray-driven forward model; replacing that simulator with a Monte Carlo or measured scatter and beam-hardening model would show whether the claimed physics fidelity survives contact with real scanners.","Prompt controllability could be quantified by training an independent anatomy classifier on real CT and checking whether images reconstructed from \"head\" and \"body\" prompts are classified correctly."],"forward_implications":["If PRO works as reported, synthetic projection data can replace real measured projections when training reconstruction networks, lowering the data-acquisition barrier for CT imaging research.","The same trained model can serve multiple downstream tasks by switching text prompts, so a single foundation model could generate training data for low-dose, sparse-view, artifact-correction, and protocol-optimization studies.","Because generation happens before reconstruction, synthetic data inherit the scanner geometry and physics of the forward model used to create training sinograms, which image-domain generators cannot do.","The two-stage arrangement means residual sinogram errors are cleaned up after reconstruction, so projection-domain synthesis does not require perfectly noise-free raw data to be useful."],"supporting_citations":[{"why":"Supplies the latent diffusion backbone that DMPD builds on for projection-domain generation.","marker":"[13]"},{"why":"Supplies the abdominal CT slices from which PRO's projection training data are simulated.","marker":"[37]"},{"why":"Gives the Siddon ray-driven algorithm used to generate projection data from the CT slices.","marker":"[38]"},{"why":"Offers the fast radiological path calculation variant of Siddon's algorithm, also used in projection simulation.","marker":"[39]"},{"why":"Defines GMSD, the prior projection-domain diffusion model used as the downstream sparse-view reconstruction testbed.","marker":"[17]"},{"why":"Defines OSDM, the prior one-sample projection-domain diffusion model used as the downstream low-dose reconstruction testbed.","marker":"[18]"},{"why":"Shows an earlier latent diffusion model in the projection domain, but for sinogram inpainting only, which PRO positions itself against.","marker":"[35]"},{"why":"Shows a masked sinogram inpainting foundation model, the other comparison point for PRO's claim of unconditional projection synthesis.","marker":"[36]"}],"fun_headline_variants":["Text-prompted diffusion generates raw CT sinograms","Projection-domain synthesis improves CT image reconstruction","First CT foundation model synthesizes projection data","Synthetic sinograms rival real data for training CT models","Generate CT projections with text prompts, no real scans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that fan-beam projections computed by Siddon's ray-driven algorithm from CT slices behave like real clinical projections, so a diffusion model trained on those simulated sinograms will transfer to genuine scanner data.","fun_headline_variants_meta":{"raw":{"variants":["Text-prompted diffusion generates raw CT sinograms","Projection-domain synthesis improves CT image reconstruction","First CT foundation model synthesizes projection data","Synthetic sinograms rival real data for training CT models","Generate CT projections with text prompts, no real scans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000234,"raw_usage":{"total_tokens":1478,"prompt_tokens":910,"completion_tokens":568,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":495}},"tokens_in":526,"tokens_out":568,"duration_ms":5610,"temperature":1.0,"reasoning_tokens":495,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:00:37.726538+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate sinograms with PRO, inject the same physical effects a real scanner has (measured detector noise, beam hardening, scatter), reconstruct and train the downstream reconstruction networks on them; if performance falls to parity with or below training on real low-dose projections, the claim that projection-domain synthesis preserves acquisition physics is refuted.","supporting_citations":[],"review_version":2}