{"id":"88a5d602-86f1-44ba-926e-7ec0d8746894","arxiv_id":"2608.10429","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LAFNO, a Fourier neural operator with CT-derived contrast and disorder proxy conditioning and lesion-aware losses, improves per-lesion activity and tumor-core radiomics reproducibility in CT-to-PSMA-PET synthesis, while peritumoral fidelity remains tracer-dependent.","lead":"This paper reports a deep learning system, LAFNO, that turns CT scans into synthetic PSMA-PET images for prostate cancer while paying special attention to tumor lesions rather than only the overall image. The method adds two simple CT-derived cues (contrast and texture disorder) plus lesion-level losses, and it reports better per-lesion activity estimates and tumor texture preservation than four baselines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proxy scales and loss weights are selected on the same test cohort, so the reported TLA and radiomics improvements may be optimistic and are not shown to be statistically robust.","rationale":"The reader's verdict is CONDITIONAL, and my stress-test does not change that verdict; it reinforces the condition. The single most load-bearing concern is that the paper's headline numbers are produced from a test cohort on which the design hyperparameters were also selected. This is a concrete statistical flaw: Section 5.1 explicitly compares loss formulations on the same 18F cohort used for final evaluation and chooses the weight that appears best, while Section 2.2 gives no validation procedure for the proxy scales. The ablation in Table 5 even shows the final peritumoral term slightly worsens the headline TLA metric and all tumor-core ICCs, which is a sign that the final architecture was selected for reported test performance rather than for a principled objective. The 68Ga results provide a partial external check because hyperparameters were not re-tuned there, but the improvement over baselines is small (patient-level TLA 64.0% vs. 63.9% for cWDM), so the central claim of advantage is not clearly established. The concern is not about internal logic or honesty; the method is coherent and the ablations are directionally sensible. The problem is that the numeric claims lack statistical grounding. A proper validation split and repeated evaluation would settle whether the reported 48.3% TLA error and radiomics ICC improvements are real or partly overfit. Because the paper otherwise presents a plausible method and useful negative results (peritumoral fidelity remains hard), conditional acceptance with a request for validation is appropriate.","tokens_in":17798,"tokens_out":9580,"duration_ms":87703,"concrete_test":"Re-run the 18F experiments with a strict three-way patient split (e.g., 60/20/20) on the 18F cohort, tuning σ, w, λ_TLA, λ_c, and λ_peri only on the validation fold and evaluating each configuration exactly once on the held-out test fold. Repeat over at least 5 random split seeds and report mean and 95% confidence interval for patient-level TLA error. If the held-out advantage of LAFNO over AFNO-L1 is substantially smaller than the reported 13.9 percentage points, or if the difference is not statistically significant, the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper tunes all key design hyperparameters on the same 18F test cohort used for final evaluation, without a held-out validation set. Section 2.2 sets the proxy scales σ=5mm and w=7.5mm 'empirically'; Section 5.1 chooses the TLA loss formulation and λ_TLA=0.05 by inspecting curves on the 18F test cohort (Figure 7); Section 5.5 sets λ_peri=0.05 after observing its effect on the same cohort. With only N=47 test patients, no confidence intervals, and no repeated seeds, the headline improvement in patient-level TLA (48.3% vs. 62.2% for AFNO-L1) could be substantially driven by selection. An internal sign of this is Table 5: adding the peritumoral term raises TLA error from 47.8% to 48.3% and lowers every tumor-core ICC, yet the term is retained in the reported final model, suggesting the configuration was chosen to match the reported test metrics. The 68Ga cohort (N=30) is less contaminated by tuning, and there LAFNO's patient-level TLA (64.0%) is not better than cWDM's (63.9%). The absence of a precise patient-level split specification also leaves open the possibility of data leakage between training and test sets.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes LAFNO, a 3D U-Net with an Adaptive Fourier Neural Operator bottleneck for synthesizing PSMA-PET from CT in prostate cancer. The main novelties are two handcrafted CT-derived proxy channels (contrast and disorder) injected at the bottleneck, and a training objective that combines whole-volume L1 with lesion-wise total lesion activity (TLA), tumor-core contrast, and distance-weighted peritumoral losses. The authors evaluate on the TCIA PSMA-PET-CT-Lesions dataset with 18F- and 68Ga-PSMA cohorts against four baselines, reporting lesion-level and patient-level TLA error, fractions of lesions within error thresholds, and tumor-core/peritumoral radiomics ICC. They report that LAFNO remains competitive on SSIM/PSNR/MAE while improving TLA error and tumor-core radiomics reproducibility, with a cumulative ablation on the 18F cohort.","tokens_in":18090,"tokens_out":6874,"duration_ms":58397,"significance":"If the results are validated, LAFNO is a useful step toward clinically meaningful CT-to-PET synthesis because it targets lesion-level activity and radiomic structure rather than only global similarity, and the proxy-conditioning idea avoids per-inference radiomics extraction. Strengths include the cumulative ablation isolating each component, the use of lesion-level and radiomics evaluation beyond global metrics, and the explicit discussion of tracer-dependent limitations. However, the absence of a held-out validation set for hyperparameter selection, the lack of statistical inference, and an inconsistency in the TLA loss formulation mean the reported quantitative claims are not yet established. The work is therefore of interest to the medical image synthesis community, but the evidence requires strengthening before publication.","major_comments":[{"comment":"The proxy scales σ=5mm and w=7.5mm are described as chosen empirically, and the TLA loss formulation, λ_TLA=0.05, and λ_peri=0.05 are selected from curves and ablations computed on the same 18F test cohort used for the final evaluation (Figure 7, Table 5). With only N=47 test patients and no held-out validation or nested cross-validation, the reported improvements, e.g., patient-level TLA 48.3% vs. 62.2%, may be inflated by selection. Please either introduce a strict validation split for all hyperparameter choices or demonstrate that the conclusions are stable under repeated random splits.","section":"§2.2, §5.1, §5.5"},{"comment":"The cumulative ablation shows that adding the peritumoral term changes patient-level TLA error from 47.8% to 48.3%, lowers the ≤25% and ≤50% fractions from 25.2/54.9 to 24.0/50.9, and decreases every tumor-core ICC class (e.g., GLSZM from 0.760 to 0.709). The text states this term leaves the metrics 'essentially unchanged', which is contradicted by the table; it also raises the question of why the term is retained in the final model. This internal inconsistency, combined with the test-set tuning, undermines the ablation-based justification of the final configuration.","section":"Table 5"},{"comment":"The TLA loss is defined with ŝ_i and s_i as 'SUV values', but the model is trained on log-normalized PET values from Eq. (9). If the loss is computed on normalized outputs, then A_k and ^A_k are sums of nonlinearly transformed values and are not proportional to physical TLA; the inverse log transform in Eq. (9) is nonlinear, so the proportionality claim does not hold. Please state explicitly whether the inverse transform is applied before computing the TLA loss, or revise the loss definition and the interpretation of the reported TLA improvements accordingly.","section":"§2.4, Eq. (5); §3.2, Eq. (9)"},{"comment":"For 68Ga-PSMA, the patient-level TLA error of LAFNO is 64.0% versus 63.9% for cWDM, so LAFNO does not reduce patient-level TLA error relative to the best baseline on that tracer. Since the 68Ga cohort is less affected by the tuning performed on 18F, this equality weakens the cross-tracer generalization claim and should be reported as a limitation rather than as a reduction.","section":"§3.1, Table 3"},{"comment":"The manuscript does not report confidence intervals, significance tests, or repeated-seed variability for the main TLA and ICC comparisons, and it does not specify the patient-level train/test split (e.g., whether all scans of a patient were confined to one split). These details are essential for assessing whether the reported differences are robust, and the split specification is needed to rule out leakage. Please add this information.","section":"§3.2, §3.3"}],"minor_comments":[{"comment":"The text refers to 'Figure 8' for peritumoral reproducibility by distance band, but the figure with that content is labeled Figure 6; the later ablation figure is Figure 8. Please fix the cross-reference.","section":"Section 4.3"},{"comment":"The sentence 'with this value, the weight decreases to half of its boundary value from the tumor boundary' is inaccurate: with τ=5mm the half-value distance is τ ln 2 ≈ 3.47mm, not 5mm, where the weight is e^{-1} ≈ 0.368 of the boundary value.","section":"Eq. (7)"},{"comment":"The abstract's phrase 'reducing per-patient TLA error to ... 64.0%' for 68Ga is misleading given that cWDM achieves 63.9% on the same metric; please qualify the claim.","section":"Abstract"},{"comment":"The paragraph describing the final choice of λ_TLA states that the overestimation rate is below 20% in the final model; the corresponding values (19.98% and 10.47%) appear only in the text and not in Table 3, making the table incomplete for this claim.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from releasing the exact train/test split and, if possible, code, because the tuning-on-test issue makes independent validation critical. Also, several 2026 references are cited; please verify that they are publicly available and correctly represented."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead LAFNO for CT-to-PSMA PET synthesis. The genuinely new piece is converting radiomics-derived CT signatures into two cheap, differentiable proxy channels (contrast and disorder) and injecting them into an AFNO bottleneck, then supervising with per-lesion TLA, tumor-core contrast, and peritumoral losses. That combination isn't in the cited literature, and the cumulative ablation is clean: each component adds something, with the TLA loss and contrast loss doing the heavy lifting. The radiomics motivation (core-to-periphery attenuation drop, entropy increase) is supported by data on 10k lesions, and the paper is honest about residual errors: TLA error is still about 48% for 18F and 64% for 68Ga, and peritumoral fidelity is tracer-dependent.\n\nWhere I'd push back: the reported gains are calibrated on the test cohort, not a held-out set. Section 5.1 selects the TLA loss weight by looking at curves on the same 18F test patients that produce Table 3's numbers; proxy scales sigma and w are also chosen empirically on that dataset. With N=47, no confidence intervals, and no repeated seeds, the 62.2% to 48.3% improvement could easily shrink under proper validation. The internal sign is Table 5: adding the peritumoral term worsens TLA (47.8% to 48.3%) and lowers every tumor-core ICC, yet it is kept. That's not disqualifying, since the authors explain the term's effect is narrow, but it does look like the configuration was picked to match a preferred narrative. On the 68Ga cohort, which was less used for tuning, LAFNO's patient-level TLA (64.0%) is not better than cWDM's (63.9%). So the cross-tracer evidence for the lesion-aware objective is thin.\n\nNone of this kills the core idea. The proxy channels are cheap, mask-free, and the ablation shows they help even without lesion supervision. The method would likely still improve TLA over plain AFNO after honest held-out tuning; the magnitude just shouldn't be taken at face value. I'd want to see code and a precise patient-level split, plus at least one validation split for choosing loss weights and proxy scales, before trusting the absolute numbers.\n\nThis paper is worth a serious referee: the design is novel enough, the writing is clear, and the limitations are discussed openly. I'd send it to review but with an explicit request for a held-out tuning story and significance testing. I wouldn't cite the quantitative claims in my own work until that exists.","headline":"A well-motivated lesion-aware synthesis method whose headline TLA gains are likely optimistic because the loss weights and proxy scales were chosen on the test cohort.","tokens_in":18671,"tokens_out":2762,"would_cite":false,"duration_ms":24845,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that conditioning a Fourier-neural-operator synthesis model on two CT-derived proxy channels and training with lesion-aware losses preserves lesion activity and tumor radiomics in synthetic PSMA-PET, where global…","keywords":["PSMA PET","Synthetic PET","Prostate cancer","Lesion-aware learning","Adaptive Fourier neural operator","Tumor microenvironment","Radiomics"],"falsifier":"On an independent cohort of PSMA-PET/CT patients from different scanners, tracers, or institutions, train LAFNO and the CT-only AFNO-L1 baseline with the proxy scales fixed at $\\sigma = 5$ mm and $w = 7.5$ mm; if per-patient TLA error and tumor-core radiomics reproducibility no longer improve over the baseline, the proxy conditioning claim is falsified. A cheaper check: replace the two proxy channels with identically shaped channels containing CT intensity or white noise, holding all losses fixed; if TLA error and ICC do not worsen, the proxies' specific encoding carries no weight.","tokens_in":17599,"feed_emoji":"🩻","tokens_out":10165,"duration_ms":86383,"temperature":0.7,"pith_summary":"Whole-body PSMA-PET carries the clinically decisive tumor signal in a tiny fraction of voxels, so models trained to maximize whole-volume similarity can look good while underestimating lesion activity. LAFNO is a CT-to-PET synthesis model that replaces expensive radiomics conditioning with two cheap CT-derived proxy channels—local density contrast and local texture disorder—and trains with lesion-level total-activity, tumor-core contrast, and peritumoral supervision. On a multi-scanner, two-tracer prostate cancer dataset, it reduces mean per-patient total-lesion-activity error to 48.3 percent for 18F-PSMA and 64.0 percent for 68Ga-PSMA, improves tumor-core radiomics reproducibility in every feature class for both tracers, and stays competitive on whole-volume quality. If the result holds, synthetic PET could support lesion quantification and radiomics without administering a tracer or delineating lesions at inference time.","feed_headline":"CT proxies cut lesion-activity error in synthetic PSMA-PET to 48%","feed_subtitle":"Tumor-level supervision keeps synthetic scans usable for quantifying PSMA lesion burden in prostate cancer.","key_machinery":"The central mechanism is the pair of differentiable CT proxy channels plus a lesion-aware loss. The contrast proxy is a Gaussian residual filter that captures the smooth density gradient from tumor core to periphery; the disorder proxy is a sliding-window local variance that captures peritumoral texture heterogeneity. They are computed from CT alone, so at inference no lesion segmentation is needed. They are pooled (average for contrast, max for disorder) and concatenated into the Adaptive Fourier Neural Operator bottleneck, which performs spectral channel mixing so the proxy-conditioned features influence long-range spatial structure. The lesion-wise total-activity loss uses a log-compressed, per-lesion-normalized summed-activity error so small lesions are not dominated by large ones; the tumor-contrast loss supervises the contrast operator on predicted versus ground-truth SUV within the tumor mask; and the peritumoral loss applies an exponential distance-decay weight in the 0–10 mm ring. Together these terms push the model to preserve lesion-level activity and tumor-adjacent structure rather than optimizing global similarity alone.","core_discovery":"The paper's central claim is that a CT-to-PSMA-PET synthesis model conditioned on radiomics-motivated CT proxies and supervised at the lesion level preserves the clinically relevant PET signal that global L1 or MSE training discards. The contrast proxy $C(\\mathbf{x}) = \\mathrm{CT}(\\mathbf{x}) - G_\\sigma * \\mathrm{CT}(\\mathbf{x})$ with $\\sigma = 5$ mm encodes the observed core-to-periphery CT attenuation gradient, and the disorder proxy $D(\\mathbf{x})$ is local variance over a $7.5$ mm window, encoding tumor-adjacent texture heterogeneity; both are injected into the Adaptive Fourier Neural Operator bottleneck of a 3D U-Net. The loss combines whole-volume L1 with a per-lesion total-activity term, a tumor-core contrast term, and an exponentially distance-weighted peritumoral term. In the reported evaluation LAFNO attains SSIM 0.960 and 0.938 and per-patient TLA errors 48.3 and 64.0 percent for 18F- and 68Ga-PSMA, and the highest tumor-core radiomics ICCs across first-order, GLCM, GLRLM, and GLSZM classes for both tracers, while peritumoral reproducibility remains tracer-dependent. The authors interpret the remaining error as inherent one-to-many CT-to-PET mapping and tracer or SUV heterogeneity rather than a failure that purely architectural changes would fix.","pith_inferences":["If the proxy scales were learned per scanner or tracer rather than fixed on the evaluation dataset, TLA gains might transfer more reliably across institutions; this is a testable extension not performed by the paper.","Because the proxies need no lesion segmentation, coupling LAFNO with an automatic lesion detector would give a fully mask-free, CT-only pipeline for PSMA lesion burden.","The 68Ga weakness and the peritumoral reproducibility peak at 3–5 mm suggest positron range and partial-volume effects are limiting; partial-volume correction before training or evaluation is a natural next test.","The peritumoral term's effect is essentially confined to GLSZM in the 0–3 mm band, implying a single distance-weighted L1 penalty is too weak to shape the whole tumor-adjacent microenvironment; adversarial or multi-scale texture losses are worth trying."],"forward_implications":["On 18F-PSMA, per-patient TLA error drops to 48.3 percent from 58.7 to 69.0 percent for the four baselines, so a trained model can estimate lesion activity from CT alone at inference.","On 68Ga-PSMA, LAFNO reaches 64.0 percent per-patient TLA error and the highest fraction of lesions within 50 and 75 percent error, though cWDM is comparable on absolute lesion-level error.","Tumor-core radiomics reproducibility is the highest of all compared models for every feature class on both tracers, indicating the synthetic volumes carry texture-level information rather than only pixel similarity.","The ablation shows proxy conditioning alone cuts TLA error from 62.2 to 56.5 percent even without lesion masks or lesion-specific losses, so the CT-derived channels contribute independent information.","Whole-volume SSIM (0.960 and 0.938) and PSNR remain competitive with baselines, so the lesion-aware objectives do not trade away global image quality."],"supporting_citations":[{"why":"Supplies the whole-body PSMA-PET/CT dataset with manually annotated lesions used for the radiomics analysis and all training and evaluation.","marker":"Jeblick et al. (2026)"},{"why":"Defines the Adaptive Fourier Neural Operator bottleneck that LAFNO uses for spectral channel mixing.","marker":"Guibas et al. (2021)"},{"why":"Defines the radiomics feature classes and extraction conventions used to establish the core-to-periphery attenuation and entropy trends.","marker":"Zwanenburg et al. (2016)"},{"why":"Demonstrates radiomics-conditioned tumor synthesis, the approach LAFNO replaces with cheap CT-derived proxies.","marker":"Kim et al. (2025)"},{"why":"Provides the Pix2Pix adversarial baseline used in the comparison.","marker":"Isola et al. (2017)"},{"why":"Provides the FlowLet flow-matching baseline used in the comparison.","marker":"Danese et al. (2026)"},{"why":"Provides the cWDM conditional diffusion baseline used in the comparison.","marker":"Friedrich et al. (2024)"},{"why":"Supplies the U-Net encoder–decoder structure on which LAFNO's backbone is built.","marker":"Ronneberger et al. (2015)"}],"fun_headline_variants":["Lesion-aware CT proxies sharpen synthetic PSMA-PET","Tumor-level supervision halves TLA error in PSMA-PET","CT proxy channels preserve PSMA lesion signal without radiomics","LAFNO: radiomics-motivated CT proxies aid PSMA-PET synthesis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the two hand-coded CT filters—a 5 mm Gaussian residual and a 7.5 mm local variance—capture CT patterns that actually predict PSMA-PET uptake in and around lesions, and that the chosen scales generalize beyond the dataset where they were selected.","fun_headline_variants_meta":{"raw":{"variants":["Lesion-aware CT proxies sharpen synthetic PSMA-PET","Tumor-level supervision halves TLA error in PSMA-PET","CT proxy channels preserve PSMA lesion signal without radiomics","LAFNO: radiomics-motivated CT proxies aid PSMA-PET synthesis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0006,"raw_usage":{"total_tokens":2935,"prompt_tokens":1210,"completion_tokens":1725,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":826,"completion_tokens_details":{"reasoning_tokens":1651}},"tokens_in":826,"tokens_out":1725,"duration_ms":13640,"temperature":1.0,"reasoning_tokens":1651,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:20:01.015097+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On an independent cohort of PSMA-PET/CT patients from different scanners, tracers, or institutions, train LAFNO and the CT-only AFNO-L1 baseline with the proxy scales fixed at $\\sigma = 5$ mm and $w = 7.5$ mm; if per-patient TLA error and tumor-core radiomics reproducibility no longer improve over the baseline, the proxy conditioning claim is falsified. A cheaper check: replace the two proxy channels with identically shaped channels containing CT intensity or white noise, holding all losses fixed; if TLA error and ICC do not worsen, the proxies' specific encoding carries no weight.","supporting_citations":[{"cited_title":"u stner, T. , author Gatidis, S. , author Fr \\","cited_arxiv_id":null,"evidence_quote":"Supplies the whole-body PSMA-PET/CT dataset with manually annotated lesions used for the radiomics analysis and all training and evaluation."},{"cited_title":", author Na, I","cited_arxiv_id":null,"evidence_quote":"Demonstrates radiomics-conditioned tumor synthesis, the approach LAFNO replaces with cheap CT-derived proxies."},{"cited_title":", author Zhu, J.Y","cited_arxiv_id":null,"evidence_quote":"Provides the Pix2Pix adversarial baseline used in the comparison."},{"cited_title":", author Lombardi, A","cited_arxiv_id":null,"evidence_quote":"Provides the FlowLet flow-matching baseline used in the comparison."}],"review_version":1}