REVIEW 3 major objections 4 minor 2 cited by
Joint Reconstruction of the Activity and the Attenuation in PET by Diffusion Posterior Sampling: a Feasibility Study
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Diffusion posterior sampling jointly reconstructs PET activity and attenuation from emission data alone, outperforming MLAA without time-of-flight information in 2D phantom tests.
desk verdict Feasibility study shows DPS can beat MLAA on non-TOF in-distribution phantoms, but the evidence is thin and the OOD non-TOF regime is untested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the joint score function that approximates the gradient of the log-density of noisy two-channel images, trained by score matching on paired activity and attenuation slices. During reconstruction, DPS approximates the conditional score by Bayes' rule, with the likelihood term evaluated at Tweedie's denoised estimate of the clean image; the reverse diffusion step is then followed by gradient updates on the activity and attenuation channels separately. The joint score is what distinguishes DPS from its independent-channel variant: the latter factorizes the score into separate activity and attenuation networks, so it assumes independence, and the paper's comparison isolates the contribution of the joint prior.
What would settle it
Run DPS on real 3D patient PET data for which a CT-derived attenuation map and an independently corrected activity image are available: if non-TOF DPS reconstructions fail to reproduce the CT-based attenuation map or introduce training-set biases on anatomies not present in the phantom data, the central claim that DPS resolves crosstalk without TOF would be refuted.
Extended reading notes
Core claim
The paper's central claim is that the activity–attenuation crosstalk that breaks non-TOF joint reconstruction can be resolved by replacing the missing joint prior with a diffusion-model score function trained on paired activity and attenuation images. In the DPS framework, the posterior over the two-channel image is sampled by alternating the reverse-diffusion update with a likelihood-gradient step computed from the forward model; the score network supplies the prior, and the joint training is what keeps the two channels consistent with each other. On 2D phantom testing slices, the paper reports that DPS without TOF outperforms MLAA with TOF on PSNR and SSIM for both activity and attenuation, and that DPS without TOF outperforms the independently trained variant with TOF. The authors interpret these results as evidence that explicitly modeling activity–attenuation dependencies is more valuable than adding TOF information, at least within the tested phantom setting.
Load-bearing premise
The method's advantage rests entirely on the learned joint prior being a faithful stand-in for real activity and attenuation maps; the paper's own out-of-distribution tumor experiments show that when an anatomy falls outside the training distribution the reconstructed shapes become piecewise-constant artifacts, so this prior fidelity is the assumption most at risk.
Editorial extensions
If this is right
- If the central claim holds, non-TOF PET scanners could perform joint activity and attenuation reconstruction from emission data alone, removing the need for CT or MR attenuation correction in some workflows.
- The observed superiority of jointly trained DPS over independently trained channels implies that cross-channel dependencies in the prior are the key mechanism for suppressing crosstalk, a lesson transferable to other joint reconstruction problems.
- The success of a learned prior on phantom data suggests that a sufficiently rich training corpus of activity and attenuation pairs could make DPS a practical alternative to MLAA in low-dose or PET-only settings.
- The out-of-distribution tumor results indicate that performance is tied to the training distribution, so deployment would require training data covering the relevant clinical variability.
- Extending the approach to 3D volumes and real patient data is the stated next step, with the paper reporting encouraging preliminary findings along that path.
Reading between the lines
- A testable extension is to quantify how much of the gain is prior strength versus joint modeling by comparing the jointly trained score against a score trained on paired images with the channel pairing scrambled, which would isolate the dependency-learning effect.
- The out-of-distribution artifacts suggest that DPS may act as a learned shape prior rather than a generic statistical prior; if so, its clinical value hinges on the coverage of the training set, much like other learned reconstruction methods.
- Since the paper ignores scatter and random coincidences, a natural next experiment is to include background terms in the forward model; a mismatched likelihood gradient could shift the crosstalk balance and test how much the prior compensates.
- The result also hints that TOF information may be partly redundant when a strong joint prior is available; a direct comparison of TOF and non-TOF DPS on the same phantom would quantify the remaining value of TOF under DPS.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a joint reconstruction of the activity and the attenuation in PET using diffusion posterior sampling (DPS). The authors train a diffusion model on pairs of 2-D XCAT activity/attenuation images, then use DPS to sample from the posterior conditioned on TOF or non-TOF emission data. Experiments on 10 test slices report higher PSNR/SSIM than MLAA and than a variant with independently trained priors (DPS2), for both TOF and non-TOF data. Out-of-distribution tests with tumors are performed only with TOF data. The paper concludes that DPS mitigates activity-attenuation crosstalk in non-TOF settings.
Significance. The central idea is sound and timely: a learned joint prior over activity and attenuation is a plausible route to resolve the cross-talk that limits non-TOF MLAA. The paper cleanly separates the joint-prior model (DPS) from an independent-prior baseline (DPS2), which helps isolate the effect of modeling dependencies. The mathematical formulation follows the standard DPS derivation and the forward model is physical. If confirmed on larger and more diverse datasets, the results would be an important step toward emission-only attenuation correction. The main weaknesses are the small test set and the absence of non-TOF out-of-distribution experiments, which are needed to support the abstract's central claim.
major comments (3)
- [Section 3.2, Figure 2] The claim that DPS 'significantly outperforms' MLAA and DPS2 is not supported by any statistical analysis. The evaluation uses 10 test slices; Figure 2 shows only point clouds and means without error bars, confidence intervals, or significance tests. Several reported differences (e.g., DPS no TOF vs DPS2 no TOF for attenuation, with PSNR 25.11 vs 24.25) appear small relative to inter-slice scatter. Please provide paired comparisons with confidence intervals, per-metric distributions, and a justified sample size, or temper the wording.
- [Section 3.2, Figure 3] The out-of-distribution evaluation is performed exclusively with TOF data. The paper's principal claim is that DPS works 'even in absence of TOF data'; in the non-TOF setting the likelihood provides no TOF information to separate activity from attenuation, so the learned prior must carry nearly the entire disambiguation burden. The OOD TOF results already show a strong prior bias (a Gaussian tumor reconstructed as piecewise-constant in Figure 3(j) and Figure 4). Without non-TOF OOD experiments, the central claim is not verified for anatomies outside the training distribution. Please add non-TOF OOD reconstructions or explicitly restrict the claim to in-distribution anatomies.
- [Section 3.1] The test set consists of only 10 slices drawn from the same XCAT distribution as the training set, with no tumors. The near-perfect in-distribution results may therefore reflect memorization rather than generalization; the paper itself raises this concern in Section 3.2, but the two OOD TOF cases are too few to resolve it. A larger held-out set with varied anatomies and pathologies is needed to substantiate the feasibility claim.
minor comments (4)
- [Section 3.2] The sentence 'We then performed applied MLAA and DPS' contains a duplicated verb; please revise.
- [Equation (9)] The norm is missing a closing parenthesis: it should read ∥s_θ(x_t,t) − ∇_{x_t} log p_t(x_t|x_0)∥₂².
- [Algorithm 1 and Section 3.1] The step sizes ζ_t and ξ_t, the number of diffusion steps T, and the noise schedule α_t are not reported; only τ = 5×10⁻¹ is given. Please provide these values to enable reproduction.
- [Figure 2] Adding per-method error bars or interquartile ranges would help the reader assess the overlap between DPS and DPS2.
Circularity Check
No significant circularity: DPS reconstruction uses a separately trained score prior and a physical forward model; reported limitations concern generalization, not derivation.
full rationale
The derivation chain is self-contained and does not reduce to its inputs. The forward model in Eqs. (1)-(3) is a physical Poisson/Beer-Lambert model, and the DPS update in Algorithm 1 follows the original DPS formulation with the standard approximation in Eq. (11); the likelihood is not fitted to the reconstructions. The prior is a score network trained by score matching on 4,000 XCAT pairs, and the reported reconstructions are evaluated on 10 held-out slices from different phantoms, so the test images are not used to fit the model. DPS2 is an independently trained variant that still uses the same forward model and sampling procedure, and MLAA is an independent classical baseline, so the claimed improvements are empirical comparisons rather than construction-identity arguments. The only self-citation, Ref. [13], is invoked for the non-load-bearing remark that similar observations were made in multi-energy CT; it does not support the central claim by itself. The paper's own limitations—training and testing on the same XCAT simulator, and out-of-distribution tests performed only with TOF data—are generalization and benchmarking caveats, not circular reasoning. No equation or fitted parameter is secretly used as the predicted quantity.
Assumptions & free parameters
free parameters (3)
- Gradient step sizes ζ_t, ξ_t =
not reported
- Noise/guidance scale τ =
5×10^-1
- Number of diffusion steps T =
not reported
assumptions (5)
- domain assumption The Poisson measurement model in Eq. (1) accurately describes PET emission counts.
- domain assumption The Beer-Lambert law in Eq. (3) correctly relates attenuation factors to the attenuation map via the Radon transform.
- domain assumption The learned score function s_θ approximates the true log-density gradient of activity-attenuation pairs.
- standard math The DPS approximation in Eq. (11), replacing ∇_xt log p(y|xt) with ∇_xt log p(y|hat_x0(xt)), is accurate enough for this inverse problem.
- domain assumption The XCAT phantom dataset provides sufficient diversity for the learned prior to generalize to unseen anatomies.
Cite this review
Pith. "Pith review of Joint Reconstruction of the Activity and the Attenuation in PET by Diffusion Posterior Sampling: a Feasibility Study." pith.science (2026). https://pith.science/paper/GCDPBHD4
@misc{pith2026241211776,
author = {Pith},
title = {Pith review of: Joint Reconstruction of the Activity and the Attenuation in PET by Diffusion Posterior Sampling: a Feasibility Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/GCDPBHD4}},
note = {Machine review of arXiv:2412.11776}
}
read the original abstract
This study introduces a novel framework for joint reconstruction of the activity and the attenuation (JRAA) in positron emission tomography (PET) using diffusion posterior sampling (DPS). By leveraging diffusion models (DMs), this approach directly addresses activity-attenuation dependencies, mitigating crosstalk issues prevalent in non-time-of-flight (TOF) settings. Experimental evaluations, conducted using 2-dimensional (2-D) XCAT phantom data, demonstrate that DPS significantly outperforms traditional maximum likelihood activity and attenuation (MLAA) methods, producing consistent and high-quality reconstructions even in the absence of TOF information. Ongoing work aims to extend our method to real 3-dimensional (3-D) data with encouraging preliminary findings.
Figures
Forward citations
Cited by 2 Pith papers
-
Joint Reconstruction of Activity and Attenuation in PET by Diffusion Posterior Sampling in Wavelet Coefficient Space
A wavelet diffusion model combined with diffusion posterior sampling enables joint 3D activity-attenuation reconstruction in PET from emission data alone, outperforming MLAA on simulated TOF data.
-
Solving Blind Inverse Problems: Adaptive Diffusion Models for Motion-corrected Sparse-view 4DCT
A diffusion-based method that jointly reconstructs motion-corrected sparse-view 4DCT images and estimates respiratory motion, outperforming three baselines on XCAT phantoms.
Reference graph
Works this paper leans on
-
[1]
Simultaneous reconstruction of activity and attenuation in time-of-flight PET
A. Rezaei, M. Defrise, G. Bal, et al. “Simultaneous reconstruction of activity and attenuation in time-of-flight PET”.IEEE transactions on medical imaging31.12 (2012), pp. 2224–2233.doi: 10.1109/ NSSMIC.2011.6153883
-
[2]
Time-of-flight PET data deter- minetheattenuationsinogramuptoaconstant
M. Defrise, A. Rezaei, and J. Nuyts. “Time-of-flight PET data deter- minetheattenuationsinogramuptoaconstant”. PhysicsinMedicine& Biology 57.4 (2012), p. 885.doi: 10.1088/0031-9155/57/4/885
-
[3]
ML-reconstruction for TOF- PET with simultaneous estimation of the attenuation factors
A. Rezaei, M. Defrise, and J. Nuyts. “ML-reconstruction for TOF- PET with simultaneous estimation of the attenuation factors”.IEEE transactions on medical imaging33.7 (2014), pp. 1563–1572.doi: 10.1109/TMI.2014.2318175
arXiv 2014
-
[4]
Attenuation correction in emission tomography using the emission data—a review
Y. Berker and Y. Li. “Attenuation correction in emission tomography using the emission data—a review”.Medical physics43.2 (2016), pp. 807–832.doi: 10.1118/1.4938264
-
[5]
Deep-learning-based methods of attenuation correction for SPECT and PET
X. Chen and C. Liu. “Deep-learning-based methods of attenuation correction for SPECT and PET”.Journal of Nuclear Cardiology30.5 (2023), pp. 1859–1878.doi: 10.1007/s12350-022-03007-3
-
[6]
Diffusion posterior sampling for general noisy inverse problems
H. Chung, J. Kim, M. T. Mccann, et al. “Diffusion posterior sampling for general noisy inverse problems”.In Proceedings of the Eleventh International Conference on Learning Representations. ICLR, 2023. doi: 10.48550/arXiv.2209.14687
-
[7]
Diffusion models for medical image reconstruction
G. Webber and A. J. Reader. “Diffusion models for medical image reconstruction”. BJR| Artificial Intelligence1.1 (2024), ubae013.doi: 10.1093/bjrai/ubae013
-
[8]
Score-basedgenerativemod- els for PET image reconstruction
I.R.Singh,A.Denker,R.Barbano,etal.“Score-basedgenerativemod- els for PET image reconstruction”.arXiv preprint arXiv:2308.14190 (2023). doi: 10.48550/arXiv.2308.14190
Show all 13 references
-
[9]
Likelihood-Scheduled Score-Based Generative Modeling for Fully 3D PET Image Recon- struction
G. Webber, Y. Mizuno, O. D. Howes, et al. “Likelihood-Scheduled Score-Based Generative Modeling for Fully 3D PET Image Recon- struction”. arXiv preprint arXiv:2412.04339(2024). doi: 10.48550/ arXiv.2412.04339
-
[10]
4D XCAT phantom for multimodality imaging research
W. P. Segars, G. Sturgeon, S. Mendonca, et al. “4D XCAT phantom for multimodality imaging research”.Medical physics37.9 (2010), pp. 4902–4915.doi: 10.1118/1.3480985
2010 doi
-
[11]
A modified expectation maximization algorithm for penalized likelihood estimation in emission tomography
A. De Pierro. “A modified expectation maximization algorithm for penalized likelihood estimation in emission tomography”.IEEE Transactions on Medical Imaging14.1 (1995), pp. 132–137.doi: 10.1109/42.370409
1995 doi
- [12]
-
[13]
Diffusionposteriorsamplingfor synergistic reconstruction in spectral computed tomography
C.Vazia,A.Bousse,B.Vedel,etal.“Diffusionposteriorsamplingfor synergistic reconstruction in spectral computed tomography”.2024 IEEE 21st international symposium on biomedical imaging (ISBI 2024). IEEE. 2024.doi: 10.1109/ISBI56570.2024.10635735. 4
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.