REVIEW 2 major objections 5 minor 34 references
FORCE-Interior: A Poisson Flow Generative Prior for Interior Tomography Reconstruction
T0 review · 2 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read FORCE-Interior shows that a Poisson-flow generative prior, anchored to the measured sinogram at every sampling step and started from a full-FOV reconstruction, improves structural and perceptual quality in interior tomography under severe R
desk verdict A competent and mostly honest extension of FORCE to interior CT, but the headline p<0.001 claims are overstated because hyperparameters were tuned on test-patient slices and the sample is small and correlated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the combination of a full-FOV OS-SART warm-start and per-step data-consistency conditioning on the ROI-truncated sinogram, inside a Poisson-flow generative sampler (an EDM-based PFGM++ that learns a denoiser/score for full-dose CT images). The full-FOV initialization accounts for out-of-ROI attenuation, while per-step OS-SART updates enforce agreement with the measured truncated projections at every denoising iteration, preventing the sampling trajectory from drifting to measurement-inconsistent anatomy. A lightweight total-variation proximal step moderates noise amplification from repeated OS-SART corrections.
What would settle it
A concrete check: retune the TV weight and starting noise level on a truly separate validation set disjoint from the evaluation slices, then re-run the comparison; if the PSNR/SSIM/LPIPS margins at radii 96 and 128 shrink to non-significance, the claim of out-of-the-box improvement is weakened. Additionally, applying the method to a different public CT dataset with varied anatomy and no structure-rich filtering would test whether the gains generalize beyond the specific test pool.
Extended reading notes
Core claim
FORCE-Interior solves the ROI-truncated inverse problem y=MAx+n by sampling from a Poisson-flow generative prior (PFGM++) while enforcing data consistency. Two design choices carry the argument: (1) initializing the sampler with a full-FOV OS-SART reconstruction rather than an ROI-restricted one, so that out-of-ROI attenuation is accounted for and cupping/bias artifacts are reduced; (2) at each denoising step, applying OS-SART updates conditioned on the truncated sinogram before the ODE step, keeping the trajectory anchored to the measurements. Experiments report that at ROI radii 96 and 128 pixels, FORCE-Interior achieves the best ROI PSNR, SSIM, and LPIPS among OS-SART, SART-TV, ROI-CT-CNN
Load-bearing premise
The reported gains rest on the assumption that the hyperparameters (TV weight lambda=0.003 and starting noise level t_start=0.6) chosen by ablations on a small random subset are not effectively tuned to the structure-rich test slices, so the improvements are not an artifact of selection on the evaluation set.
Editorial extensions
If this is right
- If the central claim holds, generative-model-based CT reconstruction can be extended to interior tomography by explicitly modeling the truncation mask in the data-consistency update.
- The full-FOV warm start implies that out-of-ROI attenuation must be accounted for even when only the ROI is of interest, challenging ROI-restricted initialization practices.
- Per-step data consistency keeps projection-domain residuals near the noise floor; removing it raises residuals by roughly 24x at radius 128, indicating that the anchoring is essential.
- The method retains small inserted lesions (false-negative rate 0) while a generative prior without data consistency misses all lesions, suggesting the anchoring preserves clinically relevant structures.
- The advantage grows as truncation becomes more severe, so the approach is most relevant for tightly targeted ROIs where the interior problem is hardest.
Reading between the lines
- The same warm-start-plus-per-step-DC recipe could likely be applied to other generative priors (e.g., score-based diffusion) to test whether the gains stem from the PFGM++ prior specifically or from the anchoring scheme generally.
- The noise-level sweep shows no per-dose retuning is needed, suggesting a path toward protocol-agnostic reconstruction, but the fixed TV weight chosen at the main dose may slightly oversmooth high-dose data, as the paper itself notes for LPIPS.
- A testable extension would evaluate the method on volumetric or cone-beam geometries and on anatomy beyond the structure-rich subset, since the reported statistics are slice-level and from two patients, which may not capture full clinical variability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FORCE-Interior, a generative reconstruction framework for interior tomography that combines a pre-trained PFGM++ (Poisson-flow generative) prior with a full-FOV OS-SART warm start and per-step data-consistency conditioning against truncated ROI measurements. The method is evaluated on the 2016 NIH-AAPM-Mayo Low-Dose CT dataset at three ROI radii (96, 128, 160 px) under Poisson noise, against OS-SART, SART-TV, ROI-CT-CNN, and DDS. The central claim is that FORCE-Interior achieves improved structural and perceptual quality (ROI PSNR, SSIM, LPIPS) at the two smaller ROI sizes with Holm-adjusted p<0.001, and competitive results at the largest, while preserving projection-domain consistency. Ablations and additional analyses (warm-start variants, TV-weight sweep, noise-level sweep, projection-residual consistency, small-lesion retention) are provided.
Significance. If the central claim holds, the paper makes a useful methodological contribution: it demonstrates how a foundation-style generative prior can be adapted to the interior-tomography setting through a measurement-anchored warm start and per-step data consistency, rather than treating generative sampling as a plug-in post-processor. The experimental design is careful in several respects: paired Wilcoxon tests with Holm correction, multiple ROI radii, a dose sweep without per-dose retuning, projection-domain consistency metrics, and a lesion-retention analysis. These features make the evaluation substantially more informative than a simple PSNR table. However, the headline significance claims are weakened by the unresolved provenance of the hyperparameter-selection subset and by slice-level statistics over two patients, as detailed below.
major comments (2)
- [Section IV-D3, Fig. 5, Table V, Appendix A] The two key hyperparameters (TV weight λ=0.003 and starting noise level t_start=0.6) are selected via ablations on a 'random 20-slice subset.' The manuscript never states that this subset is disjoint from the structure-rich test set (N=34 at r=128) on which the headline comparisons in Table I are made. Appendix A mentions a separate pool of 20 held-out slices used for the radius sweep, but does not state whether the hyperparameter ablations use that pool or a random subset of the test partition. If the latter, the test data have influenced the model configuration, making the subsequent paired Wilcoxon p<0.001 results in Table VIII partially circular for those hyperparameters. Please clarify the provenance of the ablation subset, and ideally re-select hyperparameters on a truly held-out validation set and re-run the benchmark.
- [Appendix A, Statistical analysis; Section IV-A] The significance claims rest on slice-level paired Wilcoxon tests with N=18, 34, and 16 slices drawn from only two patients. The appendix concedes that 'adjacent slices from the same patient may not be fully independent,' but no correction is made for within-patient correlation. Treating correlated slices as independent inflates the effective sample size; with two patients, the reported Holm-adjusted p<0.001 are not credible as evidence about patient-level generalization. Please provide patient-level summaries (e.g., per-patient mean differences), a mixed-effects analysis, or a conservative cluster-based correction, and temper the abstract's significance claim accordingly.
minor comments (5)
- [Abstract/Introduction] Typo: 'proposeFORCE-Interior' should be 'propose FORCE-Interior'.
- [Table I caption] The heading 'POISSON-NOISE RECONSTRUCTION ATr∈{96,128,160}PX' appears to have a formatting error ('ATr' should be 'at r'). Also, the ROI-CT-CNN caveat (trained for r=128 only, evaluated under a shifted distribution) should appear in the table caption since the table includes it at r=128.
- [Fig. 5 caption / Section IV-D3] Please clarify which 'random 20-slice subset' is used for the TV-weight and t_start ablations and whether it is the same as the 'separate pool of 20 held-out slices' described in Appendix A. This is important for the reader to assess potential leakage.
- [Algorithm 3 / Eq. (14)] The notation M_g A_{S_g} and the update equation could be defined more explicitly; currently it is not immediately clear how the subset mask M_g interacts with the projection rows A_{S_g}. A brief sentence connecting Eq. (14) to the truncation mask M in Eq. (10) would help.
- [Section IV-C2 / Appendix A] The lesion-retention analysis uses '20 structure-rich test slices.' It would be useful to confirm whether these are the same 20 slices used for hyperparameter tuning or the radius-sweep pool, and to report the overlap explicitly.
Circularity Check
No significant circularity: central claims rest on held-out empirical comparisons, not on a fitted parameter or self-citation chain; noted caveats are statistical, not constructional.
full rationale
The paper does not derive its headline result from its own inputs by construction. FORCE-Interior is defined by explicit equations (Algorithm 3; Eqs. (13)-(14)) combining a pretrained PFGM++ prior, full-FOV OS-SART warm start, per-step OS-SART data consistency, and TV proximal steps, and it is evaluated against OS-SART, SART-TV, ROI-CT-CNN, and DDS on two held-out Mayo patients. The central metric improvements are empirical outcomes, not identities. The reliance on the authors' own FORCE/PFGM++ prior is a normal self-citation: the prior's training objective (Eq. (7)) does not include the interior-tomography benchmark, and the measured gains come from the experimental protocol rather than from the prior's objective. Likewise, the projection-consistency results are not circular: they are presented as a check of the enforced constraint, with an ablation (w/o per-step DC) showing a 24x residual increase, so the analysis has independent content. The appendix explicitly notes that adjacent slices from the same two patients may not be fully independent, and the TV weight/t_start were selected on a random 20-slice subset (Section IV-D3, Table V); these are statistical validity caveats (potential dependence/selection on evaluation data) rather than a reduction of prediction to input. No equation in the paper equates a fitted parameter with the reported outcome by construction, and no load-bearing claim rests solely on an unverified self-citation. Hence no significant circularity; score 1 reflecting only minor self-citation and mild statistical caveats.
Assumptions & free parameters
free parameters (4)
- TV weight lambda =
0.003
- Starting noise level t_start =
0.6
- Per-step OS-SART settings =
G=16 subsets, 3 iterations, relaxation omega=1
- Warm-start OS-SART iterations =
30 iterations, 16 subsets
assumptions (4)
- standard math OS-SART converges as a data-consistency operator for truncated projections
- domain assumption The PFGM++/EDM denoiser approximates the score of the normal-dose CT image distribution
- ad hoc to paper A full-FOV OS-SART warm start from truncated projections captures out-of-ROI attenuation
- domain assumption Slice-level paired statistics on adjacent slices approximate independent samples
Cite this review
Pith. "Pith review of FORCE-Interior: A Poisson Flow Generative Prior for Interior Tomography Reconstruction." pith.science (2026). https://pith.science/paper/7YVMDPYI
@misc{pith2026260714320,
author = {Pith},
title = {Pith review of: FORCE-Interior: A Poisson Flow Generative Prior for Interior Tomography Reconstruction},
year = {2026},
howpublished = {\url{https://pith.science/paper/7YVMDPYI}},
note = {Machine review of arXiv:2607.14320}
}
read the original abstract
Interior tomography reconstructs a region of interest (ROI) from truncated projection measurements. However, projection truncation makes the inverse problem severely ill-posed, leading to non-unique solutions, oversmoothing, and truncation-induced artifacts when conventional reconstruction methods are directly applied to interior tomography. Existing learning-based interior CT methods have shown promising performance, but their generalization across different truncation patterns, ROI sizes, and noise levels remains an important challenge. Meanwhile, current generative model-based reconstruction methods are primarily designed for non-interior tomography settings and do not directly address ROI-based projection truncation. Moreover, without sufficient data-consistency constraints, generative sampling may yield anatomically plausible but measurement-inconsistent structures. To address these challenges, we propose FORCE-Interior, a Poisson-flow generative reconstruction framework for interior tomography. FORCE-Interior combines a full-FOV measurement-constrained initialization with per-step data consistency for ROI-truncated measurements, anchoring generative sampling to the acquired measurements throughout reconstruction. Experiments show that FORCE-Interior achieves improved structural and perceptual reconstruction quality at the two more severely truncated ROI sizes, with competitive results at the largest, while maintaining projection-domain consistency.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
The meaning of interior tomography,
G. Wang and H. Yu, “The meaning of interior tomography,”Phys. Med. Biol., vol. 58, no. 16, pp. R161–R186, 2013
2013
-
[2]
Compressed sensing based interior tomography,
H. Yu and G. Wang, “Compressed sensing based interior tomography,” Phys. Med. Biol., vol. 54, no. 9, pp. 2791–2805, 2009
2009
-
[3]
Image reconstruction for sparse-view CT and interior CT: Introduction to compressed sensing and differentiated backprojection,
H. Kudo, T. Suzuki, and E. A. Rashed, “Image reconstruction for sparse-view CT and interior CT: Introduction to compressed sensing and differentiated backprojection,”Quant. Imaging Med. Surg., vol. 3, no. 3, pp. 147–161, 2013
2013
-
[4]
A practical local tomography reconstruction algorithm based on a known sub-region,
P. Paleo, M. Desvignes, and A. Mirone, “A practical local tomography reconstruction algorithm based on a known sub-region,”Journal of Synchrotron Radiation, vol. 24, no. 1, pp. 257–268, 2017
2017
-
[5]
A. C. Kak and M. Slaney,Principles of Computerized Tomographic Imaging. New York, NY , USA: IEEE Press, 1988
1988
-
[6]
FBP and the interior problem in 2D tomography,
A. Bilgot, L. Desbat, and V . Perrier, “FBP and the interior problem in 2D tomography,” in2011 IEEE Nuclear Science Symposium Conference Record, 2011, pp. 4080–4085
2011
-
[7]
Simultaneous algebraic reconstruction technique (SART): A superior implementation of the ART algorithm,
A. H. Andersen and A. C. Kak, “Simultaneous algebraic reconstruction technique (SART): A superior implementation of the ART algorithm,” Ultrason. Imaging, vol. 6, no. 1, pp. 81–94, 1984
1984
-
[8]
Ordered-subset simultaneous algebraic recon- struction techniques (OS-SART),
G. Wang and M. Jiang, “Ordered-subset simultaneous algebraic recon- struction techniques (OS-SART),”J. X-Ray Sci. Technol., vol. 12, no. 3, pp. 169–177, 2004
2004
Show all 34 references
-
[9]
Bone-induced streak artifact suppression in sparse-view CT image reconstruction,
S. O. Jin, J. G. Kim, S. Y . Lee, and O.-K. Kwon, “Bone-induced streak artifact suppression in sparse-view CT image reconstruction,”BioMed. Eng. OnLine, vol. 11, p. 44, 2012
2012
-
[10]
Artifact reduction methods for truncated projections in iterative breast tomosynthesis reconstruction,
Y . Zhang, H.-P. Chan, B. Sahiner, J. Wei, C. Zhou, and L. M. Hadjiiski, “Artifact reduction methods for truncated projections in iterative breast tomosynthesis reconstruction,”J. Comput. Assist. Tomogr ., vol. 33, no. 3, pp. 426–435, 2009
2009
-
[11]
A diffusion-based truncated projection artifact reduction method for iterative digital breast tomosynthesis reconstruction,
Y . Lu, H.-P. Chan, J. Wei, and L. M. Hadjiiski, “A diffusion-based truncated projection artifact reduction method for iterative digital breast tomosynthesis reconstruction,”Phys. Med. Biol., vol. 58, no. 3, pp. 569– 587, 2013
2013
-
[12]
Image reconstruction in circular cone-beam computed tomography by constrained, total-variation minimization,
E. Y . Sidky and X. Pan, “Image reconstruction in circular cone-beam computed tomography by constrained, total-variation minimization,” Phys. Med. Biol., vol. 53, no. 17, pp. 4777–4807, 2008
2008
-
[13]
Prior image constrained compressed sensing (PICCS): A method to accurately reconstruct dynamic CT images from highly undersampled projection data sets,
G.-H. Chen, J. Tang, and S. Leng, “Prior image constrained compressed sensing (PICCS): A method to accurately reconstruct dynamic CT images from highly undersampled projection data sets,”Med. Phys., vol. 35, no. 2, pp. 660–663, 2008
2008
-
[14]
Solving inverse problems using data-driven models,
S. Arridge, P. Maass, O. ¨Oktem, and C.-B. Sch ¨onlieb, “Solving inverse problems using data-driven models,”Acta Numer ., vol. 28, pp. 1–174, 2019
2019
-
[15]
Low-dose ct reconstruction via edge-preserving total variation regularization,
Z. Tian, X. Jia, K. Yuan, T. Pan, and S. B. Jiang, “Low-dose ct reconstruction via edge-preserving total variation regularization,” Physics in Medicine and Biology, vol. 56, no. 18, p. 5949–5967, Aug. 2011. [Online]. Available: http://dx.doi.org/10.1088/0031- 9155/56/18/011
2011 doi
-
[16]
Deep learning interior tomography for region-of-interest reconstruction,
Y . Han, J. Gu, and J. C. Ye, “Deep learning interior tomography for region-of-interest reconstruction,” 2018. [Online]. Available: https://arxiv.org/abs/1712.10248
2018 arXiv
-
[17]
One network to solve all rois: Deep learning ct for any roi using differentiated backprojection,
Y . Han and J. C. Ye, “One network to solve all rois: Deep learning ct for any roi using differentiated backprojection,” 2019. [Online]. Available: https://arxiv.org/abs/1810.00500
2019 arXiv
-
[18]
End-to-end deep learning for interior tomography with low-dose x-ray ct,
Y . Han, D. Wu, K. Kim, and Q. Li, “End-to-end deep learning for interior tomography with low-dose x-ray ct,” 2025. [Online]. Available: https://arxiv.org/abs/2501.05085
2025 arXiv
-
[19]
Solving inverse problems in medical imaging with score-based generative models,
Y . Song, L. Shen, L. Xing, and S. Ermon, “Solving inverse problems in medical imaging with score-based generative models,” inProc. Int. Conf. Learn. Represent. (ICLR), 2022
2022
-
[20]
Dolce: A model-based probabilistic diffusion framework for limited-angle ct reconstruction,
J. Liu, R. Anirudh, J. J. Thiagarajan, S. He, K. A. Mohan, U. S. Kamilov, and H. Kim, “Dolce: A model-based probabilistic diffusion framework for limited-angle ct reconstruction,” 2022. [Online]. Available: https://arxiv.org/abs/2211.12340
2022 arXiv
-
[21]
Generative modeling in sinogram domain for sparse-view ct reconstruction,
B. Guan, C. Yang, L. Zhang, S. Niu, M. Zhang, Y . Wang, W. Wu, and Q. Liu, “Generative modeling in sinogram domain for sparse-view ct reconstruction,” 2022. [Online]. Available: https://arxiv.org/abs/2211.13926
2022 arXiv
-
[22]
Tomographic foundation model – force: Flow-oriented reconstruction conditioning engine,
W. Xia, C. Niu, and G. Wang, “Tomographic foundation model – force: Flow-oriented reconstruction conditioning engine,” 2025. [Online]. Available: https://arxiv.org/abs/2506.02149
2025 arXiv
-
[23]
Decomposed diffusion sampler for accelerating large-scale inverse problems,
H. Chung, S. Lee, and J. C. Ye, “Decomposed diffusion sampler for accelerating large-scale inverse problems,” 2024. [Online]. Available: https://arxiv.org/abs/2303.05754
2024 arXiv
-
[24]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” 2020. [Online]. Available: https://arxiv.org/abs/2006.11239
2020 arXiv
-
[25]
Score-based generative modeling through stochastic differ- ential equations,
Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differ- ential equations,” inProc. Int. Conf. Learn. Represent. (ICLR), 2021
2021
-
[26]
Elucidating the design space of diffusion-based generative models,
T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models,” inProc. Adv. Neural Inf. Process. Syst., 2022
2022
-
[27]
Poisson flow generative models,
Y . Xu, Z. Liu, M. Tegmark, and T. Jaakkola, “Poisson flow generative models,” inProc. Adv. Neural Inf. Process. Syst., 2022
2022
-
[28]
PFGM++: Unlocking the potential of physics-inspired generative mod- els,
Y . Xu, Z. Liu, Y . Tian, S. Tong, M. Tegmark, and T. Jaakkola, “PFGM++: Unlocking the potential of physics-inspired generative mod- els,” inProc. Int. Conf. Mach. Learn. (ICML), 2023, pp. 38 566–38 591
2023
-
[29]
Image quality assessment: From error visibility to structural similarity,
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,”IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, 2004
2004
-
[30]
The unreasonable effectiveness of deep features as a perceptual metric,
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 586–595
2018
-
[31]
A first-order primal-dual algorithm for convex problems with applications to imaging,
A. Chambolle and T. Pock, “A first-order primal-dual algorithm for convex problems with applications to imaging,”J. Math. Imaging Vis., vol. 40, no. 1, pp. 120–145, 2011
2011
-
[32]
Tigre v3: Efficient and easy to use iterative computed tomographic reconstruction toolbox for real datasets,
A. Biguri, T. Sadakane, R. Lindroos, Y . Liu, M. Sabat´e Landman, Y . Du, M. Lohvithee, S. Kaser, S. Hatamikia, R. Bryll, E. Valat, S. Wonglee, T. Blumensath, and C.-B. Sch ¨onlieb, “Tigre v3: Efficient and easy to use iterative computed tomographic reconstruction toolbox for ...
-
[33]
DM4CT: Benchmarking dif- fusion models for computed tomography reconstruction,
J. Shi, D. M. Pelt, and K. J. Batenburg, “DM4CT: Benchmarking dif- fusion models for computed tomography reconstruction,”arXiv preprint arXiv:2602.18589, 2026
2026
-
[2025]
Available: https://doi.org/10.1088/2631-8695/adbb3a
[Online]. Available: https://doi.org/10.1088/2631-8695/adbb3a
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.