{"id":"1768296b-f8b6-4fc2-88e1-7005b9735e69","arxiv_id":"2506.14719","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A 2.5D artifact-reduction CNN prior in plug-and-play reconstruction improves sparse-view X-ray CT quality and defect detection over a 2D prior.","lead":"The paper shows that a 2.5D neural network prior inside a plug-and-play CT reconstruction framework produces better sparse-view industrial X-ray reconstructions than the same framework with a 2D prior, recovering more small defects at nearly equal compute cost. It also demonstrates that the 2.5D prior, trained only on simulated data, suppresses beam-hardening artifacts and performs well on a real aluminum-cerium part.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All quantitative gains are measured against reconstructed references, not the true object; the synthetic CAD phantoms are known, so a ground-truth comparison is needed to validate the accuracy claim.","rationale":"I agree with the reader that the reference problem is the weakest link. It is load-bearing because every headline number in Tables 2 and 3 and every recall/precision curve is defined relative to reconstructed references. The reader's conditional verdict already captures this. I would not change the verdict: the paper is a useful empirical study, the 2.5D-vs-2D comparison is same-reference and therefore still informative for relative performance, and the synthetic/experimental results are consistent. Yet the 'accuracy' and 'eliminating BH pre-processing' claims are broader than what is measured. A secondary issue is that the BH claim is not benchmarked against standard BH correction preprocessing; adding such a baseline would also help, but the ground-truth phantom check is the more fundamental single test. The paper's own Section 7 admission of poor OOD view-sparsity generalization further supports scoping the claim, not rejection.","tokens_in":14603,"tokens_out":6408,"duration_ms":69674,"concrete_test":"Recompute the synthetic InD and OOD tables (NRMSE, SSIM, recall, precision for 75–125 µm flaws) for 2D PnP and 2.5D PnP against the CAD-derived clean volume used to generate projection data, rather than the 2132-view FDK reconstruction. If 2.5D PnP remains better on ground-truth metrics, the reference-artifact objection is substantially weakened; if the gap shrinks or reverses, the central accuracy claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central accuracy claims rest on metrics computed against reference reconstructions, not known truth. In the synthetic experiments, the training target and evaluation reference is a dense 2132-view mono-energetic FDK reconstruction (Table 1 and Fig. 1 caption); this reference can still carry cone-beam artifacts and residual streaks. In the experimental demonstration, the reference is a 580-view MBIR reconstruction (Section 5), which has its own regularization bias and resolution limits. The recall/precision analyses use Otsu segmentations of these references as ground truth (Figs. 6, 9, 12). Thus, the reported superiority of 2.5D PnP could reflect better reproduction of reference-specific artifacts, for example matching the dense-FDK target's cupping profile or MBIR's pore appearance, rather than higher fidelity to the actual object. This is not a philosophical objection: the synthetic simulation starts from known CAD phantoms, so true ground-truth volumes exist and should be the primary benchmark. Without that check, the abstract's 'fast and accurate' claim is not yet separated from 'fast and close to this particular reference.' Similarly, the beam-hardening suppression evidence is a line profile against a mono-energetic FDK reference; a ground-truth phantom would show whether cupping is truly removed rather than merely matched to the reference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a plug-and-play (PnP) reconstruction method for sparse-view cone-beam XCT in which the proximal/regularization step is a 2.5D UNet that takes five adjacent slices and denoises/artifact-corrects the center slice. The method is compared with a 2D-prior PnP counterpart on synthetic aluminum AM phantoms (in-distribution and out-of-distribution noise/view counts) and on an experimental Al-Ce part scanned with 145 views, using a 580-view MBIR reconstruction as reference. The authors report improved NRMSE/SSIM and defect recall/precision for 2.5D PnP in most conditions, plus suppression of beam-hardening cupping, at roughly equal runtime and memory to 2D PnP, and much lower cost than MBIR.","tokens_in":14790,"tokens_out":6702,"duration_ms":62865,"significance":"If the reported gains hold against true object ground truth, the paper makes a modest but useful contribution: it validates a low-cost architectural change to a practical PnP framework, demonstrates that a synthetic-trained 2.5D prior transfers to experimental data, and shows that the extra computational burden over 2D PnP is negligible. The fixed Otsu-based defect-detection evaluation and the explicit runtime/memory comparison are strengths. The central limitation is that all quantitative comparisons are made against reconstructed references rather than known object truth; because the synthetic phantoms are CAD-generated, this is addressable and should be addressed before publication.","major_comments":[{"comment":"The synthetic evaluation uses a dense 2132-view mono-energetic FDK reconstruction as both the training target and the evaluation reference (Table 1). Since the CAD phantoms are known, true ground-truth volumes exist and should be used for NRMSE/SSIM and for recall/precision with known pore labels. As written, the reported gains of 2.5D PnP over 2D PnP may partly reflect better matching of the particular FDK reference's residual artifacts (cone-beam or cupping profile) rather than higher fidelity to the true object. The Conclusion's statement that pores \"more closely match the ground truth\" is therefore stronger than the evidence supports.","section":"§3.3, Table 1; §4.1; Figs. 6, 9, 12"},{"comment":"All metrics in Table 2 are computed from a single reconstructed volume per condition; no noise realizations, bootstrap intervals, or significance tests are reported. The text in §4.1 says 2.5D PnP \"consistently achieves significantly higher recall and precision,\" but the OOD-view rows contradict this (e.g., 73 views, noise 0.5: recall 0.879 for 2D PnP vs 0.306 for 2.5D PnP; precision 0.658 vs 0.306). The claims need to be qualified to the conditions where they hold, and error bars or repeated trials are needed before calling the differences significant.","section":"§4, Table 2"},{"comment":"For the experimental demonstration, the reference is an MBIR reconstruction of a 580-view scan; all experimental image-quality and defect-detection metrics are relative to that reconstruction, which carries its own regularization bias. Moreover, Table 3 shows 2D PnP slightly outperforms 2.5D PnP on NRMSE and SSIM (0.148 vs 0.152; 0.991 vs 0.990), so the experimental advantage of 2.5D PnP rests entirely on recall/precision computed against an Otsu segmentation of the MBIR reference. The manuscript should acknowledge this reference-dependence more explicitly and ideally validate the experimental pipeline on a phantom with known defect locations.","section":"§5, Table 3"},{"comment":"The abstract claims that the 2.5D prior \"eliminates the need for artifact correction pre-processing.\" The only evidence is the line profile in Fig. 4 compared with a mono-energetic dense FDK reference. No quantitative beam-hardening metric (e.g., cupping index) is provided, and no comparison against a standard BH-correction preprocessing pipeline is shown. The claim should either be supported by such a comparison or narrowed to \"suppresses beam-hardening-like cupping relative to this reference.\"","section":"Abstract; §4.1, Fig. 4"}],"minor_comments":[{"comment":"Eq. (6) defines the network as predicting a residual (input minus target), but Algorithm 1 writes z_k ← D_θ(x_{k-1}) as if D_θ outputs the denoised center slice directly. Please clarify whether D_θ outputs the residual or the cleaned slice, and make Algorithm 1 consistent with Eq. (6).","section":"§3.1 and Algorithm 1"},{"comment":"The caption calls the reference a \"dense-view FDK reconstruction that does not contain beam hardening artifacts\"; this is accurate for the synthetic setup, but the term \"reference\" should not be equated with ground truth elsewhere in the text (e.g., §4.1's wording about matching the reference).","section":"Figure 1 caption"},{"comment":"Please report the number of defects/flaws in the reference segmentation used for recall and precision, and state how many voxels or pores contribute to the reported values; without this, the precision values for large pores are difficult to interpret.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for the journal's audience. The main risk is the reference-dependent evaluation, but I believe it is addressable in revision by adding ground-truth-based synthetic benchmarks and tempering the broader claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: it is a clean, modest A/B result. Swapping the 2D artifact-reduction CNN in the authors' own PnP framework for a 2.5D UNet improves reconstruction quality and defect detection on synthetic in-distribution data and one experimental Al-Ce part, at essentially no extra compute. The comparison is fair: both methods share the same beta grid search, CG steps, iteration count, and training recipe. The domain-transfer result—a prior trained entirely on simulated scans working on a real 145-view scan—is genuinely useful for nondestructive evaluation.\n\nThe paper is also honest. Section 7 admits the method struggles under OOD view sparsity, and Section 6 reports the precision trade-off at 73 views. That alone puts it ahead of a lot of reconstruction papers. Runtime and memory tables are there; 2.5D costs about two seconds and one extra gigabyte over 2D.\n\nSoft spots, in proportion. The main one is evaluation against reconstructed references rather than known truth. The synthetic setup starts from CAD phantoms, so true ground-truth volumes exist, but the training target and evaluation reference is a dense 2132-view mono-energetic FDK reconstruction. That reference can still carry cone-beam artifacts. The cupping figure compares a line profile to that same FDK reference, which shows the network matches the reference, not necessarily that beam-hardening is truly gone from the object. Since the phantoms are known, a ground-truth comparison would settle this, and it should be trivial to run.\n\nSecond, there are no error bars or significance tests; every condition is one volume. Some differences are small, and in a couple of conditions 2D PnP wins on NRMSE or precision. The abstract's unqualified \"improves the quality of reconstructions\" overstates what is actually a conditional result that holds across most but not all tested regimes. The body is honest about this; the abstract is not.\n\nThird, the experimental reference is itself an MBIR reconstruction, and the SNR/CNR numbers in Table 3 are absurdly high (CNR over 1300) because the PnP outputs are heavily smoothed. Those numbers do not mean the reconstruction is more faithful; they mean it is smoother. And the beam-hardening claim is never tested against a standard BH-correction baseline, so \"eliminates the need for artifact correction pre-processing\" is asserted, not demonstrated.\n\nWho this is for: people doing PnP or learned priors for industrial XCT/NDE. They get a fair, reproducible-in-spirit comparison and a clear picture of where the 2.5D prior helps and where it does not. It does not reshape the field.\n\nThis deserves a serious referee. My own verdict would be major revision: add ground-truth evaluation on the synthetic phantoms, error bars or at least variance across slices, and a BH-correction baseline. With those, it would be a solid acceptance.","headline":"A fair, honest 2.5D-versus-2D A/B in PnP XCT reconstruction; the only real weakness is that accuracy is measured against reconstructed references when known CAD ground truth exists.","tokens_in":15403,"tokens_out":3966,"would_cite":false,"duration_ms":35653,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Upgrading the artifact-reduction prior from 2D to 2.5D improves sparse-view cone-beam XCT reconstruction, preserves fine pore structure, and suppresses beam hardening directly during reconstruction.","keywords":["X-ray computed tomography","plug-and-play reconstruction","artifact reduction prior","2.5D CNN","beam hardening","sparse-view CT","defect detection","super-resolution"],"falsifier":"Recompute the experimental comparison using a reference obtained from a materially denser acquisition (for example, thousands of views at high signal) or a known ground-truth phantom, and test whether the 2.5D PnP advantage in recall and precision over 2D PnP survives; if it shrinks or reverses, the reported gains were partly matching reference artifacts instead of true pore structure.","tokens_in":14331,"feed_emoji":"🩻","tokens_out":6634,"duration_ms":60370,"temperature":0.7,"pith_summary":"Industrial cone-beam X-ray computed tomography needs many projections to produce high-quality 3D reconstructions, making scans slow and expensive. This paper argues that placing a 2.5D artifact-reduction CNN inside a plug-and-play reconstruction framework recovers near-reference quality from sparse, noisy, beam-hardened scans: the network sees five adjacent slices at once, using inter-slice context to clean the center slice while staying as cheap as a 2D model. The authors show that this 2.5D prior outperforms the 2D prior it extends, preserving fine pore shape and boosting defect-detection recall and precision, and that it removes beam-hardening artifacts inside the reconstruction so no separate artifact-correction preprocessing is needed. They also demonstrate that a prior trained only on synthetic scans reconstructs experimental aluminum-cerium parts well, which matters because deep-learned CT models typically fail on out-of-distribution data.","feed_headline":"2.5D prior improves sparse-view CT, removes beam-hardening artifacts","feed_subtitle":"Synthetic-only training transfers to real parts and matches MBIR quality at a fraction of the time and memory.","key_machinery":"The load-bearing mechanism is the alternating-minimization PnP loop with a quadratic-penalty variable split: a data-fidelity step (conjugate gradient on the cone-beam forward model) alternates with a regularization step in which the proximal operator is replaced by a 2.5D artifact-reduction UNet. The 2.5D network takes five adjacent slices as input channels and is trained with an L1 residual loss to output the cleaned center slice, giving it volumetric context at 2D memory cost; an adaptive grid search on center-slice quality sets the penalty parameter β each iteration. This combination lets the reconstruction suppress noise, cupping and streak artifacts from beam hardening, and view-sparsity artifacts while preserving pore boundaries.","core_discovery":"The paper claims that upgrading the artifact-reduction prior in a plug-and-play (PnP) reconstruction scheme from 2D to 2.5D yields a clear improvement in sparse-view cone-beam XCT. In the PnP loop, the regularization sub-problem is solved by a CNN trained to map low-quality FDK inputs (few views, Poisson-like noise, beam hardening) to dense-view artifact-free targets; the 2.5D version feeds a stack of five neighboring slices and predicts the residual for the center slice. The authors report lower NRMSE and higher SSIM on synthetic in-distribution and most out-of-distribution tests, higher recall and precision for 75–125 µm flaws, and visible recovery of pores that 2D PnP distorts or misses. On experimental 145-view scans, the 2.5D model trained entirely on synthetic data matches or beats the 2D prior on task-specific defect detection and achieves the highest SNR and CNR, at essentially the same runtime (about 48 minutes per 1356×1356×1264 volume on four GPUs) as 2D PnP and far below MBIR.","pith_inferences":["Because the 2.5D prior leverages inter-slice consistency, the same architecture could plausibly suppress other volumetric artifacts, such as metal streaks or limited-angle artifacts, if training targets expose them.","The recall and precision gains on 75–125 µm flaws suggest that PnP with 2.5D priors could lower missed-defect rates in nondestructive evaluation, but the false-positive increase under high noise and extreme sparsity implies production use would need a confidence or size filter.","A natural testable extension is to vary the input stack depth (3, 5, or 7 slices) and measure the trade-off between context and compute; the paper fixes it at five without an ablation.","If the synthetic-to-experimental transfer holds broadly, a shared artifact-reduction prior could be trained once on CAD-derived phantoms and deployed across many industrial XCT systems, avoiding per-scanner retraining."],"forward_implications":["Sparse-view industrial scans could be cut to 73–145 views and still yield defect-detection-quality volumes, with beam hardening handled inside reconstruction rather than as a preprocessing step.","Large-volume reconstruction time and memory drop by an order of magnitude versus MBIR (about 48 minutes and 35 GB on four GPUs, compared with about 6 hours and 300 GB), making high-quality XCT feasible for routine inspection.","A synthetic-only training recipe transfers to experimental parts, so new materials and geometries may not require experimental ground-truth training data.","Out-of-distribution noise levels (up to four times the training noise) degrade gracefully, while extreme view sparsity (73 views) still exposes a sensitivity and false-positive trade-off."],"supporting_citations":[{"why":"Supplies the original 2D artifact-reduction PnP framework with adaptive regularization parameter selection that this paper extends and uses as the baseline.","marker":"[23]"},{"why":"Provides the FDK algorithm used to form the low-quality input reconstructions from sparse views.","marker":"[7]"},{"why":"Defines the U-Net architecture on which both the 2D and 2.5D artifact-reduction networks are based.","marker":"[38]"},{"why":"Introduces the plug-and-play priors concept that justifies replacing the regularizer with a learned denoiser.","marker":"[13]"},{"why":"Earlier 2.5D super-resolution work that motivates using multi-slice context in XCT.","marker":"[24]"},{"why":"Defines the short-scan acquisition protocol used to generate the sparse-view scan geometry.","marker":"[40]"},{"why":"Spectrum simulation tool used to generate the polychromatic beam-hardened training data.","marker":"[43]"},{"why":"Supplies the rationale for reporting task-specific defect-detection metrics rather than NRMSE and SSIM alone.","marker":"[45]"},{"why":"Otsu thresholding is the fixed, parameter-free segmentation baseline used to compute recall and precision of detected defects.","marker":"[46]"}],"fun_headline_variants":["2.5D prior sharpens sparse CT, drops beam hardening","Synthetic-trained 2.5D prior matches MBIR, runs much faster","2.5D CNN prior beats 2D in sparse XCT, transfers sim-to-real","Better sparse CT: 2.5D prior vs 2D, no artifact preprocessing","2.5D artifact prior: sharper sparse CT, no beam hardening"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume that the reference reconstructions used for training and evaluation — a dense-view fast analytic reconstruction with a single-energy beam on synthetic data, and a slower iterative reconstruction on experimental data — faithfully represent the true object rather than carrying their own residual artifacts.","fun_headline_variants_meta":{"raw":{"variants":["2.5D prior sharpens sparse CT, drops beam hardening","Synthetic-trained 2.5D prior matches MBIR, runs much faster","2.5D CNN prior beats 2D in sparse XCT, transfers sim-to-real","Better sparse CT: 2.5D prior vs 2D, no artifact preprocessing","2.5D artifact prior: sharper sparse CT, no beam hardening"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000852,"raw_usage":{"total_tokens":3769,"prompt_tokens":1077,"completion_tokens":2692,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":693,"completion_tokens_details":{"reasoning_tokens":2587}},"tokens_in":693,"tokens_out":2692,"duration_ms":19438,"temperature":1.0,"reasoning_tokens":2587,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:47:28.382823+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the experimental comparison using a reference obtained from a materially denser acquisition (for example, thousands of views at high signal) or a known ground-truth phantom, and test whether the 2.5D PnP advantage in recall and precision over 2D PnP survives; if it shrinks or reverses, the reported gains were partly matching reference artifacts instead of true pore structure.","supporting_citations":[{"cited_title":"A Fast, Scalable, and Robust Deep Learning-based Iterative Reconstruction Framework for Accelerated Industrial Cone-beam X-ray Computed Tomography","cited_arxiv_id":"2501.13961","evidence_quote":"Supplies the original 2D artifact-reduction PnP framework with adaptive regularization parameter selection that this paper extends and uses as the baseline."},{"cited_title":"JOSA A1(6), 612–619 (1984)","cited_arxiv_id":null,"evidence_quote":"Provides the FDK algorithm used to form the low-quality input reconstructions from sparse views."},{"cited_title":"2.5D Super-Resolution Approaches for X-ray Computed Tomography-based Inspection of Additively Manufactured Parts","cited_arxiv_id":"2412.04525","evidence_quote":"Earlier 2.5D super-resolution work that motivates using multi-slice context in XCT."},{"cited_title":"Medical physics9(2), 254–257 (1982)","cited_arxiv_id":null,"evidence_quote":"Defines the short-scan acquisition protocol used to generate the sparse-view scan geometry."},{"cited_title":"0—a software toolkit for modeling x-ray tube spectra","cited_arxiv_id":null,"evidence_quote":"Spectrum simulation tool used to generate the polychromatic beam-hardened training data."},{"cited_title":"general quality assessment","cited_arxiv_id":null,"evidence_quote":"Supplies the rationale for reporting task-specific defect-detection metrics rather than NRMSE and SSIM alone."},{"cited_title":"Automatica11(285-296), 23–27 (1975)","cited_arxiv_id":null,"evidence_quote":"Otsu thresholding is the fixed, parameter-free segmentation baseline used to compute recall and precision of detected defects."}],"review_version":2}