{"id":"35124dc8-1d42-41cd-b554-d50f370a920f","arxiv_id":"2412.08350","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"A reproducible benchmark of learned CT reconstruction methods on real experimental data, covering post-processing, unrolled, learned regularizer, and plug-and-play methods across five tasks.","lead":"This paper benchmarks twelve learned CT reconstruction algorithms, in four method categories, on real-world experimental scan data from the 2DeteCT dataset across five reconstruction tasks. It also releases an open-source pipeline (LION) so future methods can be added and compared under identical conditions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All scores are computed against AGD references that are also the training targets; category-level rankings may reflect mimicry of AGD rather than reconstruction quality, and the paper does not quantify how rankings would shift under an alternative reference.","rationale":"The reader's weakest_assumption (AGD target as pseudo-ground-truth) is indeed the most load-bearing issue. It directly undermines the comparative interpretation of the benchmark, which is the paper's central value proposition. The paper is transparent about the choice and explicitly says results would change with a different regularized target, but it does not test sensitivity; therefore the CONDITIONAL verdict is appropriate. I considered the missing scripts and trained models as an alternative concern, but that is a release-timeline artifact rather than a scientific flaw in the argument; the LION toolbox is already public and the data loader and tasks are described. The target-dependence concern, by contrast, affects every reported score and the relative ranking of method categories, which is exactly what a benchmark is supposed to provide. Retraining a representative subset with an alternative reference is the decisive experiment, since evaluating existing AGD-trained models on a new target would be confounded by training-target mismatch. Recommendation: keep the verdict CONDITIONAL, with the condition being a sensitivity analysis of the reference choice (and, secondarily, artifact release).","tokens_in":23005,"tokens_out":6886,"duration_ms":73285,"concrete_test":"Retrain one representative method from each category (e.g., FBP+U-Net, LPD, ACR, DRUNet-PnP) on the Full Data, Sparse 60, and Beam-Hardening tasks with a Chambolle-Pock + TV reference (or a consensus of independent reconstructions from AGD, ChP+TV, and another iterative method) as the training and evaluation target, using the same hyperparameters and data split. Compare the category-level rankings with Tables 5-7. If the ordering of method categories changes on any task, the reported benchmark conclusions are reference-dependent; if rankings persist, the concern is bounded. A cheaper complementary check: compute PSNR/SSIM between the AGD and ChP+TV reference images on the test set; if their agreement is low (e.g., PSNR below 25 dB), the target choice is a dominant factor in the scores.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's quantitative comparisons (Tables 5-7) measure every method's PSNR/SSIM against the mode-2 AGD reconstructions, and the same AGD reconstructions are the supervision targets for all trained methods. For the Full Data task this is explicit: methods 'actually learn to mimic AGD' (Section 'Relevance and Difficulty of CT Image Reconstruction Tasks', Full Data Reconstruction paragraph), and the target is described as a choice, not ground truth. The concern is not that the target is imperfect—every real-data benchmark faces this—but that the comparison protocol conflates reconstruction quality with fidelity to one particular algorithm's output. Methods whose optimization family matches AGD (unrolled methods, learned regularizers, PnP with gradient-step or denoiser priors) are likely favored: on Full Data, LPD, DnCNN-PnP, DRUNet-PnP, and ACR reach SSIM 0.84-0.86, while FBP+U-Net (0.65) and even classical FBP (0.75) score lower—plausibly because a post-processing network operating from FBP inputs cannot reproduce AGD. The paper acknowledges that using Chambolle-Pock with TV as target 'would change the numerical results' but does not bound the change, and the Limitations section only cautions that trends may not transfer to medical CT. Since the central claim is a fair benchmark across method categories, an unquantified dependence of rankings on the reference choice is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a benchmarking study for learned computed tomography (CT) reconstruction methods using the 2DeteCT dataset, which contains real experimental X-ray measurements. Twelve learned methods are grouped into post-processing networks, learned/unrolled iterative methods, learned regularizer methods, and plug-and-play methods, alongside classical baselines (FBP, AGD, ChP). The authors define five reconstruction tasks (full data, limited-angle, sparse-angle, low-dose, and beam-hardening corrected), report PSNR and SSIM on a fixed test set, and release the pipeline in the open-source LION toolbox with trained models. The central claims are that this is the first benchmark combining real-world experimental CT data with a broad range of standardized tasks and method categories, and that the framework enables reproducible comparison and extension by the community.","tokens_in":23222,"tokens_out":3181,"duration_ms":36176,"significance":"If the benchmark is accepted as a community reference, it would fill a genuine gap: most existing CT reconstruction benchmarks use simulated data or a single task, whereas 2DeteCT provides real projection data with matching high-quality references. The open implementation in LION, the fixed train/validation/test split, the released models, and the explicit task definitions are practical strengths that lower the barrier for future method comparison. The paper is also candid about several limitations, including limited hyperparameter tuning and the fact that residual beam-hardening artifacts remain in the mode-2 references. However, the central quantitative conclusions are conditional on the choice of the AGD reconstructions as the reference for both training and evaluation, and on single training runs per method; the significance of the benchmark would be substantially strengthened by robustness checks against alternative references and repeated runs.","major_comments":[{"comment":"The rankings are based on a single training run per method with fixed hyperparameters, and the paper explicitly states that training was done \"without extensive hyperparameter tuning.\" For the post-processing and unrolled methods, training runs took tens to over a hundred hours, so repeated runs may be costly, but without them the reported differences are not statistically grounded. For instance, in Table 5 the Full Data SSIM values of LPD (0.8447), FBP+MSDNet (0.8481), ACR (0.8518), DnCNN-PnP (0.8585), and DRUNet-PnP (0.8573) differ by amounts that are comparable to the reported standard deviations over test slices; these standard deviations encode slice-to-slice variability, not run-to-run variability. I ask the authors to either supply multiple training runs per method (at least for the leading contenders) or to add a sensitivity analysis showing that the conclusions are stable to seed and hyperparameter choices. Without this, the cross-method and cross-category comparisons are vulnerable to noise in training.","section":"Training Details; Tables 5-7"}],"minor_comments":[{"comment":"On the Full Data task, FBP+U-Net has SSIM 0.6499, which is lower than plain FBP (0.7463) and well below the other post-processing methods; this is an interesting result that the text does not discuss, and it would be helpful to comment on why the U-Net post-processing degrades relative to its FBP input when the target is an AGD reconstruction.","section":"Table 5"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Max,\n\nHere is my take on Kiss et al. arXiv:2412.08350. It is a genuinely useful benchmark paper, not a flashy methods paper, and it fills a real niche: a common real-data evaluation platform for learned CT reconstruction. The authors define five tasks on 2DeteCT (full data, limited angle, sparse angle, low dose, beam-hardening), categorize twelve published methods into four families, and implement everything in the LION toolbox with classic baselines. That combination—real experimental raw data, standardized tasks, several method categories in one framework—is not something I have seen elsewhere, and it is a sensible contribution.\n\nThe paper is also honest. It says up front that the AGD reference is a choice, not ground truth; it notes that using Chambolle-Pock with TV as target would change the numbers; it flags limited hyperparameter tuning and warns that trends may not transfer to medical CT. The qualitative analysis adds value because the metrics alone mislead on limited-angle tasks. The open-source toolbox is real and useful.\n\nNow the soft spots. The stress-test concern is real, though I would not call it fatal. For Full Data, methods are literally trained to mimic AGD, so their rankings partly measure mimicry rather than reconstruction quality. The paper admits this, but it does not bound the effect. For the other tasks the same AGD reference is the target, and while the comparison is fair in the sense that all methods face the same target, absolute PSNR/SSIM values and some category-level rankings could shift under a different reference. A supplementary table with a TV-regularized target, or an experiment showing rank stability, would address this cleanly and should be requested.\n\nThe bigger practical gap is reproducibility: the benchmark tables rest on promised scripts and trained models that are not yet released with a fixed commit. For a benchmark paper, that is a concrete missing artifact, not a minor point. The single-run training and limited hyperparameter tuning also mean the numbers should be read as baselines, not as a definitive ranking; the authors say as much.\n\nCitation-wise, the self-references to 2DeteCT and LION are justified because these are the authors' own dataset and toolbox. No issue there.\n\nWho should read this? Anyone working on learned CT reconstruction who needs a real-data testbed and baseline numbers. It deserves a serious referee and, with the artifacts released and some sensitivity analysis, it will be a useful reference. I would send it out.\n\nBest,\n[Name]","headline":"A solid, useful real-data benchmark for learned CT reconstruction, with honest limitations; the AGD-target dependence is real but acknowledged, and the main fix is releasing exact scripts/models and adding sensitivity analysis.","tokens_in":23840,"tokens_out":2613,"would_cite":true,"duration_ms":29177,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that real experimental CT data, not simulations, can support a reproducible benchmark comparing learned reconstruction methods across five standardized tasks.","keywords":["computed tomography","image reconstruction","learned reconstruction","benchmarking","2DeteCT dataset","limited-angle CT","sparse-angle CT","beam-hardening correction"],"falsifier":"Recompute the full benchmark tables using a different reference, such as a total-variation-regularized primal-dual reconstruction of the same mode-2 data, and check whether any method's ranking changes on the low-dose, limited-angle, or beam-hardening tasks; a ranking change would show that the conclusions depend on the choice of reference image rather than on the algorithms alone.","tokens_in":22755,"feed_emoji":"🩻","tokens_out":8846,"duration_ms":86451,"temperature":0.7,"pith_summary":"Computed tomography research lacks a large-scale, open, real-measurement dataset on which learned reconstruction algorithms can be compared fairly; most benchmarks use simulated data or private scans. This paper attempts to close that gap by building a benchmarking framework entirely on the 2DeteCT experimental dataset, with five standardized tasks — full-data, limited-angle, sparse-angle, low-dose, and beam-hardening-corrected reconstruction — and twelve supervised methods drawn from four method families: post-processing networks, learned or unrolled iterative methods, learned regularizers, and plug-and-play methods. The authors' central claim is that this combination of real measured data, standardized tasks, and an open pipeline makes results reproducible and lets any new method be added and compared on equal footing. If the framework works as intended, it gives the field a common yardstick and reduces the sim-to-real gap that has made published CT results hard to compare.","feed_headline":"Real-data benchmark puts 12 learned CT methods on equal footing","feed_subtitle":"Five standardized tasks on the same measured scans, with open code, let any new algorithm be tested against the same baselines.","key_machinery":"The load-bearing object is the benchmark pipeline itself, organized as a sinogram-to-image experiment: each method receives a sinogram as input and must output a 1024 by 1024 image, and every task is generated by choosing which sinogram to feed it. The source of realism is the 2DeteCT dataset's three physical acquisition modes — high-dose filtered, low-dose filtered, and unfiltered — which supply the clean targets, noisy inputs, and beam-hardening inputs respectively. The standard reference target for all tasks is the mode-2 reconstruction computed by accelerated gradient descent on a 2048 by 2048 grid and cropped to 1024 by 1024; that choice makes every method's score comparable because everyone is measured against the same image. Four method categories with three methods each, plus classical filtered-backprojection, accelerated-gradient-descent, and total-variation-regularized primal-dual baselines, form the comparison grid, and structural similarity (SSIM) and peak signal-to-noise ratio (PSNR) averaged over a held-out test set quantify success.","core_discovery":"The paper's central claim is that a benchmark for learned CT reconstruction should be built on real experimental measurements rather than simulations, and that such a benchmark can cover the full range of common reconstruction tasks while keeping every method on identical data and preprocessing. Concretely, it defines five tasks from the three acquisition modes of the 2DeteCT dataset: full-data reconstruction from clean high-dose mode-2 sinograms, limited-angle and sparse-angle tasks by cropping or subsampling those sinograms, low-dose reconstruction from mode-1 sinograms, and beam-hardening correction from unfiltered mode-3 sinograms. All methods are trained in a sinogram-to-image setup and scored against the same target, the mode-2 iterative accelerated-gradient-descent reconstructions. On those tasks the paper reports that post-processing networks are consistently strong quantitatively despite lacking data-consistency guarantees, that the learned primal-dual unrolled method is the steadiest high performer, and that beam-hardening correction is the hardest task, with plug-and-play and adversarial-regularizer methods collapsing because the linear forward model cannot represent the nonlinear beam-hardening effect.","pith_inferences":["Beyond the paper: the benchmark's target choice is a sensitivity point, since all scores are distances to one algorithm's reconstruction; re-running the tables with a total-variation-regularized reference target would show whether relative rankings survive the change.","Beyond the paper: the released models and code make out-of-distribution evaluation immediate, and applying the trained models to a medical or different-geometry CT dataset would test whether the observed rankings transfer beyond the 2DeteCT scanner.","Beyond the paper: the framework omits transformer-based and diffusion-based reconstructions, so adding them is a straightforward next step, with diffusion methods mapping onto the plug-and-play slot and transformers onto the post-processing slot."],"forward_implications":["New methods can be inserted into the open pipeline and compared directly against the same twelve learned baselines on the same real sinograms, making published CT results far easier to reproduce.","Post-processing networks, despite having no data-consistency mechanism, can match or exceed more complex model-based methods on most tasks while training in less time.","Beam-hardening correction is the clearest bottleneck: methods that rely on the linear forward model or on local denoisers fail, so progress on this task requires nonlinear modeling or learned artifact correction.","Learned or unrolled iterative methods such as learned primal-dual are the most balanced performers across all five tasks, but their training cost is high relative to post-processing.","Quantitative scores alone can mislead: on 60-degree limited-angle data, plug-and-play and adversarial-regularizer methods score similarly to better-looking reconstructions, so visual inspection remains necessary."],"supporting_citations":[{"why":"Supplies the 2DeteCT real experimental dataset: all five reconstruction tasks and their reference targets are built from its three acquisition modes.","marker":"[20]"},{"why":"Defines the four-category taxonomy of data-driven reconstruction methods that organizes the benchmark's method selection.","marker":"[12]"},{"why":"Provides the Learned Gradient algorithm, one of the three unrolled iterative baselines evaluated on every task.","marker":"[5]"},{"why":"Provides the Learned Primal-Dual algorithm, the unrolled baseline that reports the best quantitative results on most tasks.","marker":"[84]"},{"why":"Supplies DnCNN, used both as a post-processing network and as the denoiser plugged into the DnCNN plug-and-play baseline.","marker":"[83]"},{"why":"Provides the Adversarial Regularizer, a learned regularizer baseline whose failure on beam-hardening correction is used to explain limits of variational methods under operator mismatch.","marker":"[85]"},{"why":"Provides the Total Deep Variation regularizer, a learned regularizer baseline trained on the variational objective.","marker":"[86]"},{"why":"Supplies the DRUNet denoiser, the higher-capacity plug-in for the DRUNet plug-and-play baseline.","marker":"[89]"},{"why":"Provides the gradient-step plug-and-play scheme used for the GS plug-and-play baseline with convergence guarantees.","marker":"[90]"}],"fun_headline_variants":["Real-data CT benchmark ranks learned reconstruction methods across five tasks","Five CT tasks, one real dataset: learned methods benchmarked openly","Benchmarking learned CT reconstruction on real measurements, not simulations","Open-source benchmark compares 12 learned CT methods on real scans","Learned CT reconstruction tested on real 2DeteCT data with reproducible code"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's weakest load-bearing premise is that the accelerated-gradient-descent reconstruction of the clean mode-2 scan is an acceptable stand-in for true ground truth for every task, because all PSNR and SSIM scores are measured relative to that single algorithm's output.","fun_headline_variants_meta":{"raw":{"variants":["Real-data CT benchmark ranks learned reconstruction methods across five tasks","Five CT tasks, one real dataset: learned methods benchmarked openly","Benchmarking learned CT reconstruction on real measurements, not simulations","Open-source benchmark compares 12 learned CT methods on real scans","Learned CT reconstruction tested on real 2DeteCT data with reproducible code"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000605,"raw_usage":{"total_tokens":2849,"prompt_tokens":1001,"completion_tokens":1848,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":1758}},"tokens_in":617,"tokens_out":1848,"duration_ms":14702,"temperature":1.0,"reasoning_tokens":1758,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:53:30.944265+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the full benchmark tables using a different reference, such as a total-variation-regularized primal-dual reconstruction of the same mode-2 data, and check whether any method's ranking changes on the low-dose, limited-angle, or beam-hardening tasks; a ranking change would show that the conclusions depend on the choice of reference image rather than on the algorithms alone.","supporting_citations":[{"cited_title":"& Öktem, O","cited_arxiv_id":null,"evidence_quote":"Provides the Learned Primal-Dual algorithm, the unrolled baseline that reports the best quantitative results on most tasks."},{"cited_title":"& Zhang, L","cited_arxiv_id":null,"evidence_quote":"Supplies DnCNN, used both as a post-processing network and as the denoiser plugged into the DnCNN plug-and-play baseline."},{"cited_title":"& Schönlieb, C.-B","cited_arxiv_id":null,"evidence_quote":"Provides the Adversarial Regularizer, a learned regularizer baseline whose failure on beam-hardening correction is used to explain limits of variational methods under operator mismatch."},{"cited_title":"& Pock, T","cited_arxiv_id":null,"evidence_quote":"Provides the Total Deep Variation regularizer, a learned regularizer baseline trained on the variational objective."},{"cited_title":"& Papadakis, N","cited_arxiv_id":null,"evidence_quote":"Provides the gradient-step plug-and-play scheme used for the GS plug-and-play baseline with convergence guarantees."}],"review_version":1}