{"id":"105c56b0-00aa-4b10-8e9e-91f8087871e2","arxiv_id":"2506.17425","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Trans2-CBCT combines a TransUNet feature extractor with a neighbor-aware Point Transformer and reports state-of-the-art PSNR/SSIM for 6-10 view sparse-view CBCT reconstruction on LUNA16 and ToothFairy.","lead":"This paper replaces the usual U-Net or ResNet image encoders in sparse-view cone-beam CT reconstruction with TransUNet and adds a neighbor-aware Point Transformer for 3D consistency. The combined model reports the best PSNR and SSIM scores on LUNA16 and ToothFairy datasets when only six to ten projection views are available.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table I's imported baseline numbers make the claimed margins unverifiable; a non-monotonic DIF-Gaussian entry suggests protocol mismatch.","rationale":"The reader's weakest assumption—imported baseline numbers under a possibly different protocol—is the same load-bearing concern I identify. I add a concrete internal inconsistency in Table I: DIF-Gaussian on ToothFairy is non-monotonic (10 views worse than 8 views), which makes protocol mismatch a live risk rather than a hypothetical one. The paper's internal ablations are coherent and the architecture is clearly described, but the central empirical claim cannot be verified without re-running the baselines or releasing code. The reader's CONDITIONAL verdict already captures this: the claim is plausible but not independently checkable. My analysis does not move that verdict, so I recommend UNCHANGED.","tokens_in":14330,"tokens_out":2787,"duration_ms":30518,"concrete_test":"Re-run DIF-Gaussian [26] and C2RV [7] on the exact LUNA16 and ToothFairy splits, DRR generation (256x256, angles in [0°, 180°), spacing), point-sampling strategy, and metric code used for Trans2-CBCT, at 6/8/10 views. If the reproduced DIF-Gaussian ToothFairy 10-view score or the C2RV ToothFairy score differs from Table I by more than 0.5 dB, the reported margins do not support the architectural-superiority claim. Also verify that Trans2-CBCT inference uses the same neighbor-construction protocol as training (e.g., KNN over 10,000 random points vs. dense volume voxels); a train/test mismatch in KNN neighborhoods would further confound the reported gains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—\"consistently outperform all prior methods in terms of both PSNR and SSIM\" (Section V)—rests entirely on Table I. The most load-bearing premise is that all methods in Table I were evaluated under one shared protocol. Section IV-A says: \"For DIF-Gaussian* ... we reproduced the experiments given the original authors' implementation. The remaining results are taken from DIF-Gaussian [26].\" Thus C2RV, the strongest LUNA16 baseline, and every ToothFairy baseline row are imported numbers, not re-run by the authors with their own data splits, DRR generation, view sampling, and metric code. A concrete red flag appears in Table I: DIF-Gaussian on ToothFairy at 10 views is reported as PSNR 29.07 / SSIM 85.17, which is sharply below the same method's 8-view numbers (29.83 / 88.67) and below its own 6-view PSNR by 0.26 dB. A more-view reconstruction should not degrade so dramatically; this is exactly the pattern expected if imported rows come from different conditions. Since the claimed margin at 10 views on ToothFairy is +4.56 dB over DIF-Gaussian, and since C2RV is missing entirely from ToothFairy, the headline superiority may be partly a protocol artifact. No code or weights are released, so these comparisons cannot be independently checked from the manuscript.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two transformer-based models for sparse-view cone-beam CT (CBCT) reconstruction from 6-10 projection views. Trans-CBCT replaces the conventional UNet/ResNet image encoder used in prior multi-view feature querying methods with TransUNet, concatenates multi-scale decoder features, queries them per 3D point via bilinear interpolation, fuses views by max pooling, and regresses attenuation with a small prediction head. Trans2-CBCT adds a Point Transformer with learnable 3D positional encoding and a KNN-restricted multi-head attention whose logits include a Gaussian distance bias. Both models are trained with point-wise MSE on LUNA16 and ToothFairy. Results in Table I report consistent PSNR/SSIM superiority over FDK, SART, NAF, NeRP, FBPConvNet, FreeSeed, BBDM, PixelNeRF, DIF-Net, DIF-Gaussian, and C2RV; ablations in Tables III-V analyze the number of sampled points, KNN size, and multi-scale feature aggregation.","tokens_in":14645,"tokens_out":6665,"duration_ms":64581,"significance":"If the comparative numbers are trustworthy, the paper makes a useful empirical contribution: it demonstrates that a CNN-transformer hybrid encoder with point-based geometric refinement yields large gains in extremely sparse-view CBCT, e.g., 1.8 dB PSNR over C2RV at 6 LUNA16 views. The internal ablations and the consistency of the Trans2-CBCT gains over Trans-CBCT support the design choices, and the efficiency analysis is helpful. However, the significance is capped by reproducibility and comparability issues: no code or weights are released, the headline claim relies on imported baseline numbers, and no uncertainty estimates are provided.","major_comments":[{"comment":"The central claim that Trans-CBCT and Trans2-CBCT 'consistently outperform all prior methods' (Section V) rests on baseline entries that were not re-run under the authors' protocol. The text states that only DIF-Gaussian* was reproduced and 'the remaining results are taken from DIF-Gaussian [26]'. If the imported rows were produced with different data splits, DRR generation, view angles, or metric computation, the reported margins (e.g., +1.8 dB over C2RV on LUNA16 at 6 views) could be protocol artifacts. A concrete red flag is the DIF-Gaussian ToothFairy row: PSNR/SSIM go from 29.83/88.67 at 8 views to 29.07/85.17 at 10 views, a sharp degradation with more views that is not seen in any other method and that suggests the imported numbers come from inconsistent experimental conditions. Please rerun all baselines under the same pipeline, or explicitly restrict the superiority claim to the conditions that were actually compared, and reconcile or remove the non-monotonic row. C2RV is also missing entirely from ToothFairy, so the 'both datasets' comparison is incomplete.","section":"Section IV-A and Table I"},{"comment":"The reported PSNR/SSIM values are single numbers without standard deviations, confidence intervals, or significance tests. Since the headline is a comparative superiority claim spanning 18 cells (two datasets x three view counts x two metrics), the lack of any uncertainty estimate makes it impossible to judge whether the margins are stable or driven by one seed. At minimum, report mean and standard deviation over multiple training runs or a paired significance test over test volumes for the key comparison (Trans2-CBCT vs C2RV on LUNA16 and vs DIF-Gaussian on ToothFairy).","section":"Section IV-B and Table I"},{"comment":"The evaluation protocol for full volumes is not described. Training samples N'=10,000 points from the ground-truth volume (Section III-D), but Table I is said to evaluate at a 'fixed resolution of 256^3'. The paper never states whether inference queries all 16.7M voxels, whether a subset of points is sampled for metric computation, or how the KNN neighborhood is defined at full resolution. This matters because PSNR/SSIM depend on the exact set of evaluated points, and the reported inference time of 53.7 s in Table II is not interpretable without this information. Please specify the full inference procedure and metric evaluation grid.","section":"Section III-D and Section IV-B"}],"minor_comments":[{"comment":"The phrase 'in term of' should be 'in terms of' in both the abstract and the conclusion.","section":"Abstract and Section V"},{"comment":"The dimensions of W1, W2, b1, and b2 in the positional encoding are not specified; please state them explicitly so that PE(p_i) can be added to the 464-dimensional feature vector F^p.","section":"Section III-C, Eq. (5)"},{"comment":"Please clarify whether sigma is learned or fixed and, if fixed, its value. The role of the log(w_ij) bias in the attention score also deserves a brief explanation, since w_ij can be very close to zero for distant neighbors.","section":"Section III-C, Eq. (6)"},{"comment":"The text says projection images are 'randomly sample[d] within the range of [0°, 180°)', but it is not stated whether the same set of projection angles is used for all methods and for all test volumes; this should be specified because random per-volume angles can affect fair comparison.","section":"Section IV-A"},{"comment":"The text says 'the forth column of Fig. 4'; this should be 'the fourth column'.","section":"Section IV-C and Fig. 4"},{"comment":"No code or pretrained weights are made available; releasing them, or at least a detailed configuration file, would substantially strengthen the reproducibility of the reported comparisons.","section":"Reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The key risk is the provenance of Table I. The non-monotonic DIF-Gaussian ToothFairy entry (10 views worse than 8 views) is a warning sign that the imported numbers are not from one consistent protocol; I would ask the authors to provide the raw tables from [26] for each cell and to rerun at least C2RV on both datasets. The paper's own contributions are plausible and the internal ablations are coherent, so this is fixable in revision rather than a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea is straightforward and defensible: swap the UNet/ResNet encoder in the DIF-Net lineage for TransUNet, which gives you multi-scale CNN features plus global attention, then refine the point features with a small neighbor-aware Point Transformer that uses KNN attention with a Gaussian distance bias. That is a legitimate new combination for sparse-view CBCT, and the authors show with their own ablations that each piece helps: multi-level feature concatenation, the Point Transformer, number of neighbors, and sampled point count. The reported deltas are internally consistent, and the efficiency table is honest about the 53.7s inference cost of the KNN module. So the engineering is credible.\n\nThe soft spot is the comparison. Section IV-A says DIF-Gaussian was reproduced from the authors' implementation, but \"the remaining results are taken from DIF-Gaussian [26].\" That means all the strong baselines on LUNA16 (C2RV) and every ToothFairy baseline row are imported numbers. The stress-test note is right: DIF-Gaussian on ToothFairy at 10 views is reported at 29.07 dB PSNR / 85.17 SSIM, which is below its own 8-view numbers (29.83 / 88.67) and below its 6-view PSNR. A more-view reconstruction should not degrade like that. That pattern is exactly what you see when imported rows come from a different protocol, and it directly undermines the +4.56 dB margin claimed at 10 views on ToothFairy. No error bars are reported, no statement on whether hyperparameters were chosen on validation or test, and no code or weights are released. So the central claim of consistently outperforming all prior methods is not independently checkable from the manuscript.\n\nThat said, the flaws are fixable. The architecture is not a paradigm shift, but it is a plausible step forward for extremely sparse-view reconstruction. If the authors re-run all baselines under their own DRR generation, view sampling, and metric code, report variance across runs, and release code, the paper would be a useful contribution. I would not desk reject it; it deserves a serious referee who can ask for those comparisons. I would bring it to reading group as a case study in how imported baseline numbers can poison a SOTA claim, but I would not cite it myself until the numbers are independently reproduced.","headline":"A sensible dual-transformer recipe for sparse-view CBCT with clean ablations, but the SOTA claim depends on imported baseline numbers, one of which is non-monotonic and likely a protocol artifact.","tokens_in":15127,"tokens_out":2773,"would_cite":false,"duration_ms":25118,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dual-transformer pipeline sets new accuracy marks for six-view CBCT.","keywords":["cone-beam CT","sparse-view reconstruction","TransUNet","Point Transformer","neighbor-aware attention","implicit neural representation","low-dose imaging","digitally reconstructed radiographs"],"falsifier":"Run C2RV, DIF-Gaussian, and Trans2-CBCT on the same LUNA16 split with identical DRR generation, the same six view angles, the same sampled points, and the same PSNR and SSIM calculation script. If the PSNR difference between Trans2-CBCT and C2RV falls below the reported 1.8 dB, or if an independent implementation of C2RV outperforms the quoted number, the central superiority claim would be falsified.","tokens_in":14124,"feed_emoji":"🩻","tokens_out":5279,"duration_ms":46258,"temperature":0.7,"pith_summary":"This paper argues that the right architecture choice can overcome the severe information loss of reconstructing 3D cone-beam CT volumes from as few as six X-ray projections. It replaces the UNet or ResNet encoders used by prior learning-based methods with TransUNet, a hybrid CNN-transformer, and then refines the resulting 3D features with a neighbor-aware Point Transformer. On the LUNA16 and ToothFairy benchmarks, the authors report that the full model exceeds the strongest previous baseline by 1.8 dB PSNR and 0.028 SSIM at six views. If the numbers hold under identical evaluation, the work would show that combining multi-scale global image features with explicit 3D positional reasoning is a practical recipe for low-dose CBCT.","feed_headline":"Dual-transformer network tops sparse-view CBCT benchmarks","feed_subtitle":"Replacing UNet encoders with TransUNet plus neighbor-aware point attention gains up to 1.8 dB PSNR at six views.","key_machinery":"The load-bearing object is the two-stage dual-transformer combination. The first stage is TransUNet, a U-shaped CNN-transformer used as a shared encoder; its four decoder feature maps are concatenated after per-view bilinear sampling and max-pooling across views, giving a 464-channel feature per 3D point. The second stage is a neighbor-aware Point Transformer with two layers, k=3 nearest neighbors, learnable positional encoding added to point features, and multi-head attention whose scores combine scaled dot-product similarity with log(w_ij), where w_ij is a Gaussian decay in Euclidean distance. This mechanism is what enforces local smoothness and spatial coherence that plain per-point MLP prediction lacks.","core_discovery":"The central claim is that sparse-view CBCT reconstruction is best modeled as a continuous attenuation field whose per-point prediction uses two complementary transformers. First, TransUNet encodes each projection at multiple scales, and view-specific features are sampled at the projected location of each 3D point and fused across views by max-pooling. Second, a Point Transformer adds learnable 3D positional encodings and a Neighbor-Aware Attention that restricts attention to k nearest neighbors, weighting scores by a Gaussian function of Euclidean distance. The paper reports that Trans-CBCT alone already beats all baselines, and Trans2-CBCT adds further gains, achieving 31.03 dB PSNR and 0.9027 SSIM on LUNA16 at six views versus 29.23 dB and 0.8747 for the strongest baseline, C2RV.","pith_inferences":["If the dual-transformer design is genuinely responsible for the margin, similar gains might appear in other inverse problems with sparse angular sampling, such as fan-beam sparse-view CT or limited-angle digital breast tomosynthesis, where global context and local smoothness are both critical.","The reported sensitivity to k suggests a testable extension: an adaptive or learned neighbor count, or a re-weighted distance kernel, could push accuracy further or shift the optimal k as voxel size or anatomy changes.","Because the paper quotes most baseline numbers from the DIF-Gaussian work rather than re-running them, the cleanest confirmation is an independent re-run under identical data splits, DRR settings, and metrics; until that happens, the architectural conclusion should be treated as conditional on protocol matching."],"forward_implications":["Replacing UNet or ResNet encoders with TransUNet yields measurable gains even before 3D refinement: +1.17 dB PSNR over the best baseline at six views on LUNA16.","The neighbor-aware Point Transformer contributes an additional 0.63 dB PSNR and 0.0117 SSIM on the same setting, showing that explicit 3D spatial reasoning adds value beyond better 2D features.","The improvements persist from six to ten views on both a chest CT dataset and a dental CBCT dataset, indicating the recipe transfers across anatomies and imaging geometries.","Using k=3 nearest neighbors beats larger neighborhoods, which means overly broad aggregation over-smooths fine anatomical boundaries and lowers both PSNR and SSIM.","Trans-CBCT keeps inference time near prior INR-based methods despite having over three times the parameters, while Trans2-CBCT's KNN search currently costs 53.7 seconds per case, identifying a concrete efficiency target."],"supporting_citations":[{"why":"Supplies the hybrid CNN-transformer encoder whose multi-scale features are the backbone of both proposed models.","marker":"[15]"},{"why":"Defines the continuous-intensity-field formulation and the point-wise feature querying pipeline that the paper extends.","marker":"[6]"},{"why":"C2RV is the strongest prior baseline that Trans2-CBCT claims to beat by 1.8 dB PSNR at six views.","marker":"[7]"},{"why":"Provides the remaining baseline numbers and the experimental protocol from which most comparisons are quoted.","marker":"[26]"},{"why":"Supplies the LUNA16 chest CT dataset used for the primary quantitative comparison.","marker":"[16]"},{"why":"Supplies the ToothFairy dental CBCT dataset used to show the method transfers across anatomy.","marker":"[17]"}],"fun_headline_variants":["Dual transformers sharpen sparse-view CBCT","Trans2-CBCT: two transformers beat UNet for CBCT","Neighbor-aware point attention lifts CBCT quality","Sparse-view CBCT gains from dual attention","Transformer duo improves low-dose CBCT reconstruction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the baseline numbers the paper quotes from DIF-Gaussian were produced under exactly the same data splits, preprocessing, view angles, and metric computation as the authors' own runs; if the protocols differ, the reported margins may reflect evaluation mismatch rather than architecture.","fun_headline_variants_meta":{"raw":{"variants":["Dual transformers sharpen sparse-view CBCT","Trans2-CBCT: two transformers beat UNet for CBCT","Neighbor-aware point attention lifts CBCT quality","Sparse-view CBCT gains from dual attention","Transformer duo improves low-dose CBCT reconstruction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001114,"raw_usage":{"total_tokens":4654,"prompt_tokens":974,"completion_tokens":3680,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":3606}},"tokens_in":590,"tokens_out":3680,"duration_ms":27831,"temperature":1.0,"reasoning_tokens":3606,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:09:13.396379+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run C2RV, DIF-Gaussian, and Trans2-CBCT on the same LUNA16 split with identical DRR generation, the same six view angles, the same sampled points, and the same PSNR and SSIM calculation script. If the PSNR difference between Trans2-CBCT and C2RV falls below the reported 1.8 dB, or if an independent implementation of C2RV outperforms the quoted number, the central superiority claim would be falsified.","supporting_citations":[{"cited_title":"Learning deep intensity field for extremely sparse-view cbct reconstruction,","cited_arxiv_id":null,"evidence_quote":"Defines the continuous-intensity-field formulation and the point-wise feature querying pipeline that the paper extends."},{"cited_title":"Cˆ2rv: Cross- regional and cross-view learning for sparse-view cbct reconstruction,","cited_arxiv_id":null,"evidence_quote":"C2RV is the strongest prior baseline that Trans2-CBCT claims to beat by 1.8 dB PSNR at six views."},{"cited_title":"Learning 3d gaussians for extremely sparse-view cone-beam ct reconstruction,","cited_arxiv_id":null,"evidence_quote":"Provides the remaining baseline numbers and the experimental protocol from which most comparisons are quoted."},{"cited_title":"Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: the luna16 challenge,","cited_arxiv_id":null,"evidence_quote":"Supplies the LUNA16 chest CT dataset used for the primary quantitative comparison."},{"cited_title":"Deep segmentation of the mandibular canal: a new 3d annotated dataset of cbct volumes,","cited_arxiv_id":null,"evidence_quote":"Supplies the ToothFairy dental CBCT dataset used to show the method transfers across anatomy."}],"review_version":1}