{"id":"f867b56b-98cd-4a27-a898-797f9623ba78","arxiv_id":"2508.09205","paper_version":2,"verdict":"REJECT","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A 14-second breath-hold MRCP reconstruction with zero-shot self-supervised learning improves over compressed sensing, and freezing early unrolled stages cuts training time up to 6.7-fold with a small PSNR penalty.","lead":"This paper tests a zero-shot deep learning method that reconstructs MRI images of the bile ducts from a single 14-second breath-hold scan, without needing a large training dataset. It also introduces a trick to speed up per-patient training time by freezing early network layers, moving the method closer to practical clinical use.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract claims comparability to successful respiratory-triggered MRCP, but Results explicitly state no breath-hold acquisition reaches that image quality; the central claim is internally contradicted.","rationale":"The strongest_claim is the abstract's conclusion. The most load-bearing condition for that claim is that zero-shot breath-hold images are comparable in quality to successful respiratory-triggered acquisitions. The paper's own Results section (Fig. 7) states the opposite: 'no breath-hold acquisition reaches the image quality of the (successful) triggered acquisition.' This is an internal inconsistency, not a matter of outside consensus, and it directly undermines the abstract's central promise. The reader's weakest_assumption focused on the PSNR reference being a reconstruction rather than ground truth and on n=2; that is also valid, but the self-contradiction is more fundamental because it does not depend on external validity. I also note the arXiv metadata abstract (explainable AI in pathology) does not match the full text (zero-shot MRCP), which is a submission-integrity issue, but it is orthogonal to the scientific claim. A blinded radiologist comparison would resolve whether the qualitative claim can be salvaged; absent that, the abstract and conclusion must be revised to state that zero-shot improves over compressed sensing but remains inferior to successful triggered acquisition. Since the reader already recommended REJECT and my concern strengthens that without changing direction, the verdict should remain as is.","tokens_in":12652,"tokens_out":3669,"duration_ms":34604,"concrete_test":"Have at least two radiologists, blinded to acquisition type, independently compare the zero-shot 14s breath-hold reconstructions (e.g., Fig. 7 cases) against the successful respiratory-triggered images on ductal delineation and overall diagnostic quality. If triggered images are rated superior in the majority of comparisons, the abstract's 'comparable' claim is falsified. Also verify the two quoted sentences in context; if both remain unqualified, the manuscript is internally inconsistent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in the abstract, is that zero-shot reconstruction 'achieved image quality comparable to that of successful respiratory-triggered acquisitions with regular breathing patterns.' The Results section, however, states explicitly (Fig. 7): 'no breath-hold acquisition reaches the image quality of the (successful) triggered acquisition.' These statements cannot both be true. The clinical-translation rationale — that a 14s breath-hold can replace a 338s triggered scan without quality loss — rests entirely on this comparability. Since the paper's own data contradict it, the headline claim is unsupported. The quantitative PSNR comparison does not rescue it: the reference is itself a CS reconstruction of a 40s R=6 breath-hold, not a fully sampled ground truth, and was obtained for only two volunteers. Thus the central quality claim depends on an internal contradiction plus an inadequate reference standard.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The full text (pages 2–24) reports a study on zero-shot self-supervised learning reconstruction for breath-hold magnetic resonance cholangiopancreatography (MRCP). The authors propose combining 2D Poisson-disk and partial-Fourier undersampling to achieve a 14-second breath-hold at R=25, then reconstruct images with a zero-shot unrolled network. To reduce training time, they split the 13 unrolled stages into frozen and trainable parts, initializing frozen stages with a pretrained SSDU backbone and training only the last stages. Experiments on 11 healthy volunteers compare zero-shot reconstructions against compressed sensing and respiratory-triggered acquisitions. The authors report that zero-shot improves over compressed sensing, and they quantify PSNR on two retrospectively undersampled breath-hold acquisitions against a CS reconstruction of a 40-s R=6 acquisition. The abstract claims image quality comparable to successful respiratory-triggered acquisitions, but the Results section explicitly states that no breath-hold acquisition reaches that quality. The paper also contains a title and abstract that are inconsistent with the body: the manuscript header describes an explainable-AI pathology system, while the full text presents the MRCP study.","tokens_in":12844,"tokens_out":5063,"duration_ms":54253,"significance":"If the central claims were established, the proposed partial-trainable zero-shot strategy could be a useful step toward practical breath-hold MRCP with high acceleration and reduced reconstruction time. The idea of freezing early unrolled stages to lower backpropagation depth and cache intermediate k-space is a reasonable technical contribution, and the qualitative figures suggest zero-shot outperforms compressed sensing at R=25. However, the contradiction between the abstract and the Results, the use of a non-ground-truth PSNR reference on only two subjects, and the selection bias in the triggered acquisitions all prevent the paper from supporting its stated comparability claim. The technical idea is potentially salvageable, but the current manuscript is not coherent enough for publication.","major_comments":[{"comment":"The manuscript is internally inconsistent at the highest level. The title and the abstract printed before the full text describe a vision-language-model explanation system for computational pathology, while the body is an MRCP reconstruction study with a different title ('Zero-shot self-supervised learning of single breath-hold magnetic resonance cholangiopancreatography (MRCP) reconstruction'). A journal submission must have a title, abstract, and body that refer to the same work. This alone is a blocking defect.","section":"Title/Abstract vs. Full Text"},{"comment":"The abstract states 'achieved image quality comparable to that of successful respiratory-triggered acquisitions with regular breathing patterns.' Yet Results (Figure 7) states: 'With the exception of residual motion artifacts ... no breath-hold acquisition reaches the image quality of the (successful) triggered acquisition.' The Discussion repeats this. These statements are directly contradictory, and the clinical-translation rationale depends on the comparability claim. The abstract must be corrected to match the actual findings, or the finding must be supported by evidence it currently lacks.","section":"Abstract vs. Results, Figure 7"},{"comment":"The PSNR reference is not a fully sampled ground truth. The 40-s R=6 breath-hold is reconstructed with ℓ1-wavelet compressed sensing, and PSNR is computed against that reconstruction after retrospective undersampling to R=25. This measures agreement with a CS reconstruction, not true image fidelity. Moreover, these numbers are reported for only two volunteers. With n=2 and a non-ground-truth reference, the quantitative claim that 'PSNR decreased only slightly' from 38.25 to 37.67 dB cannot establish image-quality equivalence, and any small differences are not statistically meaningful.","section":"Methods, Quantitative analysis; Results, Figure 4"},{"comment":"The respiratory-triggered reference scans were acquired by repeatedly scanning volunteers until a 'sufficiently regular breathing pattern' was obtained. This selection process biases the reference toward unusually high quality and does not represent the typical or clinical distribution of triggered acquisitions. Using such a cherry-picked reference as the benchmark makes the stated 'comparability' claim even more difficult to support, and the comparison is not fair to the breath-hold method. The paper should report the success rate and the range of triggered image qualities, not only the best cases.","section":"Methods, Data"},{"comment":"Many numeric values are missing or unreadable in the manuscript text: e.g., reconstruction times in Table 2 and Figure 4, PSNR values in Figure 6, the mean volunteer age in the Data section, and the equations in Methods. These placeholders make the central quantitative claims unverifiable. I understand some of this may be a typesetting artifact, but as submitted, the paper does not allow a reader to check the reported speedups (6.7-fold, 4.3-fold) or the PSNR trade-offs.","section":"Throughout Results and Methods"}],"minor_comments":[{"comment":"The compressed-sensing regularization parameter is set to 0.008 because it 'provided visually optimal reconstructions.' This is an ad hoc free parameter; a sensitivity analysis or a systematic selection criterion (e.g., L-curve) should be reported to show the comparison is not biased by an unfavorable CS setting.","section":"Methods, Conventional reconstructions"},{"comment":"The mask-split ratios Γ, Θ, and Θ_v and the number of unrolled stages (13) are introduced without justification or sensitivity analysis. Since these are free parameters of the method, the authors should explain how they were chosen and whether results are robust to them.","section":"Methods, Zero-shot learning"},{"comment":"The learning rate was reduced to 0.0001 for volunteer #11 due to instability. This subject-specific tuning is a form of peeking at the test case; its effect on the reported image quality and time should be discussed transparently.","section":"Methods, Data"},{"comment":"The note contains a typo: 'initizliaed' should be 'initialized.'","section":"Table 2"},{"comment":"The paper states 'the image quality assessment was based solely on visual inspection by the authors, without any involvement of radiologists.' This is a significant limitation for a clinical-imaging claim and should be stated earlier, ideally in the Methods.","section":"Discussion"}],"recommendation":"reject","confidential_remarks":"The mismatch between the supplied abstract (explainable AI in pathology) and the full text (MRCP reconstruction) is severe; if this is a pipeline artifact, the actual submission still needs to align its title/abstract with the body. More substantively, the central claim of comparability to triggered acquisitions is contradicted by the paper's own Results, and the quantitative support rests on two subjects and a CS reference. Even with a correction of the abstract, the evidence would need substantial strengthening (larger N, ground-truth reference, radiologist readers) to justify the clinical-translation language. I see a potentially salvageable technical contribution, but not in the current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The actual paper—the one buried under the mismatched metadata—is a solid engineering study. The core idea is simple and useful: instead of backpropagating through all 13 unrolled stages during zero-shot SSDU training, freeze the early stages from a pretrained network, cache their output, and only train the last few. That cuts reconstruction time by up to 6.7x with a small PSNR drop (about 0.6 dB), and it holds across the configurations tested. The application to 14-second breath-hold MRCP at R=25 is new, and the authors go out of their way to compare against compressed sensing and a pretrained model. Credit where due: they ship code, release sample data, and the limitations section is unusually honest. They explicitly say no radiologist reader study was done, that the PSNR reference is itself a CS reconstruction, and that only two volunteers were used for quantitative numbers.\n\nNow the soft spots, and one of them is a load-bearing crack. The abstract claims image quality comparable to successful respiratory-triggered acquisitions, but the Results state flatly that “no breath-hold acquisition reaches the image quality of the (successful) triggered acquisition.” Those cannot both be true. The Discussion also concedes the point. That contradiction is not a minor wording issue; it is the clinical-translation claim. On the quantitative side, the PSNR numbers come from two retrospectively undersampled subjects, against an R=6 CS reconstruction, not a fully sampled ground truth. No error bars, no reader study, and one volunteer needed a special learning rate. These are all acknowledged in the text, which helps, but they still cap how much the evidence can support.\n\nThe freeze-train splitting and caching is a genuine methodological contribution, and the comparison against CS with the proposed sampling pattern is informative. But the paper as submitted oversells its headline result. A serious referee should see it; with a revised abstract and Results framing that matches the actual quality gap, plus more subjects and ideally a reader study, it could be a useful clinical imaging paper. The metadata mismatch is strange and should be flagged to the authors—it looks like a submission error, not an intentional bait-and-switch, but it makes the arXiv index wrong and undermines trust.","headline":"A real MRCP reconstruction engineering contribution with a head-scratcher of a submission: the arXiv metadata abstract is about pathology explainability, while the actual paper is about zero-shot MRCP, and the abstract's quality claim is contradicted by the paper's own Results.","tokens_in":13378,"tokens_out":954,"would_cite":false,"duration_ms":11989,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a partially trainable zero-shot self-supervised network can reconstruct 14-second breath-hold MRCP at R=25 with image quality close to triggered acquisitions and reconstruction times up to 6.7 times shorter than full","keywords":["breath-hold MRCP","zero-shot self-supervised learning","deep learning-based MRI reconstruction","MR cholangiopancreatography","self-supervised training","compressed sensing","unrolled network","partial trainable"],"falsifier":"Scan a phantom with a known fully sampled ground truth using the same R=25 Poisson-disk/partial-Fourier pattern; if the 12/1 zero-shot reconstruction loses more than the reported ~0.6 dB against that truth, or leaves partition-encoding aliasing behind, the claim that partial training preserves fidelity is falsified. Independently, timing the 0/13 through 12/1 configurations on one GPU would check the claimed approximately linear relationship between reconstruction time and number of trainable stages.","tokens_in":12534,"feed_emoji":"🩻","tokens_out":11794,"duration_ms":106445,"temperature":0.7,"pith_summary":"This paper is trying to establish that zero-shot self-supervised reconstruction—training a network on the single undersampled scan being reconstructed—can make breath-hold MRCP practical at 14 seconds under 25-fold acceleration, without a large training dataset of fully sampled images. The payoff is clinical: current breath-hold MRCP needs around 20 seconds, and respiratory-triggered alternatives can fail or stretch over several minutes when breathing is irregular. The proposed mechanism is to split the unrolled reconstruction network into frozen early stages, initialized from a pretrained backbone, and a small number of trainable late stages; this cuts reconstruction time up to 6.7-fold while the image-quality metric PSNR (peak signal-to-noise ratio) drops only from 38.25 dB to 37.67 dB. If true, this offers a route from scan-specific deep learning into time-constrained clinical workflows.","feed_headline":"Partial training makes zero-shot MRCP 6.7x faster","feed_subtitle":"Freezing early network stages barely dents image quality while cutting reconstruction time, bringing breath-hold scans closer to clinics.","key_machinery":"The central mechanism is a frozen/trainable split of a 13-stage unrolled reconstruction network, where each stage alternates a residual-network regularization block with a conjugate-gradient data consistency block that enforces the MRI encoding equation. The first $f$ stages are initialized from a self-supervised pretrained backbone and frozen; their input-to-output map is cached after the first epoch, so backpropagation passes only through the $t$ trainable stages. This cuts gradient depth and memory, and the pretrained frozen part supplies a strong starting point that keeps quality high when only a few stages are trained.","core_discovery":"The paper's central claim is that zero-shot self-supervised learning can reconstruct 14-second breath-hold MRCP at R=25 with image quality clearly better than compressed sensing and close to successful respiratory-triggered acquisitions. It further claims that the practical blocker—multi-hour zero-shot training—can be lifted by freezing early unrolled stages initialized from a pretrained backbone and training only the last stages; this yields up to a 6.7-fold reduction in reconstruction time with a PSNR change of roughly half a decibel. Results also show that using a pretrained network for the frozen stages beats an $\\ell_1$-wavelet compressed-sensing initialization, raising PSNR by more tha","pith_inferences":["Editorial inference: the frozen-stage caching strategy is not MRCP-specific; it should transfer to other scan-specific unrolled reconstructions, such as diffusion-weighted or cardiac MRI, wherever a fully sampled readout dimension can be decoupled, since the speedup comes from reducing backpropagation depth rather than from biliary anatomy.","Editorial inference: the residual partition-encoding aliasing seen with the 12/1 configuration suggests the safe number of trainable stages depends on the sampling geometry; with R=2 equidistant undersampling along the partition direction, one trainable stage may be too few to fully suppress aliasing, so the optimal split should be tuned per sampling pattern.","Editorial inference: the clinical-translation claim would be directly testable in a blinded reader study on patients with strictures or dilations; the paper's own visual assessment by the authors and the two-volunteer PSNR numbers are not strong enough evidence for diagnostic equivalence."],"forward_implications":["A 14-second breath-hold MRCP at R=25 becomes feasible, within the 10–15 s range considered tolerable for patients who cannot hold their breath longer.","Freezing early stages and training only the last stages cuts reconstruction time up to 6.7-fold with random initialization and 4.3-fold with pretrained initialization, while PSNR moves from 38.25 dB (0/13) to 37.67 dB (12/1).","Initializing the frozen stages from a pretrained network instead of from compressed sensing raises PSNR by roughly 6.3–6.7 dB and also shortens reconstruction time.","At R=25, zero-shot reconstruction shows better ductal delineation and less aliasing than $\\ell_1$-wavelet compressed sensing, though successful respiratory-triggered acquisitions still produce the sharpest images.","Because the regularization weight is learned during training, the method avoids per-scan regularization tuning beyond occasional learning-rate adjustment."],"supporting_citations":[{"why":"Supplies the zero-shot self-supervised reconstruction method that the paper adapts to breath-hold MRCP.","marker":"29"},{"why":"Provides the self-supervised physics-guided loss and k-space partition strategy used to pretrain the backbone and to define the zero-shot objective.","marker":"28"},{"why":"Supplies the 3T raw MRCP dataset and previous reconstruction setup used to train the pretrained network.","marker":"27"},{"why":"Defines $\\ell_1$-wavelet compressed sensing, the conventional reconstruction baseline that zero-shot must beat.","marker":"21"},{"why":"Provides the combined compressed-sensing/parallel-imaging sampling design used to achieve R=25 in a 14-second acquisition.","marker":"24,25"},{"why":"Supplies the ESPIRiT coil sensitivity estimation used in the encoding operator for all reconstructions.","marker":"38"},{"why":"Provides GRAPPA, used to reconstruct the respiratory-triggered acquisitions that serve as the image-quality reference.","marker":"42"}],"fun_headline_variants":["From explainable to explained AI via falsification","Quantify and falsify AI explanations","Human-VLM loop tests and quantifies explanations","Sliding-window tests verify pathology AI explanations","Explained AI: quantify and falsify model claims"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that a compressed-sensing reconstruction of a 40-second breath-hold at R=6 is an acceptable stand-in for ground truth when measuring PSNR at R=25, and that the PSNR numbers come from just two volunteers; if that reference is biased, the reported quality advantage of zero-shot reconstruction is not established.","fun_headline_variants_meta":{"raw":{"variants":["From explainable to explained AI via falsification","Quantify and falsify AI explanations","Human-VLM loop tests and quantifies explanations","Sliding-window tests verify pathology AI explanations","Explained AI: quantify and falsify model claims"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1119,"prompt_tokens":709,"completion_tokens":410,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":340}},"tokens_in":453,"tokens_out":410,"duration_ms":4270,"temperature":1.0,"reasoning_tokens":340,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:26:04.775665+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Scan a phantom with a known fully sampled ground truth using the same R=25 Poisson-disk/partial-Fourier pattern; if the 12/1 zero-shot reconstruction loses more than the reported ~0.6 dB against that truth, or leaves partition-encoding aliasing behind, the claim that partial training preserves fidelity is falsified. Independently, timing the 0/13 through 12/1 configurations on one GPU would check the claimed approximately linear relationship between reconstruction time and number of trainable stages.","supporting_citations":[{"cited_title":"Zero-Shot Self-Supervised Learning for MRI Reconstruction","cited_arxiv_id":null,"evidence_quote":"Supplies the zero-shot self-supervised reconstruction method that the paper adapts to breath-hold MRCP."},{"cited_title":"Deep Learning��Based Accelerated MR Cholangiopancreatography Without Fully��Sampled Data","cited_arxiv_id":null,"evidence_quote":"Supplies the 3T raw MRCP dataset and previous reconstruction setup used to train the pretrained network."}],"review_version":1}