{"id":"5e2d4da8-e50d-42f8-87ef-f53fb25a78a9","arxiv_id":"2412.07590","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"PFAD removes MRI motion artifacts without paired clean data by guiding a pretrained diffusion model with low-frequency k-space information and alternating complementary masks in pixel and frequency domains.","lead":"This paper presents PFAD, an unsupervised method that removes MRI motion artifacts by guiding a pretrained diffusion model with the low-frequency parts of the corrupted image and alternating masks in both pixel and frequency domains. It reports better artifact removal than several GAN- and diffusion-based baselines on simulated brain, knee, and abdominal MRI, and higher radiologist ratings on real clinical images.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The π/10 high-frequency assumption is load-bearing: simulated artifacts are generated with the same cutoff the method uses to freeze low-frequency guidance, so the quantitative benchmark may confirm the premise rather than test the method.","rationale":"The reader's weakest_assumption identifies the same point: the method's correctness hinges on artifacts being high-frequency, and the simulation uses the same cutoff, making the quantitative evaluation circular with respect to that assumption. This is the most load-bearing concern because the central claim of 'superior performance' is supported primarily by Table 1, and that table is generated under the method's own premise. It is not an internal inconsistency: the paper explicitly labels the assumption as lenient and even tests cutoff sensitivity, so the genuine risk is external validity. The proposed test directly disrupts the premise by injecting artifacts into low frequencies; if the method still wins, the central claim survives at least for that perturbation class; if it does not, the reported superiority is an artifact of benchmark construction. I considered whether the omission of Oh et al. 2023 as a baseline is more load-bearing, but that is a comparison-fairness issue; even if resolved, it would not test whether the method's mechanism works when its core assumption fails. The hyperparameter tuning on the test set is secondary because the ablation shows qualitative stability, and the real-image radiologist study provides some independent evidence, though with only 40 images. Overall, the appropriate verdict remains CONDITIONAL: the method is plausible and reproducible in structure, but the quantitative claim should not be accepted without testing the method on artifacts that violate the π/10 assumption.","tokens_in":16394,"tokens_out":5348,"duration_ms":52219,"concrete_test":"Using the released code and the same hyperparameters and random seeds, rerun the brain and knee experiments with a modified simulation: set k0 = 0 in Supplementary Eq. 13 (and analogously remove the high-pass restriction in Eq. 14), so phase perturbations are applied across all of k-space including low frequencies. Compare PFAD against the same baselines on PSNR, SSIM, and LPIPS. If PFAD's margins over the best baseline shrink substantially or reverse, the π/10 assumption is load-bearing; if the gains persist, the concern is largely answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The method's central mechanism relies on the assumption that motion artifacts are confined to high-frequency k-space, |ky| > π/10 (Related Work). Equations 4-5 then freeze the corrupted image's low-frequency component into every reverse step as guidance, so if real artifact energy leaks below π/10, the guidance itself is corrupted and the method cannot remove that part of the artifact. The quantitative support for the central claim comes almost entirely from simulated images, and the simulation shares the same premise: Supplementary Eq. 13-14 impose phase perturbations exp(-jΦ(ky)) only for |ky| > k0, with k0 fixed to π/10. Thus the brain, knee, and abdominal benchmarks are in-distribution for the method's key assumption. The cutoff study in Table 7 further selects π/10 as optimal on images generated with exactly that cutoff, so the optimum is expected rather than evidence about real artifacts. Real motion can occur during the initial acquisition of central k-space, and respiratory or bulk motion can produce low-frequency components; in those cases the frozen low-frequency guidance in Eq. 4 would reintroduce artifact structure. The paper is honest about calling the assumption 'lenient,' but the consequence is that the reported PSNR/SSIM/LPIPS gains over baselines may reflect alignment between the artifact simulator and the method's filter, not a demonstrated ability to remove clinically realistic motion artifacts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes PFAD (Pixel-Frequency Alternate Masks Diffusion), an unsupervised method for removing MRI motion artifacts. A diffusion model pre-trained on clean images is guided during reverse diffusion by the low-frequency k-space of the corrupted image, while alternating complementary masks in the pixel and frequency domains are used to destroy artifact structure and exploit usable information. A time-varying hyperparameter a balances frequency- and pixel-domain guidance. The method is evaluated on simulated motion-corrupted brain (HCP), knee (fastMRI), and abdominal MRI, and on 40 real clinical abdominal images, with quantitative metrics and radiologist Likert scoring.","tokens_in":16779,"tokens_out":5460,"duration_ms":47369,"significance":"If the reported gains are robust, PFAD would be a useful contribution to unsupervised MRI motion-artifact correction: it needs no paired data, preserves a clear role for k-space information, and provides a concrete inference algorithm with released code. The paper is honest about its 'lenient' assumption that artifacts are high-frequency, and it includes ablation studies and a small radiologist study. However, the central quantitative claim is currently supported mainly by a simulation pipeline that embeds the method's own cutoff assumption, and one key hyperparameter is selected on the evaluation set. These issues make the superiority claim (Abstract, Table 1) not yet convincing; the real-image evaluation is too small and too subjective to carry the claim on its own.","major_comments":[{"comment":"The quantitative benchmark is circular with respect to the method's central assumption. In Related Work the authors adopt the assumption that motion artifacts are confined to |ky| > pi/10, and Eq. (4) freezes all low-frequency content of the corrupted image as guidance. The simulated artifacts in Supplementary Eqs. (13)-(14) are generated by applying phase perturbations only for |ky| > k0, with k0 fixed to pi/10. Thus the test data are produced under exactly the premise that the method's filter uses, and the large gains over baselines in Table 1 may reflect this alignment rather than an ability to remove clinically realistic artifacts. The cutoff experiment in Table 7 does not resolve this because it is also conducted on images simulated with k0 = pi/10; selecting pi/10 as optimal under that generator is an expected outcome, not evidence about real artifact spectra. Please add experiments with artifact perturbations extending below pi/10 (e.g., k0 = pi/20 or with low-frequency leakage), or provide real artifact data with paired ground truth or quantitative spectral characterisation.","section":"Supplementary A.2, Eqs. (13)-(14); Related Work"},{"comment":"The balance parameter a is selected on the evaluation set, which inflates the reported metrics. Table 5 lists PSNR/SSIM/LPIPS/... for different a values and states that 'we choose the value of a for the best case of the total metric, where a is equal to 0.7.' These appear to be the same test-set metrics reported in Table 1 for the brain dataset, so the method's hyperparameter has been tuned to the test data. Please either (i) derive a from a separate validation split and report test metrics only for the fixed value, or (ii) clearly state that a was selected on a validation set and provide those validation results. The ad hoc 'Total' metric (sum of some metrics minus others, with no normalisation or justification) should also be justified or replaced with a standard criterion.","section":"Table 5 and 'Hyperparameter Study'"},{"comment":"The central mechanism is not robust to low-frequency artifact energy. Because Eq. (4) copies the entire low-frequency component of the corrupted image into every reverse step, any real artifact energy below the pi/10 cutoff is frozen into the output and cannot be removed. The paper explicitly acknowledges the assumption is 'lenient', but it never tests the failure mode. Real patient motion during the initial acquisition of central k-space, or respiratory/bulk motion with low-frequency components, violates this assumption. Please provide an experiment that quantifies performance degradation as artifact energy is progressively moved below pi/10, or a spectral analysis of the real clinical images showing that their artifact energy is indeed confined to |ky| > pi/10. Without such evidence, the claim that PFAD works on 'real clinical images' (Figure 5, Table 2) is not strongly supported.","section":"Eq. (4); Related Work"},{"comment":"The closest prior work is not compared. The paper adopts its pi/10 assumption from Oh et al. (2023) and describes that work as using score-based diffusion with measured k-space values, but Oh et al. is absent from Table 1 and from the qualitative comparisons. Since Oh et al. is a recent diffusion-based method designed for the same task (MR motion artifact reduction), omitting it undercuts the claim of 'superior performance.' Please include Oh et al. (and, if practical, one more recent k-space-aware baseline) under the same evaluation protocol.","section":"Table 1; Comparison Approaches"},{"comment":"Statistical reporting is incomplete. Table 1 reports only means for each metric, with a footnote that Mann-Whitney U tests were used, but no variances, confidence intervals, or effect sizes are given for the quantitative metrics. With the reported differences being small in several cases (e.g., brain PSNR 27.60 for PFAD vs 26.76 for Pix2pix), the reader cannot judge whether the differences are practically meaningful. Please report standard deviations or 95% confidence intervals, and specify the number of test images for each dataset. For the radiologist study, only 40 real images were scored; this is a small sample and the inclusion of variance in Table 6 is helpful, but the protocol should state whether the two radiologists' scores were averaged or adjudicated per-image, and how inter-rater agreement was quantified.","section":"Table 1; Supplementary B.2, Table 6"}],"minor_comments":[{"comment":"In the fourth contribution the phrase 'demonstrates the of our method in metrics' is missing a word ('effectiveness' or 'superiority'); please fix this typo.","section":"Introduction, list of contributions"},{"comment":"The modulus operation in Eq. (6) discards phase information, and the text states that pixel-domain processing compensates for this. Please make this explanation more precise: what exactly is lost, and how does the pixel-domain branch recover it? A short worked illustration would help.","section":"Eq. (6) and surrounding text"},{"comment":"The sentence describing the difference heatmap ('The color ranges from blue to red, indicating differences from small to large, with deeper colors representing smaller differences') is self-contradictory as written; please clarify whether deeper colors mean larger or smaller errors.","section":"Figure 4 caption"},{"comment":"The sentence 'Noting that most previous works are based on GANs and remove artifacts only in pixel domain' is inaccurate for UDDN, which is described earlier in the same paragraph as removing artifacts in the frequency domain; please rephrase to acknowledge UDDN's k-space operation.","section":"Related Work, paragraph on Motion Artifact Removal"},{"comment":"The definition of omega_i in line 3 could be confused with the noise schedule; please add a short comment that omega_i is the standard DDPM posterior weight (1 - sqrt(alpha_bar_i)) used in RePaint-style guidance.","section":"Algorithm 1, line 3"}],"recommendation":"major_revision","confidential_remarks":"The main concern is structural rather than adversarial: the evaluation pipeline reproduces the method's own high-frequency assumption, and the hyperparameter a is selected on the test set. Both issues are fixable with additional experiments (variable k0, hold-out validation, and a comparison with Oh et al. 2023). I recommend major revision rather than rejection because the core algorithm is clearly described and the real-image qualitative results, though small, suggest the method may have practical value. Please do not accept without the additional experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one for the mechanism, not for the headline numbers. PFAD's combination—freezing low-frequency k-space of the corrupted image, flipping complementary checkerboard masks each reverse step, and blending pixel and frequency guidance with a time-varying weight—is genuinely new relative to Repaint, UDDN, DR2, and GDP. The algorithm is clearly specified, the ablations make sense, and the authors are upfront that the π/10 high-frequency assumption is borrowed from Oh et al. They also release code.\n\nThe main soft spot is that the simulation generator and the method share the same premise. Supplementary Eqs. 13–14 apply phase perturbations only for |ky| > π/10, and Eq. 4 freezes the corrupted image's low-frequency content below exactly that cutoff. So the reported PSNR/SSIM/LPIPS gains on brain, knee, and abdominal data are partly in-distribution for the method's core assumption. If real motion perturbs central k-space, the frozen low-frequency guidance reintroduces that artifact and the method has no way to remove it. The cutoff study in Table 7 only deepens this: π/10 wins on images generated with π/10. That isn't a knock on the authors' honesty—they call the assumption 'lenient'—but it means the quantitative claim is conditional on a premise the evaluation never tests.\n\nA couple of smaller issues: the balance parameter a is chosen on the evaluation set using an ad hoc total metric (Table 5), and the closest SOTA, Oh et al. 2023, is cited but not benchmarked. The real-image study is small (40 images, two junior readers plus a senior adjudicator), though the scoring protocol is reasonable.\n\nNone of this sinks the paper. The method is clearly designed, the ablations show each component earns its place, and the clinical evaluation, while small, is a genuine attempt. The proper fix is a robustness experiment: simulate or collect artifacts with energy below π/10, or at least show sensitivity of the guidance to the cutoff. That's a revision, not a rejection.\n\nWho this is for: anyone working on unsupervised MRI artifact removal or diffusion-based inverse problems. Worth a serious referee; the authors should be pushed on the low-frequency robustness question.","headline":"Clever diffusion-based recipe for MRI motion artifact removal, but its quantitative benchmark leans on the same high-frequency assumption the method makes, so the reported gains are less decisive than the tables suggest.","tokens_in":17208,"tokens_out":2607,"would_cite":false,"duration_ms":23259,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PFAD removes MRI motion artifacts at inference without paired data by locking low-frequency k-space and alternating complementary masks across pixel and frequency domains, outperforming supervised and unsupervised baselines on brain…","keywords":["MRI motion artifact removal","diffusion model","unsupervised image restoration","k-space guidance","alternate complementary masks","pixel-frequency domain","DDPM","medical image quality"],"falsifier":"Simulate rigid-motion phase corruption applied only to low-frequency k-space lines with $|k_y| < \\pi/10$ (below the cutoff the method treats as clean), with the same amplitude used for high-frequency corruption. If PFAD's output then retains visible ghosting or its PSNR and SSIM drop far below the high-frequency-corruption case, the low-frequency-lock premise is falsified. Alternatively, on retrospectively motion-tracked real data, measure the k-space phase-error spectrum: energy below $\\pi/10$ would directly violate the assumption.","tokens_in":16134,"feed_emoji":"🩻","tokens_out":9417,"duration_ms":85600,"temperature":0.7,"pith_summary":"The paper proposes PFAD, an unsupervised inference-time method for removing MRI motion artifacts with a pretrained diffusion model. Its central claim is that a motion-corrupted scan can be purified without paired clean data by copying the corrupted image's low-frequency k-space into every reverse-diffusion step and using alternating complementary masks to blend corrupted and generated content in both frequency and pixel domains. If this claim holds, artifact cleanup becomes practical in the clinic, because a model trained once on clean images of an anatomy can be applied to whatever corrupted scan arrives. The paper reports that PFAD beats supervised (Pix2pix) and unsupervised (CycleGAN, UDDN, DR2, GDP) baselines on quantitative metrics for simulated brain, knee, and abdominal motion artifacts, and receives the highest radiologist Likert scores on real clinical abdominal images.","feed_headline":"No paired data needed to erase MRI motion artifacts","feed_subtitle":"Freezing low-frequency k-space and alternating masks, PFAD lets a diffusion model rebuild clean tissue on real scans","key_machinery":"The central mechanism is the alternate complementary mask pair $M_t$ and $1-M_t$, with $M_t = \\omega_t m_t$ and $m_t$ flipped at every reverse step ($m_{t-1}=1-m_t$). It splits both domains into alternating checkerboard areas, so the artifact-corrupted high-frequency and pixel content is never used wholesale, while useful structure from the actual scan is still fed to the diffusion model; the time-varying $\\omega_t=1-\\sqrt{\\bar{\\alpha}_t}$ weakens the corrupted guidance in late reverse steps. The other load-bearing piece is the low-frequency lock $\\Phi_l(f_{x'_{t-1}})=\\Phi_l(f_{x_{\\mathrm{ori}}})$, which freezes k-space below cutoff $\\pi/10$ from the corrupted input so tissue texture stays anchored. These are combined by $x'_{t-1}$ in the frequency domain and $x''_{t-1}$ in the pixel domain, then blended by $\\gamma_t$ into $\\tilde{x}_{t-1}$.","core_discovery":"On its own terms, the discovery is that the diffusion reverse process can be reorganized in two coupled domains to remove motion artifacts without training on corrupted images. At each reverse step the method first lets the pretrained diffusion model predict the previous timestep, then overwrites low-frequency k-space of that prediction with the low-frequency k-space of the corrupted input, and overwrites part of the high-frequency k-space with high-frequency content of the corrupted input through a mask that flips every step. In the pixel domain, it mixes the diffusion prediction with a forward-noised version of the corrupted image under the complementary mask. A step-dependent weight $\\omega_t = 1-\\sqrt{\\bar{\\alpha}_t}$ gradually reduces how much corrupted guidance enters as the image becomes cleaner, and a scalar $\\gamma_t$ shifts emphasis from frequency-domain texture anchoring to pixel-domain sharpness. The paper claims this preserves tissue textures and destroys artifact structure, yielding top quantitative results on simulated brain, knee, and abdominal data and the best radiologist ratings on real abdominal images.","pith_inferences":["A natural next step the paper does not take is replacing the fixed $\\pi/10$ cutoff with a learned or estimated corruption map, since real motion can in principle corrupt low-frequency k-space and would then be locked into the output by the low-frequency guidance.","The method still needs a clean-image pretraining set for each anatomy and contrast; a hospital with only corrupted archives would have to borrow clean data or pretrain on a compatible public set before PFAD could run.","Because the mechanism attacks structured phase errors rather than image content, it may carry over to other k-space phase artifacts such as respiratory ghosting or EPI Nyquist ghosts, with the same alternate-mask design.","The grid size of the checkerboard mask is a free parameter with visible quality effects, so clinical deployment would likely tune it per anatomy rather than using a single size."],"forward_implications":["Hospitals can clean new motion-corrupted scans with no paired rescans: the same pretrained diffusion model, trained once on clean images of that anatomy, purifies each incoming corrupted image at inference.","Because the corrupted image's own low-frequency k-space is preserved, the output stays anchored to the actual scan, reducing the risk that the generative model invents plausible but wrong anatomy.","Alternating complementary masks distribute artifact removal over the whole image across reverse steps, so no fixed region is left either fully corrupted or fully hallucinated.","The time-varying mask weight and domain-balance parameter give a principled schedule for leaning on the corrupted guidance early and the generated content late, which the ablation studies show is needed.","Radiologist evaluation on real clinical images indicates the benefit transfers beyond the simulated benchmark, which matters because simulated artifacts share the method's own high-frequency assumption."],"supporting_citations":[{"why":"Establishes the high-frequency (above $\\pi/10$) assumption for motion artifacts and provides an unpaired score-based diffusion baseline.","marker":"(Oh et al. 2023)"},{"why":"Supplies the masked-guidance scheme (RePaint) that PFAD turns into alternating complementary masks.","marker":"(Lugmayr et al. 2022)"},{"why":"Defines the DDPM forward/reverse process and training objective that PFAD runs at inference.","marker":"(Ho, Jain, and Abbeel 2020)"},{"why":"Provides the guided-diffusion architecture and training recipe for the pretrained clean-image model.","marker":"(Dhariwal and Nichol 2021)"},{"why":"Gives the k-space phase-perturbation motion model used to simulate corrupted images.","marker":"(Tamada et al. 2020)"},{"why":"Supplies the unpaired motion-artifact simulation and evaluation precedent used for the benchmark.","marker":"(Oh, Lee, and Ye 2021)"},{"why":"UDDN, the frequency-domain unsupervised baseline that PFAD compares against and claims to surpass.","marker":"(Wu et al. 2023)"},{"why":"Pix2pix, the supervised paired baseline PFAD claims to beat on simulated and real images.","marker":"(Isola et al. 2017)"},{"why":"CycleGAN, the unpaired GAN baseline used in the comparisons.","marker":"(Zhu et al. 2017)"}],"fun_headline_variants":["Erase MRI motion artifacts without paired training data","Frequency-domain guide cleans MRI scans unsupervised","No clean pairs needed: diffusion purifies MRI motion artifacts","Alternating masks let diffusion repair MRI artifacts from noisy images"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes motion artifacts live almost entirely in the high-frequency part of k-space above a cutoff of $\\pi/10$ along the phase-encoding direction, and the simulated benchmark is generated with phase perturbations placed only above that same cutoff, so the main quantitative evidence inherits the assumption rather than testing it.","fun_headline_variants_meta":{"raw":{"variants":["Erase MRI motion artifacts without paired training data","Frequency-domain guide cleans MRI scans unsupervised","No clean pairs needed: diffusion purifies MRI motion artifacts","Alternating masks let diffusion repair MRI artifacts from noisy images"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000716,"raw_usage":{"total_tokens":3216,"prompt_tokens":942,"completion_tokens":2274,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":2213}},"tokens_in":558,"tokens_out":2274,"duration_ms":15765,"temperature":1.0,"reasoning_tokens":2213,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:40:51.573723+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate rigid-motion phase corruption applied only to low-frequency k-space lines with $|k_y| < \\pi/10$ (below the cutoff the method treats as clean), with the same amplitude used for high-frequency corruption. If PFAD's output then retains visible ghosting or its PSNR and SSIM drop far below the high-frequency-corruption case, the low-frequency-lock premise is falsified. Alternatively, on retrospectively motion-tracked real data, measure the k-space phase-error spectrum: energy below $\\pi/10$ would directly violate the assumption.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the k-space phase-perturbation motion model used to simulate corrupted images."}],"review_version":1}