{"id":"afe39df8-220f-431e-8282-9b8820367507","arxiv_id":"2505.05137","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"The paper claims a diffusion-model anomaly detection framework with wavelet and attention modules outperforms prior methods on images and time series, but supplies no code, data, or audio results to support the claim.","lead":"This paper proposes a diffusion-model anomaly detector that combines wavelet decomposition, attention, and reconstruction error scoring, and reports high AUC on image and time-series benchmarks. The experimental support is thin: no code, no data, no error bars, and the abstract promises audio experiments that never appear.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed audio and SOTA superiority is unsupported: no UrbanSound8K experiments and no comparison with the diffusion-based SOTA methods cited in the paper.","rationale":"The reader's REJECT verdict is correct and should stand. My stress-test identifies a more concrete failure than the reconstruction-separation assumption: the paper's own evidence does not cover the audio modality named in the abstract, and the 'SOTA' baselines are too weak. A positive test would be to evaluate the method on UrbanSound8K and against a cited SOTA; absent such evidence, the central claim is not merely unproven but internally contradicted (abstract promises audio; results have none). I also note the ablation table's Image AUC (0.941) differs slightly from the average of Table I (0.943), and Table II lists five time-series subsets while Section IV-A enumerates six; these inconsistencies reinforce the need for code or additional detail, though they are secondary. The reader's weakest assumption (separation property) is related but not identical; I focus on the missing audio/SOTA evaluation, which is more directly decisive.","tokens_in":8410,"tokens_out":5860,"duration_ms":59553,"concrete_test":"Run the proposed method on UrbanSound8K using the CWT front-end described in Section III-A and evaluate AUC per anomaly class, comparing against a standard autoencoder and at least one diffusion-based SOTA method (e.g., DDAD) on the same evaluation protocol; if no audio result is produced or the method does not exceed these baselines, the abstract's multimodal claim must be retracted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract—that the method 'outperforms state-of-the-art anomaly detection techniques' on 'both image and audio data,' validated on MVTec AD and UrbanSound8K—is not supported by the experimental sections. Section IV-A describes only image (MVTec AD, six categories) and time-series (NAB, UCR) datasets; UrbanSound8K is never defined, and no audio result, audio-specific table, or audio metric appears anywhere in Section IV. Section V's stated limitations (manual lambda tuning, preprocessing dependence) do not acknowledge this missing modality. Additionally, Tables I and II compare only against weak baselines (VAE, AnoGAN, PatchSVDD, DDPM, Transformer-AD, Autoformer); the diffusion-based SOTA methods cited in Section II (e.g., DDAD [22], Masked Diffusion Posterior Sampling [21]) are absent, so the 'state-of-the-art' claim is not tested. Because the claim is empirical and modality-spanning, the absence of the audio evaluation is load-bearing: without it, the paper has not shown what it claims.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an unsupervised anomaly detection framework based on diffusion probabilistic models, combining reconstruction error and semantic discrepancy computed with a pretrained perceptual network. The architecture adds a wavelet pyramid module, multi-head attention, a modality-shared input adapter, and a hybrid time embedding, with a training loss that mixes noise prediction and perceptual feature preservation. The authors claim state-of-the-art performance on both image and audio data, validated on MVTec AD and UrbanSound8K, with additional time-series experiments on NAB and UCR. The experimental section, however, contains only image (MVTec AD, six categories) and time-series (NAB/UCR subsets) results; no audio experiment or UrbanSound8K result appears anywhere in the manuscript.","tokens_in":8595,"tokens_out":4647,"duration_ms":46265,"significance":"If the reported results were reproducible and the audio evaluation existed, the framework would be a plausible contribution to diffusion-based anomaly detection, and the ablation study is sensibly designed to attribute gains to the proposed modules. The paper also gives explicit equations for the forward/reverse diffusion processes and the anomaly score, which is helpful. However, the central multimodal claim is not supported: the abstract promises image and audio validation, but the experiments cover only images and time series, and no state-of-the-art diffusion baselines are compared. The manuscript ships no code or data, and key hyperparameters are unreported. As a result, the claimed superiority over state-of-the-art is neither demonstrated nor independently checkable, and the contribution cannot currently be assessed.","major_comments":[{"comment":"The abstract and introduction state that the method is validated on both image and audio data (UrbanSound8K) and outperforms state-of-the-art anomaly detection techniques, but Section IV-A describes only MVTec AD (images) and NAB/UCR (time series) datasets; UrbanSound8K is never defined, and no audio experiment, table, or metric appears anywhere in Section IV. This omission is load-bearing because the paper's central multimodal claim rests on it, and Section V's stated limitations (manual lambda tuning, preprocessing dependence) do not acknowledge this missing modality.","section":"Abstract and Section IV-A"},{"comment":"The 'state-of-the-art' claim is not tested: the comparison baselines are VAE, AnoGAN, PatchSVDD, and DDPM, while the diffusion-based methods cited as state of the art in Section II (e.g., DDAD [22], Masked Diffusion Posterior Sampling [21]) are absent from all experiments. In addition, the six MVTec AD categories are selected without stated criteria, and Tables I and II report no error bars, standard deviations, or number of independent runs, so the claimed average improvement of 2.9% over DDPM is not established as statistically significant.","section":"Section IV-B and Table I"},{"comment":"The anomaly score A = λE_recon + (1−λ)E_feat and the training loss L = L_MSE + γL_feat depend on hyperparameters λ, γ, the number of diffusion steps T, the noise schedule β_t, the wavelet family/level, and the network architecture, but none of these values is reported. Section V acknowledges that λ requires manual tuning, yet no selected values or sensitivity analysis are given; without these details the experiments are not reproducible and the comparison with baselines is not interpretable.","section":"Section III-C and Section IV"},{"comment":"The feature-level error E_feat uses a 'frozen lightweight perceptual network (e.g., pretrained MobileNet)' f(·), but the paper does not specify how an image-oriented network such as MobileNet is applied to 1D time-series inputs or to the CWT spectrograms described in Section III-A. This is a reproducibility gap that directly affects the time-series results in Tables II and III and would also affect any intended audio evaluation.","section":"Section III-C"},{"comment":"The 'Full Model' row in Table III reports aggregate AUCs of 0.941 (image) and 0.919 (time-series), but Tables I and II give only per-category values and no averaging procedure is defined; the percentage drops in Table III are stated without confidence intervals or per-seed variation, so it is impossible to determine whether the ablation differences are within run-to-run variability.","section":"Section IV-D and Table III"}],"minor_comments":[{"comment":"The abstract claims validation on UrbanSound8K and audio data, but the experiments cover only images and time series; the abstract and introduction should be aligned with the actual experimental content.","section":"Abstract vs. paper body"},{"comment":"The displayed formula for the reverse step has a formatting problem: the fraction and the √(1−β_t) term are not properly typeset, making the equation hard to read.","section":"Section III-B, Eq. for xt−1"},{"comment":"The reported 'average improvement' percentages (2.9% and 1.7%) are ambiguous: it is unclear whether they are absolute percentage-point differences or relative improvements; please define the measure.","section":"Tables I and II"},{"comment":"The dataset name is inconsistent: Table II uses 'Pems-Bay' while the text in Section V uses 'PEMS-Bay'; please unify the spelling.","section":"Section V and Table II"},{"comment":"Reference [35] is a Medium.com post and is not a peer-reviewed source; consider replacing it with a stable archival citation.","section":"References"},{"comment":"The paper includes no data or code availability statement; given the manual tuning and missing hyperparameters, such a statement is essential for reproducibility.","section":"General"}],"recommendation":"reject","confidential_remarks":"The mismatch between the abstract (which promises audio experiments on UrbanSound8K) and the actual experimental section (which contains none) is a fundamental problem that should have been caught before submission. Even setting aside the missing audio results, the absence of hyperparameters, error bars, and comparisons to the cited state-of-the-art diffusion baselines means the central claims are currently unsupported. A revision would require substantial new experiments and a full reporting overhaul rather than a local fix, so I do not think a major-revision recommendation would be fair to the authors or the reviewing process."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this one before you skim it: the abstract says the method beats state-of-the-art on both image and audio data, but Section IV has no audio experiment at all, and the comparison tables only include weak baselines like VAE, AnoGAN, and plain DDPM. So the main advertised result is not actually in the paper.\n\nWhat is genuinely there: the architecture is a reasonable combination of known pieces — a wavelet pyramid, multi-head attention, a shared input adapter, hybrid time embeddings, and an anomaly score that blends reconstruction error with a MobileNet feature discrepancy. That specific assembly is not in the cited references, and the idea of using perceptual features alongside pixel error with a diffusion backbone is sensible. The ablation study, though coarse, does suggest each module contributes a little. The writing is clear and the author clearly knows the diffusion anomaly detection literature.\n\nThe soft spots are not minor. The missing audio evaluation is load-bearing because the abstract and method section explicitly claim audio support. The tables report AUC on six MVTec categories and five time-series subsets, but give no error bars, no code, no training details, and no hyperparameter values (lambda, gamma, number of diffusion steps, wavelet level). The \"state-of-the-art\" claim is not tested: the paper cites DDAD and Masked Diffusion Posterior Sampling in related work but does not compare against them. The limitations section honestly admits manual tuning of lambda and preprocessing dependence, but it does not acknowledge the missing audio modality, which the abstract clearly promises. The core assumption — that reconstruction and feature errors separate normal from anomalous inputs — is never independently validated per modality.\n\nFor a researcher working on diffusion-based anomaly detection, this reads like an early draft or workshop note, not a complete paper. It deserves a desk reject in the current form. A serious referee could not verify the central claims without the audio experiments, proper SOTA comparisons, and reproducibility details. The right next step is for the author to fill those gaps and resubmit.","headline":"The paper's central empirical claim is unsupported: the abstract promises audio and SOTA results, but the experiments contain no audio evaluation and the baselines are not state-of-the-art.","tokens_in":9101,"tokens_out":1671,"would_cite":false,"duration_ms":17906,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a normal-only diffusion model with wavelet and attention modules outperforms current anomaly detectors, raising average MVTec AD AUC by 2.9% over DDPM.","keywords":["anomaly detection","diffusion models","multi-scale feature extraction","attention mechanisms","wavelet transform","reconstruction error","semantic discrepancy","time-series anomaly detection"],"falsifier":"Train the model on normal samples from one MVTec AD category, compute the anomaly score $A$ for every test image, and plot its distribution for normal versus anomalous samples. If many normal images with ordinary texture variation score as high as true defects, or if fixing $\\lambda=0.5$ removes the reported 2.9% average gain over DDPM, the separation assumption the method rests on fails.","tokens_in":8180,"feed_emoji":"🔍","tokens_out":12926,"duration_ms":113366,"temperature":0.7,"pith_summary":"The paper tries to establish that a diffusion model trained only on normal samples can serve as an effective anomaly detector once the reconstruction score is combined with a semantic feature distance and the denoiser is augmented with wavelet and attention modules. It reports that this framework beats five baselines on six MVTec AD object categories, averaging 2.9% higher AUC than DDPM, and beats four baselines on time-series subsets from NAB and UCR by 1.7% on average. If those results hold, industrial visual inspection and streaming monitoring systems could use a single normal-only diffusion model to flag defects without adversarial training. The abstract also promises audio detection through wavelet spectrograms, but the paper reports no audio experiments.","feed_headline":"Normal-trained diffusion detector lifts anomaly AUC by 2.9%","feed_subtitle":"Wavelet pyramid and attention let one normal-only model flag image defects and time-series anomalies.","key_machinery":"The load-bearing object is the anomaly score $A = \\lambda E_{\\text{recon}} + (1-\\lambda)E_{\\text{feat}}$, computed after one reverse-diffusion reconstruction $\\tilde{x}_0$ of $x_0$; $E_{\\text{recon}}$ is squared pixel error and $E_{\\text{feat}}$ is squared distance between frozen MobileNet features. The reconstruction is produced by a U-Net noise predictor enhanced with a wavelet pyramid module (multilevel wavelet decomposition into approximation and detail subbands), multi-head self-attention between encoder and decoder stages, a modality-shared convolutional input adapter, and a hybrid sinusoidal-plus-convolution time embedding. The paper's argument is that these components make the normal-only model reconstruct normal inputs faithfully while leaving anomalous structure unrecovered, so both error terms rise for anomalies.","core_discovery":"The central claim is that anomalies reveal themselves through reconstruction failure: after training a diffusion model on normal samples only, the score $A = \\lambda\\|x_0 - \\tilde{x}_0\\|_2^2 + (1-\\lambda)\\|f(x_0)-f(\\tilde{x}_0)\\|_2^2$, with $f$ a frozen MobileNet feature extractor, separates normal from anomalous inputs better than pixel error alone. The paper argues that a wavelet pyramid in the U-Net encoder resolves fine edges and frequency shifts, multi-head self-attention captures long-range dependencies, and a shared input adapter lets one architecture handle images and CWT audio spectrograms under the same noise-prediction objective. In the reported MVTec AD experiments the full model reaches AUCs of 0.938–0.948 across six categories, an average 2.9% improvement over DDPM; on selected NAB and UCR time-series subsets it improves 1.7% over DDPM. Ablations assign the largest drops to removing the wavelet module (3.5% image, 5.1% time-series AUC) and to removing attention (2.8% image, 3.9% time-series).","pith_inferences":["Inference: the paper reports only image-level AUC, but the same reconstruction and feature-distance maps could be thresholded spatially; testing whether the wavelet pyramid improves per-pixel anomaly localization would be a direct extension of its claims.","Inference: the abstract promises audio results on UrbanSound8K, yet no audio experiment appears; if the CWT-to-image pipeline works as claimed, audio AUC would be the decisive test of modality generality.","Inference: because $\\lambda$ is manually tuned, the reported gains may partly reflect dataset-specific weighting; fixing or learning $\\lambda$ and rerunning the comparisons would reveal how much of the 2.9% improvement is architectural rather than score tuning.","Inference: applying the same weighted score to VAE and GAN reconstructions would isolate whether the diffusion backbone or the wavelet and attention modules drive the gain; the paper does not run that cross-model comparison."],"forward_implications":["On the six MVTec AD categories tested, the proposed score reaches AUCs of 0.938–0.948, so a defect detector for these industrial objects could operate without any anomalous training examples.","On the selected NAB and UCR time-series streams, the same framework reports a 1.7% average AUC gain over DDPM, suggesting the approach transfers from images to sensor, traffic, and ECG monitoring.","The ablation study implies the wavelet pyramid is the largest single contributor, with 3.5% image and 5.1% time-series AUC lost when removed, while multi-head attention matters most for time series, with a 3.9% drop.","Because training uses only the standard noise-prediction loss plus a perceptual feature loss, no adversarial training is required, which the paper argues avoids mode collapse and stabilizes normal-data modeling."],"supporting_citations":[{"why":"Supplies the DDPM training objective and reverse-sampling rule the framework builds on, and serves as the main baseline the proposed method reports beating.","marker":"[13]"},{"why":"Establishes the reconstruction-error criterion for diffusion-based anomaly detection that the paper extends by adding a semantic feature distance.","marker":"[18]"},{"why":"Provides a conditioned diffusion anomaly-detection model used as state-of-the-art context for image detection.","marker":"[22]"},{"why":"Introduces masked diffusion posterior sampling, the conditional-sampling approach cited for improving anomaly localization.","marker":"[21]"},{"why":"Provides the DDMT time-series diffusion baseline that motivates modeling long-range temporal information with attention.","marker":"[23]"},{"why":"Presents the DiffAD stepwise-weighting diffusion baseline for time-series anomaly detection.","marker":"[24]"},{"why":"Supplies the diffusion-time-estimation scoring alternative from the density-estimation branch of diffusion anomaly detection.","marker":"[26]"},{"why":"Supports the stability and expressiveness advantages of score-based diffusion models that the paper invokes for normal-data modeling.","marker":"[14]"}],"fun_headline_variants":["Diffusion models detect anomalies by reconstruction failure","Wavelet pyramid boosts diffusion anomaly detection by 2.9%","Attention and wavelet U-Net lift anomaly AUC to 0.948","One diffusion model handles image and audio anomalies","Reconstruction error plus semantics: diffusion anomaly detector"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that anomalous inputs are consistently harder for a normal-only diffusion model to reconstruct, and that this shows up in both raw pixel error and pretrained feature distance under the manually chosen weight $\\lambda$; the paper does not test this separation independently for images, time series, or audio.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion models detect anomalies by reconstruction failure","Wavelet pyramid boosts diffusion anomaly detection by 2.9%","Attention and wavelet U-Net lift anomaly AUC to 0.948","One diffusion model handles image and audio anomalies","Reconstruction error plus semantics: diffusion anomaly detector"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1471,"prompt_tokens":994,"completion_tokens":477,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":399}},"tokens_in":610,"tokens_out":477,"duration_ms":4702,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:11:57.295401+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the model on normal samples from one MVTec AD category, compute the anomaly score $A$ for every test image, and plot its distribution for normal versus anomalous samples. If many normal images with ordinary texture variation score as high as true defects, or if fixing $\\lambda=0.5$ removes the reported 2.9% average gain over DDPM, the separation assumption the method rests on fails.","supporting_citations":[{"cited_title":"Denoising Diffusion Pro babilistic Mod- els","cited_arxiv_id":null,"evidence_quote":"Supplies the DDPM training objective and reverse-sampling rule the framework builds on, and serves as the main baseline the proposed method reports beating."},{"cited_title":"DiffAD: Denoising diffusion-based anoma ly detection for time series","cited_arxiv_id":null,"evidence_quote":"Presents the DiffAD stepwise-weighting diffusion baseline for time-series anomaly detection."},{"cited_title":"Score-based Generative Modelin g through Stochastic Differential Equations","cited_arxiv_id":null,"evidence_quote":"Supports the stability and expressiveness advantages of score-based diffusion models that the paper invokes for normal-data modeling."}],"review_version":1}