{"id":"857a06b2-9219-45e8-87b1-f390d7747385","arxiv_id":"2505.21228","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Projecting pre-trained image features into hyperbolic space and classifying them with a learned hyperplane improves medical anomaly detection AUROC over Euclidean baselines on BMAD benchmarks.","lead":"This paper tests whether mapping medical image features into hyperbolic space, a curved space for tree-like relationships, improves detection of anomalies like tumors. The method beats Euclidean feature-based baselines on image-level accuracy across five medical datasets and works well when only a few healthy images are available.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing Euclidean twin: the paper never varies geometry alone, so the reported gains cannot be attributed to hyperbolic space.","rationale":"The reader correctly flags the synthetic-anomaly proxy as a threat to clinical transfer, and I agree with the CONDITIONAL verdict. My stress-test identifies a more basic internal-validity problem: even on the benchmarks actually reported, the paper does not establish that hyperbolic geometry causes the improvement. The only trainable components are the hyperbolic linear layer, curvature, and hyperplane; every comparison is against Euclidean methods that differ in feature handling, aggregation, and supervision. Thus the headline 'hyperbolic space consistently outperforms Euclidean-based frameworks' rests on a confounded comparison. A Euclidean twin using the same synthetic anomalies and classifier would settle whether geometry is the active ingredient. I also note the abstract overstates pixel-level performance, as Table 2 shows two datasets where a Euclidean baseline has higher PAUROC. These issues are addressable and do not require rejecting the paper, hence CONDITIONAL rather than REJECT or ACCEPT.","tokens_in":7424,"tokens_out":4040,"duration_ms":44497,"concrete_test":"Run the identical architecture with a Euclidean twin: same frozen WideResNet50 features, same patchification, same synthetic anomalies (CutPaste, Gaussian intensity, source deformation), same weighted centroid aggregation, and same learned hyperplane classifier, but replace the Lorentz exponential map with the identity and the Lorentz hyperplane distance (Eq. 4) with the Euclidean hyperplane distance, and remove the trainable curvature. Compare image-level and pixel-level AUROC on all five BMAD datasets over the same five seeds. If the Euclidean twin matches or exceeds the hyperbolic pipeline, the central claim is unsupported; if the hyperbolic pipeline consistently wins, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that hyperbolic space, not the surrounding machinery, drives the anomaly-detection gains. Yet the experimental design in Sections 2.3-2.4 and Table 2 never isolates geometry. The method trains a hyperplane classifier on synthetic anomalies (CutPaste, Gaussian intensity, source deformation, Section 2.1) and compares against Euclidean baselines such as PatchCore, RD4AD, STFPM, PaDiM, and CFA. Those baselines differ in many ways beyond geometry: memory banks, teacher-student distillation, hypersphere fitting, and Gaussian modeling, and none use the same synthetic anomalies or a comparable centroid-plus-hyperplane classifier. Consequently, the reported image-level AUROC improvements could come from the anomaly-synthesis procedure or from the learned hyperplane rather than from hyperbolic curvature. The abstract's pixel-level claim is also stronger than Table 2 supports: on BraTS2021, RD4AD achieves PAUROC 96.36 vs. Ours 95.56; on RESC, PatchCore achieves 95.87 vs. Ours 95.32. The paper would need a Euclidean counterpart of the same pipeline, with identical features, synthetic anomalies, aggregation, and classifier, to substantiate that hyperbolic geometry is what matters.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a hyperbolic anomaly detection and localization framework for medical images. The method generates synthetic anomalies (CutPaste, Gaussian intensity, source deformation), extracts patchified features from a frozen WideResNet50, projects them into the Lorentz model of hyperbolic space, aggregates layer-wise features using confidence weights based on Euclidean L2 norms, and trains a hyperbolic hyperplane classifier with binary cross-entropy. It is evaluated on five BMAD datasets (BraTS2021, BTCV+LiTs, RESC, OCT2017, RSNA) using image-level and pixel-level AUROC, with comparisons to five Euclidean baselines (RD4AD, STFPM, PaDiM, PatchCore, CFA). The paper also reports ablations on curvature, patch size, and dimensionality, plus few-shot experiments with 1 to 25 normal images.","tokens_in":7701,"tokens_out":9266,"duration_ms":86787,"significance":"If the headline claim were fully supported, this would be a practical empirical contribution: a simple, frozen-backbone method that beats established Euclidean anomaly-detection benchmarks on medical images would be useful, especially in low-data regimes. The paper has clear strengths: it uses the standardized BMAD benchmark, reports mean/min/max over five random seeds, includes ablations of key hyperparameters, and compares against five standard baselines. The image-level results are favorable and the experimental setup is described in enough detail to be reproduced. However, the core attribution of the gains to hyperbolic geometry is not established by the current design, and the pixel-level and few-shot claims are stronger than the reported evidence. The paper would be materially improved by adding a Euclidean control of the same pipeline and by toning down the abstract and conclusions to match the quantitative results.","major_comments":[{"comment":"The abstract states that hyperbolic space achieves higher AUROC scores at both image and pixel levels across multiple datasets, but Table 2 does not support the pixel-level part: on BraTS2021 the proposed method reaches PAUROC 95.56 versus RD4AD's 96.36, and on RESC it reaches PAUROC 95.32 versus PatchCore's 95.87. The text in Section 4 correctly describes the pixel-level results as competitive, but the abstract and Section 5 repeat the stronger claim. Please revise the abstract and conclusions to state that image-level IAUROC is the consistently superior metric and that pixel-level PAUROC is competitive but not always best.","section":"Abstract and Table 2"},{"comment":"The central claim that hyperbolic geometry causes the observed improvement is not tested. The proposed pipeline differs from every Euclidean baseline in several components at once: the synthetic anomaly generation (Section 2.1), the confidence-weighted aggregation (Eq. 2), and the hyperplane classifier (Eqs. 3-6). None of the five baselines shares these components, so the higher IAUROC could come from the anomaly synthesis or the classifier rather than from curvature. To justify the title's claim, the authors should add a Euclidean counterpart of the same pipeline with identical frozen features, synthetic anomalies, aggregation weights, and an analogous Euclidean hyperplane classifier, and report whether hyperbolic projection still improves over that control. Without this comparison, the statement that hyperbolic space is what matters remains an attribution hypothesis rather than a demonstrated result.","section":"Sections 2.3-2.4 and Table 2"},{"comment":"The few-shot claim is supported only by Figure 3, with no numeric table, confidence intervals, or significance tests. The text says the hyperbolic model 'significantly outperforms' PaDiM and PatchCore, but no Mann-Whitney U results are reported for these comparisons, and only two baselines are included. Please provide a table with mean/min/max IAUROC and PAUROC for each dataset and each sample count (1, 3, 5, 10, 25), and report p-values for the comparisons against both baselines.","section":"Section 4.2 and Figure 3"},{"comment":"The confidence weights w_{i,l} are a distinctive part of the method, but there is no ablation that removes them or replaces them with equal weights. As a consequence, the contribution of the hyperbolic projection is conflated with the contribution of the weighting scheme. Please add an ablation with equal weights and, ideally, with weights computed from Euclidean norms, to show which component drives the reported gains.","section":"Eq. (2) and Section 2.3"}],"minor_comments":[{"comment":"Please clarify the notation f_{i,l} and the aggregation dimensions after upsampling; the text says 'to give a feature map f_{i,l} ∈ R^C' but Eq. (2) sums over l in a way that suggests each level contributes a centroid.","section":"Section 2.2"},{"comment":"The caption should state explicitly that PAUROC is not reported for OCT2017 and RSNA, or the missing entries should be marked; the current table header is ambiguous about which columns are image-level and which are pixel-level for these datasets.","section":"Table 2"},{"comment":"Please report the learned curvature values for each dataset and add a log-scale label to the x-axis of the curvature plot, since the text claims the learnable curvature adaptively optimizes the geometry.","section":"Section 4.1"},{"comment":"The x-axis values 1, 3, 5, 10, 25 are not evenly spaced; use a consistent linear or logarithmic scale and ensure the legend is identical across subplots.","section":"Figure 3"},{"comment":"Minor typos and heading issues: 'Synthesis Anomalies' in Section 2.1 should be 'Anomaly Synthesis' or 'Synthesizing Anomalies'; 'the the exponential map' appears in Section 2.3; 'T able 1' appears before Table 1.","section":"Throughout"},{"comment":"The conclusions mention reconstruction-based methods [24, 25] and gradient-based methods [15] without introducing them earlier; add a sentence of context where these categories first appear.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable empirical study, but the title and abstract outrun the evidence. The missing Euclidean control of the same pipeline is the key technical risk; if the authors add it and the geometry still helps, this could be a solid contribution for a medical-imaging or anomaly-detection venue. I would not reject, but the revision needs substantive additional experiments and a more careful framing of the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The empirical headline holds up: a frozen WideResNet with a learned hyperbolic projection, confidence-weighted aggregation, and a hyperplane classifier beats five established Euclidean baselines on image-level AUROC across all five BMAD datasets, and the few-shot plots show a plausible advantage. The paper is a genuine contribution to medical anomaly detection. But the central attribution is not actually tested. There is no Euclidean twin of the same pipeline; every baseline differs in features, aggregation, classifier, and anomaly generation, so the reported gains could come from any of those components rather than hyperbolic curvature. The ablation varies curvature within hyperbolic space, which is not the same as comparing geometry. This is the paper's real weak spot, and it is addressable.\n\nWhat is new here is the specific combination—synthetic anomalies (CutPaste, Gaussian intensity, source deformation), multi-layer feature projection into the Lorentz model, norm-based confidence weighting, and a hyperbolic hyperplane—applied systematically to medical anomaly detection on BMAD. The close relation to the authors' prior hyperbolic outlier work [14] is fine; the benchmark comparison is new and reasonably rigorous.\n\nAlso, the abstract overstates localization: it says higher AUROC at both image and pixel levels, but Table 2 shows RD4AD ahead on BraTS PAUROC (96.36 vs 95.56) and PatchCore ahead on RESC (95.87 vs 95.32). The paper's own conclusion admits 'competitive' for pixel-level, so the abstract needs correction. Few-shot results are plot-only, with no tables or per-seed values; minor but annoying for reproducibility. No code is released, which makes the few-shot numbers harder to verify.\n\nWhat I like: the choice of BMAD, five seeds with min/max ranges, Mann-Whitney testing, and ablations of curvature, patch size, and dimensionality. The synthetic anomaly proxy is an assumption about clinical transfer, but that is standard for self-supervised anomaly detection and not a fatal issue.\n\nBottom line: this deserves a serious referee, but I would send it back asking for a Euclidean counterpart of the same pipeline, a corrected abstract, and tables for the few-shot results. The core contribution is solid enough that these are fixable rather than fatal.","headline":"Solid empirical win for a hyperbolic anomaly-detection pipeline on BMAD, but the title and abstract overclaim attribution to geometry; needs a Euclidean twin and an honest abstract before I'd trust the 'all you need' framing.","tokens_in":8171,"tokens_out":2647,"would_cite":true,"duration_ms":26890,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that hyperbolic space, not a larger model or dataset, is the active ingredient in medical anomaly detection: projecting frozen pre-trained features onto a Lorentz hyperboloid and separating them with a learned hyperplane…","keywords":["hyperbolic space","medical anomaly detection","anomaly localization","Lorentz model","few-shot learning","pre-trained features","synthetic anomalies","BMAD benchmark"],"falsifier":"The cleanest test is to keep every component identical but replace the Lorentz exponential map with the identity so features stay Euclidean; if this Euclidean twin matches or exceeds the reported image-level AUROC on the five BMAD datasets, the central claim that hyperbolic geometry is responsible for the gains is falsified.","tokens_in":7252,"feed_emoji":"🩺","tokens_out":13957,"duration_ms":136602,"temperature":0.7,"pith_summary":"This paper tries to establish that the geometry of the representation space is a practical lever for medical anomaly detection: projecting a frozen pre-trained network's features into hyperbolic space, rather than leaving them in Euclidean space, improves separation between healthy and anomalous images. The method synthesizes anomalies (CutPaste patches, Gaussian intensity changes, source deformations), extracts multi-layer features, maps them onto the Lorentz hyperboloid, aggregates them by confidence, and classifies with a learned hyperbolic hyperplane. On the five BMAD medical benchmarks, it reports the highest image-level AUROC on every dataset, with pixel-level AUROC competitive but not always highest. The paper further claims that the hyperbolic pipeline is robust to parameter variations and dominates Euclidean baselines in few-shot settings with one to twenty-five healthy images. If these results hold, anomaly detection in medicine can be improved without a new backbone, more data, or annotation effort.","feed_headline":"Hyperbolic space beats Euclidean rivals on all five benchmarks","feed_subtitle":"No new backbone or large training set: the gain comes from the geometry.","key_machinery":"The load-bearing object is the Lorentz hyperboloid $\\mathbb{L}_c^n$, the standard model of hyperbolic space with constant negative curvature $c$, used as the feature space for every image. The argument runs through three ingredients: the exponential map of eq. (1) carries Euclidean patch features onto the hyperboloid; the weighted Lorentzian centroid of eq. (2) aggregates features from different network layers into a single point, with weights given by the Euclidean norm of each feature after mapping to the Poincaré ball, interpreted as confidence; and the hyperbolic hyperplane of eqs. (3)–(5) classifies a point by the signed geodesic distance to a learned separating hyperplane, trained with binary cross-entropy. The curvature $c$ is itself trainable, so the geometry adapts to each dataset, and the ablations show that fixing the curvature hurts performance.","core_discovery":"The paper's central claim is that hyperbolic geometry itself, not a larger model or richer training set, is what buys the improvement. Starting from a frozen WideResNet-50, the authors generate synthetic anomalies and push patchified features from layers 2 and 3 onto the Lorentz hyperboloid via the exponential map; a hyperbolic linear layer adapts the ImageNet-pretrained features to the medical domain, and a weighted Lorentzian centroid pools the hierarchical levels into one point per image, with weights set by each feature's distance from the origin. A learned hyperplane in the Lorentz model then separates normal from anomalous samples. Across the five BMAD datasets (brain MRI, liver CT, retinal OCT, chest X-ray), the paper reports the best image-level AUROC for every dataset—for example, 92.49 versus 92.02 for PatchCore on BraTS2021—while pixel-level AUROC is competitive but not uniformly the best. The paper also claims strong few-shot behaviour, beating PaDiM and PatchCore when only a handful of healthy images are available, and graceful degradation when the hyperbolic embedding dimension is cut to as low as two.","pith_inferences":["A natural test the paper does not run is to train the same hyperplane on real, annotated lesion patches instead of synthetic anomalies; if the reported advantage shrinks, the synthetic proxy is carrying part of the result.","The same recipe—frozen features, confidence-weighted hyperbolic centroid, learned hyperplane—should transfer beyond medicine to industrial or satellite anomaly detection, since none of the components is modality-specific.","The table suggests the geometric gain is larger for whole-image decisions than for pixel-level localization, since different Euclidean baselines win PAUROC on specific datasets; testing aggregation weights or adding earlier-layer features could show where the localization ceiling comes from.","Because the backbone is frozen, the method could be composed with larger or medical-specific feature extractors as they appear, likely preserving the geometric advantage."],"forward_implications":["Medical anomaly detection can be improved by a geometric swap alone: no new backbone, no extra labels, and no data augmentation are needed.","Few-shot deployment becomes feasible: with as few as one to twenty-five healthy images, the hyperbolic detector is reported to outperform PaDiM and PatchCore, which matters for rare diseases and new imaging modalities.","The framework tolerates low-dimensional embeddings, so memory-constrained settings can use compact hyperbolic representations without losing much localization accuracy.","Because the curvature is learned per dataset, the method effectively tunes the shape of the representation space rather than assuming a fixed geometry."],"supporting_citations":[{"why":"It supplies the five BMAD benchmark datasets, their train/validation/test splits, and the evaluation protocol used in every comparison.","marker":"[2]"},{"why":"It provides the CutPaste operation used to create one of the synthetic anomaly types for training.","marker":"[23]"},{"why":"It is the source of the Gaussian intensity variation synthetic anomaly type.","marker":"[37]"},{"why":"It provides the source-deformation synthetic anomaly type and the multi-task synthetic anomaly approach.","marker":"[3]"},{"why":"It supplies the Lorentz model formulas, including the exponential map and the weighted centroid used for projection and aggregation.","marker":"[21]"},{"why":"It motivates the distance-from-origin confidence weighting that sets the aggregation weights in eq. (2).","marker":"[18]"},{"why":"It provides the hyperbolic linear layer that adapts ImageNet-pretrained features to the medical target domain.","marker":"[4]"},{"why":"It is the strongest Euclidean baseline on most datasets and the main comparison point in the few-shot experiments.","marker":"[28]"},{"why":"It is the Euclidean patch-distribution baseline used alongside PatchCore in the few-shot evaluation.","marker":"[8]"},{"why":"It defines the frozen WideResNet-50 feature extractor used for all feature extraction.","marker":"[36]"}],"fun_headline_variants":["Hyperbolic space tops all five medical anomaly benchmarks","Hyperbolic space outperforms Euclidean for anomaly detection","Simple hyperbolic projection beats complex Euclidean methods","Geometry, not data, gives medical anomaly detection a boost","Hyperbolic geometry wins anomaly detection without extra data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's claims rest on synthetic anomalies—CutPaste patches, Gaussian intensity changes, and source deformations—being a faithful proxy for the real lesions in the BMAD test sets, since the hyperbolic hyperplane is trained only on those synthetic examples.","fun_headline_variants_meta":{"raw":{"variants":["Hyperbolic space tops all five medical anomaly benchmarks","Hyperbolic space outperforms Euclidean for anomaly detection","Simple hyperbolic projection beats complex Euclidean methods","Geometry, not data, gives medical anomaly detection a boost","Hyperbolic geometry wins anomaly detection without extra data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.002051,"raw_usage":{"total_tokens":7968,"prompt_tokens":912,"completion_tokens":7056,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":6985}},"tokens_in":528,"tokens_out":7056,"duration_ms":55021,"temperature":1.0,"reasoning_tokens":6985,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:32:03.498415+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The cleanest test is to keep every component identical but replace the Lorentz exponential map with the identity so features stay Euclidean; if this Euclidean twin matches or exceeds the reported image-level AUROC on the five BMAD datasets, the central claim that hyperbolic geometry is responsible for the gains is falsified.","supporting_citations":[{"cited_title":"In: CVPRW (Apr 2024)","cited_arxiv_id":null,"evidence_quote":"It supplies the five BMAD benchmark datasets, their train/validation/test splits, and the evaluation protocol used in every comparison."},{"cited_title":"In: CVPR (2021)","cited_arxiv_id":null,"evidence_quote":"It provides the CutPaste operation used to create one of the synthetic anomaly types for training."},{"cited_title":"In: MICCAI (2024)","cited_arxiv_id":null,"evidence_quote":"It is the source of the Gaussian intensity variation synthetic anomaly type."},{"cited_title":"In: MICCAI (Jul 2023)","cited_arxiv_id":null,"evidence_quote":"It provides the source-deformation synthetic anomaly type and the multi-task synthetic anomaly approach."},{"cited_title":"In: ICML (2019)","cited_arxiv_id":null,"evidence_quote":"It supplies the Lorentz model formulas, including the exponential map and the weighted centroid used for projection and aggregation."},{"cited_title":"In: CVPR (Mar 2020)","cited_arxiv_id":null,"evidence_quote":"It motivates the distance-from-origin confidence weighting that sets the aggregation weights in eq. (2)."},{"cited_title":"In: ICLR (Oct 2023)","cited_arxiv_id":null,"evidence_quote":"It provides the hyperbolic linear layer that adapts ImageNet-pretrained features to the medical target domain."},{"cited_title":"In: CVPR (May 2022)","cited_arxiv_id":null,"evidence_quote":"It is the strongest Euclidean baseline on most datasets and the main comparison point in the few-shot experiments."},{"cited_title":"In: ICPRW (2021)","cited_arxiv_id":null,"evidence_quote":"It is the Euclidean patch-distribution baseline used alongside PatchCore in the few-shot evaluation."}],"review_version":1}