{"id":"7a3a8d03-c007-4fa1-8e53-7868f196cd10","arxiv_id":"2506.14322","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"FRIDU refines functional maps by treating them as images and applying a conditional diffusion model with point-to-point and geometric guidance at inference.","lead":"FRIDU trains an image diffusion model to clean up functional maps, the matrix representation of shape correspondences, by treating the matrices as images. The key extra idea is using test-time guidance to push the output toward valid point-to-point maps, and the method is competitive with established refinement algorithms.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-dataset evidence for the central claim is compromised: SHREC19 rows in Table 1 trail DiffZO by roughly 2x, and the reported numbers used guidance strength s=2000 selected on that test set, contrary to the paper's default s=500.","rationale":"The reader identified sensitivity to the initial map as the weakest assumption; I partially agree, but locate the sharper problem in the empirical support. On SHREC19, the only cross-dataset test, FRIDU trails DiffZO by roughly a factor of two, and the reported result was obtained after increasing guidance strength s from 500 to 2000 specifically for cross-dataset evaluation, i.e., hyperparameters were selected on the test set. This makes the headline 'competitive with state-of-the-art' claim non-informative as generalization evidence. It does not overturn the in-distribution refinement result (e.g., FAUST from 2.5 to 1.5/1.7), which is credible and consistent with the qualitative no-guidance improvement shown in Figure 3. But the method's value as a generalizable refiner is exactly what the cross-dataset rows should demonstrate, so the verdict remains CONDITIONAL: code, hyperparameters fixed before seeing the test set, and variance reports are needed before the stronger claim can be accepted.","tokens_in":22124,"tokens_out":10403,"duration_ms":116451,"concrete_test":"Release code and trained models, then rerun the SHREC19 rows of Table 1 with guidance strength s=500 (default) and s=2000, under the same DiffZO initialization, and report per-pair mean and standard deviation. If the s=500 SHREC19 error is substantially above the reported 9.4/7.1, or if the s=2000 improvement disappears when the setting is fixed before seeing the test set, the cross-dataset competitiveness claim is not supported. As a second check, degrade the SHREC19 initial maps by adding controlled noise and plot error versus initial-map quality to directly test the Section 4 sensitivity limitation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that a diffusion prior trained on ground-truth functional maps can correct noisy initial maps, and remain competitive with existing refinement methods, is only tested fairly where the initial maps come from the training distribution. In the one held-out benchmark (SHREC19), Table 1 shows FRIDU at 9.4 / 7.1 versus DiffZO at 4.2 / 3.6 for Train F / Train F+S. Section 3.3 states that for cross-dataset evaluations 'we find that increasing s to 2000 improves performance', meaning the guidance strength used in those rows was selected after looking at the SHREC19 test set; the default and ablated value is s=500. Section 4 admits the method is 'somewhat sensitive to the initial map' and that for a very bad initialization it 'does not fully reach the optimal solution.' Thus the load-bearing generalization claim is supported only in the favorable regime of in-distribution initial maps, and is not established for the cross-dataset regime where a learned refinement prior would add value; the reported SOTA-competitive numbers on SHREC19 may be a tuning artifact rather than evidence of generalization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes FRIDU, a method that treats functional map matrices as 2D images and trains a conditional image diffusion model to denoise/refine an initial, inaccurate functional map. At inference, a point-to-point (P2P) guidance term, optionally combined with orthogonality or Laplacian-commutativity regularizers, steers the diffusion process. Training is performed purely on functional maps (only the preprocessing uses ground-truth P2P maps), and a patch-based diffusion strategy is used for efficiency. Experiments cover descriptor-based initial maps (WKS, SHOT) on the Michael/TACO dataset, and feature-based initial maps from DiffZO on FAUST, SCAPE, and SHREC19. The paper reports improved accuracy over the initial maps in most settings and competitive in-distribution results with DiffZO, but substantially worse cross-dataset results on SHREC19.","tokens_in":22200,"tokens_out":5061,"duration_ms":48982,"significance":"Treating a functional map matrix as an image and applying a conditional diffusion model is a novel and nontrivial reformulation; the patch-based training on small datasets is a practical contribution, and the plug-and-play guidance framework is flexible. The in-distribution evidence is credible: on FAUST, FRIDU with orthogonality guidance reaches 1.5 mean geodesic error versus 1.9 for DiffZO, and the Michael/WKS and SHOT experiments show consistent refinement over initial maps. The paper is also honest about its limitations. However, the only held-out benchmark in the main comparison, SHREC19, shows FRIDU at 9.4/7.1 versus DiffZO at 4.2/3.6, and the reported cross-dataset numbers use a guidance strength selected on that test set. The cross-dataset generalization claim is thus not established, and the strength of the paper rests on the in-distribution refinement results.","major_comments":[{"comment":"The SHREC19 rows of Table 1 are obtained with guidance strength s=2000, not the default s=500 used elsewhere. The text states: 'For cross-dataset evaluations in Section 3.2, we find that increasing s to 2000 improves performance.' Because SHREC19 is the only held-out dataset in Table 1, selecting s after observing its test errors is a form of test-set tuning. Please report the SHREC19 results with the default s=500 (and with orthogonality guidance, if applicable), or provide a validation-based selection protocol that does not use the SHREC19 test set. This is load-bearing because the cross-dataset numbers are already the weakest part of the comparison.","section":"Section 3.3"},{"comment":"The paper's own limitation statement ('somewhat sensitive to the initial map... for a very bad initialization... does not fully reach the optimal solution') applies precisely to the SHREC19 setting, where the initial maps are very noisy (initial error 12.4 for Train F) and come from a different feature extractor than the training distribution. With those initial maps, FRIDU reduces the error to 9.4 whereas DiffZO, starting from the same initial condition, reaches 4.2. The central claim that the learned prior refines maps between 'arbitrary shape pairs' is therefore only supported in the in-distribution regime; the cross-dataset regime, which is the most valuable case for a learned refinement prior, is not established. I would like the authors to discuss this distribution-shift limitation explicitly and temper the 'comparable' wording in Section 3.2 accordingly.","section":"Section 4 / Table 1"}],"minor_comments":[{"comment":"The appendix references 'Section 4.2' and 'Section 4.1' (e.g., 'except in Section 4.2, where we match the resolution used in DiffZO' and 'for the experiments in Section 4.1'), but the paper's sections are numbered 1-5 and the experiments are in Section 3. Please correct these cross-references.","section":"Appendix A and B"},{"comment":"The baseline name is spelled inconsistently as both 'DiffZO' (Section 3.2, Table 1) and 'DiffZo' (Figure 9 caption, Appendix C). Please standardize.","section":"Throughout"},{"comment":"The sentence 'Our method consistently improves the initial mapping and outperforms DiffZO in intra-dataset evaluations, while maintaining comparable results in cross-dataset evaluations relative to supervised methods' is ambiguous. Since the primary comparison is DiffZO, the SHREC19 gap (9.4 vs 4.2) should be stated explicitly rather than qualified as 'comparable.'","section":"Section 3.2"},{"comment":"The stop-gradient operation for Π21 is described in prose but not shown in the equation. Please indicate the stop-gradient (e.g., using a ⊥ symbol or an explicit 'detach' notation) to make the guidance loss self-contained.","section":"Section 2.6, Eq. (7)"},{"comment":"There is a typo: 'the number of recurrent stepsk' should read 'the number of recurrent steps k.'","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The authors are well-known in the field, and the paper's reframing of functional map refinement as image diffusion is interesting. The main concern is methodological: the test-set selection of s for the SHREC19 rows undermines what is already the weakest evidence for the central claim. The appendix referencing nonexistent sections suggests the manuscript was revised quickly; the authors should be asked to fix these and to re-run the SHREC19 experiments with the default parameters before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my read on FRIDU.\n\nThe core idea is genuinely new: treat a functional map matrix as an image and train a conditional diffusion model to denoise it, with the initial map as condition and inference-time guidance from P2P consistency and classical regularizers. Prior diffusion work on functional maps [ZLG25] predicts to a fixed template; FRIDU instead refines arbitrary given maps. That is a real distinction, and training purely in spectral space is efficient. The patch-based training is a sensible adaptation to small datasets.\n\nThe paper does well on its home turf. On FAUST and SCAPE intra-dataset, FRIDU consistently lowers mean geodesic error relative to the initial maps, and with orthogonality guidance it edges out DiffZO (1.5 vs 1.9 on FAUST). The P2P guidance trick—extract the nearest-neighbor map, stop gradients, use it as a loss—is simple and works. The authors also state their limitations honestly: sensitivity to the initial map, square maps only, and the cross-dataset weakness.\n\nThe soft spots are in the SOTA comparison. On SHREC19, FRIDU is 9.4 vs DiffZO's 4.2 (train F) and 7.1 vs 3.6 (train F+S). That is a 2x gap, not \"comparable.\" The paper says they found \"increasing s to 2000 improves performance\" for cross-dataset evaluations, with the default being s=500. That is test-set tuning on the one benchmark that would demonstrate generalization, so the headline cross-dataset numbers are not clean evidence. This is the main issue. The sensitivity to the initial map also means the method is most useful for same-class batch refinement rather than as a universal refiner.\n\nThe math and citations are solid; no circularity, standard regularizers. The missing piece is code, plus variance over runs and a separated validation set for s.\n\nBottom line: this deserves a serious referee. The idea is novel and the in-distribution refinement is convincing, but the paper needs revision on the SOTA claim and on reproducibility. I'd send it to review, with the expectation that the authors either close the cross-dataset gap with properly tuned and honestly reported hyperparameters, or state clearly where the method does and doesn't compete.","headline":"Novel diffusion-on-functional-maps idea with solid in-distribution refinement, but the SOTA-competitive claim is undermined by test-set tuning and a clear SHREC19 miss.","tokens_in":22900,"tokens_out":3285,"would_cite":true,"duration_ms":31788,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FRIDU treats a functional map matrix as an image and trains a guided diffusion model to refine noisy shape correspondences, with the refined maps competitive with state-of-the-art methods on FAUST, SCAPE, and SHREC19.","keywords":["functional maps","shape correspondence","diffusion models","image diffusion","map refinement","point-to-point guidance","spectral geometry","geodesic error"],"falsifier":"Take a trained model and feed it initial maps whose quality is systematically degraded (e.g., by increasing descriptor noise, reducing the spectral basis, or using a distant shape class), then plot refined-map geodesic error against initial-map error. If there is a corruption level beyond which refinement stops improving the map or makes it worse, the distributional assumption is violated. A cleaner test: train only on WKS-based maps and apply the model to SHOT-based initial maps on a held-out shape class, and check whether refinement still consistently lowers error.","tokens_in":21747,"feed_emoji":"🗺️","tokens_out":6459,"duration_ms":57027,"temperature":0.7,"pith_summary":"The paper sets out to prove that functional map refinement can be solved by image diffusion: a functional map matrix, normalized to pixel values, is a valid 2D image, and a diffusion model trained to denoise such images conditioned on a noisy initial map produces accurate refined maps. Training is done entirely in spectral space, so no point-to-point maps are needed during learning. At inference, the pointwise map recovered from the current functional map is injected as guidance, optionally alongside orthogonality or Laplacian-commutativity objectives. The authors report that the method is competitive with dedicated refinement algorithms, improving FAUST mean geodesic error from 2.5 initially to 1.5 with orthogonality guidance. If correct, it makes diffusion models a flexible, plug-and-play tool for functional-map processing that adapts to new tasks through guidance alone.","feed_headline":"Guided diffusion cuts FAUST matching error from 2.5 to 1.5","feed_subtitle":"Functional maps become images; a diffusion prior repairs noisy correspondences at inference time.","key_machinery":"The central object is the functional map matrix itself, treated as an image patch. The machinery is a patch-based EDM-DDPM++ diffusion model that denoises corrupted functional maps, conditioned on the initial map and on extra position channels encoding where each patch lies in the full matrix. At inference, guidance from the point-to-point map—solved by nearest-neighbor search under the smoothness prior and held fixed during backpropagation—injects the geometric requirement that the functional map come from a valid dense map; orthogonality and Laplacian-commutativity losses add optional spectral regularizers.","core_discovery":"The central claim is that a functional map—a change-of-basis matrix between two shapes' spectral embeddings—can be refined by treating it as an image and running a conditioned image diffusion model. The model learns, from pairs of noisy initial maps and ground-truth maps, to denoise a corrupted functional map while conditioning on the initial map and on the spatial position of each patch. At test time, the point-to-point map corresponding to the current diffusion sample is computed by nearest-neighbor search and used as a guidance objective, with gradients stopped through that extraction; the paper argues this indirect signal steers generation toward maps consistent with valid dense correspondences. The result is a refinement pipeline that works for initial maps from different descriptors and learned features, generalizes across shape classes, and matches or beats existing refinement methods on FAUST, SCAPE, and SHREC19.","pith_inferences":["One consequence the authors leave implicit: the same image-diffusion treatment could extend to other spectral operators, such as shape difference operators or functional vector fields, since those are also matrices on the spectral basis; the paper lists this as future work.","If a diffusion prior were trained on a broad distribution of noisy correspondences across many shape classes, guidance alone might enable zero-shot refinement for new datasets and pipelines, avoiding per-dataset training—this is the paper's stated foundation-model vision.","The patch-based, position-encoded training suggests the model's effective receptive field is limited; a testable extension is whether training with larger patches or full maps on larger datasets removes the remaining sensitivity to very bad initializations."],"forward_implications":["The same trained model can refine functional maps from different descriptor sources (WKS, SHOT) and from learned feature extractors, including zero-shot across descriptor types.","Because training is purely in spectral space and uses patches, refinement can be learned from very small datasets, such as 190 correspondence pairs.","Inference-time guidance for orthogonality or Laplacian commutativity changes the output without retraining, letting one model serve multiple tasks.","A model trained on one shape class (Michael) improves maps on unseen human shapes (FAUST) and a non-human animal (wolf), indicating cross-category generalization.","One recursive refinement iteration improves accuracy further, while additional iterations degrade it."],"supporting_citations":[{"why":"Introduces the functional map representation that the paper refines.","marker":"[OBCS∗12]"},{"why":"ZoomOut, the classical refinement baseline whose pointwise-map insight motivates the P2P guidance.","marker":"[MRR∗19]"},{"why":"Provides the smoothness-based point-to-point extraction objective used in guidance.","marker":"[EBC17]"},{"why":"Supplies the EDM formulation and sampler the diffusion model is built on.","marker":"[KAAL22]"},{"why":"Patch-based diffusion training, adapted for data efficiency on small map datasets.","marker":"[WJZ∗23]"},{"why":"Universal guidance algorithm used at inference to inject geometric losses.","marker":"[BCS∗23]"},{"why":"DiffZO, the learned refinement baseline and source of the feature extractor producing initial maps.","marker":"[MO24]"},{"why":"Concurrent diffusion approach to functional maps, contrasted for its template-based rather than refinement-based setup.","marker":"[ZLG25]"},{"why":"Wave Kernel Signature descriptors used to compute the initial functional maps in training and evaluation.","marker":"[ASC11]"}],"fun_headline_variants":["Diffusion refines functional maps by treating them as images","Functional maps as images: guided diffusion improves shape correspondence","Image diffusion model repairs functional maps for shape matching","From matrices to pixels: diffusion-based functional map refinement","Guided diffusion in functional space boosts map accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The noisy initial maps seen at test time must lie close enough to the distribution of descriptor-based maps used during training that the learned diffusion prior can correct them; the paper states the method is 'somewhat sensitive to the initial map'.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion refines functional maps by treating them as images","Functional maps as images: guided diffusion improves shape correspondence","Image diffusion model repairs functional maps for shape matching","From matrices to pixels: diffusion-based functional map refinement","Guided diffusion in functional space boosts map accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000714,"raw_usage":{"total_tokens":3164,"prompt_tokens":853,"completion_tokens":2311,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":2236}},"tokens_in":469,"tokens_out":2311,"duration_ms":15111,"temperature":1.0,"reasoning_tokens":2236,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:52:34.719348+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a trained model and feed it initial maps whose quality is systematically degraded (e.g., by increasing descriptor noise, reducing the spectral basis, or using a distant shape class), then plot refined-map geodesic error against initial-map error. If there is a corruption level beyond which refinement stops improving the map or makes it worse, the distributional assumption is violated. A cleaner test: train only on WKS-based maps and apply the model to SHOT-based initial maps on a held-out shape class, and check whether refinement still consistently lowers error.","supporting_citations":[],"review_version":1}