{"id":"a0221ab3-9b08-45b7-9741-6d184cfba3ef","arxiv_id":"2501.13597","paper_version":2,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey organizes spectral clustering methods around graph structure learning, covering fixed and adaptive pairwise, anchor, and hypergraph graphs in single- and multi-view settings.","lead":"This paper is a survey that reviews spectral clustering methods through the lens of graph structure learning, organizing methods by graph type, adaptivity, and single- or multi-view setting. A generalist would read it to get a structured map of a mature clustering field and to see where open problems remain.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first/most extensive' survey claim is unsupported and already internally contradicted: Table VIII classifies DSC as fixed-Gram while §IV.B.2 describes it as adaptive-bipartite, with no coverage protocol to resolve which is right.","rationale":"The reader identified the completeness assumption as the weakest point: the 'most extensive and detailed' claim depends on representative selection and correct classification, but no search protocol or baseline is supplied. I agree, and I found concrete internal evidence that strengthens the concern: Table VIII's entry for DSC directly conflicts with the prose in Section IV.B.2, and Table VI contains uncited rows with no corresponding discussion. These are not just formatting issues; they make the taxonomy impossible to verify from the manuscript itself. The survey otherwise has a coherent structure and covers many relevant methods, so I do not regard the paper as worthless or the central idea as invalid. The appropriate verdict remains UNVERDICTED, because a survey's value cannot be adjudicated through the normal accept/reject scientific-claim tests, and the observed inconsistencies reinforce rather than overturn that judgment. A revision could address the concern by adding a search protocol, a coverage comparison against [34], and a table-by-table source audit.","tokens_in":33501,"tokens_out":4947,"duration_ms":46491,"concrete_test":"Perform a source-level audit: for every row of Tables II–VIII, compare the stated graph-construction category and partitioning process against the cited paper's own abstract or stated model, and compute the mismatch rate. In parallel, build a coverage baseline by merging the reference list of [34] with a systematic keyword search (e.g., 'spectral clustering' plus 'graph structure learning', 'anchor graph', or 'hypergraph', 2015–2024) and listing GSL-spectral methods absent from the survey. If DSC is in fact adaptive/bipartite in [66], or if any known GSL-spectral method is missing, the 'first/most extensive' claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section I is: 'For the first time, we present the most extensive and detailed survey on spectral clustering, with a particular emphasis on GSL.' For that claim to hold, the selected methods must be representative and correctly classified. The paper provides no search protocol, inclusion/exclusion criteria, time window, or coverage comparison against the acknowledged prior survey [34], so the completeness assertion is untestable. This is not merely an omissions issue: internal evidence indicates the taxonomy is unreliable. Table VIII row 2 classifies DSC [66] as 'K-means / Fixed / Gram' for anchor selection, graph construction, and similarity matrix, whereas Section IV.B.2 states that DSC 'adopts k-means for anchor selection and constructs adaptive anchor graphs' and 'generates similarity matrices via bipartite graphs.' At least one of the table or the prose is wrong, and no source-level audit is included to let the reader verify. Table VI additionally contains two uncited rows, MVGL (2018) and OMSC (2019), with no corresponding method descriptions in the text. If the central enumerations contain such contradictions, the claim of being the most extensive and detailed survey cannot be supported, and the taxonomy's value as a reliable reference is materially reduced. This is a correctness risk for the paper's primary contribution, not a stylistic nit.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey organizes spectral clustering methods around graph structure learning (GSL), classifying graph construction into pairwise, anchor, and hypergraph approaches in both fixed and adaptive forms, and grouping algorithms into single-view versus multi-view and one-step versus two-step frameworks. It provides background on graph cuts, Laplacians, spectral embedding, and partitioning, and it offers comparative tables of methods. The stated contribution is a first-of-its-kind, most extensive survey with GSL at its center.","tokens_in":33704,"tokens_out":5698,"duration_ms":45062,"significance":"If the taxonomy were reliable, the survey would be a useful reference for researchers entering spectral clustering with an emphasis on GSL. The mathematical background is mostly standard and reproduced without major distortion, and the organizational tables condense a large literature. However, the centrality of GSL is asserted rather than empirically demonstrated, and the paper's 'most extensive' claim is not backed by a reproducible selection protocol. Several internal inconsistencies in the tables and prose currently undermine the utility of the taxonomy.","major_comments":[{"comment":"The classification of DSC is internally contradictory. Table VIII, row 2, lists DSC (2019) with anchor selection 'K-means', anchor graph 'Fixed', and similarity matrix 'Gram'; Section IV.B.2 states that DSC 'adopts k-means for anchor selection and constructs adaptive anchor graphs' and 'generates similarity matrices via bipartite graphs.' At least one of these is wrong. Because the survey's central contribution is a classification of GSL methods, this contradiction materially reduces confidence in the taxonomy. The authors should correct the entry and add a source-level audit, or at least an explicit paper-by-paper justification, for each table row.","section":"Section IV.B.2 and Table VIII"},{"comment":"Rows 6 (MVGL, 2018) and 7 (OMSC, 2019) have no citation keys and no corresponding descriptions in Section IV.B.1. Without references or text descriptions, the reader cannot verify the classification, and the completeness claim 'most extensive and detailed survey' is untestable for exactly the kind of entries that should be traceable. Add citations and short descriptions, or remove the rows.","section":"Table VI"},{"comment":"The claim 'For the first time, we present the most extensive and detailed survey on spectral clustering, with a particular emphasis on GSL' is not supported by a methodology. The paper reports no search protocol, inclusion/exclusion criteria, time window, or coverage comparison against the acknowledged prior survey [34]. A survey's value depends on reproducible coverage; either add a methodology subsection or soften the claim to a scope statement.","section":"Section I"}],"minor_comments":[{"comment":"The constraint in Eq. (13) is written as \\sum_{j=1}^n w_{ij}=1 while w_{ij} is defined over the m anchors; this should be \\sum_{j=1}^m w_{ij}=1.","section":"Section III.A.2.b, Eq. (13)"},{"comment":"The heading 'Adaptive Neighbor Methods:' is repeated twice, and the text says 'PTAG [73]' while Table II row 12 lists 'CTAG (2023) [73]'. Unify the name and citation.","section":"Section IV.A.1 and Table II"},{"comment":"There is a typo 'Thye previous survey' that should read 'The previous survey'.","section":"Section I"},{"comment":"In the paragraph on IMVSC, 'alternating optimizationoptimization' contains a duplicated word.","section":"Section IV.B.1"},{"comment":"Reference [66] is used for two different methods: in Section III.C.2.b it is cited as 'Discrete Spectral Clustering (DSC)' for Eq. (29), while in Table VIII it is used for the multi-view anchor-graph method 'DSC (2019)'. The reference list identifies [66] as Luo et al., 'Discrete multi-graph clustering' (TIP 2019), a multi-view method; one of these usages is likely a misattribution. Clarify which method Eq. (29) is taken from.","section":"Section III.C.2.b and Table VIII"}],"recommendation":"major_revision","confidential_remarks":"The survey has a useful organizational structure and covers a broad literature, but the internal inconsistencies (especially the DSC contradiction between Table VIII and Section IV.B.2) and the uncited rows in Table VI make it unsuitable for acceptance in its current form. The 'first/most extensive' claim should be toned down or supported by a reproducible coverage protocol. These issues are fixable within the paper's scope, hence major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a survey, not a research preprint, and that is fine. The GSL-centered organizing scheme — fixed versus adaptive graphs, pairwise versus anchor versus hypergraph, single-view versus multi-view, one-step versus two-step — is sensible and gives practitioners a genuine map of the field. The background math on graph cuts, Laplacians, and spectral embedding is standard and reproduced without serious distortion. The tables are a convenient index to a large body of work, especially the multi-view fusion categories, and the coverage of methods up to 2024 is broad. If you work in this area, you will probably want a copy on your desk.\n\nThat said, the central claim in Section I — 'For the first time, we present the most extensive and detailed survey' — is not supported. The paper gives no search protocol, no inclusion criteria, no time window, and no coverage comparison against the prior survey [34], so the completeness claim is untestable. Worse, there are concrete reliability problems that undercut the taxonomy itself. Table VIII classifies DSC as fixed-Gram, while Section IV.B.2 describes DSC as constructing adaptive anchor graphs with bipartite similarity matrices. At least one of those is wrong, and there is no source-level audit to let the reader check. Table VI has two uncited rows, MVGL and OMSC, with no corresponding descriptions in the text. Equation (13) sums the anchor-weight constraint over n instead of m anchors. There are also many typos — 'Thye', 'similairy', 'optimizationoptimization' — which are minor but suggest the manuscript was not carefully proofread.\n\nI want to be clear about proportions. These are fixable issues, not a broken core. The taxonomy is coherent, the standard math is correct, and the survey does fill a gap by putting graph structure learning at the center. The 'first and most extensive' language is the main offender; it should be softened to something like 'a survey focused on GSL' or the authors should add a methodology section and a coverage comparison against [34]. The DSC contradiction and the uncited rows need to be resolved before this can be trusted as a reference.\n\nWould I bring it to a reading group? Probably not as a paper to discuss, but it could be useful background for someone entering the area. I would not cite it in my own work until the internal inconsistencies are fixed. But it deserves a serious referee: a careful reviewer who knows the spectral clustering literature could help the authors fix the errors and turn this into a genuinely useful reference. I would send it to review rather than desk-reject it.","headline":"A useful but carelessly assembled survey of spectral clustering with graph structure learning; the 'first and most extensive' claim is unsupported and the taxonomy has internal contradictions.","tokens_in":34255,"tokens_out":2588,"would_cite":false,"duration_ms":23871,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T10","05C50","62H30"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that graph structure learning is the central and decisive stage of spectral clustering, and offers a taxonomy organizing the field around graph families, learning modes, and partitioning frameworks.","keywords":["spectral clustering","graph structure learning","adaptive neighbors","anchor graph","hypergraph","multi-view clustering","one-step clustering","graph fusion"],"falsifier":"A systematic literature search for spectral clustering papers with graph structure learning published before 2025, using a defined query and inclusion criteria, would either reproduce the survey's coverage or find omissions; likewise, a benchmark comparing clustering accuracy across methods while holding data fixed would show whether graph construction indeed dominates other design choices. Any of these can be checked independently.","tokens_in":33285,"feed_emoji":"🕸️","tokens_out":3478,"duration_ms":30589,"temperature":0.7,"pith_summary":"This survey argues that graph structure learning (GSL) is the pivotal stage in spectral clustering, and that how the similarity graph is built matters more than any later step. It claims to be the first comprehensive review devoted specifically to GSL within spectral clustering, organizing roughly sixty methods into a taxonomy with two axes: graph type (pairwise, anchor, hypergraph) and learning mode (fixed vs. adaptive). It further splits clustering pipelines by partition strategy (one-step vs. two-step) and by view count (single vs. multi-view), with multi-view fusion decomposed into consensus, complementary, and weighting mechanisms. The intended payoff is a map that lets researchers position new methods, choose graph construction strategies, and identify open problems.","feed_headline":"Survey: how the graph is learned drives spectral clustering","feed_subtitle":"New taxonomy orders pairwise, anchor, and hypergraph methods in fixed and adaptive settings.","key_machinery":"The load-bearing object is the GSL taxonomy: pairwise, anchor, and hypergraph graph architectures, each split into fixed and adaptive construction, crossed with one-step versus two-step partitioning and single- versus multi-view fusion. Within it, the adaptive neighbor model, the anchor-graph formulation, the hypergraph Laplacian, and the self-expressive subspace representation are the canonical exemplars that anchor each cell of the taxonomy. The survey's classification rules — for instance, whether the graph is predetermined or optimized inside a clustering objective — do the argumentative work of making GSL the organizing principle.","core_discovery":"The paper's central claim is that the entire spectral clustering literature can be organized around how the graph is constructed and learned. It partitions graph construction into three families — pairwise graphs, anchor graphs, and hypergraphs — each with fixed and adaptive variants, and shows that adaptive methods such as adaptive neighbors, self-expressive subspace recovery, and adaptive anchor or hypergraph learning have progressively replaced fixed protocols. A second axis separates one-step clustering, which jointly learns the spectral embedding and discrete cluster assignments, from the classical two-step relax-and-discretize pipeline. For multi-view data the survey identifies fusion as the fourth dimension: shared or consensus structure, complementary view-specific structure, and view weighting. If the taxonomy is right, it supplies the first unified terminology for a field that previously grew method-by-method.","pith_inferences":["If the taxonomy becomes standard, the \"most extensive\" claim will eventually be superseded as new methods appear; the durable contribution would be the naming and ordering of the design space itself.","The same fixed-versus-adaptive axis could be applied to deep clustering and graph neural networks, where graph construction is often treated as a fixed preprocessing step rather than part of the learned objective.","A quantitative test measuring how much of the variance in clustering accuracy across surveyed methods is explained by graph construction versus partitioning would either support or weaken the paper's premise that GSL is the dominant factor."],"forward_implications":["New spectral clustering papers can be positioned in the taxonomy by specifying graph family, fixed or adaptive mode, partition strategy, and fusion method, making method comparison more systematic.","The survey's emphasis on adaptive graph learning implies that future gains in clustering accuracy will come substantially from better graph construction rather than from better partitioning alone.","For multi-view data, the consensus–complementary–weighting decomposition gives a checklist for designing fusion strategies and for deciding which view information to preserve.","The one-step versus two-step distinction clarifies a design tradeoff: joint optimization avoids information loss but restricts the objective, while two-step pipelines allow modularity at the cost of discretization errors."],"supporting_citations":[{"why":"Supplies the foundational definitions of graph Laplacians, graph cuts, and KNN-based fixed graph construction that the taxonomy builds on.","marker":"[19]"},{"why":"The prior spectral clustering survey whose graph-cut and Laplacian focus this paper claims to extend by foregrounding graph structure learning.","marker":"[34]"},{"why":"Introduces the adaptive neighbors objective that anchors the adaptive pairwise graph category.","marker":"[42]"},{"why":"Provides the anchor graph construction and discrete clustering model that anchors the anchor graph family.","marker":"[48]"},{"why":"Defines the normalized hypergraph Laplacian that underlies the entire hypergraph category.","marker":"[52]"},{"why":"Establishes self-expressive subspace representation, the core mechanism behind the self-expressive pairwise and hypergraph methods.","marker":"[45]"},{"why":"Provides the auto-weighted multi-view fusion framework that the multi-view consensus and weighting analysis builds upon.","marker":"[116]"},{"why":"Exemplifies one-step discrete multi-view clustering, supporting the one-step versus two-step partitioning axis.","marker":"[66]"}],"fun_headline_variants":["Spectral clustering reorganized around graph learning","Graph learning is the new lens for spectral clustering","Survey maps spectral clustering by graph construction","How you build the graph defines spectral clustering","The key to spectral clustering: graph structure learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey assumes that the papers it selected are representative of the field and that its chosen categories—graph type, fixed versus adaptive, one-step versus two-step, single- versus multi-view—are the right organizing axes; if important methods are missing or misclassified, the claim of being the most extensive and accurate survey fails.","fun_headline_variants_meta":{"raw":{"variants":["Spectral clustering reorganized around graph learning","Graph learning is the new lens for spectral clustering","Survey maps spectral clustering by graph construction","How you build the graph defines spectral clustering","The key to spectral clustering: graph structure learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000751,"raw_usage":{"total_tokens":3310,"prompt_tokens":882,"completion_tokens":2428,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":2361}},"tokens_in":498,"tokens_out":2428,"duration_ms":13816,"temperature":1.0,"reasoning_tokens":2361,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:47:05.263093+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search for spectral clustering papers with graph structure learning published before 2025, using a defined query and inclusion criteria, would either reproduce the survey's coverage or find omissions; likewise, a benchmark comparing clustering accuracy across methods while holding data fixed would show whether graph construction indeed dominates other design choices. Any of these can be checked independently.","supporting_citations":[{"cited_title":"Auto-weighted multi-view learning for image clustering and semi-supervised classification,","cited_arxiv_id":null,"evidence_quote":"Provides the auto-weighted multi-view fusion framework that the multi-view consensus and weighting analysis builds upon."}],"review_version":1}