{"id":"6685f07a-229a-4b7c-8ba8-ddfa0040dd55","arxiv_id":"2608.05336","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A fixed-length, 3D-aware molecular fingerprint based on graph Laplacian eigenvalues separates conformers and stereoisomers and scales to millions of molecules.","lead":"This paper introduces a fixed-length 3D molecular fingerprint built from eigenvalues of graph Laplacians over four physics-inspired interaction channels. It gives a cheap, alignment-free way to tell apart conformers and stereoisomers that 2D fingerprints collapse, and it scales to millions of molecules.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Default M=16 eigenvalue truncation is dataset-specific; SI Figure S2 shows large-system similarity depends on M, so the generalizability claim rests on an untuned hyperparameter.","rationale":"The paper's core construction is mathematically sound: eigenvalues of a normalized graph Laplacian are permutation- and E(3)-invariant, and the four physics-inspired channels provide a plausible 3D encoding at low per-molecule cost. The demonstrations on conformers, stereoisomers, and small-molecule datasets are credible and independently supported by the open-source code and Zenodo archive. The most load-bearing weakness is the default truncation M=16, exactly the reader's weakest assumption. This is not an internal inconsistency; it is a correctness risk because the central generalization claim spans systems (proteins, MOFs, GEMS) where N is far larger than M, and the authors' own SI Figure S2 shows M-dependence for such systems. The proposed sweep over M would settle whether the reported benchmark rankings and pairwise similarity values are stable or artifacts of a single empirical hyperparameter choice. Since the reader already assigned CONDITIONAL on this basis, my independent read does not change the verdict; the appropriate action is to require the M-robustness check or an explicit per-domain M-selection protocol before accepting the broad generalizability claim.","tokens_in":32291,"tokens_out":6382,"duration_ms":58052,"concrete_test":"Recompute the QMOF band-gap clustering, GEMS electronic-energy clustering (Table 2), and the P450/WelO5 pairwise similarities (Table S6) for M ∈ {8, 16, 32, 64, 128} with all other settings fixed. If the U-test/DBI/CHI rankings relative to CD-RACs and Morgan change, or if protein similarity values shift by more than ~0.05, then the default-M results are not robust and the generalizability claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that spectral fingerprints are a generalizable, geometry-aware similarity measure across organic, inorganic, biological, and reticular chemistry depends on the default M=16 per-channel eigenvalue truncation (Eq. 8, Section 3a). For any molecule with N>16 atoms, the fingerprint retains only 16 eigenvalues per channel, discarding N−16 eigenvalues; for a 3,600-atom protein this is under 0.5% of the spectrum. The authors state M was chosen empirically and explicitly note in Supporting Information Figure S2 that similarity for large proteins depends on the number of eigenvalues, while Figure S1 shows the plateau used to select M=16 was established on small molecules. Thus the reported protein similarities (e.g., 0.976 for Gly-to-Ala, 0.436 for Gly-to-Phe) and the QMOF/GEMS clustering results are conditional on an M value that may be suboptimal for large systems. Because the abstract claims maintenance of chemical resolution and low cost for 'vast chemical spaces' including macromolecular and reticular chemistry, the lack of an M-robustness study is a load-bearing gap: a different M could change relative method rankings in Tables 1 and 2, or collapse the pairwise separations that motivate the geometry-sensitivity claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 'spectral fingerprints': for each molecule, four weighted complete graphs are constructed from 3D atomic coordinates using heuristic electrostatics, bonding, sterics, and dispersion interaction functions; the eigenvalues of each channel's graph Laplacian are truncated to a fixed length (M=16 per channel) and concatenated into a length-64 vector. Similarity is defined as an inverse Euclidean distance. The representation is argued to be permutation-invariant, E(3)-invariant, alignment-free, and cheap, and is demonstrated on conformers, stereoisomers, reactions, MOFs, and proteins. The authors benchmark Leiden-based clustering against scaffolds, Morgan, RACs, and CD-RACs on six datasets (QM9, tmQMg, QCell, Transition1x, QMOF, GEMS), and evaluate k-NN regression and applicability-domain estimation for QM9 and tmQMg. The central claim is that spectral fingerprints provide a generalizable, interpretable, geometry-aware similarity measure that overcomes 2D limitations at low computational cost.","tokens_in":32586,"tokens_out":4474,"duration_ms":41570,"significance":"If the claims hold, this is a practically useful contribution: a fixed-length, alignment-free 3D descriptor with an interpretable construction and no fitted target-property model. The paper ships an open-source implementation, a Zenodo repository, and machine-checked case studies; the representation is not trained on target properties, and the evaluation is largely external, which is a strength. The clearest value is in settings where 2D connectivity is insufficient (stereoisomers, conformers, coordination geometry) and where alignment-based 3D methods are too costly. However, the main benchmarking and the large-system demonstrations rest on empirically tuned hyperparameters and on comparisons across different parsable subsets and cluster counts, so the strength of the generalizability claim is currently not fully supported.","major_comments":[{"comment":"The default truncation M=16 is a tuned hyperparameter, not a derived guarantee. The text states that M 'was determined empirically' and that 'larger chemical systems beyond the scope of small molecules would benefit from dataset-specific tuning'; SI Figure S2 explicitly shows that protein similarities depend on the number of eigenvalues, while the plateau used to select M=16 (Figure S1) was established on small systems. Because the abstract and §3b claim generalizability to macromolecules and reticular chemistry, the reported protein similarities (0.976 and 0.436) and the QMOF/GEMS clustering results are conditional on an M value that may be suboptimal for large systems. Please provide an M-robustness analysis: report the P450/WelO5 similarities and the key Table 1–2 metrics for M = 8, 16, 32, 64, and show that method rankings are stable, or explicitly temper the generalizability claim to the small-molecule regime.","section":"§3a, Eq. (8); SI Figures S1–S2"},{"comment":"The benchmark comparisons are made across different parsable subsets and different numbers of clusters, which undermines the 'best overall performance' conclusion. For example, on QM9 the number of clusters is 76 for spectral, 16 for Morgan, and 1,132 for scaffolds; on QCell, scaffold parsing succeeds for only 0.4% of the dataset. Since the U-test percentage, DBI, and CHI depend on both the population and the number of clusters, the values in Tables 1–2 are not directly comparable across representations. The k-NN analysis already includes a parsability-controlled comparison (Tables S12–S14); please perform the analogous common-subset evaluation for the clustering benchmarks and report cluster counts alongside all metrics.","section":"Tables 1 and 2; §3c"},{"comment":"The primary extrinsic metric is the percentage of cluster pairs with statistically significant Mann–Whitney U tests at α=0.05, reported without any multiple-testing correction or confidence interval. With 76 clusters there are 2,850 pairwise tests, and at α=0.05 one expects roughly 5% false positives even under the null; the expected number of significant pairs grows quadratically with the number of clusters. This makes the reported percentages (e.g., 98.0% for QM9 spectral, 87.3% for tmQMg) hard to interpret across methods that produce very different numbers of clusters. Please report adjusted p-values (e.g., Benjamini–Hochberg), or effect sizes with confidence intervals, or at least state the number of pairwise tests and the expected false-positive rate for each row.","section":"§3c, Tables 1–2; Text S3"},{"comment":"The k-NN results are reported for k=5, but the text states that k=5 'is observed to maximize the mean rank correlation across all tasks and representations' after evaluating k ∈ {1, 5, 25}. Selecting k on the same evaluation data can inflate the reported Spearman correlations and may favor the method for which the selected k is most beneficial. The SI does contain results for k=1 and k=25, which is helpful; please state explicitly that k was tuned on the evaluation set and show the sensitivity of the main conclusions to k (e.g., whether spectral fingerprints still achieve the highest rank correlations at k=1 and k=25 on the same tasks).","section":"§3d; SI Tables S9–S11"}],"minor_comments":[{"comment":"The bonding-channel formula is not rendered clearly; the denominator in the exponential appears to be missing in the typeset equation. Please define it explicitly, e.g., as exp(−r_ij / (r_cov,i + r_cov,j)).","section":"§3a, Eq. (4)"},{"comment":"The table headers state 'Bold indicates the best result for a given metric, and the arrow indicates what direction corresponds to a better value,' but no arrows appear in the printed headers. Adding ↑/↓ to the metric names would make the tables self-explanatory.","section":"Tables 1 and 2"},{"comment":"The text says RACs vectors 'consistently require over ten times longer' than spectral fingerprints; for QCell the ratio is over 40×. The statement is technically true but understates the observed gap; consider giving the range of speedups.","section":"§3c, Supporting Information Table S8"},{"comment":"In the paragraph on Transition1x reactions, the text reports similarities of 0.434 and 0.429 between the ring-opening reaction and the two oxidation/reduction reactions, but Figure 2c is described as showing 'reaction spectral fingerprints between (reactant, product) pairs'; please clarify in the figure caption whether the displayed values correspond to pairwise reaction similarities or to the concatenated fingerprints themselves.","section":"§3b"}],"recommendation":"major_revision","confidential_remarks":"The central construction is sound and the paper is generally well written, but the two load-bearing issues are the M=16 truncation for large systems and the comparability/statistical rigor of the clustering benchmarks. Both are fixable within the manuscript's scope: an M-sensitivity study and a common-subset, multiple-testing-aware evaluation would substantially strengthen the claims. The paper would be a good fit for the journal if these are addressed. I do not see evidence of circularity or missing prior work; the distinction from BCUT descriptors and Coulomb-matrix eigenvalues is adequately drawn."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe core idea is solid and worth taking seriously. The authors define a complete graph over 3D coordinates with four physics-based edge weights (electrostatics, bonding, sterics, dispersion), take graph Laplacian eigenvalues per channel, and concatenate a fixed-length vector. That specific combination is new and it delivers on the main promise: it separates conformers and stereoisomers (chair vs boat cyclohexane at 0.961 similarity, cisplatin vs transplatin at 0.897, fac/mer at 0.956), and it is cheap enough to run on 2.7M GEMS fragments. The reaction fingerprint trick, concatenating reactant and product spectra, is also clever and seems to track activation energy differences in their examples. They ship code and data on GitHub and Zenodo, which counts for a lot.\n\nThe soft spots are real but addressable. The default M=16 eigenvalues per channel is empirically chosen on small molecules, and SI Figure S2 shows the large-protein similarities depend on the number of eigenvalues. The authors acknowledge this in the text, so it is not hidden, but the abstract's \"vast chemical spaces\" claim rests on an untuned hyperparameter for exactly the macromolecular and reticular systems where the method is supposed to shine. A proper M-robustness study for the QMOF and GEMS clustering would settle the question.\n\nThe benchmarking also has confounds. Cluster counts differ wildly across methods (76 for spectral on QM9 vs 1132 for scaffolds), so the pairwise Mann-Whitney U-test percentages are not directly comparable. No confidence intervals or multiple-testing corrections accompany those percentages. k=5 for k-NN was chosen by maximizing mean rank correlation, which is a mild tuning leak. And the \"best overall\" claim leans on parsability, which conflates RDKit's limitations with chemical generality. That said, the parsability argument is legitimate for practical screening.\n\nOverall, the central geometry-sensitivity claim holds. The paper deserves a serious referee. I would send it to review, but ask for an M-robustness analysis and a more careful treatment of the clustering metrics before acceptance.\n\nWho should read it: anyone building 3D similarity baselines or working on dataset curation and ML preprocessing for chemical spaces. I would cite it as a baseline and bring it to the next reading group.","headline":"A genuinely useful new 3D fingerprint that distinguishes conformers and stereoisomers cheaply, but its generalizability claim rests on a dataset-specific M=16 truncation and the clustering benchmarks need tightening.","tokens_in":33057,"tokens_out":3154,"would_cite":true,"duration_ms":29765,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that graph-Laplacian eigenvalues of a complete graph over 3D atomic coordinates form a fixed-length, permutation-invariant fingerprint that distinguishes conformers and stereoisomers that 2D connectivity collapses, at a…","keywords":["spectral graph theory","molecular fingerprint","chemical similarity","graph Laplacian","3D molecular representation","stereoisomer discrimination","cheminformatics","community detection"],"falsifier":"A direct falsifier would be a pair of stereoisomers or conformers with identical 2D connectivity, known different properties, and nearly identical spectral fingerprints (similarity above about 0.99), while unrelated molecules score lower; that would show the truncation discards the discriminating geometry. A quantitative version is to scan a large conformational ensemble and check whether pairwise spectral distances order the molecules the same way as RMSD or energy differences; any systematic inversion would contradict the claimed locality.","tokens_in":32112,"feed_emoji":"🧪","tokens_out":8342,"duration_ms":71352,"temperature":0.7,"pith_summary":"The paper introduces a molecular fingerprint built from the eigenvalues of graph Laplacians on a complete graph whose edge weights are four physics-inspired pairwise interactions derived from 3D coordinates and tabulated atomic properties. The authors aim to show that this fixed-length fingerprint is geometry-aware without the cost of pairwise alignment, and that it distinguishes structures such as conformers, stereoisomers, and mutated proteins that 2D connectivity representations collapse together. If correct, the method would give cheminformatics an interpretable, cheap, and domain-general similarity measure for screening large chemical spaces, with immediate uses in clustering, nearest-neighbor property estimation, and applicability-domain analysis.","feed_headline":"A 64-number fingerprint separates 3D molecular shapes","feed_subtitle":"Graph-Laplacian eigenvalues of atomic-interaction graphs separate conformers and stereoisomers without alignment.","key_machinery":"The central object is the multichannel graph Laplacian of a complete graph whose vertices are atoms. For each of four physical channels, the Laplacian is constructed from a distance-weighted adjacency matrix, and its eigenvalues are collected; the smallest nonzero eigenvalue (the algebraic connectivity) is placed first and the remaining extreme eigenvalues follow, truncated to 16 per channel after Frobenius normalization of the adjacency matrices. Because eigenvalues are invariant to vertex relabeling and the interaction weights depend only on interatomic distances, the resulting 64-component vector is permutation-invariant and E(3)-invariant, and it is fixed-length regardless of molecular size.","core_discovery":"The paper's central claim is that the spectrum of a graph Laplacian built from a complete graph over 3D atomic coordinates encodes molecular similarity in a way that is simultaneously geometry-aware, permutation-invariant, alignment-free, and fixed-length. The fingerprint uses four channels—electrostatics, bonding, sterics, and dispersion—whose edge weights are heuristic physical functions of interatomic distance and tabulated atomic properties; the channel-wise graph Laplacians are formed as $L^{(k)} = D^{(k)} - A^{(k)}$, and for each channel the smallest nonzero eigenvalue is placed first and the largest eigenvalues are kept down to a length of 16, giving a 64-dimensional vector. The authors report that this vector separates chair from boat cyclohexane, cis from trans platinum complexes, fac from mer iridium isomers, and proteins differing by side-chain mutations, all cases where 2D connectivity fingerprints coincide. They further report that clusters formed from fingerprint neighborhoods track property distributions across organic, inorganic, biological, reaction, and framework datasets, and that nearest-neighbor predictions in fingerprint space rank-correlate strongly with several quantum-chemical properties.","pith_inferences":["A natural next test is whether the default 16-eigenvalue cutoff can be replaced by an adaptive per-molecule cutoff chosen from the eigenvalue decay, which would remove the main tuned hyperparameter.","Because the fingerprint needs only 3D coordinates, it could serve as a cheap geometry-aware uncertainty signal for machine-learned property models even when those models are trained on 2D inputs.","The same complete-graph construction could extend to periodic solids or supramolecular assemblies by using cutoffs and supercell coordinates, provided the channel formulas are rescaled consistently.","If the reported locality is robust, spectral fingerprint distances could be used to design training sets that deliberately span underrepresented geometries rather than random splits."],"forward_implications":["Molecule libraries can be screened without computing pairwise alignments, because similarity is just Euclidean distance between fixed-length vectors.","Stereoisomers and conformers that are invisible to 2D fingerprints become distinguishable from coordinates alone.","The fingerprint gives a training-free property estimate: nearby molecules in fingerprint space have similar size-extensive properties such as atomization energy and polarizability.","It offers a model-agnostic applicability-domain estimate that can be computed before any model is trained.","Reactions can be compared by concatenated reactant and product fingerprints, and reaction similarity tracks activation-energy differences without transition-state geometries."],"supporting_citations":[{"why":"Supplies the spectral graph theory of graph Laplacian eigenvalues on which the fingerprint is built.","marker":"[71]"},{"why":"Defines the algebraic connectivity, the smallest nonzero Laplacian eigenvalue that the fingerprint places first.","marker":"[72]"},{"why":"Provides the extended-connectivity 2D fingerprint baseline that cannot distinguish structures differing only in 3D geometry.","marker":"[6]"},{"why":"Supplies the revised autocorrelation descriptor baseline for continuous physics-based fingerprints.","marker":"[15]"},{"why":"Introduces the Coulomb matrix, a prior eigenvalue-based 3D representation whose electrostatics channel this work modifies.","marker":"[24]"},{"why":"Provides the QM9 organic-molecule dataset used for clustering and nearest-neighbor benchmarks.","marker":"[85]"},{"why":"Supplies the tmQMg transition-metal dataset and the list of erroneous structures excluded from training.","marker":"[86]"},{"why":"Contributes the community detection algorithm used to partition fingerprint similarity graphs for large-scale screening.","marker":"[98]"},{"why":"Provides the 2.7-million-fragment GEMS dataset used for the large-scale scaling test.","marker":"[94]"}],"fun_headline_variants":["Spectral fingerprints reveal 3D structure at 2D speed","64 eigenvalues encode 3D molecular shape","Eigenvalue fingerprints distinguish 3D isomers","Graph Laplacian spectra give alignment-free 3D fingerprints","3D molecular similarity from graph eigenvalues"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a fixed 16-eigenvalue truncation, combined with Frobenius normalization, preserves enough chemical information to compare molecules that differ greatly in size; the paper states this cutoff was chosen empirically and that results for large proteins depend on it.","fun_headline_variants_meta":{"raw":{"variants":["Spectral fingerprints reveal 3D structure at 2D speed","64 eigenvalues encode 3D molecular shape","Eigenvalue fingerprints distinguish 3D isomers","Graph Laplacian spectra give alignment-free 3D fingerprints","3D molecular similarity from graph eigenvalues"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001062,"raw_usage":{"total_tokens":4505,"prompt_tokens":1050,"completion_tokens":3455,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":3380}},"tokens_in":666,"tokens_out":3455,"duration_ms":27159,"temperature":1.0,"reasoning_tokens":3380,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:03:37.400482+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct falsifier would be a pair of stereoisomers or conformers with identical 2D connectivity, known different properties, and nearly identical spectral fingerprints (similarity above about 0.99), while unrelated molecules score lower; that would show the truncation discards the discriminating geometry. A quantitative version is to scan a large conformational ensemble and check whether pairwise spectral distances order the molecules the same way as RMSD or energy differences; any systematic inversion would contradict the claimed locality.","supporting_citations":[{"cited_title":"The Shark Integral Generation and Digestion System","cited_arxiv_id":null,"evidence_quote":"Supplies the revised autocorrelation descriptor baseline for continuous physics-based fingerprints."},{"cited_title":"J.; Zunker, M.; Wolf, J","cited_arxiv_id":null,"evidence_quote":"Introduces the Coulomb matrix, a prior eigenvalue-based 3D representation whose electrostatics channel this work modifies."},{"cited_title":"P.; Kondor, R.; Csányi, G","cited_arxiv_id":null,"evidence_quote":"Contributes the community detection algorithm used to partition fingerprint similarity graphs for large-scale screening."}],"review_version":1}