{"id":"b3a80dd8-6887-4e47-9fcd-38d11096c67d","arxiv_id":"2507.00408","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Sort-based featurization plus autoencoders with an orthogonality loss yields permutation-symmetric collective variables whose reduced-model committor serves as a reaction coordinate for Lennard-Jones cluster transition rates.","lead":"A numerical pipeline learns symmetry-preserving collective variables for atomic clusters by sorting pairwise distances or coordination numbers, then uses the reduced model's committor as a reaction coordinate for forward flux sampling. On Lennard-Jones-7 (2D) and Lennard-Jones-8 (3D), the pipeline's escape-rate estimates mostly match brute-force simulation, though several reverse-rate estimates and the underlying smoothness theory remain weak points.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LJ8 'ML CV' success skips the autoencoder/orthogonality step, so the 3D case does not actually test Algorithm 1; the central 3D rate claim rests on diffusion-map eigenfunctions, not the proposed CV learner.","rationale":"The strongest claim bundles two assertions: Algorithm 1 (with the autoencoder and orthogonality loss) yields useful symmetry-respecting CVs, and the reduced-model committor produces rates and residence times matching brute force. Reading the workflow carefully, the LJ8 'ML CV' case—the only 3D demonstration—is obtained by stopping after the diffusion map; Section 6.1 explicitly says there is no need to learn V1 or the CVs. The tables labelled 'ML CV–sortrcs' accordingly test diffusion-map coordinates, not Algorithm 1. This is a more direct attack on the central claim than the piecewise-smoothness gap: the claimed method's key step is absent from the flagship example. The numerical comparisons also show the abstract overstates agreement even for the CVs that were used: Table 6 at β = 15 gives kB = (2.6 ± 0.3)×10^-3 by FFS versus (4.6 ± 0.4)×10^-3 by brute force, a roughly 4σ discrepancy, and Table 5 has similar factor-of-two gaps for (μ2, μ3). These issues are not fatal to the paper's engineering value—LJ7 demonstrates the full pipeline and the LJ8 kA estimates are good—but they are exactly the caveats the abstract should state. A concrete check running full Algorithm 1 on LJ8 would settle whether the 3D success belongs to the proposed innovation or to diffusion maps on sorted coordination numbers. The reader's CONDITIONAL verdict remains appropriate, so no verdict change is recommended.","tokens_in":44238,"tokens_out":8013,"duration_ms":90781,"concrete_test":"Run the LJ8-in-3D workflow exactly as Algorithm 1 prescribes: after the diffusion map, train a diffusion net and confining potential V1 (Sec. 4.2), then train the encoder/decoder pair with loss (4.6), and use the two learned CVs for the free-energy, diffusion-tensor, committor, and FFS pipeline (Sec. 4.4). Compare kA, kB, νAB, ρA, ρB at β = 10, 15, 20 with the brute-force values in Tables 5–7. If the AE-learned CVs do not separate Minima 1 and 2, or if their rates agree with brute force substantially worse than the reported diffusion-map CVs, the claim that Algorithm 1 produces the agreeing 3D rates is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central 3D claim is carried by the LJ8 'ML CV–sortrcs' case, but that case does not execute Algorithm 1. In Section 6.1, after computing the target-measure diffusion map, the authors state: 'Therefore, there is no need to learn the confining potential V1 and the CVs: ψ1pϕpxqq and ψ2pϕpxqq are the desired CVs in this case.' These ψ's are diffusion-map eigenvectors, not outputs of the autoencoder trained with the orthogonality loss (4.6). Thus the load-bearing novelty of the paper—learning CVs by enforcing Dξ∇V1 = 0—is untested on the 3D example used to support the abstract's claim that reduced-model rates agree with brute force. The LJ7 case does exercise the autoencoder, but only in 2D, and the sortrd2s variant even required extra MEP constraints (Eq. 5.1). This is independent of the smoothness issue: even setting aside the Sec. 4.1 admission that sorted features violate the regularity assumption behind Sec. 3.3, the empirical 3D success attributes to Algorithm 1 a result obtained by skipping its distinctive step. The conclusion's phrasing that CVs were 'learned by Algorithm 1' for LJ8 is therefore not supported by the reported workflow.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework, Algorithm 1, for learning collective variables that are invariant under translation, rotation, and permutation of identical particles, by combining sorted feature maps, diffusion-map/diffusion-net residence manifold learning, and an autoencoder whose loss enforces the Legoll–Lelièvre orthogonality condition. The reduced-model committor is then used as a forward-flux-sampling reaction coordinate and as a stochastic control for sampling transition paths. Two Lennard-Jones case studies are presented: LJ7 in 2D and LJ8 in 3D, with rates obtained by FFS compared against brute-force all-atom simulations. The paper reports that the reduced-model rates agree with brute-force rates, and it makes code and detailed implementation data available.","tokens_in":44615,"tokens_out":4439,"duration_ms":50139,"significance":"If the method works as claimed, it is a useful contribution to collective-variable discovery for systems with permutational symmetry, an important class that is poorly served by traditional physically motivated CVs. The manuscript is particularly strong in its engineering: it provides reproducible code, detailed appendices on neural-network architectures, Jacobian derivatives of the sorted feature maps, and extensive FFS versus brute-force comparisons. However, the theoretical grounding is weakened by the acknowledged non-smoothness of the sorted features, and the 3D case study does not actually execute the autoencoder/orthogonality step that is the paper's main algorithmic novelty. The rate agreement is also more mixed than the abstract suggests. The 2D LJ7 demonstration is credible, but the central 3D claim currently overstates what has been tested.","major_comments":[{"comment":"The LJ8 'ML CV–sortrcs' case does not execute Algorithm 1. After computing the target-measure diffusion map, the authors state: 'Therefore, there is no need to learn the confining potential V1 and the CVs: ψ1pϕpxqq and ψ2pϕpxqq are the desired CVs in this case.' These ψ's are diffusion-map eigenvectors, not outputs of the autoencoder trained with the orthogonality loss (4.6). Consequently, the conclusion's statement that the LJ8 CVs were 'learned by Algorithm 1' (Section 8) is not supported by the reported workflow. The distinctive component of the proposed framework—learning CVs by enforcing Dξ∇V1 = 0—remains untested on the 3D example used to support the abstract's rate-agreement claim.","section":"Section 6.1"},{"comment":"The orthogonality condition (3.19) is motivated by the relative entropy bound (3.17), which is derived under a regularity assumption (Duong et al., Assumption 2.3). Section 4.1 explicitly acknowledges that the sorted feature maps 'sortrd2s and sortrcs are continuous but only piecewise smooth, which violates the regularity assumption.' Since the CV ξ(ϕ(x)) and its Jacobian Dϕξ pass through sorting boundaries, the theoretical justification for the orthogonality-based loss does not directly apply to the feature maps actually used. The paper should either provide a relaxation of the regularity condition for sorting maps, or supply quantitative evidence (e.g., orthogonality residuals and sensitivity of the computed rates to small perturbations that change sorting order) that the lack of smoothness does not materially affect the bound.","section":"Sections 3.3 and 4.1"},{"comment":"The abstract's claim that 'the transition rates and residence times computed with the aid of the reduced models agree with those obtained via brute-force methods' is stronger than the data support. For the LJ8 ML CV–sortrcs case at β = 10, Table 6 reports FFS νAB = 3.1 ± 2.2 × 10⁻² versus brute-force νAB = 6.1 ± 0.1 × 10⁻², a factor-of-two discrepancy that is not within one standard deviation, and kB = 3.4 ± 2.6 × 10⁻² versus 7.2 ± 0.2 × 10⁻², also a factor of about two. The LJ7 kB estimates in Tables 1 and 3 likewise deviate by more than one standard deviation at several temperatures. The kA estimates generally agree well, but the overall rate and residence-time claim needs to be qualified, with relative errors reported systematically.","section":"Section 6.2, Tables 5–6"},{"comment":"The successful LJ7 result with the sortrd2s feature map uses the minimum-energy-path arc-length constraint (5.1), which is not part of Algorithm 1. The unmodified Algorithm 1 is therefore validated only by the LJ7 sortrcs case and by the LJ8 diffusion-map-eigenvector case, which, as noted above, skips the autoencoder step. This limitation should be stated in the conclusion so that the reader does not attribute the sortrd2s success to the baseline algorithm.","section":"Section 5.2"}],"minor_comments":[{"comment":"The label 'ML CV' for the LJ8 sortrcs CVs is misleading because the CVs are diffusion-map eigenvectors, not outputs of the machine-learned autoencoder. A name such as 'diffusion-map CV' would be clearer.","section":"Section 6.1"},{"comment":"Table 2 appears to be an exact duplicate of Table 1. This duplication should be removed or the two tables should be given distinct content.","section":"Appendix D"},{"comment":"There are numerous typos and garbled expressions: 'vice versus' in Table 1, 'Mimimum 2' in Section 7.3, 'committtor' in several Appendix headings, and corrupted figure labels such as 'vs , ML CV, , LJ8log(kA) β 𝗌𝗈𝗋𝗍[c]' in Figure 24. These should be corrected in a careful revision.","section":"Throughout"},{"comment":"The quantities ρA and ρB are last-hit probabilities, not residence times in the usual sense of mean residence times in the metastable sets. The paper's use of 'residence times' should be defined or replaced with 'last-hit probabilities' to avoid confusion.","section":"Sections 4.4 and 5.4"},{"comment":"In the reconstruction loss, the decoder's target is Ψ_DNet(ϕ), the diffusion-net embedding, rather than the original feature vector ϕ. This is an unusual choice and should be clearly motivated, since the autoencoder is then not reconstructing the input feature space but a lower-dimensional embedding of it.","section":"Equation (4.6)"}],"recommendation":"major_revision","confidential_remarks":"The skeptical point about the LJ8 case is accurate: the paper's central 3D claim is carried by diffusion-map eigenvectors, not by the proposed autoencoder with orthogonality loss. This is a load-bearing inconsistency that should be fixed by either running Algorithm 1 on LJ8 or substantially reframing the claims. The rate comparisons in Tables 5–6 also show factor-of-two discrepancies for νAB and kB that contradict the abstract's blanket agreement statement. The 2D LJ7 sortrcs demonstration is a solid positive result, and the paper is likely salvageable with a revision that narrows its claims and addresses the theoretical gap for non-smooth feature maps."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper is useful: sortrcs is a sensible new feature map for permutationally symmetric atomic clusters, and the workflow of diffusion-map manifold, orthogonality-loss autoencoder, reduced-model committor, and FFS is implemented carefully with open code and comparisons to brute force. Second, the headline claim is too strong. The LJ8 “ML CV–sortrcs” case, the centerpiece 3D demonstration, skips the autoencoder: Section 6.1 says the diffusion map eigenvectors directly are the CVs, so there was no need to learn the confining potential or train the orthogonality-loss autoencoder. The novel learning step is only actually tested in 2D. That does not sink the paper, but it changes what has been demonstrated.\n\nThe paper earns credit. It ships code, documents architectures, computes free energies and diffusion tensors, and runs FFS and brute force with repeated trials. The sortrcs featurization is simple and effective, and the comparisons with pµ₂,µ₃ and LDA CVs are informative. The authors are also honest in the discussion that kB agreement is worse than kA.\n\nThe soft spots are real but manageable. The rate agreement is uneven, and the abstract overstates it: for LJ8 at β=10, FFS gives νAB = 3.1±2.2×10⁻² versus brute force 6.1±0.1×10⁻², and kB similarly disagrees. The paper’s claim that rates “agree” leans heavily on kA. There is also a theory gap: the orthogonality condition is motivated by relative entropy bounds derived under smoothness assumptions, and the sorted feature maps are only piecewise smooth, as the paper admits. So the loss is applied heuristically. That is a modest issue because the method’s value here is empirical, but it deserves a clear statement rather than a silent skip.\n\nBottom line: a specialist in rare-event sampling or CV discovery should read this. It deserves peer review, not desk rejection, but the revision should narrow the abstract, re-frame the LJ8 result as a diffusion-map CV success rather than an Algorithm 1 success, and either extend the regularity argument or state plainly that the orthogonality condition is used heuristically.","headline":"Useful, honest engineering for permutationally symmetric CVs, but the headline overclaims: the LJ8 3D success skips the autoencoder step and the rate agreement is mainly kA.","tokens_in":45107,"tokens_out":2383,"would_cite":true,"duration_ms":27619,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["82.20.Db","05.10.Gg"],"model":"deepseek-v4-flash","headline":"Sorting atom descriptors before manifold learning and autoencoding yields collective variables that respect permutational symmetry, and the reduced-model committor used as a reaction coordinate reproduces brute-force transition rates.","keywords":["collective variables","permutational symmetry","coarse-graining","autoencoders","diffusion maps","Lennard-Jones clusters","committor","forward flux sampling"],"falsifier":"Extend the brute-force all-atom estimates with about ten times more statistics in the channel where the paper already reports the worst agreement—the back-rate $k_B$ for LJ7 at $\\beta = 7$ with the learned CVs, where the FFS and brute-force means differ by roughly a factor of two, and the $\\nu_{AB}$ channel for LJ8 at $\\beta = 20$, where the FFS mean is about three times smaller than the brute-force mean with comparatively tight error bars. If either discrepancy persists beyond three standard deviations at full statistics, the claim that reduced-model rates agree with brute force is false for that channel.","tokens_in":43997,"feed_emoji":"⚛️","tokens_out":15942,"duration_ms":162160,"temperature":0.7,"pith_summary":"The paper tries to establish a workable route from raw atomic coordinates to low-dimensional collective variables for clusters of identical particles, where swapping atoms leaves the system unchanged. Its proposal is to sort a symmetry-invariant descriptor—pairwise distances squared or coordination numbers—so that every permutation of a configuration maps to the same feature vector, learn the residence manifold in that feature space, and train an autoencoder whose loss enforces the Legoll–Lelièvre orthogonality condition between the collective-variable level sets and that manifold. The reduced model's committor, the hitting probability of one metastable state before the other, is then lifted back to atomic coordinates and used as the reaction coordinate for forward flux sampling. The paper reports that for Lennard-Jones-7 in 2D and Lennard-Jones-8 in 3D the resulting transition rates and residence times match brute-force all-atom simulation, which matters because it offers a way to estimate rare-event rates from cheap reduced models and to visualize metastable states in two dimensions.","feed_headline":"Sorted-particle autoencoders match brute-force cluster rates","feed_subtitle":"Lennard-Jones clusters in 2D and 3D get rate estimates from reduced models that agree with full all-atom simulation.","key_machinery":"The load-bearing identity is the orthogonality condition (3.19), $D\\xi(x)\\nabla V_1(x) = 0$, which demands that the level sets of the collective variable be orthogonal to the residence manifold where the invariant measure concentrates; Legoll and Lelièvre (2010) and Duong et al. (2018) proved that this condition suppresses the relative-entropy bound (3.17) measuring the gap between the reduced dynamics and the projected original dynamics. The paper turns this analytic condition into a training term in the autoencoder loss (4.6), alongside reconstruction and linear-independence terms. The complementary mechanism is the rate-versus-committor identity of Zhang, Hartmann, and Schütte (2016): the transition rate in any collective-variable space is never smaller than the true rate, with equality when the collective variable is the committor, which is why the reduced-model committor, not the raw learned variables, is lifted to atomic coordinates and used as the reaction coordinate for forward flux sampling and for the control $2\\beta^{-1}\\nabla\\log q(x)$ for sampling transition paths.","core_discovery":"The paper's central claim is that Algorithm 1—sort-based featurization, residence-manifold learning by diffusion maps and a diffusion net, and an autoencoder whose loss contains an orthogonality term—produces collective variables that respect translational, rotational, and permutational symmetry, and that the committor of the reduced model in those variables, composed with the learned features to give $q(x) = \\tilde{q}(\\xi(\\varphi(x)))$, is a good reaction coordinate for forward flux sampling and for the stochastic control used to sample transition paths. For both case studies the reported escape rates $k_A$ and the transition rates $\\nu_{AB}$ obtained with the learned variables agree with brute-force all-atom results within statistical error; the standard coordination-moment variables $(\\mu_2, \\mu_3)$ fare comparably, while linear-discriminant-analysis variables show notably larger rate errors. The paper further claims that its proposed feature map, the sorted vector of coordination numbers $\\mathrm{sort}[c_s]$, separates the relevant metastable basins in both systems, whereas the sorted squared-distance map $\\mathrm{sort}[d^2]$ works for LJ7 in 2D but fails to separate the two main minima of LJ8 in 3D.","pith_inferences":["The sorting trick is generic: any permutation-invariant ordering of a symmetric descriptor collapses the symmetry group, so the same pipeline should transfer to other symmetric clusters, and the paper's own named next target, LJ38 with its double-funnel landscape, is a direct test of whether $\\mathrm{sort}[c_s]$-style features separate competing funnels as cleanly as they separate the LJ8 minima.","The piecewise-smoothness gap suggests a discriminating experiment the authors did not run: replace $\\mathrm{sort}[d^2]$ and $\\mathrm{sort}[c_s]$ with a smooth permutation-invariant embedding, for instance an equivariant graph-network descriptor, and re-run the pipeline; if rates stay the same, the orthogonality loss's theoretical basis is incidental to the empirical success, and if they change, th","A weaker reading of the paper's message is that the orthogonality loss does not have to preserve rates directly; it only has to produce collective variables whose level-set foliation lets the lifted committor approximate the true committor closely enough for interface-crossing probabilities, and measuring the overlap of the reduced-model committor's level sets with true isocommittor surfaces for L"],"forward_implications":["Rate estimation for permutationally symmetric clusters can be done with forward flux sampling driven by a two-dimensional reduced-model committor instead of an all-atom committor, and the reported $k_A$ and $\\nu_{AB}$ values match brute-force all-atom estimates within the combined statistical error bars at essentially every temperature tested.","The $\\mathrm{sort}[c_s]$ feature map separates the free-energy basins of the relevant minima in both LJ7 in 2D and LJ8 in 3D, so the learned collective variables can serve both to define metastable sets and to visualize where transition trajectories concentrate.","The same lifted committor defines an approximately optimal control for the transition path process, allowing the bias vector $2\\beta^{-1}\\nabla\\log q(x)$ to be added to the drift so that reactive trajectories can be sampled in bulk and their probability density estimated by binning in the CV space.","Because the reduced-model rates follow Arrhenius behavior across $\\beta = 5,7,9$ for LJ7 and $\\beta = 10,15,20$ for LJ8, in agreement with brute force, the framework inherits the temperature dependence of the underlying dynamics rather than imposing an ad hoc one."],"supporting_citations":[{"why":"Supplies the orthogonality condition (3.19) and the relative-entropy bound (3.17) that the autoencoder loss turns into a training objective.","marker":"[52]"},{"why":"Extends the orthogonality condition and entropy bound to vector collective variables; its smoothness Assumption 2.3 is the regularity assumption the sorted feature maps violate.","marker":"[28]"},{"why":"Proves that the collective-variable-space transition rate equals or exceeds the true rate, with equality when the collective variable is the committor, motivating the reduced-model committor as the forward flux sampling reaction coordinate.","marker":"[73]"},{"why":"The parallel work whose five-step framework (feature map, residence manifold, orthogonality-conditioned autoencoder, free energy and diffusion tensor, committor-based FFS) is here adapted to permutational symmetry.","marker":"[63]"},{"why":"Introduces the coordination-number central moments $(\\mu_2, \\mu_3)$ used both as the comparison collective variables and as the biasing variables for generating the input datasets.","marker":"[64]"},{"why":"The forward flux sampling method used to compute escape rates from interface-crossing probabilities with the committor as reaction coordinate.","marker":"[2]"},{"why":"Diffusion nets, used to extend the diffusion map beyond the point cloud so that the orthogonality loss can be differentiated in feature space.","marker":"[55]"},{"why":"Supplies the finite-element committor solver and the stochastic control for the transition path process that the rate-estimation step reuses.","marker":"[71]"}],"fun_headline_variants":["Symmetry-aware autoencoders reproduce cluster escape rates","Sort-based collective variables match exact cluster transition rates","Permutation-invariant learning hits brute-force cluster rates","Autoencoders with sorted inputs recover cluster rates","Respectful of swaps, learned variables mirror cluster rates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The theoretical guarantee behind the training loss is proven only for smooth feature maps, while the sorted feature maps actually used are only piecewise smooth, so the orthogonality condition the loss enforces is not rigorously justified for the true inputs, a limitation the paper itself acknowledges in Section 4.1.","fun_headline_variants_meta":{"raw":{"variants":["Symmetry-aware autoencoders reproduce cluster escape rates","Sort-based collective variables match exact cluster transition rates","Permutation-invariant learning hits brute-force cluster rates","Autoencoders with sorted inputs recover cluster rates","Respectful of swaps, learned variables mirror cluster rates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1528,"prompt_tokens":981,"completion_tokens":547,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":472}},"tokens_in":597,"tokens_out":547,"duration_ms":7494,"temperature":1.0,"reasoning_tokens":472,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:19:15.474833+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Extend the brute-force all-atom estimates with about ten times more statistics in the channel where the paper already reports the worst agreement—the back-rate $k_B$ for LJ7 at $\\beta = 7$ with the learned CVs, where the FFS and brute-force means differ by roughly a factor of two, and the $\\nu_{AB}$ channel for LJ8 at $\\beta = 20$, where the FFS mean is about three times smaller than the brute-force mean with comparatively tight error bars. If either discrepancy persists beyond three standard deviations at full statistics, the claim that reduced-model rates agree with brute force is false for that channel.","supporting_citations":[{"cited_title":"Legoll and T","cited_arxiv_id":null,"evidence_quote":"Supplies the orthogonality condition (3.19) and the relative-entropy bound (3.17) that the autoencoder loss turns into a training objective."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends the orthogonality condition and entropy bound to vector collective variables; its smoothness Assumption 2.3 is the regularity assumption the sorted feature maps violate."},{"cited_title":"Zhang, C","cited_arxiv_id":null,"evidence_quote":"Proves that the collective-variable-space transition rate equals or exceeds the true rate, with equality when the collective variable is the committor, motivating the reduced-model committor as the forward flux sampling reaction coordinate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The parallel work whose five-step framework (feature map, residence manifold, orthogonality-conditioned autoencoder, free energy and diffusion tensor, committor-based FFS) is here adapted to permutational symmetry."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the coordination-number central moments $(\\mu_2, \\mu_3)$ used both as the comparison collective variables and as the biasing variables for generating the input datasets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The forward flux sampling method used to compute escape rates from interface-crossing probabilities with the committor as reaction coordinate."},{"cited_title":"Mishne, U","cited_arxiv_id":null,"evidence_quote":"Diffusion nets, used to extend the diffusion map beyond the point cloud so that the orthogonality loss can be differentiated in feature space."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the finite-element committor solver and the stochastic control for the transition path process that the rate-estimation step reuses."}],"review_version":1}