{"id":"445e6d6f-e44c-4b93-ab04-ae5e1661a671","arxiv_id":"2506.01222","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A neural network collective variable built by enforcing the Legoll-Lelievre orthogonality condition on a diffusion-map surrogate manifold reproduces butane's anti-gauche transition rate to within 10 percent.","lead":"The paper learns a low-dimensional collective variable for butane that reproduces the anti-gauche transition rate to within 10 percent, by enforcing an orthogonality condition from quantitative coarse graining on a learned surrogate manifold. It also reports that rank-deficient diffusion tensors can still yield faithful rates, and that hydrogen atoms matter for the learned embedding.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Latent-space (OC) is never shown to imply true (OC): Algorithm 1 enforces ∇bξ·∇bΦ=0, but this gives no control on Dξ∇V1 for ξ=bξ∘Ψ; the 1.25e-2 rate is 10.6% above reference, not 'less than ten percent'.","rationale":"The reader's weakest assumption focuses on the quality of the learned surrogate manifold; my concern is narrower and is testable even if the surrogate is perfect: the condition actually imposed in Algorithm 1 is not the condition used in the theory. The paper states in Sec. 6 that OC is imposed on the surrogate rather than in all-atom space, but it gives no proposition or bound connecting ∇bξ·∇bΦ=0 to Dξ∇V1=0 for ξ=bξ∘Ψ. Because the diffusion net is trained only on configurations on M, its derivatives in the true normal direction are unconstrained, so latent orthogonality could hold while true OC fails badly. The empirical butane result remains plausible, and the availability of code is a real point in favor; however, the central claim that rate preservation is achieved through the orthogonality mechanism lacks a direct check. A single Jacobian computation on the existing trajectory and force field would settle it. I am not recommending rejection: the required test is straightforward, the reported rate is close, and the reader's CONDITIONAL verdict already asks for additional validation. The only additional, minor correction is that the learned rate of 1.25e-2 is 10.6% above the reference 1.13e-2, so the abstract's 'less than ten percent relative error' is a slight overstatement. Prop. 2 itself is sound under the full-rank assumption; the gap is in the transfer of the surrogate condition to the true potential, not in the algebra of the equivalence proof.","tokens_in":30192,"tokens_out":8093,"duration_ms":99503,"concrete_test":"Evaluate the normalized true orthogonality residual R = Eμ[|Dξ∇V1|/(||Dξ|| ||∇V1||)] for the learned PlaneAlign CV, using the same force field and trajectory and computing Dξ by automatic differentiation through ξ=bξ∘ΨF∘F. Compare R with the same residual for the dihedral angle θ and for (sinθ, cosθ). If R for the learned CV is not an order of magnitude smaller than for θ, then the rate reproduction is not explained by (OC); if it is, the latent-to-true transfer concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Algorithm 1 replaces the theoretical condition (OC), Dξ∇V1=0, by ∇bξ·∇bΦ=0 in latent coordinates. The error estimates of Sec. 2 (eqs. 32, 40, 48) and Prop. 2 apply to the true V1 in the all-atom space RN. For ξ=bξ∘Ψ, Dξ=(Dbξ∘Ψ)DΨ, so latent orthogonality controls only the component of DΨ∇V1 that is aligned with ∇bΦ. Since the SE(3)-invariant feature map F is non-injective, and the diffusion net Ψ is trained only on M, DΨ∇V1 is essentially unconstrained; ∇bΦ is the SDF normal to Ψ(M), not the pushforward of ∇V1. The paper acknowledges in Sec. 6 that OC is imposed on the surrogate rather than in RN, but no bound transfers the surrogate condition to the true coarse-graining error. Consequently the reported 1.25×10−2 ps−1 rate for PlaneAlign could be generated by arctan2 providing a smooth reparameterization of the dihedral angle, not by satisfaction of the advertised orthogonality condition. As a numerical aside, the headline comparison is 1.25 vs 1.13×10−2, i.e. 10.6% relative error, slightly above the claimed 'less than ten percent.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper revisits quantitative coarse-graining for overdamped Langevin dynamics, proves the equivalence of the orthogonality condition (OC) and the projected orthogonality condition (POC), and proposes a data-driven pipeline (Algorithm 1) for learning collective variables: featurize MD data, embed the residence manifold as a hypersurface via diffusion maps/diffusion nets/LAPCAE, learn a surrogate signed-distance potential, and enforce orthogonality in the latent space. The method is tested on butane, where the PlaneAlign-based CV is reported to reproduce the anti-gauche transition rate at 1.25×10−2 ps−1 versus the reference 1.13±0.08×10−2 ps−1. The paper also provides empirical evidence that a rank-deficient diffusion tensor can still reproduce transition rates and emphasizes the role of hydrogen atoms in featurization.","tokens_in":30562,"tokens_out":4094,"duration_ms":46144,"significance":"If the central claim were fully supported, the paper would make a useful contribution: Proposition 2 gives a clean identification of two conditions from different error estimates, the surrogate-manifold framework is a plausible route to imposing geometric conditions without knowing V1, the butane study is detailed, and the code is public. The empirical demonstration with (sin θ, cos θ) showing that a rank-deficient diffusion tensor can reproduce rates is interesting and relevant. However, the quantitative headline claim is contradicted by the authors' own table, and the connection between the latent-space condition actually imposed and the theoretical condition used in the error estimates is not established. The approach is defensible, but the paper currently overstates what is proven and what is demonstrated.","major_comments":[{"comment":"The paper repeatedly claims a 'less than ten percent' relative error for the PlaneAlign CV, but Table 2 gives 1.25×10−2 ps−1 against the reference 1.13±0.08×10−2 ps−1. Relative to the point estimate this is (1.25−1.13)/1.13 ≈ 10.6%, slightly above ten percent. The Introduction's phrase 'nearly 10%' is accurate, but the Abstract and §5.4 are not. This is a concrete, checkable numerical claim and should be corrected or carefully qualified (e.g., 'within 1.5 standard deviations of the reference').","section":"Abstract; §5.4, Table 2"},{"comment":"Algorithm 1 enforces the condition ∇bξ·∇bΦ = 0 in latent coordinates, but the theoretical estimates (32) and (40), as well as Proposition 2, concern the true condition Dξ∇V1 = 0 in the all-atom space R^N. For ξ = bξ ∘ Ψ, the chain rule gives Dξ = (Dbξ ∘ Ψ)DΨ, so latent orthogonality controls only the component of DΨ∇V1 aligned with ∇bΦ; no quantitative bound is supplied that transfers the surrogate condition to the true coarse-graining error. The acknowledgment in §6 that (OC) is imposed on the surrogate rather than in R^N does not by itself close this gap. Without such a transfer estimate, the reported PlaneAlign rate improvement cannot be attributed to satisfaction of the advertised orthogonality condition rather than to a favorable reparameterization of the dihedral angle.","section":"Algorithm 1; §5.2; §6 (Limitations)"},{"comment":"The successful case study is selected by visual inspection, and the successful CV is then also chosen by inspection. Section 4.4 states that PlaneAlign 'clearly visually' gives the best 2D embedding despite its low HyperSurface score, and §5.2 sets bξ(x) = arctan2(x2, x1) after observing that the embedding is a circle parameterized by the dihedral angle. This weakens the demonstration that the proposed algorithm automatically discovers a good CV: the favorable result could be obtained by any smooth reparameterization of the dihedral angle. The paper should either present these two choices as human-guided and moderate the claims accordingly, or provide an automated selection criterion that recovers PlaneAlign without visual inspection.","section":"§4.4; §5.2"}],"minor_comments":[{"comment":"The transition rate for ξ3 = (sin θ, cos θ) is printed as '1.19 ps−1' in the text; it should read 1.19×10−2 ps−1 to match Table 1.","section":"§3.2.1"},{"comment":"The inequality is attributed to 'Cziszar-Kullback'; the standard spelling is Csiszár-Kullback. The same spelling appears in the proposition label and should be corrected.","section":"§2.2.1"},{"comment":"The notation 'Leikonal' is used for the eikonal loss in equation (71); consider renaming to 'L_eikonal' for consistency with the rest of the text.","section":"§4.5; §5.1"},{"comment":"The header 'F eature map' contains a stray space; also, the table would be easier to read if the hyperparameter column distinguished αZero, αnormals, and αOC explicitly, since αnormals is reported as 0.0 for both rows.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The headline numerical claim is contradicted by the paper's own Table 2, and the latent-space OC condition is not connected to the true condition by any quantitative estimate; both issues are fixable in revision, but the current version substantially overstates the strength of the demonstration. The paper otherwise fits the journal's scope and the code availability is a positive feature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe cleanest thing here is Proposition 2: (OC) and (POC) are equivalent. The SVD proof is short and correct. That is a real, citable contribution. The pipeline around it is also genuinely assembled - diffusion maps, LAPCAE, an eikonal surrogate potential, and an OC-style loss for a learned CV - and the code is available, which makes the butane experiment reproducible. That deserves credit.\n\nNow the soft spots, in order of size.\n\nFirst, the algorithm enforces ∇bξ·∇bΦ=0 on the surrogate manifold, not Dξ∇V1=0 in the all-atom space. For ξ=bξ∘Ψ, the chain rule gives Dξ∇V1=(Dbξ∘Ψ)DΨ∇V1, and nothing transfers the latent orthogonality to a statement about the true coarse-graining error. The paper admits this in Section 6, but the abstract and Section 3 present the method as if it realizes (OC). It doesn't, strictly. That gap is load-bearing for the 'principled' framing, though it doesn't kill the numerical result.\n\nSecond, the headline rate is 1.25e-2 vs 1.13±0.08e-2, which is 10.6% relative error, not 'less than ten percent'. The introduction's 'nearly 10%' is fair; the abstract overstates. There's also no uncertainty on the learned rate, so the comparison is a single number against an interval. Minor but real sloppiness.\n\nThird, and related to circularity: the PlaneAlign embedding is a circle parameterized by the dihedral angle, and the learned CV is arctan2 of two coordinates on that circle - effectively a smooth reparameterization of θ. The paper says it 'correlates significantly' with θ. So it's not clear the OC loss is doing the work; the rate improvement might just come from using a different parameterization and different boundary intervals. A head-to-head with (sinθ, cosθ) in the same pipeline, plus sensitivity to feature map choice, would settle this. Without it, the empirical claim is suggestive, not demonstrative.\n\nThe rank-deficiency observation is interesting but only empirical; no theory is offered. The citation pattern is fine - [ZHS16] and [Mis+19] are the right anchors.\n\nBottom line: this deserves a serious referee, not a desk reject. The referee should ask for a precise statement of what the surrogate OC actually controls, error bars on the learned rate, and the (sinθ, cosθ) baseline. I'd bring it to the reading group as a good case of a principled idea meeting the messiness of learned representations.\n\nRecommendation: peer review, major revision expected.","headline":"A clean equivalence proof and a well-described CV discovery pipeline, but the surrogate-space OC is not the advertised theorem and the headline rate is 10.6%, not <10%.","tokens_in":31107,"tokens_out":5883,"would_cite":true,"duration_ms":58878,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65C30","60J60"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that enforcing the orthogonality condition $D\\xi\\nabla V_1 = 0$ on a learned surrogate manifold yields collective variables that preserve transition rates, and demonstrates it on butane with under ten percent rate error.","keywords":["collective variables","transition rates","orthogonality condition","effective dynamics","diffusion maps","manifold learning","butane","overdamped Langevin dynamics"],"falsifier":"Use the GramMatrixCarbon feature map in the same pipeline: its embedding is parameterized by $\\cos\\theta$, the collective variable that overestimates the butane rate at $1.52\\times 10^{-2}\\,\\mathrm{ps}^{-1}$. If an orthogonality-respecting collective variable learned from that embedding also lands near $1.52\\times 10^{-2}$ rather than near the reference $1.13\\times 10^{-2}$, then enforcing (OC) on a surrogate that misrepresents the manifold does not preserve kinetics.","tokens_in":29952,"feed_emoji":"⚛️","tokens_out":8521,"duration_ms":78533,"temperature":0.7,"pith_summary":"This paper tries to show that one geometric condition from effective-dynamics theory—the orthogonality condition, which asks that the collective variable’s level sets run along the slow directions of a scale-separated potential $V = V_0 + \\epsilon^{-1} V_1$—is a practical criterion for learning reaction coordinates that preserve kinetics, not just metastable state separation. It proves that this condition is equivalent to the projected orthogonality condition used for pathwise error estimates, so the same constraint controls both relative entropy and pathwise distance between the projected and effective dynamics. The paper turns the condition into a numerical recipe: learn a surrogate hypersurface for the residence manifold from simulation data, learn a surrogate confining potential whose gradient is normal to that surface, and train a neural network collective variable whose gradient is orthogonal to that normal field. On butane, the learned collective variable reproduces the anti-gauche transition rate at $1.25\\times 10^{-2}\\,\\mathrm{ps}^{-1}$ against a reference of $1.13\\pm 0.08\\times 10^{-2}\\,\\mathrm{ps}^{-1}$ (under ten percent relative error), while the dihedral angle gives 24 percent error. The paper also gives evidence that the diffusion tensor need not be uniformly positive definite for the rate to come out right.","feed_headline":"Orthogonality condition yields butane transition rates within 10%","feed_subtitle":"A learned collective variable reproduces butane's anti-gauche rate at 1.25×10−2 ps−1; the dihedral angle is off by 24%.","key_machinery":"The load-bearing object is the orthogonality condition (OC), $D\\xi\\nabla V_1 = 0$: the gradient of the collective variable is perpendicular to the gradient of the stiff confining potential, so level sets of $\\xi$ lie along the slow directions of $V = V_0 + \\epsilon^{-1}V_1$. Proposition 2 identifies (OC) with the projected orthogonality condition (POC), $(I-\\Pi)^{\\top}\\nabla V_1 = 0$, where $\\Pi$ is the projection induced by $\\xi$, making the same condition serve both the relative-entropy estimate and the pathwise-distance estimate. Computationally, the condition is enforced in a latent space rather than all-atom space: group-invariant features are embedded with diffusion maps, independent eigencoordinate selection and a hypersurface search choose coordinates so the embedded residence manifold is a hypersurface, diffusion nets extend the embedding out of sample, and a Laplacian conformal autoencoder removes spurious self-intersections; a surrogate potential is learned as a signed distance whose gradient is the surface normal, and the collective variable is trained so its gradient is orthogonal to that normal.","core_discovery":"The central claim is that a collective variable satisfying the orthogonality condition $D\\xi\\nabla V_1 = 0$ reproduces the statistical properties of the original overdamped Langevin dynamics under scale separation, and that this condition is equivalent to the projected orthogonality condition $(I-\\Pi)^{\\top}\\nabla V_1 = 0$. The equivalence means one geometric constraint controls both relative entropy and pathwise error estimates in coarse graining. The paper implements this by learning a surrogate manifold from featurized simulation data, learning a signed-distance surrogate potential on it, and training an encoder whose gradient is orthogonal to the surrogate potential’s gradient; the resulting collective variable is the composition of that encoder with the manifold embedding. In the butane case study, the learned variable separates the anti and gauche states and reproduces the anti-gauche transition rate within ten percent relative error, whereas the conventional dihedral angle overestimates it by 24 percent. The paper further claims that a rank-deficient diffusion tensor—as in the $(\\sin\\theta, \\cos\\theta)$ variable—does not prevent faithful transition rates.","pith_inferences":["The same pipeline should carry over to any molecule whose stiff degrees of freedom produce a codimension-one residence manifold, but only if the feature map yields a topologically faithful embedding; the paper’s own HyperSurface score gives false negatives, so this is the main transfer risk.","The rank-deficiency evidence invites a theoretical extension: error estimates for effective dynamics with degenerate diffusion tensors, replacing uniform positive definiteness with a weaker condition that the diffusion tensor does not vanish on its support.","The Laplacian conformal autoencoder’s ability to undo self-intersections is a standalone manifold-learning contribution that could be tested on other spectral embeddings independent of collective-variable construction.","One testable extension is to compare the feature-map route to Haar-averaged group-invariant diffusion-map kernels on the same butane data, since the paper notes the two routes have not been compared."],"forward_implications":["Transition rates for rare conformational changes can be computed from the low-dimensional effective dynamics once the learned collective variable satisfies the orthogonality condition, avoiding brute-force all-atom simulation.","The equivalence of (OC) and (POC) means a single constraint improves both relative-entropy and pathwise estimates of coarse-graining error.","The butane experiments suggest that requiring $D\\xi D\\xi^{\\top}$ to be uniformly positive definite is too strong; rank-deficient diffusion tensors still yield correct rates as long as the collective variable does not collapse distinct metastable states.","Group-invariant featurization that retains hydrogen coordinates can change the learned manifold enough to make the difference between a collective variable that separates metastable states and one that does not."],"supporting_citations":[{"why":"Foundation: effective dynamics conditioned on the invariant measure, the orthogonality condition (OC), and the relative-entropy estimate that the CV design is built on.","marker":"[LL10]"},{"why":"Supplies the generalized relative-entropy error estimates and assumptions that the paper reviews and relaxes when it drops uniform positive definiteness.","marker":"[Duo+18]"},{"why":"Supplies the pathwise-distance estimate and the projected orthogonality condition (POC) that Proposition 2 proves equivalent to (OC).","marker":"[LZ19b]"},{"why":"Provides the diffusion-map construction used to embed the featurized data into low-dimensional coordinates.","marker":"[Coi+08]"},{"why":"Provides the diffusion-net out-of-sample extension that turns the discrete embedding into a neural-network map.","marker":"[Mis+19]"},{"why":"Supplies independent eigencoordinate selection, which the hypersurface search extends to choose the embedding coordinates.","marker":"[CM19]"},{"why":"Provides the string-method and spring-force estimator used to compute the diffusion tensor in the rate calculation.","marker":"[Mar+06]"},{"why":"Provides well-tempered metadynamics, used to compute the free energy for the transition-rate formula.","marker":"[BBP08]"},{"why":"Supplies the transition-rate comparison formula and the observation that the committor is the optimal collective variable, the baseline for the butane experiments.","marker":"[ZHS16]"}],"fun_headline_variants":["Orthogonal CVs preserve rare-event rates","Butane rate matched within 10% by learned CV","Design rule for CVs: orthogonality beats dihedral","Surprising: rank-deficient diffusion still works"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The construction assumes the learned surrogate manifold faithfully represents the true residence manifold, with its normal directions matching the true stiff fast directions; if the diffusion-map embedding is distorted, self-intersecting, or mis-dimensioned, the orthogonality condition enforced in latent space need not correspond to the fast subspace of the original dynamics.","fun_headline_variants_meta":{"raw":{"variants":["Orthogonal CVs preserve rare-event rates","Butane rate matched within 10% by learned CV","Design rule for CVs: orthogonality beats dihedral","Surprising: rank-deficient diffusion still works"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1458,"prompt_tokens":968,"completion_tokens":490,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":428}},"tokens_in":584,"tokens_out":490,"duration_ms":5093,"temperature":1.0,"reasoning_tokens":428,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:46:50.530924+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use the GramMatrixCarbon feature map in the same pipeline: its embedding is parameterized by $\\cos\\theta$, the collective variable that overestimates the butane rate at $1.52\\times 10^{-2}\\,\\mathrm{ps}^{-1}$. If an orthogonality-respecting collective variable learned from that embedding also lands near $1.52\\times 10^{-2}$ rather than near the reference $1.13\\times 10^{-2}$, then enforcing (OC) on a surrogate that misrepresents the manifold does not preserve kinetics.","supporting_citations":[],"review_version":1}