Pith. sign in

REVIEW 4 major objections 5 minor 3 references

WeTICA: A directed search weighted ensemble based enhanced sampling method to estimate rare event kinetics in a reduced dimensional space

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A binless weighted-ensemble algorithm guided by a target-directed distance objective recovers protein unfolding times from simulations roughly an order of magnitude shorter than the events themselves.

desk verdict Plausible binless WE variant with a real idea worth testing, but the paper must pin down the merge rule and be honest about parameter tuning before its rate claims can be trusted. read the letter →

arxiv 2501.08926 v1 pith:QP56KC6A submitted 2025-01-15 cond-mat.soft physics.bio-phphysics.chem-phphysics.comp-ph

classification cond-mat.softphysics.bio-phphysics.chem-phphysics.comp-ph
keywords weightedensemblesimulationbinlessresamplingTICAcollectivevariablesproteinunfoldingkineticsmeanfirstpassagetimeenhancedsamplingtarget-directed
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is trying to establish that a binless weighted-ensemble algorithm can estimate rare-event kinetics without optimizing bins or even knowing slow collective variables in advance. The proposed method, WeTICA, points every resampling decision at a target conformation: it clones the walker closest to the target and merges distant pairs, using the inverse distance to the target as the quantity to maximize. In tests on three proteins with reported unfolding times of 3 to 40 microseconds, WeTICA returns unfolding times of 3.77 +/- 0.15, 12.87 +/- 0.09, and 47.35 +/- 0.39 microseconds from cumulative simulations of roughly 0.7, 1.2, and 1.35 microseconds, meaning the rate is recovered with about an order of magnitude less simulation time than the event itself. The method also works when the TICA eigenvectors are trained only on folded and unfolded end-state trajectories rather than on a long trajectory containing transitions. If right, it gives a practical route to kinetic rates for systems where binning and good reaction coordinates are the main obstacles.

What carries the argument

The central object is the target-directed trajectory variation $V = \sum_i 1/d_{\mathrm{target}}(x_i)$, evaluated on fixed low-dimensional projections of the walkers. Distances are Euclidean distances between walker projections and the target-state projection on the first two TICA eigenvectors, where TICA is a linear dimensionality-reduction method that extracts the slowest independent motions from time-lagged molecular data. Each cycle, WeTICA clones the walker closest to the target and merges the walker farthest from the target with a nearby walker within a cutoff distance $d_{\mathrm{merge}}$ whose combined weight stays below a cap, randomly selecting one survivor from the pair; warping occurs when a walker enters the target hypersphere of radius $d_{\mathrm{warp}}$, and the mean first passage time is computed from the weights of warped walkers. The variation-maximization objective is what replaces binning: no bins partition the collective-variable space, and cloning and merging decisions are made entirely from the distance geometry on the projection plane.

What would settle it

For a model with an analytically known mean first passage time (e.g., a one-dimensional double well), run WeTICA while varying $d_{\mathrm{warp}}$ across the basin-size range suggested by the distance-distribution minima; an unbiased first-passage estimator should return the same mean first passage time, while a systematic drift with $d_{\mathrm{warp}}$ would show that the boundary, not the dynamics, sets the rate. Also check exact conservation of total walker weight through the random survivor selection in each merging event.

Watch

Extended reading notes

Core claim

WeTICA claims that rare-event kinetics can be estimated by a binless weighted-ensemble simulation in which every resampling decision is driven by a single target-directed objective. On a fixed projection plane spanned by the first two TICA eigenvectors, each walker carries a variation value $V_i = 1/d_{\mathrm{target}}(x_i)$; the walker with the largest variation (closest to target) is cloned, while the walker with the smallest variation is merged with a nearby walker if their combined weight stays below 0.20. A walker entering the hypersphere of radius $d_{\mathrm{warp}}$ around the target projection is warped, its weight is recorded, and the walker restarts from the folded state; the mean first passage time is total simulation time divided by the summed warped weights. Applied to TC10b, TC5b, and Protein G, the method yields 3.77 +/- 0.15, 12.87 +/- 0.09, and 47.35 +/- 0.39 microseconds against reported values of 3 +/- 1, 12.7, and 37 +/- 10 microseconds, using total simulations of roughly 0.70, 1.20, and 1.35 microseconds. The Protein G test uses TICA eigenvectors trained on two short end-state trajectories only, showing that transition-containing training data are not required.

Load-bearing premise

The load-bearing premise is that the merging step chooses the surviving walker in proportion to its weight, so total statistical weight is conserved, and that the per-system warping distance $d_{\mathrm{warp}}$ marks a true first-passage boundary; if either gives way, the computed unfolding time is biased.

Editorial extensions

If this is right

  • Unfolding times for TC5b, TC10b, and Protein G can be estimated from cumulative weighted-ensemble runs of about one microsecond or less, removing the need to simulate the full unfolding event.
  • Because the collective variables are fixed linear projections, the same algorithm can be run with other linear dimensionality-reduction coordinates and is not restricted to TICA.
  • Training TICA on short end-state trajectories, as in the Protein G case, makes the method usable when no long transition-containing trajectory is available.
  • The merging and warping scheme may extend binless weighted-ensemble simulation to other rare events, such as ligand unbinding or conformational transitions, once a target-state conformation and a projection space are chosen.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next test is to apply WeTICA to a model with an exactly calculable mean first passage time, where the sensitivity of the reported rate to the warping distance and to the merging survivor rule can be checked without force-field uncertainty.
  • The paper tunes $d_{\mathrm{warp}}$ per system, choosing the largest value that still allows escape from the folded basin; this suggests the reported rates may inherit a parameter-selection bias, and a principled way of setting or averaging over this boundary would strengthen the method.
  • Because merging decisions use only the current projection distances, the algorithm's efficiency should degrade gracefully as the target basin becomes broad or multimodal; a test would be unfolding to a target defined by an ensemble of unfolded conformations rather than a single representative structure.
  • If the resampling is truly unbiased, the same directed-search principle should work in non-protein contexts, such as crystal nucleation or conformational switching, with the target state specified in whatever linear collective-variable space separates the endpoints.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript introduces WeTICA, a binless weighted-ensemble enhanced-sampling method that drives walkers toward a predefined target state using a variation function V = sum_i 1/d_target(x_i), where d_target is the Euclidean distance in a fixed low-dimensional TICA projection plane. The method is implemented by modifying the Wepy codebase and is tested on the unfolding kinetics of three proteins: the TC10b Trp-cage mutant, the TC5b Trp-cage mutant, and Protein G. The authors report computed unfolding times of 3.77 +/- 0.15 microseconds, 12.87 +/- 0.09 microseconds, and 47.35 +/- 0.39 microseconds, respectively, matching reported values of 3 +/- 1, 12.7, and 37 +/- 10 microseconds with more than one order of magnitude less cumulative simulation time than the unfolding time scales. The paper also includes a set of heuristic guidelines for choosing simulation parameters such as the number of walkers, the merging distance, and the warping distance.

Significance. If the method is statistically exact and the reported validations are reliable, WeTICA offers a practical binless alternative to conventional WE binning. A notable strength is that the paper provides code and data via a public GitHub repository, and the methodology is demonstrated on atomistic explicit-solvent simulations of three proteins with microsecond unfolding times. The extension beyond TICA to other linear CVs and the proposed resampling scheme for nonlinear CV spaces are also of interest. However, the significance of the central claim is currently limited by ambiguities in the resampling rule and by the sensitivity of the reported rates to parameter choices and reference-state definitions.

major comments (4)
  1. [Section II, merging step] The text states that 'a walker from that merging pair (j, k) is randomly selected for continuation with the total weight (w0+w1) and the other walker is discontinued.' This phrasing is ambiguous about the selection probability. For the weighted-ensemble resampling to be unbiased, the surviving configuration must be chosen with probability proportional to its weight, i.e., configuration j survives with probability w_j/(w_j+w_k). If the selection is instead uniform (a fair coin), the expected weight in each configuration is not conserved, and the MFPT computed by Eq. (2) becomes biased. Because every reported rate constant depends on this step, the manuscript must specify the weight-proportional rule explicitly and confirm that the provided code implements it. Without this, the central claim of unbiased kinetics estimation is not established.
  2. [Section V.B and Table 1] The WeTICA simulation for TC5b is performed at 300 K, while the experimental value used for comparison (12.7 microseconds, from reference 90) was measured at 296 K. The manuscript does not account for this 4 K temperature difference or provide an argument that its effect on the unfolding time is negligible. Since protein unfolding rates typically depend strongly on temperature, a 4 K mismatch could lead to a non-negligible shift in the reference value. The comparison in Figure 4d and Table 2 should either correct for the temperature difference using an assumed Arrhenius behavior, or the simulation should be repeated at 296 K, or the uncertainty in the comparison should be expanded to reflect the temperature sensitivity.
  3. [Section V.C, Protein G d_warp selection] For Protein G, the warping distance d_warp is not set by the distance-distribution criterion of Section III D. Instead, the authors scan three values (1.0, 1.2, 1.4) and select d_warp = 1.4 because the walkers still leave the folded basin (Q > 0.8). Since d_warp defines the boundary of the target state, and the MFPT is inversely proportional to the probability flux into that boundary, the reported rate depends directly on this choice. The manuscript does not report how the computed unfolding time varies with d_warp for the three tested values, nor does it provide an independent, a priori criterion for choosing the largest acceptable value. As written, the selection is post hoc and could bias the result toward the reference value. The authors should report a sensitivity analysis of the MFPT versus d_warp and justify the choice on grounds other than agreement with the benchmark.
  4. [Sections IV and V, CV training and target-state construction] For TC10b and TC5b, the TICA eigenvectors are trained on full-length trajectories that contain multiple folding/unfolding transitions, and the target unfolded conformations are selected from those same trajectories. The reported reference unfolding times for TC10b come from the same trajectory, and for TC5b from the same force field used in the training trajectory. This circularity means that the claim of working 'without a priori knowledge of the CVs' is not fully supported by these two cases. For Protein G, the TICA eigenvectors are trained on two 2 microsecond end-state segments from the same long Anton trajectory that also provides the 37 +/- 10 microsecond reference value, so the independence of the benchmark is only partial. A more convincing demonstration would use a reference from an independent simulation or experiment, with CVs trained solely on short end-state simulations that do not contain the transition. The authors should at least clearly state this limitation and discuss how it affects the generality of the conclusions.
minor comments (5)
  1. [Eq. (2) and surrounding text] The definition of the set U in Eq. (2) is not fully explicit in the main text; it would help to state clearly that U contains all warping events from all walkers up to time T, and that the sum is over the weights of those events. The notation should also be cleaned up, since the equation appears garbled in the manuscript.
  2. [Section III D and Figure 5b] The advice in Section III D to 'choose the largest value' of d_warp when the target basin is flat is potentially circular, because the choice is made by observing which runs appear 'productive.' This guideline should be presented as a practical heuristic rather than a validated selection rule, and its effect on the reported rate should be examined.
  3. [Figure 5b] The caption states that the three curves are from 'three productive WeTICA unfolding trajectories' generated with different d_warp values, but it is not clear whether these are independent runs or one run per d_warp. Please clarify the number of independent runs and how productivity was judged.
  4. [Table 1 and Section IV.B] The table lists the TC10b temperature as 290 K, which matches the original Anton simulation, but the text in Section IV.B does not explicitly state that the WeTICA simulation was performed at 290 K, only that the NVT ensemble was used at this temperature in the table. A sentence in the main text would improve clarity.
  5. [Figure 1] The workflow diagram contains the abbreviation 'Clst walk. dist' without a definition in the main text; it should be spelled out or defined in the caption for clarity.

Circularity Check

1 steps flagged · score 4.0 of 10

Trp-cage demonstrations are partially circular because the TICA CVs and target structures come from the benchmark trajectories; the Protein G end-state-trained case is an independent test.

  1. self definitional [Section V.B (Trp-cage discussion); cf. Section IV.B TC10b setup]
    "However, we calculated the TICA eigenvectors (CVs) for these two systems using long enough unbiased trajectories where several folding-unfolding events have been observed. Thus, by construction, the TICA eigenvectors are embedded with all the necessary information related to the slow folding-unfolding processes."

    For TC10b and TC5b, the fixed TICA eigenvectors used as CVs are fitted to the same long folding-unfolding trajectories from which the benchmark unfolding times and the target unfolded conformations are taken. The WeTICA objective V=sum(1/d_target) and the warping boundary are both defined on those CVs, so the reported MFPTs (3.77 us and 12.87 us) are first-passage times to a target basin on a coordinate system already informed by the very transition whose kinetics are being predicted. The paper itself concedes this by saying the eigenvectors are 'by construction' embedded with the slow folding-unfolding information. The Protein G case, where TICA is trained only on two end-state segments, is an independent validation and prevents the entire claim from collapsing.

full rationale

The only circular element found is in the two Trp-cage demonstrations: their TICA eigenvectors and target structures are constructed from the same long trajectories that supply or define the benchmark kinetics, so the agreement for TC10b (3.77 vs 3±1 us) and TC5b (12.87 vs 12.7 us) is partly a consistency check of the WE resampling on a reaction coordinate already informed by the target process. This is explicitly acknowledged in the paper. The central protocol still has independent content: Protein G uses TICA trained on two short end-state segments with no inter-state hopping, and its 47.35 us result against the external 37±10 us benchmark provides a genuine out-of-sample check. The ambiguous 'randomly selected' merge rule and the per-system d_warp tuning are correctness and parameterization risks, not circular-equivalence reductions under the strict definition. There is no load-bearing self-citation chain.

Assumptions & free parameters 9 free parameters · 5 assumptions · 0 invented entities

The method introduces no new physical entities, but it does depend on several heuristically chosen parameters and domain assumptions. The most significant are the warping distance d_warp, which directly sets the target boundary and is tuned per system, and the assumption that weight-conserving resampling with the described selection rules yields unbiased kinetics. The TICA CVs are derived from the benchmark trajectories themselves for the two Trp-cage cases, which weakens the 'no a priori CV knowledge' claim for those tests.

free parameters (9)
  • d_merge (walker merging distance cutoff) = 0.50 (TC5b, TC10b), 0.25 (Protein G)
    Set from the position of the minimum of the pairwise projection distance distribution of the initial state ensemble (Figures S2, S4, S6). Determines which walker pairs can be merged in resampling.
  • d_warp (walker warping distance) = 0.75 (TC5b, TC10b), 1.40 (Protein G)
    Set from the minimum of the target state distance distribution; for Protein G, values 1.0, 1.2 and 1.4 were tested and 1.4 selected based on observed Q values of productive trajectories. Directly defines the target boundary for MFPT and is post hoc.
  • p_max (maximum walker weight for merging) = 0.20
    Chosen by hand as a cap on merged walker weight (Section II).
  • p_min (minimum walker weight for cloning) = 1e-5
    Chosen by hand; walkers with weight below this are prohibited from cloning (Section II).
  • Number of walkers N_w = 24
    Chosen as a balance between search speed and computational cost (Section III A).
  • Resampling interval = 20 ps
    Fixed for all systems (Table 1); not optimized.
  • TICA lag time = 10 ns (Trp-cage mutants), 20 ns (Protein G)
    Chosen for TICA training in Section IV.
  • Number of TICA eigenvectors = 2
    Chosen to separate the two end states while minimizing dimensionality (Section III B).
  • Target state conformation = A snapshot with Q ~ 0.2 from the reference trajectory
    Selected for each system as the representative unfolded state; all d_target distances and the warping sphere are defined relative to this conformation.
assumptions (5)
  • domain assumption Weighted ensemble resampling with weight-conserving clone/merge is unbiased for estimating MFPTs.
    The MFPT formula (Eq. 2) and the algorithm rely on standard WE unbiasedness; invoked implicitly in Section II.
  • domain assumption TICA eigenvectors trained on the specified trajectories capture the slow unfolding modes separating folded and unfolded states.
    Section IV uses TICA eigenvectors as CVs; Figures S1, S3, S5 validate separation on the projection plane.
  • domain assumption Distance distribution minima between projection points characterize the sizes of the initial and target state basins.
    Section III C/D prescribes this heuristic; used to set d_merge and d_warp.
  • domain assumption The Q ~ 0.2 snapshot is a valid representative of the unfolded state for defining the warping boundary.
    Section IV selects one snapshot per system; the computed MFPT depends on this choice.
  • domain assumption Comparing the TC5b simulation at 300 K with the experimental unfolding time at 296 K is valid without temperature correction.
    Section V B compares values at different temperatures without correction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of WeTICA: A directed search weighted ensemble based enhanced sampling method to estimate rare event kinetics in a reduced dimensional space." pith.science (2026). https://pith.science/paper/QP56KC6A

@misc{pith2026250108926,
  author       = {Pith},
  title        = {Pith review of: WeTICA: A directed search weighted ensemble based enhanced sampling method to estimate rare event kinetics in a reduced dimensional space},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QP56KC6A}},
  note         = {Machine review of arXiv:2501.08926}
}
read the original abstract

Estimating rare event kinetics from molecular dynamics simulations is a non-trivial task despite the great advances in enhanced sampling methods. Weighted Ensemble (WE) simulation, a special class of enhanced sampling techniques, offers a way to directly calculate kinetic rate constants from biased trajectories without the need to modify the underlying energy landscape using bias potentials. Conventional WE algorithms use different binning schemes to partition the collective variable (CV) space separating the two metastable states of interest. In this work, we have developed a new "binless" WE simulation algorithm to bypass the hurdles of optimizing binning procedures. Our proposed protocol (WeTICA) uses a low-dimensional CV space to drive the WE simulation toward the specified target state. We have applied this new algorithm to recover the unfolding kinetics of three proteins: (A) TC5b Trp-cage mutant, (B) TC10b Trp-cage mutant, and (C) Protein G, with unfolding times spanning the range between 3 and 40 {\mu}s using projections along predefined fixed Time-lagged Independent Component Analysis (TICA) eigenvectors as CVs. Calculated unfolding times converge to the reported values with good accuracy with more than one order of magnitude less cumulative WE simulation time than the unfolding time scales with or without a priori knowledge of the CVs that can capture unfolding. Our algorithm can be used with other linear CVs, not limited to TICA. Moreover, the new walker selection criteria for resampling employed in this algorithm can be used on more sophisticated nonlinear CV space for further improvements of binless WE methods.

Figures

Figures reproduced from arXiv: 2501.08926 by the authors.

Figure 1
Figure 1. Workflow diagram illustrating the WeTICA algorithm. III. Guidelines for choosing WeTICA simulation parameters Setting up a WeTICA simulation requires careful optimisation of certain input parameters to ensure efficient sampling of the productive trajectories. Here we provide some specific guidelines for choosing these parameters such as the number of walkers, generation and dimensionality of the CV space, walker mer… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

3 extracted references · 3 canonical work pages

  1. [1]

    The goal is to generalize the scope of REVO-based binless methods to a diverse class of problems, not limited to ligand binding/unbinding

  2. [2]

    2) Cloning and merging are decided based on projections of the walkers on a lower dimensional CV space, where we calculate the Euclidean distance d!

    Cloning and merging of walkers are decided based on ligand RMSD between a pair of walkers d!". 2) Cloning and merging are decided based on projections of the walkers on a lower dimensional CV space, where we calculate the Euclidean distance d!" between those projections of walkers and also their distance from the target conformation d#$%& ()$*+(!. 3) Vari...

  3. [4]

    How fast-folding proteins fold

    Clone that walker which is farthest from all other walkers. 4) Clone that walker which is closest to the target state conformation. S2 Table S2. Information related to all the independent WeTICA simulations. System Run Number of warping events Total number of cycles Total simulation time (ns) Trp-cage (TC10b) 1 311 1461 701.28 2 264 1460 700.80 3 113 1465...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.