{"id":"4ecf57dd-151c-4085-8e7c-9dde3b4f2f27","arxiv_id":"2412.18275","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Protein flexibility can be predicted from sequence or structure, and a fine-tuned inverse folding model can be steered toward generating sequences with increased predicted flexibility.","lead":"This paper builds machine-learning predictors that estimate the flexibility of each residue in a protein, and then adapts a protein design model to generate sequences with requested flexibility in selected regions. The significance is a step toward making flexibility, not just shape, a controllable property in computational protein design.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The design result is evaluated with the same Flexpert-3D predictor used to build its training signal; without independent MD or experimental validation, the 1.52x enrichment may reflect predictor bias rather than physical flexibility.","rationale":"The reader's weakest assumption identifies exactly the load-bearing fragility: the Flexpert-Design evaluation is circular with respect to Flexpert-3D, because the same predictor provides the training pseudolabels (Eq. 1), the training loss (Eq. 2), and the evaluation metric (Eq. 6). The paper's own Appendix K explicitly concedes a possible overprediction bias, and the vanilla ProteinMPNN baseline already shows 61% flexibility-increasing mutations under this metric. Appendix I further shows structural destabilization in engineered regions, which is consistent with the predicted 'flexibility' being an artifact of misfolding rather than enhanced native dynamics. These self-reported limitations are in-scope evidence and should be weighed alongside the positive Table 4 results.\n\nThe prediction contributions are not in question: Flexpert-Seq and Flexpert-3D outperform baselines on ATLAS and are externally evaluated on mdCATH, with a reasonable performance drop at higher temperatures. The design contribution, however, is currently demonstrated only against a surrogate. This does not require rejection, because the concern is addressable: an independent MD-based evaluation, or a careful reframing of the claim to 'increased predicted flexibility,' would resolve it. The reader's CONDITIONAL verdict is therefore appropriate, and this stress-test does not move it.","tokens_in":20360,"tokens_out":3055,"duration_ms":28932,"concrete_test":"Run short MD simulations on a sample of designed and baseline sequences. Take, e.g., 100 CATH4.3 test proteins from Table 4; generate sequences with Flexpert-Design and vanilla ProteinMPNN under identical flexibility-increasing instructions (|S|=50, τ=5); for each, run 3×100 ns MD replicas with the same protocol as ATLAS, compute per-residue Cα RMSF, and recompute the median enrichment ratio and proportion of flexibility-increasing mutations using the MD RMSF of the native sequence as the reference. If the MD-based enrichment of Flexpert-Design is not significantly greater than that of ProteinMPNN, the headline claim should be weakened to 'increased predicted flexibility.' A secondary check: correlate Flexpert-3D predictions with MD RMSF on the same designed sequences to test for systematic overprediction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 1 — that inverse folding can be steered toward increased protein flexibility — rests on the Flexpert-Design evaluation in Section 5.2, where enrichment is measured by Flexpert-3D (Eq. 6). But Flexpert-3D is the same model used to generate native pseudolabels (Eq. 1) and to define the training loss (Eq. 2). The evaluation therefore tests whether the model can satisfy its own surrogate, not whether the designed sequences are physically more flexible.\n\nThe paper's own Appendix K acknowledges that Flexpert-3D 'might be slightly biased toward overpredicting flexibility since it predicts a flexibility increase even for the vanilla ProteinMPNN model'; Table 4 shows 61% of ProteinMPNN mutations are classified as flexibility-increasing, which is consistent with a systematic positive bias. Appendix I adds a second red flag: Flexpert-Design sequences have higher backbone RMSD (3.36 Å vs 2.00 Å) and much lower pLDDT in the engineered region (0.58 vs 0.81), so the predicted flexibility increase may reflect local unfolding or misfolding rather than native-like enhanced dynamics.\n\nUnless Flexpert-3D transfers to designed sequences, the 1.52 median enrichment ratio is not evidence about real protein flexibility.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the problem of engineering protein flexibility in computational protein design. The authors first compare several flexibility quantification methods (MD RMSF, PDB B-factors, AlphaFold2/ESMFold pLDDT, GNM, ANM) on the ATLAS dataset, concluding that MD-derived RMSF is the most reliable learning target and that ANM/GNM are the strongest structure-based baselines. They then introduce two flexibility predictors: Flexpert-Seq, which uses a ProtTrans protein language model with LoRA fine-tuning and a linear regression head, and Flexpert-3D, which additionally incorporates ANM features through a CNN adaptor. Both predictors are evaluated on held-out ATLAS test proteins and on the mdCATH dataset across simulation temperatures; Flexpert-3D achieves a Pearson correlation of 0.83 with MD RMSF, outperforming ANM (0.76) and approaching an estimated upper bound of 0.88. Finally, the paper proposes Flexpert-Design, a method that fine-tunes ProteinMPNN to accept per-residue flexibility instructions, using Flexpert-3D to generate pseudolabels for native sequences (Eq. 1), a flexibility-matching loss (Eq. 2), and a flexibility enrichment ratio for evaluation (Eq. 6). On CATH4.3, Flexpert-Design is reported to achieve a median enrichment ratio of 1.52 and 83% flexibility-increasing mutations, compared to 1.07 and 61% for vanilla ProteinMPNN.","tokens_in":20628,"tokens_out":5166,"duration_ms":44120,"significance":"If the design claim holds, the paper makes a useful contribution by demonstrating that inverse folding models can be conditioned on per-residue flexibility instructions, which is a novel capability with potential impact on enzyme engineering and protein design. The predictor contribution is solid and independently meaningful: the comparison of flexibility quantification methods is informative, and Flexpert-Seq and Flexpert-3D are carefully evaluated on ATLAS and mdCATH, showing favorable correlations to MD against several baselines. The authors also ship code and trained weights, which supports reproducibility. However, the central claim about steering inverse folding toward increased physical flexibility is currently supported only by the same surrogate model used to create the training signal, and the paper's own appendices reveal a likely positive bias and structural-destabilization confound. These issues are load-bearing for the paper's main advertised capability, so the design claim requires additional independent validation or a substantially more cautious framing.","major_comments":[{"comment":"The evaluation of Flexpert-Design is circular: the pseudolabels used to construct training instructions (Eq. 1), the flexibility-matching loss (Eq. 2), and the enrichment metric (Eq. 6) all use the same Flexpert-3D model. The reported median enrichment ratio of 1.52 therefore demonstrates that the fine-tuned model produces sequences that Flexpert-3D scores as more flexible, not necessarily that the sequences are physically more flexible. This concern is amplified by the vanilla ProteinMPNN baseline, which also yields 61% flexibility-increasing mutations under the same metric, and by the authors' acknowledgment in Appendix K that Flexpert-3D 'might be slightly biased toward overpredicting flexibility since it predicts a flexibility increase even for the vanilla ProteinMPNN model.' To support the central claim in Section 1 that inverse folding can be steered toward increased protein flexibility, the paper needs independent validation of designed sequences, such as MD simulations on a subset of designs, or agreement with a structurally orthogonal flexibility measure (e.g., ANM fluctuations on the designed structures or experimental B-factors).","section":"Section 4.2, Eqs. (1)-(3) and Section 5.2, Eq. (6)"},{"comment":"The structure preservation analysis raises a serious confound. Flexpert-Design sequences have substantially higher Cα RMSD to the ground-truth backbone (3.36 Å vs 2.00 Å for ProteinMPNN) and a pronounced pLDDT drop in the engineered region (0.58 vs 0.81). This pattern is consistent with local unfolding or misfolding rather than native-like enhanced dynamics, which would confound the interpretation of the enrichment ratio. The authors should either stratify the evaluation by structural integrity (e.g., pLDDT or RMSD thresholds) to show that the flexibility increase is not an artifact of destabilization, or substantially temper the claim that the method 'engineers flexibility' in the sense of native dynamics rather than partially unfolding the engineered region.","section":"Section 5.2 and Appendix I, Table 9"},{"comment":"The median enrichment ratios and proportions are reported without confidence intervals or significance tests. Given the small difference between the full Flexpert-Design model (1.52) and its without-loss variant (1.43), and the high baseline proportion for ProteinMPNN (61% flexibility-increasing mutations), it is unclear whether the differences are statistically meaningful. The authors should report bootstrap confidence intervals across CATH4.3 test proteins or paired per-protein tests, and should also discuss whether the improvement over the no-loss variant justifies the additional complexity of the flexibility-matching loss.","section":"Section 5.2, Table 4"}],"minor_comments":[{"comment":"The abstract and introduction state that the method 'demonstrate[s] that inverse folding models can be steered toward' increased flexibility without the caveat 'as measured by our protein flexibility predictor,' which appears only in the conclusion. Please align these statements to avoid overclaiming in the abstract.","section":"Abstract and Section 1 vs. Section 6"},{"comment":"Please specify why 7 of the 1390 ATLAS proteins were skipped, since the text says 'some were skipped due to missing pieces of structure resulting in NaNs from ENMs' but does not quantify the number or give criteria.","section":"Section 3.2, Table 1"},{"comment":"The notation 'LF lexpert' is inconsistent and appears with various spacing; please use a single consistent subscript, e.g., L_flex, throughout the paper.","section":"Section 4.2, Eq. (3) and surrounding text"},{"comment":"The term 'topology splitting' is used without definition or reference; please provide details in the experimental setup or point to an appendix that explains how topologies are split and how leakage is prevented.","section":"Section 5.1"},{"comment":"The right panel of Figure 8, which shows the effect of segment length, does not have labeled axes in the text description; please clarify the x-axis and y-axis in the caption.","section":"Appendix J, Figure 8"}],"recommendation":"major_revision","confidential_remarks":"The circularity concern is the main issue and it is substantial, but the paper's predictor contribution is concrete and the code is available, so I believe the design claim can be fixed within revision scope by adding independent MD validation on a subset of designed sequences or by explicitly reframing the claim as surrogate-level steerability with supporting evidence that the surrogate transfers to designed sequences. The paper is likely to be salvageable as a meaningful contribution to the protein design literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a read for the flexibility predictor work; the Flexpert-Design claim, as the stress-test note says, is only as good as its own surrogate. The predictors Flexpert-Seq and Flexpert-3D are solid: they beat ANM/GNM and pLDDT baselines on ATLAS, generalize reasonably to mdCATH at near-physiological temperature, and the systematic comparison of B-factor, pLDDT, GNM, ANM against MD is a good contribution. The architecture makes sense — PLM embeddings plus ANM correction, trained on MD RMSF — and the code and weights are released.\n\nThe soft spot is exactly where the authors make their headline claim. Flexpert-Design fine-tunes ProteinMPNN using Flexpert-3D pseudolabels as ground truth, and then evaluates the designed sequences with the same Flexpert-3D. The 1.52x median enrichment ratio and 83% flexibility-increasing mutations may partly be the model learning to satisfy its own surrogate. The paper's own numbers support this worry: vanilla ProteinMPNN already produces 61% flexibility-increasing mutations under the same metric (Table 4), and Appendix K concedes Flexpert-3D 'might be slightly biased toward overpredicting flexibility.' Appendix I shows the engineered sequences have higher backbone RMSD (3.36 Å vs 2.00 Å) and much lower pLDDT in the engineered region (0.58 vs 0.81), which is consistent with local unfolding rather than native-like flexibility. So the physical claim is not established.\n\nThat said, the circularity is not fatal for the whole paper. The prediction contributions stand on their own, and the design pipeline is a plausible method for steering inverse folding; what is missing is independent validation — MD on a sample of designed sequences, or at least careful correlation against a held-out predictor. The authors should also report error bars on the enrichment metrics and clarify whether the claim is about the surrogate or about physical flexibility.\n\nFor a reading group, this is a useful case study in evaluation design. I would not cite the design claim without a caveat, but I would cite the predictor work. A serious referee should engage with this; the prediction half is sound, and the design half is a testable hypothesis that the authors can strengthen.","headline":"Useful flexibility predictors, but the design steering claim is only validated against its own surrogate.","tokens_in":21175,"tokens_out":1891,"would_cite":true,"duration_ms":16251,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that protein inverse folding models can be steered toward increased flexibility by conditioning on per-residue flexibility predictions from a learned predictor.","keywords":["protein flexibility","inverse folding","protein design","molecular dynamics","RMSF","ProteinMPNN","Flexpert","deep learning"],"falsifier":"Take a sample of Flexpert-Design-generated sequences from CATH4.3, run new atomistic molecular dynamics simulations on both native and designed proteins under identical conditions, and compute per-residue RMSF; if the MD-based median enrichment ratio in the engineered regions falls to roughly 1, the steering effect is an artifact of the predictor rather than real flexibility change.","tokens_in":20171,"feed_emoji":"🧬","tokens_out":5262,"duration_ms":45468,"temperature":0.7,"pith_summary":"The paper aims to make protein flexibility an explicit, controllable input to generative protein design. It first benchmarks ways of quantifying flexibility and settles on molecular-dynamics-derived root mean square fluctuations as the learning target. It then trains fast predictors, Flexpert-Seq from sequence alone and Flexpert-3D from sequence plus backbone structure, using a pre-trained protein language model to overcome limited data. Finally, it introduces Flexpert-Design, which fine-tunes an inverse folding model to accept flexibility instructions and generate sequences with increased flexibility in specified regions. If correct, this would let protein engineers request flexible loops, tunnels, or other regions directly, without slow molecular dynamics simulations in the design loop.","feed_headline":"Protein design can be steered by flexibility instructions","feed_subtitle":"Fine-tuned inverse folding follows flexibility targets, lifting median enrichment to 1.52 on CATH4.3.","key_machinery":"The load-bearing component is Flexpert-3D: a protein language model (ProtTrans) with LoRA fine-tuning and a linear regression head, augmented by a CNN adaptor that injects ANM-computed flexibility values into the embedding space, so the model learns to correct crude ANM estimates toward MD ground truth. Flexpert-Design then wraps this predictor around ProteinMPNN: flexibility instructions are added as zero-initialized node features, sequences are sampled with a straight-through Gumbel-Softmax estimator, and the sampled sequence is passed through Flexpert-3D so that a flexibility-matching loss can be backpropagated while sequence cross-entropy loss keeps the inverse folding ability intact.","core_discovery":"The central claim is that per-residue protein flexibility can be predicted quickly and then used as a conditioning signal for inverse folding. Flexpert-3D, which combines the sequence-based Flexpert-Seq with an Anisotropic Network Model corrected by a small convolutional adaptor, reaches a Pearson correlation of 0.83 to MD-derived flexibility on the ATLAS test set. Flexpert-Design takes a ProteinMPNN inverse folding model, adds flexibility instructions as input node features, and fine-tunes it with a loss that matches the flexibility of sampled sequences to the instructions using Flexpert-3D as the evaluator. On CATH4.3, the resulting model achieves a median flexibility enrichment ratio of 1.52 and 83% flexibility-increasing mutations, compared with 1.07 and 61% for the vanilla ProteinMPNN baseline.","pith_inferences":["Inference: The steering result is measured by the same Flexpert-3D predictor that generated the training pseudolabels, so an independent check with new molecular dynamics simulations of designed sequences would be needed to confirm that the enrichment reflects true conformational flexibility rather than predictor bias.","Inference: The increased proportion of glycine and alanine in engineered segments is biochemically plausible but also points to a possible shortcut; testing whether the model still raises flexibility when these small residues are disallowed would clarify whether the signal is structural or residue-type-driven.","Inference: The same conditioning-and-loss loop could plausibly be applied to other inverse folding backbones, such as flow-matching or diffusion-based design models, and to other per-residue properties such as stability or solubility."],"forward_implications":["Protein engineers can request increased flexibility in a designated contiguous region while keeping the backbone fixed, and the redesigned sequences show higher predicted flexibility in that region.","The flexibility predictors are fast enough to be embedded in iterative design pipelines, unlike the molecular dynamics simulations used to generate their training labels.","Fine-tuning with the flexibility loss retains sequence recovery within about one percentage point of the ProteinMPNN baseline, so the steering does not come at a large cost to inverse folding quality.","The same procedure is not effective for decreasing flexibility, so the method currently provides one-directional control over protein flexibility.","The flexibility signal is continuous and per-residue, which makes the training loop adaptable to other per-residue design objectives beyond flexibility."],"supporting_citations":[{"why":"Supplies the ATLAS dataset of molecular dynamics trajectories and per-residue RMSF values used as the flexibility learning target.","marker":"Vander Meersche et al., 2023"},{"why":"Provides the ProteinMPNN inverse folding model that Flexpert-Design modifies and fine-tunes.","marker":"Dauparas et al., 2022"},{"why":"Supplies the ProtTrans pre-trained protein language model whose embeddings form the base of Flexpert-Seq and Flexpert-3D.","marker":"Elnaggar et al., 2022"},{"why":"Provides LoRA low-rank adaptation, the parameter-efficient fine-tuning method used for ProtTrans.","marker":"Hu et al., 2021a"},{"why":"Provides the Gumbel-Softmax straight-through estimator that keeps the Flexpert-Design training pipeline differentiable through discrete sequence sampling.","marker":"Jang et al., 2017"},{"why":"Provides the CATH4.3 dataset used for training and evaluating the inverse folding models and for the flexibility engineering experiments.","marker":"Pearl et al., 2003"},{"why":"Provides the ProDy implementation of elastic network models used to compute ANM-based flexibility inputs for Flexpert-3D.","marker":"Bakan et al., 2011"},{"why":"Provides the mdCATH dataset used to test how the flexibility predictors generalize to different simulation temperatures and protein domains.","marker":"Mirarchi et al., 2024"}],"fun_headline_variants":["Flexpert guides protein design to follow flexibility goals","Flexibility instructions steer inverse folding in protein design","Predict flexibility, then design proteins to match it","Flexpert-Design tunes inverse folding to boost flexibility"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire steering result rests on the assumption that Flexpert-3D's flexibility predictions stay accurate for mutated, designed sequences rather than systematically overpredicting flexibility for any design change.","fun_headline_variants_meta":{"raw":{"variants":["Flexpert guides protein design to follow flexibility goals","Flexibility instructions steer inverse folding in protein design","Predict flexibility, then design proteins to match it","Flexpert-Design tunes inverse folding to boost flexibility"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000241,"raw_usage":{"total_tokens":1491,"prompt_tokens":887,"completion_tokens":604,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":544}},"tokens_in":503,"tokens_out":604,"duration_ms":6076,"temperature":1.0,"reasoning_tokens":544,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:50:39.840207+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of Flexpert-Design-generated sequences from CATH4.3, run new atomistic molecular dynamics simulations on both native and designed proteins under identical conditions, and compute per-residue RMSF; if the MD-based median enrichment ratio in the engineered regions falls to roughly 1, the steering effect is an artifact of the predictor rather than real flexibility change.","supporting_citations":[],"review_version":1}