claims depot shelf
Protein Folding
Formal claims (Lean)
Stated claims
-
The paper's load-bearing assertion is that 'the objective of protein structure prediction must shift from recovering the coordinates of a single conformation to inferring a conformational state space' (Section I). If the paper is correct, the field's success metric changes: a predictor is judged by whether it recovers accessible states, their populations, transition pathways, and their dependence on biological context, not by coordinate accuracy on a dominant structure.
-
On its own terms, the paper establishes a modeling architecture for unified scientific intelligence: a vision-language backbone forms task-conditioned hidden states from an instruction and a native scientific object; structured reasoning supervision ('natural-world knowledge alignment') shapes those states so they encode evidence about composition, symmetry, contacts, peaks, or spatial relations; and a task token selects a decoder that produces an output verifiable in the domain's native form. The paper's empirical claim is that this single model outperforms leading closed general-purpose models on most of 66 evaluated tasks and matches or exceeds specialist models on a meaningful subset—not
-
On 101 experimentally resolved binding-pocket peptides (5–18 residues), fully executed on IBM Heron R2, non-iterative multi-β sampling of a fixed amino-acid-level tetrahedral-lattice Hamiltonian yields median Cα RMSD near 2.7 Å and improves accuracy by 27–71% over evaluated AI and quantum baselines, while recovering ground-state energy under noise several times typical hardware rates and reducing mean quantum execution time by about 27× relative to VQE.
-
Using lattice-based quantum protein structure prediction as a case study, the contact-energy cost Hamiltonian is not sufficiently aligned with structural accuracy as measured by RMSD against experimentally determined structures. For small peptides and on average, the energy landscape of the considered cost Hamiltonian is not correlated well enough to the actual error to provide meaningful predictions. The correlation increases for larger problem instances and when more interaction shells are considered, as estimated through Monte-Carlo sampling.
-
EasyNano optimizes CDR residue logits via gradient descent through the ESMFold2 pairwise distance distogram, using the lightweight ESMFold2-Fast model as a differentiable oracle guided by a composite loss including a dedicated epitope proximity term. A full ESMFold2 CA-coordinate structure prior prevents framework pose drift. Across six target-framework pairs the procedure improves ipTM by up to +0.559 while preserving ipTM on already-strong binders.
-
EpiFormer is a general encoder-decoder framework whose central mechanism is interleaved cross-attention placed inside GNN encoding layers. This design produces bidirectional antigen-antibody information flow at every stage of representation learning instead of only at the output. The same early-fusion principle works across different GNN backbones and becomes especially effective when combined with sparsity-aware objectives. On standard benchmarks the resulting model exceeds the previous best method by more than 40 percent in F1 score, shows cross-dataset transfer, and produces attention patterns and feature preferences that match known biology without any explicit supervision on those prope
-
AIMS-Fold is an inference-time guided-diffusion framework that converts XL-MS spatial restraints and HDX-MS solvent accessibility profiles into differentiable physical potentials derived from structural proteomics measurements; these potentials steer the generative sampling trajectory of pretrained diffusion models and yield higher accuracy on induced proximity targets than purely computational unguided models.
-
DCFold is a single-step generative model that attains AlphaFold3-level accuracy on all-atom protein structures by means of a Dual Consistency training framework that incorporates a Temporal Geodesic Matching scheduler, delivering 15 times faster inference while preserving predictive fidelity on structure prediction and binder design benchmarks.
-
The central claim is that iterative, feedback-guided refinement of the Hamiltonian's three constraint-penalty coefficients—chirality, backbone-geometry, and local-overlap terms—produces directed improvement that independent sampling does not. The agent pipeline never sees ground-truth RMSD; it adjusts penalties based on VQE energy trajectories and structural-validation metrics, and over three to five cycles the normalized energy-to-penalty ratio E/P decreases significantly in most agent configurations while remaining flat for LLM-only controls. Structural validity on unseen sequences rises from 87.5% to 98.7%, and 87% of initially invalid candidates are recovered by best-of-three selection;
-
ConforNets consist of channel-wise affine transforms applied specifically to the pre-Pairformer pair latents inside the AF3 architecture. These transforms globally modulate the model's internal representations in a manner reusable across different proteins. The method attains state-of-the-art success rates for unsupervised generation of alternate conformational states on all existing multi-state benchmarks. When trained in a supervised setting on one source protein, the same transforms induce a conserved conformational change across an entire protein family while preserving overall structure prediction accuracy.
-
On the paper's own terms: ESMFold folds a beta hairpin in two causally separable stages. During early blocks (0–7), residue identity and biochemical features such as charge flow from the sequence representation into the pairwise representation through the seq2pair pathway; sequence patching is effective only there. During late blocks (roughly 25 onward), the pairwise representation accumulates distance and contact information, modulates sequence attention through the pair2seq bias, and directly controls the structure module; pairwise patching is effective only there. The causal interventions are charge steering, distance steering, attention redirection, and pairwise scaling. The abstract ext
-
In twisted dipolar clusters of magnetic rods forming polygons, the relative twist angle induces noncollinear chiral magnetic phases ranging from vortex-like flux closure to radial hedgehog configurations. Chirality quantified by a bond order parameter behaves as an Ising variable, while a clock index rooted in the C_N symmetry of the polygons distinguishes different chiral textures within the same sector. As twist increases, the competition between continuous phase shift and discrete anisotropy creates a tilted N-fold energy landscape whose minimum switches discontinuously between clock sectors, with the response becoming nearly U(1)-invariant for large site numbers.
-
The central claim is that generative AI has moved from a niche technique to a broadly applicable tool in bioinformatics, with the review's evidence showing that specialized model architectures generally outperform general-purpose models on biological sequence and structure tasks. Across six research questions, the authors find that GenAI supports sequence analysis, molecular design, and integrative data modeling; that domain-specific pretraining and context-aware tokenization drive performance gains; that protein structure prediction, functional annotation, and synthetic data generation have advanced substantially; and that a diverse set of molecular, cellular, and textual datasets enables t
-
The central claim is that the Evoformer's attention-based operations can be faithfully expressed as the vector field of an ordinary differential equation, so that the 48 discrete layers become a discretization of a continuous-time model solved by Neural ODE techniques; this yields constant memory cost in depth, a tunable runtime-accuracy tradeoff via solver choice, and protein predictions that remain structurally plausible while capturing certain secondary structures such as alpha-helices.
-
The central claim is that AlphaFold 3 embodies a paradigm shift toward differentiable simulation. Its multi-scale transformer architecture, biologically informed cross-attention, and geometry-aware optimization make every predicted coordinate a differentiable function of the input, so gradients computed through the network can be used to move a structure around its predicted landscape. The authors argue this turns structure prediction into a foundation for dynamic molecular simulation: the same model that predicts a folded state can, through its gradients, suggest how that state responds to perturbation or evolves in time. The paper presents this as a reframing—structure as a point on a diff
-
The central discovery is a sampler reconfiguration, not a retrained model: keeping $\gamma_0=0$ and setting $\eta=1.0$ turns the AF3 EDM sampler into a pure ODE, and the resulting two-step trajectory yields complex LDDT of 0.822 versus 0.820 for the 200-step baseline on RecentPDB proteins under 768 tokens. The claim is that AF3-style models, whether trained with EDM or flow matching, are inherently robust to drastically reduced sampling steps once the noise injection is removed and the step scale is corrected. The paper then shows the same robustness carries over to a compact architecture, Protenix-Mini, which drops redundant early pairformer blocks and uses one MSA block, producing 1-5% lower performance on benchmarks while cutting FLOPs from 93 to 20.
-
On the paper's own terms, the central claim is that fold-switching proteins challenge the classical expectation that a globular protein's sequence encodes a single fixed fold. The review reports roughly one hundred experimentally characterized fold switchers, estimates that up to 4% of proteins in the Protein Data Bank and up to 5% of E. coli proteins may switch folds, and argues that dual-fold coevolutionary signals show both conformations of many switchers are under selection. It also argues that fold switching can be an evolutionary end point, as in XCL1 and RfaH, or an evolutionary intermediate, as in the stepwise helix-turn-helix to winged-helix transition seen in bacterial response regulators. The review treats the emergence of these proteins as evidence that fold space is more fluid than the one-sequence-one-structure doctrine implies, and that AI-based structure predictors often fail on fold switchers because they memorize training-set structures rather than inferring alternative folds from coevolution.
-
On the paper's own terms, the discovery is that the FCC lattice can be encoded for quantum protein structure prediction without slack variables, and that the resulting Hamiltonian's ground state — located by two distinct variational routes on noisy hardware and decoded back to coordinates — matches the lowest-energy conformation found by classical exhaustive search for the six-residue peptide KLVFFA. The authors argue the FCC lattice is not just a bigger version of earlier lattices but a qualitatively better model of biology: across eleven test proteins spanning helices, sheets, and loops, FCC best fits reach RMSDs below 2.0 Å in every case and several near sub-angstrom, whereas tetrahedral fits stretch helices to a Cα pitch of roughly 9 Å against nature's 5.4 Å. The quantum construction is a four-term Hamiltonian built from turn-indicator polynomials — backtracking penalties, penalties on the four redundant bitstrings, a nearest-neighbor interaction term scored with Miyazawa–Jernigan contact energies through ancilla flags, and an overlap term — where the overlap constraint is enforced either by fitting a polynomial to the ideal penalty functional (PolyFit) or by converting the constraint into a Lagrangian saddle-point problem (VQEC). Both schemes keep the configuration encoding at $4N-10$ qubits (24 total for KLVFFA, including contact ancillas), and on hardware both assign the highest sampled probability to a ground-state turn sequence, a claim the authors verify by comparing the decoded conformations against the classical ground-truth ensemble.
-
On the paper's own terms, the discovery is that a coarse-grained quantum optimization—each residue mapped to a tetrahedral lattice node, the conformational energy encoded as a four-term Hamiltonian, and the ground state found with a variational quantum eigensolver on a real superconducting processor—produces fragment structures that beat the two dominant deep-learning predictors on the two metrics that matter for docking. Compared with experimentally determined X-ray structures, the quantum fragments give lower root-mean-square deviation (RMSD) of backbone carbon positions in 51 of 55 cases against the older deep-learning model and 40 of 55 against the newest one; compared with the same deep-learning models, docking against native ligands gives lower (more favorable) binding-affinity scores in 53 of 55 and 50 of 55 cases, respectively. The paper reads these results as evidence that quantum-first modeling, grounded in physical energy minimization rather than training-data statistics, can handle short ligand-binding fragments better than data-driven approaches.
-
The paper's central discovery is that the sum-of-pairs score of an alignment can be evaluated exactly through a classical query function $f_i(k)$ defined on the qubit bitstring, which maps each 1 to the index of the corresponding letter in the original sequence by counting preceding 1s, and maps 0s and excess 1s to a dummy index. Substituting this query into the SP-score yields the hybrid loss $L(x;p)$ of Eq. (8), whose ground state corresponds to the optimal alignment. Because the encoding is positional rather than one-hot, the required qubits fall to $O(NL)$. On a 37-qubit ion-trap quantum computer, 2000-shot measurements of the variational state reproduce the optimal alignment with significant probability at 8, 12, and 16 qubits, and the experimental time per iteration grows polynomially, while a noisy classical simulation grows exponentially, placing a crossover around 22 qubits. The authors state that this is the largest digital simulation on a trapped-ion quantum computer for a life-science problem at the time of writing.
-
The paper's central claim is that the Transformer's application to protein informatics has reached a state where a comprehensive, domain-oriented synthesis is both possible and needed, and that this synthesis shows the architecture reshaping the field. It surveys more than 100 studies and organizes them into four application domains: protein structure prediction, protein function prediction, protein–protein interaction analysis, and drug discovery/target identification. It further claims that the Transformer's self-attention mechanism explains the gains—because protein sequence, structure, and function depend on distal residue interactions—and that pre-trained protein language models such as AlphaFold, ESM-Fold, ProtTrans, and ProteinBERT are the vehicles through which those gains appear. The paper also asserts that curating datasets and code repositories is an essential part of the contribution, since reproducibility and benchmarking depend on them, and that future progress will come from multimodal integration, hybrid physics-informed modeling, efficiency improvements, and interpretability.
-
AlphaFold's potential energy function, parameterized by deep models, implements probability kinematics by using distance information as uncertain evidence to update a prior over structures. This process explicitly defines a posterior distribution, generalizing standard Bayesian updating to cases where evidence is not certain. The synthetic angular random walk example shows how the update works in a tractable setting without the complexity of real proteins.
-
The discovery is an application-level integration: a PyQt5-based graphical user interface that calls the Python libraries of Boltz-1, Chai-1, and Protenix for local folding, and the ESM3 Forge API for remote folding, then loads the returned coordinate files as ordinary PyMOL objects. For file-based models the plugin prepares temporary FASTA or JSON inputs with chain and sequence metadata; for library-based models it passes the sequence directly to the folding function. With Boltz-1 and Chai-1 a SMILES string can be included so the small molecule is placed in the predicted complex. Once loaded, the structure is interactive and customizable like any PDB-derived structure, and can be colored by AlphaFold-style confidence.
-
The paper's central claim is that protein language models “skillfully grasp the foundational knowledge of proteins and can be effectively generalized to solve a diversity of sequence-structure-function reasoning problems.” The evidence it assembles is a taxonomy: sequence-only pLMs (ESM-2, ProtGPT2, xTrimoPGLM) capture evolutionarily favored amino acid patterns; structure- and function-enhanced pLMs (SaProt, ESM-3) add explicit 3D and annotation knowledge; multimodal pLMs (ProLLaMA, BioT5) bridge protein sequences with natural language and molecule languages. On top of these foundations, the paper reports pLM-based single-sequence structure prediction comparable to MSA-based methods, zero-shot fitness and mutation-effect prediction, text-guided and condition-tagged protein generation, and ChatGPT-like protein question answering. The intended conclusion is that pLMs, not bespoke per-task models, now carry the main line of computational protein science.
-
The paper's central claim is that the forward noising process of a diffusion model should be identified with an exact renormalization-group coarse-graining. For data modeled as a field $\phi$, a cutoff function $K_t(k)$ with a chosen regulator $r(x)$ determines how much each Fourier mode is erased at time $t$: $\bar{\alpha}_{tk}=K_t(k)$ and $\bar{\beta}_{tk}=k^{-2}(1-K_t(k))$. Because the exact RG guarantees scale separation, the erased high-wavenumber modes become Gaussian and independent of the retained modes, so the model can discard them entirely during training and sampling. Reversing this flow generates data coarse-to-fine, and the empirical claim is that this consistently outperforms the standard DDPM, which denoises white noise on all modes at once, on protein structure prediction (RMSD, TM-score, GDT-TS, GDT-HA) and image generation (FID on CIFAR-10 and FFHQ), often improving quality and/or cutting the required steps by an order of magnitude. Since the schedule follows from the RG once a regulator is specified, the model removes the heuristic, data-dependent tuning of the noise schedule.
-
The paper's central assertion is that the three canonical problems of protein bioinformatics form a directed cycle: sequence determines structure, structure and sequence determine function, and design is the inverse map from function or structure back to sequence. Within that frame, the paper claims that the strongest recent design results did not come from better energy functions or fragment libraries but from reusing deep-learning predictors—most notably trRosetta and AlphaFold2—either as scoring oracles in generative models or as differentiable networks that can be run backward by gradient ascent to 'hallucinate' sequences for a desired structure or motif. It treats this inversion as the main explanation for why design methods have improved since CASP13, while acknowledging that in silico performance has repeatedly failed to survive in vitro synthesis. The survey also claims the structure-first, function-second ordering is useful because functional labels are sparser than structural data, making function the weakest link in the cycle.
-
The central claim is that the two architectures make Deep Q-Networks competitive with dedicated HP-model solvers: FFNN-R, a fully connected network fed by a fixed random reservoir, reaches best-known energies on most short benchmark sequences (A1–A5 and A8–A10) with roughly 25% fewer episodes than the vanilla FFNN baseline, and LSTM-A, an LSTM with multi-head attention, matches best-known energies on the 3d1–3d4 sequences and comes within a few hydrophobic contacts of best-known values on 3d5–3d9. The paper takes this as evidence that implicit temporal memory (reservoir) and long-range attention (LSTM-A) are the ingredients that let DQN scale on the 3D HP model.
-
Generative AI, led by diffusion models and transformer architectures, has enabled significant breakthroughs in medical imaging (including image reconstruction, image-to-image translation, generation, and classification), protein structure prediction, clinical documentation, diagnostic assistance, radiology interpretation, clinical decision support, medical coding, and billing, as well as drug design and molecular representation. These innovations have enhanced clinical diagnosis, data reconstruction, and drug synthesis.
-
The suggested hybrid algorithm is assessed on the quality of produced solutions and compared with state-of-the-art algorithms on well-studied benchmark instances for the PSP problem.