REVIEW 6 cited by
ProteinBench: A Holistic Evaluation of Protein Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent years have witnessed a surge in the development of protein foundation models, significantly improving performance in protein prediction and generative tasks ranging from 3D structure prediction and protein design to conformational dynamics. However, the capabilities and limitations associated with these models remain poorly understood due to the absence of a unified evaluation framework. To fill this gap, we introduce ProteinBench, a holistic evaluation framework designed to enhance the transparency of protein foundation models. Our approach consists of three key components: (i) A taxonomic classification of tasks that broadly encompass the main challenges in the protein domain, based on the relationships between different protein modalities; (ii) A multi-metric evaluation approach that assesses performance across four key dimensions: quality, novelty, diversity, and robustness; and (iii) In-depth analyses from various user objectives, providing a holistic view of model performance. Our comprehensive evaluation of protein foundation models reveals several key findings that shed light on their current capabilities and limitations. To promote transparency and facilitate further research, we release the evaluation dataset, code, and a public leaderboard publicly for further analysis and a general modular toolkit. We intend for ProteinBench to be a living benchmark for establishing a standardized, in-depth evaluation framework for protein foundation models, driving their development and application while fostering collaboration within the field.
Forward citations
Cited by 6 Pith papers
-
Improving Protein Sequence Design through Designability Preference Optimization
A residue-level preference optimization method that uses AlphaFold pLDDT scores as rewards triples the in silico design success rate of LigandMPNN on enzyme benchmarks.
-
Designing Cyclic Peptides via Harmonic SDE with Atom-Bond Modeling
CpSDE generates cyclic peptides of all four cyclization types for a protein pocket by alternating a harmonic-SDE structure denoiser with a residue-type predictor.
-
Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design
AAMFM combines ESM3, an antigen-geometry adapter, and Cal-DPO preference optimization rewarded by AlphaFold3-style scores to design antibody CDRs and structures, reporting higher predicted binding scores than prior methods.
-
Protein-SE(3): Benchmarking SE(3)-based Generative Models for Protein Structure Design
Protein-SE(3) is a unified benchmark that retrains six SE(3) protein backbone generative models on the same data and evaluates them with identical metrics for designability, diversity, and novelty.
-
PFMBench: Protein Foundation Model Benchmark
A comprehensive benchmark of 17 protein foundation models across 38 tasks yields task correlations, a streamlined protocol, and identifies ProTrek as the strongest general performer.
-
AffinityFlow: Guided Flows for Antibody Affinity Maturation
AffinityFlow guides AlphaFlow structure generation toward low Rosetta binding energy, then inverse-folds the structures to propose antibody mutations, and reports top scores on a computational affinity maturation benchmark.
Discussion (0). Continue with ORCID to comment.