REVIEW 1 cited by
Exploring Large Protein Language Models in Constrained Evaluation Scenarios within the FLIP Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this study, we expand upon the FLIP benchmark-designed for evaluating protein fitness prediction models in small, specialized prediction tasks-by assessing the performance of state-of-the-art large protein language models, including ESM-2 and SaProt on the FLIP dataset. Unlike larger, more diverse benchmarks such as ProteinGym, which cover a broad spectrum of tasks, FLIP focuses on constrained settings where data availability is limited. This makes it an ideal framework to evaluate model performance in scenarios with scarce task-specific data. We investigate whether recent advances in protein language models lead to significant improvements in such settings. Our findings provide valuable insights into the performance of large-scale models in specialized protein prediction tasks.
Forward citations
Cited by 1 Pith paper
-
EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?
A new closed-book, sequence-only benchmark, EpiBench, measures epitope reasoning in LLMs and finds them near chance on residue-level localization and escape assessment, with only coarse region-level signal.
Discussion (0). Continue with ORCID to comment.