REVIEW 1 major objections 4 minor 34 references
Is Retrieval All You Need? Assessment and Emergence of Novelty in Protein Structure Generation
T0 review · 1 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Low full-chain novelty does not prove a protein generator invented a new fold.
desk verdict A solid evaluation paper that makes a real point about full-chain novelty, but its calibrated-threshold claim is over-generalized from domain-level to chain-level queries. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is RetFold, a zero-training retrieval-and-assembly baseline that constructs backbones by retrieving three fragments from CATH S40 (one core domain at 40-60 percent of target length and two extensions at 15-35 percent each), using an ensemble of ESM-2, Foldseek 3Di, and ProstT5 embeddings, then connecting them with idealized alpha-helix linkers optimized for geometry and diversity. Because RetFold's composition is known by construction, it serves as a positive control for any retrieval criterion. The paper also uses DRR as the diagnostic metric: it decomposes generated backbones with Merizo and asks whether any domain matches a CATH S40 domain at aligned-length TM above a threshold. The argument hinges on calibrating three TM-score normalizations: query-normalized qTM (which carries a length ceiling), aligned-length alnTM (which has high false positive rates under leave-one-topology-out), and target-normalized tTM (the only one that passes both positive and negative controls at threshold 0.7).
What would settle it
A direct test would be to build a chain-sized negative control: take RetFold's 500 chains, remove the specific CATH domains used to construct each chain from the reference set by structural exclusion (not by topology label), and measure how often each score retrieves the remaining structure above threshold. If reference-normalized tTM at 0.7 shows a false positive rate well above 5.38 percent on such chains, the paper's claim that only tTM at 0.7 supports a novelty reading would be weakened.
Extended reading notes
Core claim
The paper's central claim is that apparent full-chain novelty coexists with high domain-level retrievability, and that the conventional protocol of reporting low full-chain TM-score against a structural database does not distinguish a genuinely new fold from a novel assembly of known domains. It demonstrates this with the Domain Retrieval Rate (DRR), which finds that a majority of generated backbones from eight models contain domains aligned to known CATH S40 entries, even when full-chain retrieval rates are near zero. The paper then introduces RetFold, a training-free pipeline that retrieves CATH domains and assembles them into chains with helix linkers; because its composition is known by construction, any criterion applied to it has a known answer. The conventional full-chain query-normalized protocol retrieves RetFold at only 20.0 percent, while the reference-normalized, coverage-controlled protocol retrieves it at 100.0 percent even at a threshold of 0.9. Against a leave-one-topology-out negative control, only the reference-normalized TM-score reaches a false positive rate below 10 percent (5.38 percent at threshold 0.7), while aligned-length TM retrieves 90.04 percent of queries whose fold class has been removed from the searched set. Therefore, the paper concludes, the retrieval envelope reported for learned generators does not by itself certify fold invention, and novelty claims should be calibrated against positive and negative controls.
Load-bearing premise
The key assumption is that removing a query's CATH topology label from the searched reference set removes all true matches, so that any hit above threshold in that control is a false positive; but because the exclusion is a discrete label rather than a structural region, cross-topology hits at TM above 0.5 include genuine similarity, and the authors state their false positive rates are upper bounds.
Editorial extensions
If this is right
- Novelty rates in protein structure generation should be reported with the calibration of the score used, including false positive rates on a negative control whose fold class is absent from the reference set.
- A generated backbone that fails full-chain retrieval is not evidence of a new fold; it may be a recombination of known domains or a distorted version of a known fold.
- The domain–full-chain gap, DRR minus FC, should not be read as a measure of structural content unless the reference value under a true-negative control is subtracted.
- Reference-normalized TM-score with coverage control, at a threshold of 0.7, separates a recombined baseline from learned unconditional generators, suggesting a protocol that can distinguish recombination from generated novelty.
- Designability screens can pass near-identical retrieved copies if the reference database contains close homologs, so reported pass rates should be accompanied by a matched trivial-baseline comparison.
Reading between the lines
- A testable extension is to run a chain-sized negative control that excludes near matches by structural exclusion rather than by discrete CATH topology labels, which would directly calibrate the false positive rates for full-chain queries.
- The paper implicitly suggests a 'novelty calibration report' standard: every novelty rate should be accompanied by the rate the same protocol assigns to a known-recombined positive control and a known-absent negative control.
- Because RetFold's 20 percent full-chain retrieval comes from its large verbatim core domain, the result may extend to other fragment-assembly methods: any generator that uses long unmodified reference fragments will be misclassified as novel by qTM for length reasons alone.
- A natural next experiment is to apply the calibrated tTM criterion to conditional binder generation in documented inference mode, where the paper's current task-matched cohorts were run under target-free ablations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that low full-chain similarity to known proteins is insufficient evidence of fold-level novelty in protein backbone generation. It introduces the Domain Retrieval Rate (DRR), which decomposes generated backbones into structural domains and measures retrieval against the CATH S40 domain database, and proposes RetFold, a zero-training baseline that assembles known CATH domains into recombined chains via geometry-based helix-linker optimization. The authors calibrate three structural-similarity scores (alnTM, qTM, tTM) against RetFold as a positive control with known composition and against a leave-one-topology-out negative control, concluding that only reference-normalized tTM at a threshold of 0.7 supports reading a retrieval rate as a novelty estimate. Across eight backbone generators, they report that most outputs contain retrievable domains even when full-chain retrieval is low, and that RetFold lands inside the retrieval envelope of learned generators at roughly two orders of magnitude lower cost.
Significance. If the chain-level calibration is established, the paper makes a significant methodological contribution. It demonstrates a granularity mismatch in existing novelty evaluation, provides a strong positive control with known composition, and quantifies score-specific false-positive rates. The paper is unusually thorough: it includes multiple segmentations (Merizo and Chainsaw), quality filtering, per-length analysis, sensitivity scans over thresholds and coverage rules, and explicit self-flagged limitations. The RetFold baseline is a creative and useful diagnostic, and the paper's core message—that full-chain novelty is not evidence of fold invention—is well supported even independently of the specific calibration threshold.
major comments (1)
- [Appendix G, Table 3, Abstract] The central prescriptive claim—that reference-normalized tTM at τ=0.7 is the only score that supports reading a retrieval rate as a novelty estimate—rests on a negative control measured on domain-sized queries. Table 2 reports a 5.38% FPR for tTM at τ=0.7 on CATH S40 domains after leave-one-topology-out exclusion, and Appendix G explicitly states: 'we therefore do not use them to judge whether a given chain-sized retrieval rate exceeds chance.' Yet Table 3 applies the criterion to full-chain queries, and the abstract and conclusion state the 5.38% figure and the 'only one of the three' conclusion without the domain/chain qualifier. A full chain has more surface area and more opportunities to partially cover a reference domain, so its tTM@0.7 FPR could be substantially higher than the domain-level 5.38%. The positive control (RetFold, 100% saturated at tTM≥0.7) demonstrates sensitivity on chains but says nothing about specificity on chains. Without a chain-level negative control—for example, assembling CATH domains into chimeric chains with known non-matching topologies—the calibrated separation of RetFold from the unconditional generators in Table 3 is not established. The authors should either add such a control or restrict the conclusion to domain-level queries.
minor comments (4)
- [Abstract and Section 'A zero-training baseline exposes the resolution limit'] The 5.38% figure is described as a false positive rate, but Appendix G defines it as an upper bound because the leave-one-topology-out exclusion removes a discrete label rather than a region of structure space; the text should consistently say 'at most 5.38%' or 'upper bound of 5.38%'.
- [Appendix G, 'Noise Floor Details'] The sentence 'the true floor is correspondingly higher' is ambiguous; it should clarify that the reference Δ value is correspondingly higher because the qTM false-positive rate on chains is even lower than the domain-level 21.06%.
- [Table 2] AUC is reported but not defined in the caption or text; please define it as the area under the ROC curve for distinguishing same-topology from cross-topology queries.
- [Figures 3 and 7] The label 'RFD3' is used for RFDiffusion without prior definition; please introduce the abbreviation at first use in the text or figure caption.
Circularity Check
No significant circularity: the paper's controls, calibrations, and negative-control FPRs are externally grounded and explicitly qualified.
full rationale
The paper's load-bearing steps are empirical calibrations against external resources rather than reductions to their own inputs. DRR uses CATH S40 with Merizo segmentation, and the high aligned-length domain retrieval is then bounded by the leave-one-topology-out negative control, whose ground-truth labels come from CATH topology rather than from the scores under test. RetFold is a constructed positive control: its composition is known by construction, but it is used as a sensitivity probe and counterexample, not as a fitted predictor of learned generators. The central 'envelope does not certify invention' argument is an existence proof, a deliberately recombined chain shown to land inside the reported retrieval envelope, and it does not derive the generators' behavior from RetFold. The choice of tTM at tau=0.7 is transparently reported with its 5.38% FPR and 100% sensitivity on the positive control; no fitted parameter is renamed as a prediction. The paper explicitly flags the one place where circular reasoning would arise and avoids it: in Appendix G it states that using the class of a generated backbone's best hit for model-specific FPR corrections 'is circular when the hit itself may be the false positive.' The acknowledged limitation that the negative-control FPRs are measured on domain-sized queries and do not transfer directly to chain-sized retrieval rates is a scope caveat about the evidence, not a definitional equivalence. There are no self-citations used as load-bearing support, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in by citation; all external tools (CATH, Merizo, Foldseek, ProteinMPNN, AlphaFold3) are independent. The paper is self-contained against external benchmarks, so no significant circularity is present.
Assumptions & free parameters
free parameters (5)
- RetFold ensemble embedding weights =
23% ESM-2, 38.5% Foldseek 3Di, 38.5% ProstT5
- RetFold fragment length windows =
core 40-60% L, extensions 15-35% L
- RetFold helix-linker parameters =
6-12 residue linkers, rise 1.5 A/res, radius 2.25 A, turn 100 deg/res, axis within +/-65 deg
- RetFold filtering thresholds =
Calpha distance >3 A, Rg <2.1 x expected, 40 variants per linker length
- Evaluation thresholds and coverage rule =
tau 0.5/0.7, tcov>=0.7, pLDDT>=70, iPTM>=0.6
assumptions (5)
- domain assumption CATH S40 (34,653 non-redundant domains) represents known fold space for this evaluation.
- domain assumption Merizo domain segmentation recovers the fold units of generated backbones.
- domain assumption Leave-one-topology-out exclusion removes true matches for a query.
- standard math TM-score conventions from Xu and Zhang 2010, TM>0.5 as same-fold, and the query-length bound hold.
- domain assumption AlphaFold3 pLDDT/iPTM and ProteinMPNN constitute a valid designability screen.
Cite this review
Pith. "Pith review of Is Retrieval All You Need? Assessment and Emergence of Novelty in Protein Structure Generation." pith.science (2026). https://pith.science/paper/S3KEU3NR
@misc{pith2026260810598,
author = {Pith},
title = {Pith review of: Is Retrieval All You Need? Assessment and Emergence of Novelty in Protein Structure Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/S3KEU3NR}},
note = {Machine review of arXiv:2608.10598}
}
read the original abstract
Protein backbone generation models are often credited with exploring novel fold space based solely on low full-chain similarity to known proteins, yet this cannot distinguish a genuinely new fold from a novel assembly of known structural units. We first ask whether this granularity mismatch alone explains the reported rates, and introduce the Domain Retrieval Rate (DRR), the fraction of generated backbones for which any constituent domain matches a known domain in CATH S40. Applied to eight backbone generation models spanning diffusion and flow-matching paradigms, DRR finds locally alignable known structure in most outputs, while the fraction containing a substantially covered complete domain is considerably smaller and depends on the scoring convention. To calibrate what retrieval alone can achieve, we propose RetFold, a zero-training baseline that constructs backbones by retrieving CATH domains and refining inter-domain connections through geometry-based helix-linker optimization, at two orders of magnitude lower cost on CPU alone.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Journal of molecular biology , volume=
Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions , author=. Journal of molecular biology , volume=. 1997 , publisher=
work page 1997
-
[2]
Methods in enzymology , volume=
Protein structure prediction using Rosetta , author=. Methods in enzymology , volume=. 2004 , publisher=
work page 2004
-
[3]
Nature structural biology , volume=
Protein building blocks preserved by recombination , author=. Nature structural biology , volume=. 2002 , publisher=
work page 2002
-
[4]
Proceedings of the National Academy of Sciences , volume=
An all-atom protein generative model , author=. Proceedings of the National Academy of Sciences , volume=. 2024 , publisher=
work page 2024
-
[5]
Boltzgen: Toward universal binder design , author=. bioRxiv , pages=. 2025 , publisher=
work page 2025
-
[6]
Nature , volume=
De novo design of protein structure and function with RFdiffusion , author=. Nature , volume=. 2023 , publisher=
2023
-
[8]
arXiv preprint arXiv:2302.02277 , year=
SE(3) Diffusion Model with Application to Protein Backbone Generation , author=. arXiv preprint arXiv:2302.02277 , year=
-
[9]
arXiv preprint arXiv:2310.05297 , year=
Fast protein backbone generation with se (3) flow matching , author=. arXiv preprint arXiv:2310.05297 , year=
Show all 34 references
-
[10]
Nature biotechnology , volume=
Fast and accurate protein structure search with Foldseek , author=. Nature biotechnology , volume=. 2024 , publisher=
2024
-
[11]
Nature Communications , volume=
Merizo: a rapid and accurate protein domain segmentation method using invariant point attention , author=. Nature Communications , volume=. 2023 , publisher=
2023
-
[12]
Nature , volume=
Accurate structure prediction of biomolecular interactions with AlphaFold 3 , author=. Nature , volume=. 2024 , publisher=
2024
-
[13]
bioRxiv , pages=
PXDesign: Fast, modular, and accurate de novo design of protein binders , author=. bioRxiv , pages=. 2025 , publisher=
2025
-
[14]
De novo design of protein structure and function with
Watson, Joseph L and Juergens, David and Bennett, Nathaniel R and Trippe, Brian L and Yim, Jason and Eisenach, Helen E and Ahern, Woody and Borber, Andrew J and Ragotte, Robert J and Milles, Lukas F and others , journal=. De novo design of protein structure and function with. ...
2023
-
[15]
Yim, Jason and Trippe, Brian L and De Bortoli, Valentin and Mathieu, Emile and Doucet, Arnaud and Barzilay, Regina and Jaakkola, Tommi , journal=
-
[16]
Nature , volume=
Illuminating protein space with a programmable generative model , author=. Nature , volume=. 2023 , publisher=
2023
-
[17]
arXiv preprint arXiv:2310.16624 , year=
Yim, Jason and Campbell, Andrew and Foong, Andrew Y K and Gastegger, Michael and Jim. arXiv preprint arXiv:2310.16624 , year=
-
[18]
arXiv preprint arXiv:2310.02391 , year=
SE(3)-Stochastic Flow Matching for Protein Backbone Generation , author=. arXiv preprint arXiv:2310.02391 , year=
-
[19]
Proceedings of the National Academy of Sciences , volume=
An all-atom protein generative model , author=. Proceedings of the National Academy of Sciences , volume=
-
[20]
Nature Machine Intelligence , volume=
Predicting equilibrium distributions for molecular systems with deep learning , author=. Nature Machine Intelligence , volume=
-
[21]
Science , volume=
Scaffolding protein functional sites using deep learning , author=. Science , volume=. 2022 , publisher=
2022
-
[22]
Robust deep learning--based protein sequence design using
Dauparas, Justas and Anishchenko, Ivan and Bennett, Nathaniel and Baek, Minkyung and Juergens, David and Ragotte, Robert J and Milles, Lukas F and Wicky, Basile I M and Galber, Alexis and Baker, David , journal=. Robust deep learning--based protein sequence design using. 2022 ...
2022
-
[23]
Highly accurate protein structure prediction with
Jumper, John and Evans, Richard and Pritzel, Alexander and Green, Tim and Figurnov, Michael and Ronneberger, Olaf and Tunyasuvunakool, Kathryn and Bates, Russ and. Highly accurate protein structure prediction with. Nature , volume=. 2021 , publisher=
2021
-
[24]
Accurate structure prediction of biomolecular interactions with
Abramson, Josh and Adler, Jonas and Dunger, Jack and Evans, Richard and Green, Tim and Pritzel, Alexander and Ronneberger, Olaf and Willmore, Lindsay and Ballard, Andrew J and others , journal=. Accurate structure prediction of biomolecular interactions with. 2024 , publisher=
2024
-
[25]
Fast and accurate protein structure search with
van Kempen, Michel and Kim, Stephanie S and Tumescheit, Charlotte and Mirdita, Milot and Lee, Jeongjae and Gilchrist, Cameron L M and S. Fast and accurate protein structure search with. Nature Biotechnology , volume=. 2024 , publisher=
2024
-
[26]
Science , volume=
Evolutionary-scale prediction of atomic-level protein structure with a language model , author=. Science , volume=. 2023 , publisher=
2023
-
[27]
2024 , publisher=
Heinzinger, Michael and Weissenow, Konstantin and Littmann, Maria and Steinegger, Martin and Rost, Burkhard , journal=. 2024 , publisher=
2024
-
[28]
arXiv preprint arXiv:2407.04967 , year=
Generative Augmentation Flows , author=. arXiv preprint arXiv:2407.04967 , year=
-
[29]
Nature , volume=
One thousand families for the molecular biologist , author=. Nature , volume=. 1992 , publisher=
1992
-
[30]
1997 , publisher=
Orengo, Christine A and Michie, Alex D and Jones, Susan and Jones, David T and Swindells, Mark B and Thornton, Janet M , journal=. 1997 , publisher=
1997
-
[31]
Journal of Molecular Biology , volume=
Estimating the number of protein folds and families from complete genome data , author=. Journal of Molecular Biology , volume=. 2000 , publisher=
2000
-
[32]
Proceedings of the National Academy of Sciences , volume=
Nature of the protein universe , author=. Proceedings of the National Academy of Sciences , volume=. 2009 , publisher=
2009
-
[33]
2021 , publisher=
Sillitoe, Ian and Bordin, Nicola and Dawson, Natalie and Waman, Vaishali P and Ashford, Paul and Scholes, Harry M and Pang, Chan and Woodridge, Laura and Rauer, Clemens and Sen, Neera and Abbasian, Maryam and LeCornu, Steven and Lam, Su Datt and Berka, Karel and Varekova, Radk...
2021
-
[34]
Bioinformatics , volume=
Chainsaw: protein domain segmentation with fully convolutional neural networks , author=. Bioinformatics , volume=. 2024 , publisher=
2024
-
[35]
Bioinformatics , volume=
How significant is a protein structure similarity with TM-score= 0.5? , author=. Bioinformatics , volume=. 2010 , publisher=
2010
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.