Pith. sign in

REVIEW 3 major objections 5 minor 21 references

HelixDesign-Antibody: A Scalable Production-Grade Platform for Antibody Design Built on HelixFold3

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Antibody design improves with computational scale: on five antigen–antibody systems, generating and folding more candidate sequences reliably lifts the predicted quality of the top-ranked binders.

desk verdict A competent engineering platform whose central scaling-law claim is an order-statistic artifact; the only real evidence is a modest KD correlation benchmark. read the letter →

arxiv 2507.02345 v1 pith:KBGI2HBM submitted 2025-07-03 q-bio.BM cs.AI

classification q-bio.BMcs.AI
keywords antibodydesignlarge-scalehigh-throughputscreeningHelixFold3scalinglawinversefoldingbindingaffinitypredictionipTM
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

HelixDesign-Antibody is a computational pipeline that takes a known antibody–antigen complex and produces thousands of redesigned antibody sequences: an inverse-folding model proposes variants in user-chosen regions, HelixFold3 folds every candidate against the antigen, and a multi-dimensional score ranks them. The paper's central claim is that this scaling works: on five test antigens, the designed antibodies match or beat the wild-type antibody on predicted interface confidence and computed binding energy, and as the number of sampled sequences grows, the top-ranked candidates keep improving. Finding this would matter because antibody discovery is normally gated by slow, expensive wet-lab screening, and the result suggests that compute can do part of the filtering before anything is made in the lab. The paper also reports that HelixFold3's interface confidence score correlates with experimentally measured influenza-antibody affinities, supporting the use of such scores as a ranking signal.

What carries the argument

The carrying mechanism is the sampling-and-scoring loop. An inverse-folding model (ESM-IF1, adapted here to design partial antibody regions) generates 1,000 candidate sequences per run from the template backbone; HelixFold3, a high-accuracy complex-structure predictor that outputs interface predicted TM-score (ipTM), per-residue pLDDT, and predicted aligned error, folds each candidate against the antigen; and a multi-dimensional filter combines sequence fitness, structural confidence, and FoldX binding free energy to rank survivors. The scaling result—top-100 quality rising with sample count—is the load-bearing evidence that computational breadth substitutes for experimental screening.

What would settle it

Clone, express, and measure binding of the top-ranked and lower-ranked designed antibodies for one of the five antigens (for example, 6U6U) with surface plasmon resonance or biolayer interferometry. If the top-ranked designs do not bind as well as or better than the wild-type antibody, or if their ordering does not track the predicted ipTM and $\Delta G$ rankings, the platform's quality and scaling claims fail; a minimum version is computing the correlation between model rankings and measured dissociation constants on the same panel.

Watch

Extended reading notes

Core claim

On five antigen systems (ACVR2B, TNFRSF9, FXI, IL-36R, IL17A), the pipeline redesigns only the heavy-chain CDRs, keeping the framework and light chain fixed. The designed candidates achieve HelixFold3 interface scores comparable to the wild-type complexes—four of five above 0.7 ipTM, with CDR pLDDT generally above 80—and lower computed binding free energy ($\Delta G$) than the wild-type antibody in all five cases, while retaining moderate-to-high overlap with the native epitope. The paper reports a scaling law: as sampling size increases, the mean ipTM and binding free energy of the top-100 ranked designs improve across all targets, which it takes as evidence that large-scale sequence exploration raises the chance of finding optimal binders. In a separate benchmark on influenza broadly neutralizing antibodies, the same interface score achieves Pearson correlations of 0.40, 0.64, and 0.52 against experimental dissociation constants on three datasets, outperforming two other sequence/structure predictors.

Load-bearing premise

The load-bearing premise is that the model's predicted interface score and computed binding energy faithfully represent real antibody binding and stability; the new designs were never tested at the bench, so this is assumed rather than shown.

Editorial extensions

If this is right

  • If the reported correlations hold, interface confidence from HelixFold3 can pre-rank thousands of candidates before any wet-lab work, cutting the number of antibodies that need experimental testing.
  • Because top-100 quality improves monotonically with sampling size on all five test systems, allocating more HPC compute to a design campaign should directly produce better top-ranked candidates.
  • The platform's support for user-defined design regions, IgGs, and nanobodies means the same workflow can be pointed at different antibody formats and epitopes without retraining the underlying models.
  • Retaining moderate-to-high epitope overlap with the wild-type antibody suggests the platform explores sequence diversity while preserving the native antigen-recognition site, a useful property for affinity maturation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The scaling law is demonstrated inside the model, not at the bench: no designed antibody was expressed or affinity-tested, so the practical payoff rests on a follow-up experiment that has not yet been reported.
  • The affinity correlations (0.40–0.64) are moderate, so the platform is best read as a high-throughput coarse filter that shrinks a large sequence space to a shortlist for experimental validation, rather than as a final affinity predictor.
  • A natural testable extension is to measure whether wet-lab hit rates from the top-100 predictions rise with sample count the way ipTM and $\Delta G$ do; that would move the scaling claim from predicted quality to true binding.
  • Because the design starts from a fixed reference backbone and mutates only CDR side chains, it samples sequence space within known scaffolds; adding a backbone-generating step would let the same scaling argument apply to novel loop conformations.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript describes HelixDesign-Antibody, a computational antibody design pipeline that combines ESM-IF1 inverse folding with HelixFold3 structure prediction and FoldX binding-energy scoring, integrated with HPC infrastructure. The authors claim the platform generates diverse, high-quality antibody candidates across five antigen targets, exhibits a scaling law whereby larger sequence sampling improves top-ranked candidates, and that HelixFold3's ipTM score correlates with experimental dissociation constants for influenza antibodies. Validation is entirely computational: quality is assessed by ipTM, pLDDT, and FoldX ΔG computed on model-predicted structures, with one external benchmark correlating ipTM to KD for known natural antibodies.

Significance. If the claims were supported, the platform could be a useful practical tool: it integrates existing models into an accessible workflow, and the KD benchmark in Section 3.4 offers some non-tautological evidence that HelixFold3 ipTM carries affinity information for natural antibody variants. However, the central scientific claims are not backed by the evidence presented. The 'scaling law' is an order-statistic artifact, and the design-quality evaluation is self-referential because the same computational pipeline both generates and scores the candidates, with no experimental validation of any designed sequence. The paper is better framed as a software system description with a limited predictive-correlation benchmark than as a validated antibody optimization method.

major comments (3)
  1. [§3.3, Figure 4] The scaling-law claim is an order-statistic artifact. For any fixed distribution of ipTM or FoldX scores, the expected value of the top-100 order statistics strictly increases with the number of sampled sequences N; Figure 4 therefore reproduces a property of random sampling and cannot support the manuscript's claim that larger sequence exploration 'raises the likelihood of identifying candidates with favorable biophysical characteristics.' To make this claim informative, the authors need a null baseline—e.g., applying the identical scoring pipeline to randomly generated CDR sequences or a shuffled-identity control—and should show that HelixDesign-Antibody's improvement curve is steeper than that null.
  2. [§3.1–§3.3] The evaluation of design quality is circular with respect to the pipeline. 'High-quality' is defined through HelixFold3 ipTM/pLDDT and FoldX energies computed on HelixFold3-predicted structures, so the same in silico framework that produces the designs also supplies the scores that certify them. The manuscript contains no experimental binding, stability, or expression data for any designed sequence, and Section 3.4's KD benchmark uses natural antibodies (CR9114, CR6261) rather than HelixDesign-Antibody outputs. The abstract's claim that the platform 'generate[s] diverse and high-quality antibodies' is therefore unsupported. Either an experimental validation or an external benchmark on previously characterized designed antibodies is required.
  3. [§3.4, Figure 5] The binding-discrimination evidence is weakened by the ad hoc 5:1 weighted sampling used for the CR9114-H3N2 dataset. Re-weighting a heavily imbalanced affinity distribution before computing Pearson correlation can inflate the reported coefficient, and the manuscript provides no sensitivity analysis, no confidence intervals, and no exact sample sizes. Report the correlations on unweighted data and with bootstrap or jackknife intervals, and justify the weighting factor.
minor comments (5)
  1. [§2, first paragraph] There are repeated typos, including 'candidate candidate antibodies' and the ungrammatical 'these candidates are subsequently undergo structure prediction'; Figure references are also inconsistent ('figure 6' versus 'Figure 1').
  2. [Abstract and §4] The platform is said to be available via PaddleHelix, but the manuscript gives no URL, version, or access instructions; this impedes reproducibility and makes the 'production-grade' claim difficult to assess.
  3. [§3.4, Figure 5] Figure 5 shows only a qualitative heatmap with no numerical correlation values or uncertainty estimates; report the PCC values in text or a table with errors.
  4. [§3.1, Figure 2] Figure 2 panels lack statistical detail such as the number of candidates, distribution shapes, and outliers; the claims of 'comparable' ipTM and 'more favorable' ΔG are not accompanied by any statistical test.
  5. [§3.1] The 8Å epitope-contact cutoff is introduced without justification, and epitope overlap is interpreted as functional retention despite being only a geometric proxy for the true paratope-epitope interaction.

Circularity Check

1 steps flagged · score 7.0 of 10

The §3.3 'scaling law' is an order-statistic tautology: top-100 scores must improve with N for any scoring function, so it does not demonstrate that larger sampling finds better antibodies.

  1. renaming known result [Section 3.3 (Computational Scale Enables Antibody Optimization), Figure 4]
    "As the sampling size increases, both the average predicted binding free energies and the ipTM scores of the top-ranked 100 designs consistently improve across all targets. This trend highlights the pivotal role of large-scale design: by extensively exploring the antibody sequence space, we substantially raise the likelihood of identifying candidates with favorable biophysical and structural characteristics."

    The plotted quantity is the mean score of the top-100 candidates as a function of total sample size N. For any fixed score distribution, the expected top-100 order statistics are monotonically nondecreasing in N; this is true by construction for any scoring function, even pure noise. The paper presents this sampling-theoretic identity as an empirical 'scaling law' that validates the platform. Since 'quality' is defined by the same ipTM and FoldX scores used to rank the candidates, the observed improvement is a mathematical consequence of selecting extremes from a larger pool, not evidence that the designed sequences are better binders. No experimental binding data on the designed sequences is provided to close the gap between improved model scores and real antibody quality.

full rationale

Most of the pipeline is engineering: ESM-IF generates candidate sequences, HelixFold3 predicts complex structures, and FoldX estimates binding energies; these choices are not circular. The central quantitative claim, however, is the 'scaling law' in Section 3.3. Figure 4 plots the mean ipTM and binding free energy of the top-ranked 100 designs as a function of total sampling size. For any fixed score distribution, the expected value of the top-100 order statistics increases with N; this is a known mathematical property and holds regardless of whether the scores come from a good structural model or from random noise. The paper labels this identity as 'computational scale enables antibody optimization,' so the prediction reduces by construction to the selection of top-ranked candidates. The quality metrics themselves (ipTM, pLDDT, FoldX ΔG) are generated by the same computational pipeline that designs and ranks the sequences, and no experimental binding data for the designed antibodies are provided, leaving the loop closed within the model. Section 3.4's KD correlations (Pearson 0.40–0.64) are an independent external check that HelixFold3 ipTM carries some affinity signal for known natural influenza antibodies; this is real evidence for HelixFold3 as a scoring tool, but it does not validate the scaling-law claim for designed sequences. Self-citations to HelixFold-Multimer and HelixFold3 are normal model citations and are partly corroborated by the KD benchmark, so they are not treated as load-bearing. The score reflects one central 'prediction' that is an order-statistic artifact, while acknowledging the external KD anchor that prevents the whole paper from being purely definitional.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The paper does not introduce new physical entities or forces. It relies on existing models and scoring functions, and the main free parameters are threshold and weighting choices that shape the reported validation metrics. The core assumption that in silico scores equate to real antibody quality is the most significant epistemic load.

free parameters (3)
  • 5:1 weighting factor for H3N2 affinity sampling = 5:1
    Chosen ad hoc to rebalance the imbalanced H3N2 affinity data (Section 3.4). This choice affects the composition of the test set and the reported correlation.
  • 8Å epitope contact cutoff = 8 Å
    Used to define antigen-contacting residues for epitope overlap calculation (Section 3.1). This threshold is conventional but arbitrary and influences the overlap metric.
  • ipTM > 0.7 threshold for 'high confidence' = 0.7
    Used in Section 3.1 to describe interfaces as high confidence. The threshold is not justified and affects the interpretation of design quality.
assumptions (3)
  • domain assumption ipTM, pLDDT, and FoldX binding energy are valid proxies for antibody binding affinity and stability.
    The entire validation relies on these model-generated scores representing real biological quality, but no experimental binding data is provided to support this assumption (Sections 2.2, 3.1).
  • domain assumption ESM-IF1 can generate diverse sequences that preserve the reference backbone and epitope specificity.
    The platform assumes that redesigning CDRs with ESM-IF1 while keeping the wild-type backbone as template yields meaningful antibody variants without disrupting the binding mode (Section 2.1).
  • ad hoc to paper Using the wild-type antibody's CDR backbone as a template is sufficient to retain the original epitope and functional activity.
    The design protocol fixes all non-CDR regions and the CDR backbone, implicitly assuming that the original binding mode is preserved, which is unverified experimentally (Section 3.1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of HelixDesign-Antibody: A Scalable Production-Grade Platform for Antibody Design Built on HelixFold3." pith.science (2026). https://pith.science/paper/KBGI2HBM

@misc{pith2026250702345,
  author       = {Pith},
  title        = {Pith review of: HelixDesign-Antibody: A Scalable Production-Grade Platform for Antibody Design Built on HelixFold3},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KBGI2HBM}},
  note         = {Machine review of arXiv:2507.02345}
}
read the original abstract

Antibody engineering is essential for developing therapeutics and advancing biomedical research. Traditional discovery methods often rely on time-consuming and resource-intensive experimental screening. To enhance and streamline this process, we introduce a production-grade, high-throughput platform built on HelixFold3, HelixDesign-Antibody, which utilizes the high-accuracy structure prediction model, HelixFold3. The platform facilitates the large-scale generation of antibody candidate sequences and evaluates their interaction with antigens. Integrated high-performance computing (HPC) support enables high-throughput screening, addressing challenges such as fragmented toolchains and high computational demands. Validation on multiple antigens showcases the platform's ability to generate diverse and high-quality antibodies, confirming a scaling law where exploring larger sequence spaces increases the likelihood of identifying optimal binders. This platform provides a seamless, accessible solution for large-scale antibody design and is available via the antibody design page of PaddleHelix platform.

Figures

Figures reproduced from arXiv: 2507.02345 by the authors.

Figure 1
Figure 1. HelixDesign-Antibody is a structure-based antibody design pipeline that takes as input the complex structure [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Comparison and evaluation of designed antibody sequences across multiple metrics. The top row shows [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Structural comparison of wild-type and designed antibodies for two targets. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Computational scale enables antibody optimization. Performance improvement as a function of sampling [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Binding discrimination capability of different scoring methods across datasets. Heatmap showing Pearson [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Example input interface of the HelixDesign-Antibody Server. [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Example output display page of the HelixDesign-Antibody Server. [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

21 extracted references · 15 canonical work pages

  1. [1]

    Protein complex prediction with alphafold-multimer

    Richard Evans, Michael O’Neill, Alexander Pritzel, Natasha Antropova, Andrew Senior, Tim Green, Augustin Žídek, Russ Bates, Sam Blackwell, Jason Yim, et al. Protein complex prediction with alphafold-multimer. biorxiv, pages 2021–10, 2021

  2. [2]

    Accurate structure prediction of biomolecular interactions with alphafold 3

    Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature, 630(8016):493–500, 2024

  3. [3]

    Precise Antigen-Antibody Structure Predictions Enhance Antibody Development with HelixFold-Multimer

    Jie Gao, Jing Hu, Lihang Liu, Yang Xue, Kunrui Zhu, Xiaonan Zhang, and Xiaomin Fang. Precise antigen-antibody structure predictions enhance antibody development with helixfold-multimer. arXiv preprint arXiv:2412.09826, 2024

  4. [4]

    Technical report of helixfold3 for biomolecular structure prediction

    Lihang Liu, Shanzhuo Zhang, Yang Xue, Xianbin Ye, Kunrui Zhu, Yuxin Li, Yang Liu, Jie Gao, Wenlai Zhao, Hongkun Yu, et al. Technical report of helixfold3 for biomolecular structure prediction. arXiv preprint arXiv:2408.16975, 2024

  5. [5]

    Learning inverse folding from millions of predicted structures

    Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives. Learning inverse folding from millions of predicted structures. ICML, 2022

  6. [6]

    Robust deep learning–based protein sequence design using proteinmpnn

    Justas Dauparas, Ivan Anishchenko, Nathaniel Bennett, Hua Bai, Robert J Ragotte, Lukas F Milles, Basile IM Wicky, Alexis Courbet, Rob J de Haas, Neville Bethel, et al. Robust deep learning–based protein sequence design using proteinmpnn. Science, 378(6615):49–56, 2022

  7. [7]

    The foldx web server: an online force field

    Joost Schymkowitz, Jesper Borg, Francois Stricher, Robby Nys, Frederic Rousseau, and Luis Serrano. The foldx web server: an online force field. Nucleic acids research, 33(suppl_2):W382–W388, 2005

  8. [8]

    Prodigy: a web server for predicting the binding affinity of protein–protein complexes

    Li C Xue, João Pglm Rodrigues, Panagiotis L Kastritis, Alexandre Mjj Bonvin, and Anna Vangone. Prodigy: a web server for predicting the binding affinity of protein–protein complexes. Bioinformatics, 32(23):3676–3678, 2016

Show all 21 references
  1. [9]

    Protein data bank

    Protein Data Bank. Protein data bank. Nature New Biol, 233(223):10–1038, 1971

  2. [10]

    Atomically accurate de novo design of antibodies with rfdiffusion

    Nathaniel R Bennett, Joseph L Watson, Robert J Ragotte, Andrew J Borst, DéJenaé L See, Connor Weidle, Riti Biswas, Yutong Yu, Ellen L Shrock, Russell Ault, et al. Atomically accurate de novo design of antibodies with rfdiffusion. bioRxiv, pages 2024–03, 2025

  3. [11]

    Helixfold-multimer: Elevating protein complex structure prediction to new heights

    Xiaomin Fang, Jie Gao, Jing Hu, Lihang Liu, Yang Xue, Xiaonan Zhang, and Kunrui Zhu. Helixfold-multimer: Elevating protein complex structure prediction to new heights. arXiv preprint arXiv:2404.10260, 2024

  4. [12]

    Proteingym: Large-scale benchmarks for protein design and fitness prediction

    Pascal Notin, Aaron W Kollasch, Daniel Ritter, Lood van Niekerk, Steffanie Paul, Hansen Spinner, Nathan Rollins, Ada Shaw, Ruben Weitzman, Jonathan Frazer, et al. Proteingym: Large-scale benchmarks for protein design and fitness prediction. bioRxiv, 2023

  5. [13]

    Unsupervised evolution of protein and antibody complexes with a structure-informed language model

    Varun R Shanker, Theodora UJ Bruun, Brian L Hie, and Peter S Kim. Unsupervised evolution of protein and antibody complexes with a structure-informed language model. Science, 385(6704):46–53, 2024

  6. [14]

    De novo design of high-affinity protein binders with alphaproteo

    Vinicius Zambaldi, David La, Alexander E Chu, Harshnira Patani, Amy E Danson, Tristan OC Kwan, Thomas Frerix, Rosalia G Schneider, David Saxton, Ashok Thillaisundaram, et al. De novo design of high-affinity protein binders with alphaproteo. arXiv preprint arXiv:2409.08022, 2024

  7. [15]

    De novo design of protein structure and function with rfdiffusion

    Joseph L Watson, David Juergens, Nathaniel R Bennett, Brian L Trippe, Jason Yim, Helen E Eisenach, Woody Ahern, Andrew J Borst, Robert J Ragotte, Lukas F Milles, et al. De novo design of protein structure and function with rfdiffusion. Nature, 620(7976):1089–1100, 2023

  8. [16]

    Flex ddg: Rosetta ensemble-based estimation of changes in protein–protein binding affinity upon mutation

    Kyle A Barlow, Shane Ó Conchúir, Samuel Thompson, Pooja Suresh, James E Lucas, Markus Heinonen, and Tanja Kortemme. Flex ddg: Rosetta ensemble-based estimation of changes in protein–protein binding affinity upon mutation. The Journal of Physical Chemistry B , 122(21):5389–5399, 2018

  9. [17]

    X-ray crystal structure localizes the mechanism of inhibition of an il-36r antagonist monoclonal antibody to interaction with ig1 and ig2 extra cellular domains

    Eric T Larson, Debra L Brennan, Eugene R Hickey, Raj Ganesan, Rachel Kroe-Barrett, and Neil A Farrow. X-ray crystal structure localizes the mechanism of inhibition of an il-36r antagonist monoclonal antibody to interaction with ig1 and ig2 extra cellular domains. Protein Scien...

  10. [18]

    The influenza hemagglutinin stem antibody cr9114: Evidence for a narrow evolutionary path towards universal protection

    Anna L Beukenhorst, Jacopo Frallicciardi, Clarissa M Koch, Jaco M Klap, Angela Phillips, Michael M Desai, Kanin Wichapong, Gerry AF Nicolaes, Wouter Koudstaal, Galit Alter, et al. The influenza hemagglutinin stem antibody cr9114: Evidence for a narrow evolutionary path towards...

  11. [19]

    Antibody recognition of a highly conserved influenza virus epitope

    Damian C Ekiert, Gira Bhabha, Marc-André Elsliger, Robert HE Friesen, Mandy Jongeneelen, Mark Throsby, Jaap Goudsmit, and Ian A Wilson. Antibody recognition of a highly conserved influenza virus epitope. Science, 324(5924):246–251, 2009

  12. [20]

    Binding affinity landscapes constrain the evolution of broadly neutralizing anti-influenza antibodies

    Angela M Phillips, Katherine R Lawrence, Alief Moulana, Thomas Dupic, Jeffrey Chang, Milo S Johnson, Ivana Cvijovic, Thierry Mora, Aleksandra M Walczak, and Michael M Desai. Binding affinity landscapes constrain the evolution of broadly neutralizing anti-influenza antibodies. ...

  13. [21]

    Language models of protein sequences at the scale of evolution enable accurate structure prediction

    Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Sal Candido, et al. Language models of protein sequences at the scale of evolution enable accurate structure prediction. bioRxiv, 2022. 12

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.