{"id":"0d131e58-09ed-47a7-841a-d1b89e2b9b2b","arxiv_id":"2508.02834","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Per-target tuned multi-expert guidance improves in silico antibody design success from 36.7% to 41.1% on six benchmark targets.","lead":"This paper adds four physics-inspired guidance modules to an antibody diffusion pipeline and tunes each target's guidance schedule with Bayesian optimization. It is a compact example of how online adaptation can be fitted to the very metrics the authors report.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Per-target Bayesian optimization selects Beta schedules on the same success metrics used in Table 2, with no held-out split; the reported 41.1% success rate may be a selection artifact, not a generalizable gain.","rationale":"The paper's central claim is an empirical one: adaptive multi-expert guidance with per-target Bayesian optimization yields better antibody designs. For that claim to be supported, the reported metrics must be an unbiased estimate of what the pipeline produces. The paper's own description of the BO loop (Eqs. 23–26, Appendix I) shows that the (α, β) parameters are selected by maximizing Expected Improvement on a GP whose training targets are CDR-H3 RMSD, pAE, and ipAE—the same metrics that define the success criterion and that are reported in Table 2. The appendix even states that the success rate climbs to 41% during optimization. With no held-out split, the final numbers are the result of fitting to the evaluation metric on the evaluation targets. Baseline methods do not receive this per-target fitting, so the comparison is not method-vs-method but method-vs-method-plus-oracle-tuning. This is a correctness risk that is independent of the hotspot annotation issue the reader highlighted: even with perfect hotspots, the reported gain could be a selection artifact. The reader's hotspot concern is real and partially overlaps with the circularity concern, but it is secondary. A concrete held-out test would settle the matter. If the held-out gain persists, the claim is credible; if it vanishes, the paper's current evidence does not support the abstract. Because the required re-analysis is straightforward but absent, and no code or data is provided for independent verification, I recommend UNVERDICTED rather than accepting the current conditional based on reported numbers.","tokens_in":25101,"tokens_out":10128,"duration_ms":111127,"concrete_test":"For each of the six targets, split the 1,000 sampled designs into a tuning set (e.g., first 700) and a held-out set (last 300) that the GP never queries. Run the online Bayesian optimization on the tuning set, freeze the selected (α, β), then generate a fresh batch of 300 designs per target with those frozen parameters and evaluate success rate plus all Table 2 metrics only on this fresh batch. Repeat for five random split orders. Compare against RFAntibody evaluated on equally fresh batches. If the mean held-out success-rate gain over RFAntibody is below roughly 2 points, or if it is negative on multiple targets, the reported 41.1% versus 36.7% advantage is attributable to selection-on-test-set and the central claim fails.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The most load-bearing flaw is not the guidance mechanism itself but the evaluation protocol. Eq. (25) defines the GP acquisition objective as a weighted sum of CDR-H3 RMSD, pAE, and ipAE, and the paper's success criterion is thresholds on exactly these same three metrics. Appendix I states that during online optimization the success rate rises from 36.7% to 39–40% and then exceeds 41%, so the headline 'Ours 41.1%' is the value obtained after selecting per-target (α, β) using the evaluation metrics on those very targets. No validation/holdout split is described anywhere. The baseline RFAntibody receives no equivalent per-target fitting. Thus the comparison conflates method quality with test-set selection: any optimizer that can query the evaluation metric will appear to improve that metric. This holds even if hotspot annotations are perfect, so it is more load-bearing than the reader's hotspot-sensitivity concern. Per-target tables also show one target (5J13) where Ours is below baseline, and standard deviations overlap, so the 4.4-point pooled gain is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an adaptive multi-expert guidance framework for SE(3)-equivariant diffusion-based antibody design. Four physics-inspired experts (van der Waals, molecular recognition, energy balance, and interface geometry) provide guidance gradients during reverse diffusion, with per-target temporal schedules parameterized by Beta distributions whose shape parameters (α, β) are selected by Bayesian optimization with Gaussian processes. The method is evaluated on six antibody-antigen targets against DiffAb and RFAntibody, reporting a 41.1% success rate versus 36.7% for RFAntibody, along with improvements in CDR-H3 RMSD, hotspot coverage, interface pAE, and shape complementarity. The central claim is that per-target online adaptation of guidance improves multi-objective antibody design without retraining the base diffusion model.","tokens_in":25302,"tokens_out":4929,"duration_ms":47992,"significance":"If the reported results were established by a sound evaluation protocol, the paper would make a useful contribution: it extends diffusion-model guidance with target-specific online adaptation, provides an explicit equivariance argument for the guidance gradients, includes an ablation separating fixed expert guidance from online learning, and broadens the benchmark to six targets. The appendix discussion of hotspot sensitivity is unusually candid and is a credit to the authors. However, the current evidence for the central claim is undermined by the circular evaluation protocol: the Bayesian optimizer is run on the same metrics that define the success criterion, on the same targets that are then reported in the headline tables, with no validation split. The hotspot-coverage result is also partially self-referential because hotspot coverage is both a guided objective and an evaluation metric. The significance of the contribution is therefore contingent on a re-evaluation that separates parameter selection from performance reporting.","major_comments":[{"comment":"The per-target Bayesian optimization selects (α, β) by maximizing L(θ) = Σ_m ω_m·metric_m/ν_m, where the metrics are CDR-H3 backbone RMSD, pAE, and ipAE, and the paper's success criterion is thresholds on exactly these three metrics. Appendix I states that during online optimization the success rate rises from 36.7% to 39–40% and then exceeds 41%, so the headline \"Ours 41.1%\" in Table 2 is the value obtained after per-target selection using the evaluation metrics on the tested targets. No held-out validation or nested evaluation is described anywhere in the paper. Because RFAntibody receives no equivalent per-target fitting, the comparison conflates method quality with test-set selection; any optimizer that can query the evaluation score would appear to improve that score. This is load-bearing for the central claim and must be addressed by reporting fixed-hyperparameter results, a validation split, or a statistically disciplined selection protocol.","section":"Online Parameter Learning (Eq. 25), Appendix I, Tables 2 and 3"},{"comment":"The appendix states that varying hotspot selections on the same target can change success rates by up to 15%. Since the Molecular Recognition Expert actively attracts CDR residues toward the manually annotated hotspots and Hotspot Coverage is both a guided objective and an evaluation metric, the reported gains (e.g., hotspot coverage increasing from 48.5% to 57.5%, and the overall success improvement) may be largely determined by the choice of hotspot annotation rather than by the guidance framework itself. The main text gives no robustness analysis over hotspot sets, so the transferability of the headline results to new targets is not established.","section":"Appendix \"Impact of Hotspot Residue Selection\" and Eqs. (14)–(15)"},{"comment":"The pooled improvement is not statistically established. For 5J13, Ours (29.4% important pass rate) is below RFAntibody (30.8%), and for several targets the per-target differences are within the reported spreads; the pooled success rates 41.1±10.5 vs 36.7±14.0 overlap at one standard deviation. The paper should provide per-target significance tests or effect sizes, and should report the number of independent design batches used to compute the standard deviations.","section":"Tables 6 and 7, per-target results"},{"comment":"The DiffAb baseline numbers are implausible as structural RMSD values: CDR-H3 RMSD = 25.56±0.22 Å for 5NGV, 20.57±0.58 Å for 6PPG, and 17.03±0.62 Å for 5J13, with near-zero standard deviations, while DiffAb CDR Int. Pass is 0.0 for every target. This suggests either a different evaluation protocol (e.g., alignment to a different reference or inclusion of failed designs) or a computational error. Because DiffAb is one of only two baselines, this undermines the comparative claim; please clarify the exact protocol and report success rates separately for designs that produced valid structures.","section":"Tables 6 and 7, DiffAb baseline"}],"minor_comments":[{"comment":"RFAntibody is cited inconsistently as \"Adolf-Bryfogle, Toth, and Bahl 2024\" in Related Work and as \"Bennett et al. 2024\" in the Experiments section; please unify the citation and specify the exact model version used.","section":"Related Work and Experiments"},{"comment":"Table 3 uses \"RFdiffusion\" as the label for the no-guidance baseline while the text calls the baseline \"RFAntibody\"; clarify whether the baseline is RFdiffusion alone or the full RFAntibody pipeline.","section":"Table 3"},{"comment":"The Figure 3 caption says \"across four antibody targets,\" but the experiments report six targets; please update the caption to be consistent with the main text.","section":"Figure 3 caption"},{"comment":"The normalizers ν_m for RMSD, pAE, and ipAE are listed as 1.5, 7.0, and 10.0, but the success thresholds are 3.0, 10.0, and 10.0; the relationship between the normalizers and the thresholds should be explained.","section":"Table 1 and Eq. (25)"},{"comment":"The claim of being the \"first biologically-motivated framework\" is too strong given existing affinity-maturation-inspired design work cited in Appendix A; please soften or qualify the novelty claim.","section":"Abstract"},{"comment":"The Beta parameter ranges are listed as [0.5, 10.0] for both α and β, but the initial value is (2.0, 2.0) and Appendix I reports convergence to α∈[1.5,3.5], β∈[1.5,4.0]; clarify whether the ranges are search bounds or expected ranges.","section":"Appendix E, Table 1"},{"comment":"No data or code availability statement is provided; to support reproducibility, please release the hotspot annotations, evaluation code, and per-target parameter traces.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The circular evaluation protocol is the main obstacle to acceptance. I would not require the authors to abandon per-target adaptation, but they must demonstrate that the reported gains survive a protocol where the parameters are not selected on the same targets that are reported. A fixed-hyperparameter variant, a validation-target split, or a properly nested Bayesian optimization loop with a separate evaluation set would suffice. The DiffAb baseline inconsistency should also be corrected before the comparison is taken at face value. The hotspot sensitivity analysis is honest and should be moved into the main text if the hotspot-guided metric remains a headline result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper has a real idea—per-target Beta-shaped guidance schedules for diffusion-based antibody design, tuned online with Bayesian optimization—but the evaluation protocol undercuts the headline numbers. The optimizer's objective (Eq. 25) is a weighted sum of CDR-H3 RMSD, pAE, and ipAE, and the success criterion is thresholds on the same three metrics. The appendix confirms success climbs from 36.7% to 39–40% to just over 41% during optimization. So the reported 41.1% is at least partly a selection artifact; the baseline RFAntibody gets no equivalent per-target fitting. That is the load-bearing flaw, and it is addressable: hold out targets or folds, and give the baseline the same tuning budget.\n\nWhat is genuinely new: combining four hand-designed expert losses (VDW, recognition, energy balance, interface geometry) with problem-driven routing and per-target temporal schedules. That is a plausible extension of guided diffusion, and the ablation separating 'expert only' from 'expert + online' is a nice touch. The equivariance argument is standard but correct.\n\nSoft spots beyond the circularity: the appendix admits up to 15% success-rate variation from hotspot selection, and hotspot coverage is both a guided objective and an evaluation metric—so biased annotations would inflate the comparison. The DiffAb baseline numbers in Table 6 look implausible (e.g., CDR-H3 RMSD 25.6 Å for 5NGV), which suggests a metric mismatch or a broken baseline. The text lists six targets but the per-target tables switch one PDB ID (6A3W vs 6MI2), and RFAntibody is cited two different ways. No code or data are provided, so none of this can be independently checked.\n\nOverall: the idea is worth a serious referee, but the central claim as stated is not established. I would send it to review with a request for a proper validation split, a tuned baseline, and code/data or a much more careful ablation. A reader looking for a cautionary tale about optimizing evaluation metrics will get value from it.","headline":"Per-target Bayesian optimization tunes the exact metrics the paper reports as gains, so the headline success-rate improvement is largely a selection artifact.","tokens_in":25912,"tokens_out":3372,"would_cite":false,"duration_ms":30518,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adaptive multi-expert physics guidance with online per-target tuning improves antibody design success rates.","keywords":["antibody design","diffusion models","SE(3)-equivariant generation","multi-objective optimization","physics-based guidance","online Bayesian optimization","hotspot coverage","affinity maturation"],"falsifier":"Run the same pipeline on one target with three different hotspot sets—the paper's curated five, five high-buried-surface-area residues chosen independently, and five random residues—and compare success rates. If the curated set beats the random set by less than the 4.4-point advantage over the undirected baseline, or if random selection closes that gap, then hotspot annotation, not adaptive guidance, explains the result.","tokens_in":24830,"feed_emoji":"🧬","tokens_out":9736,"duration_ms":90651,"temperature":0.7,"pith_summary":"Antibody design by diffusion models fails when one metric lags far behind the rest, so the paper argues that generation should be steered per antigen rather than by a fixed protocol. It adds four physics-based guidance experts—van der Waals balance, epitope-hotspot recognition, contact-density balance, and interface geometry—to an SE(3)-equivariant backbone diffusion pipeline, and an online Bayesian optimizer learns each expert's temporal activation profile for each target. The paper reports that this adaptive guidance lifts the fraction of designs passing its success thresholds from 36.7% to 41.1%, with gains in CDR-H3 RMSD, hotspot coverage, interface pAE, and shape complementarity, while reducing variance. If correct, the practical consequence is that better antibody designs can be obtained by tuning guidance per target without retraining or replacing the base generative model. The biological inspiration is B cell affinity maturation, where antibodies are refined by competing selection pressures acting over iterative rounds.","feed_headline":"Per-target tuning lifts antibody design success to 41.1%","feed_subtitle":"Adaptive physics experts improve every antibody design metric, from RMSD to hotspot coverage.","key_machinery":"The load-bearing object is the guidance gradient $g(T_t,t)=\\sum_{i=1}^4 w_i(t)\\,\\lambda_i(t)\\,\\nabla \\mathcal{L}_i(T_t)$ injected into every reverse-diffusion step. Each $\\mathcal{L}_i$ is one expert's loss: $L_{\\text{vdw}}$ penalizes atom overlaps, $L_{\\text{hotspot}}$ pulls CDR residues toward annotated epitope hotspots, $L_{\\text{contact}}$ keeps interface contact counts in a target range, and $L_{\\text{geom}}$ enforces distance uniformity and penalizes cavities. The temporal strength $\\lambda_i(t)=\\lambda_{\\mathrm{base},i}\\,f_{\\text{temporal}}(t,\\alpha_i,\\beta_i)$ is a Beta-distribution profile whose shape parameters $(\\alpha,\\beta)$ are learned per antigen by Bayesian optimization; small values of $\\alpha$ give early-peaking guidance for global structure, large $\\alpha$ give late-peaking atomic refinement. Because all gradients are built from pairwise distances and relative orientations, the combined guidance remains SE(3)-equivariant.","core_discovery":"On its own terms, the paper's central claim is that the “weakest link” problem in computational antibody design can be attacked during generation. The method wraps an existing SE(3) diffusion backbone in a multi-expert guidance layer: each expert computes a gradient from a physical loss—steric clashes, hotspot proximity, contact density, interface uniformity—a routing module weights experts by current structural severity, and a Gaussian-process Bayesian optimizer chooses the shape parameters of each expert's Beta-distributed activation schedule for the target at hand. The reported outcome is a balanced improvement across all evaluation metrics, with success rate 41.1% versus 36.7% for the state-of-the-art undirected baseline, a 7% reduction in CDR-H3 RMSD, 9% higher hotspot coverage, 12% better interface pAE, and 5% higher shape complementarity. The paper further claims that fixed expert guidance alone reaches only 38.9% success, so the online adaptation step, not merely the physics losses, drives the gain.","pith_inferences":["A natural test is to use the learned Beta schedules as a prior for a new antigen: if schedules cluster by epitope type, they could seed Bayesian optimization at a good starting point instead of the neutral (2,2) profile.","Because hotspot coverage is both the steering signal and one of the headline metrics, removing the recognition expert or replacing manual hotspots with an automated predictor would reveal how much of the gain is adaptive scheduling versus annotation information.","The online-optimization loop is generic: the same Gaussian-process search over schedule parameters could tune guidance for other generative structure models, including flow-matching or sequence co-design pipelines, rather than only diffusion backbones.","The current adaptation happens between batches; a truly B-cell-like version would update the schedule within a single denoising run based on real-time structural metrics, a variation the paper's routing mechanism only partially implements."],"forward_implications":["Per-target adaptation is the main driver of improvement: fixed physics guidance alone raises success from 36.7% to 38.9%, and online learning adds a further 2.2 points while shrinking standard deviations.","Different antigen classes need different temporal schedules, so a universal guidance protocol leaves performance on the table; the paper finds early-peaking activation for small epitopes and late-peaking activation for large interfaces.","The reported success-rate gain means fewer designs have to be generated and validated per viable candidate, which can offset the roughly 15% extra compute per design in a full campaign.","The same guided-diffusion pipeline generalizes across six therapeutically relevant targets spanning small epitopes and large protein interfaces, suggesting the method can be transferred to new antigens without retraining the base model."],"supporting_citations":[{"why":"Supplies the SE(3)-equivariant diffusion backbone whose reverse process this paper guides.","marker":"(Watson et al. 2023)"},{"why":"Provides the DiffAb baseline, a torsion-space CDR diffusion method with much lower success rates.","marker":"(Luo et al. 2022)"},{"why":"Defines the RFAntibody pipeline that serves as the main state-of-the-art baseline.","marker":"(Adolf-Bryfogle, Toth, and Bahl 2024)"},{"why":"Provides ProteinMPNN, the sequence-design step that converts guided backbones into antibodies.","marker":"(Dauparas et al. 2022)"},{"why":"Underlies AlphaFold2-Multimer confidence metrics (pAE, ipAE, pLDDT) used for evaluation and success criteria.","marker":"(Jumper et al. 2021)"},{"why":"Establishes SE(3) diffusion and the IGSO3 rotation noise model that the guidance is built on.","marker":"(Yim et al. 2023)"},{"why":"Motivates the molecular recognition expert via geometric deep learning of interaction fingerprints.","marker":"(Gainza et al. 2020)"},{"why":"Motivates the interface geometry expert from structure-based antibody screening.","marker":"(Schneider et al. 2022)"},{"why":"Gives the denoising diffusion probabilistic framework underlying the reverse process.","marker":"(Ho, Jain, and Abbeel 2020)"}],"fun_headline_variants":["B cell-inspired adaptive experts lift antibody design to 41.1%","Per-antigen expert routing boosts antibody success to 41.1%","Adaptive multi-expert diffusion tailors antibody designs per target","Online-optimized physics experts improve antibody design metrics","Multi-expert diffusion adapts per antigen, hits 41.1% success"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the five manually chosen hotspot residues per target are the right ones: they steer the recognition expert and are also the metric used to score designs, and the paper reports that swapping hotspot sets can change success rates by up to 15 percentage points, so noisy annotation could produce the observed gain.","fun_headline_variants_meta":{"raw":{"variants":["B cell-inspired adaptive experts lift antibody design to 41.1%","Per-antigen expert routing boosts antibody success to 41.1%","Adaptive multi-expert diffusion tailors antibody designs per target","Online-optimized physics experts improve antibody design metrics","Multi-expert diffusion adapts per antigen, hits 41.1% success"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000453,"raw_usage":{"total_tokens":2291,"prompt_tokens":968,"completion_tokens":1323,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":1231}},"tokens_in":584,"tokens_out":1323,"duration_ms":10951,"temperature":1.0,"reasoning_tokens":1231,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:05:34.381225+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same pipeline on one target with three different hotspot sets—the paper's curated five, five high-buried-surface-area residues chosen independently, and five random residues—and compare success rates. If the curated set beats the random set by less than the 4.4-point advantage over the undirected baseline, or if random selection closes that gap, then hotspot annotation, not adaptive guidance, explains the result.","supporting_citations":[],"review_version":1}