{"id":"1c1e36df-90c2-4a4e-b841-9ceab4ffb4bb","arxiv_id":"2412.05776","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"ProtGO, a fusion of three fine-tuned ProtBert transformers, reports higher GO-term prediction accuracy than Proteinfer on a SwissProt benchmark restricted to 100 frequent terms per aspect.","lead":"This paper presents ProtGO, a model that combines three fine-tuned ProtBert transformers to predict Gene Ontology terms from protein sequences. The authors report accuracy and F1 improvements over Proteinfer, but only on a reduced set of the 100 most common GO terms per category.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The alleged state-of-the-art advantage over Proteinfer rests on an unspecified baseline protocol: the baseline numbers may come from a different label set or evaluation setup, making the comparison apples-to-oranges.","rationale":"I read the paper in good faith: the model architecture is described in reasonable detail, and the fusion of three ProtBert modules is a plausible design. The load-bearing claim, however, is comparative: ProtGO achieves state-of-the-art accuracy by outperforming Proteinfer and ProteinferEN. That claim can only hold if the comparison is like-for-like. The paper uses a reduced top-100 label set while baselines are not described as being adapted to that same protocol. The reader's weakest assumption identifies exactly this: the Proteinfer numbers may come from its original full-GO evaluation rather than from an identical re-run. I agree with that assessment. This is not an internal mathematical inconsistency in ProtGO, but it is an externally unverified benchmark equivalence, and it is decisive for the central claim. The concrete check—re-evaluating Proteinfer under the exact ProtGO protocol—would settle the matter. Until that is done, the verdict remains REJECT; no new concern moves it to a different judgment.","tokens_in":10396,"tokens_out":2110,"duration_ms":25794,"concrete_test":"Obtain the released Proteinfer and ProteinferEN checkpoints and evaluation code, construct the exact top-100 GO label set per aspect from the SwissProt data described in Section 2.2, and evaluate the baselines on the same random and clustered test splits with the same multi-label accuracy metric and the same default threshold. If the reproduced baseline accuracies match Tables 3 and 4, the comparison is validated; if they do not, the reported ProtGO gains are unsupported. As a secondary check, train or evaluate ProtGO on the full GO label space to test whether the 'full-scale protein sequence' claim survives outside the reduced top-100 label set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that ProtGO outperforms Proteinfer and ProteinferEN by 3–7% accuracy on all three GO aspects (Tables 3 and 4). However, Section 2.2 states that only the top 100 most populous GO terms per aspect are used for training and evaluation, while the paper never states whether the Proteinfer baselines were retrained or re-evaluated on the same top-100 label set, the same SwissProt random/clustered splits, the same accuracy definition, or the same decision threshold. Section 2.4 only says that a default threshold was used 'for all algorithms,' but no threshold value or evaluation code is provided, and Proteinfer's original evaluation protocol is not described anywhere in the manuscript. If the baseline numbers are taken from Proteinfer's original full-GO evaluation, then Tables 3 and 4 compare ProtGO on a reduced 100-label task against Proteinfer's numbers from a much harder full ontology task; the reported gains would then reflect task narrowing rather than superior modeling. This is not a minor methodological detail: the abstract, introduction, and conclusion all rest on this numerical superiority. Without a like-for-like comparison, the state-of-the-art claim has no evidentiary support. The absence of released code, checkpoints, and error bars compounds the problem, but the decisive gap is the unstated baseline protocol.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ProtGO, a transformer-based fusion model that combines three ProtBert-derived modules to predict Gene Ontology (GO) terms for the biological process, molecular function, and cellular component aspects from full-length protein sequences. The model is trained and evaluated on SwissProt random and clustered splits using only the top-100 most frequent GO terms per aspect, and is compared against Proteinfer and ProteinferEN in terms of accuracy, F1 score, precision, and recall. The authors report state-of-the-art performance, with gains of roughly 3-7% in accuracy over the benchmarks, plus ROC curves and a sequence-length robustness analysis.","tokens_in":10658,"tokens_out":3765,"duration_ms":38501,"significance":"If the central comparison were like-for-like and fully specified, a single lightweight transformer-based model that outperforms Proteinfer and ProteinferEN on three GO aspects across two challenging splits would be a practically valuable contribution to automated protein annotation. The paper has some strengths: it evaluates on an external benchmark, addresses computational efficiency, and provides a sequence-length analysis. However, as written, the central empirical claim is not adequately supported because the baseline protocol, evaluation metric, and several architectural details are unspecified or inconsistent. The results may still be of interest, but the current manuscript does not provide enough information to judge or reproduce them.","major_comments":[{"comment":"The comparison to Proteinfer and ProteinferEN is not specified as a like-for-like evaluation. The paper never states whether the baselines were retrained on the same top-100 GO label set per aspect, the same random and clustered SwissProt splits, the same accuracy definition, or the same decision threshold. If the baseline numbers are taken from Proteinfer's original full-GO evaluation, then Tables 3 and 4 compare ProtGO on a reduced 100-label task against results from a much harder full-ontology task, and the reported gains would reflect task narrowing rather than superior modeling. This is load-bearing because the abstract, introduction, and conclusion all rest on the state-of-the-art claim.","section":"Section 2.4, Tables 3 and 4"},{"comment":"The accuracy metric is never defined. Since GO term prediction is a multi-label problem, 'accuracy' could mean exact match accuracy, per-label accuracy, micro- or macro-averaged accuracy, or Jaccard-style overlap. Without a definition, the reported percentages and the 3-7% gains are not interpretable. The paper also provides no confidence intervals, error bars, or repeated-run variability, so even the point estimates are not statistically grounded.","section":"Section 3, Tables 3 and 4"},{"comment":"The abstract and title claim prediction of GO terms from 'full-scale protein sequences' and state-of-the-art accuracy, but the evaluation uses only the top 100 most populous GO terms per aspect, and sequences are truncated at 1000 tokens. This restricts the task to a small subset of the GO vocabulary and to a limited sequence-length range. The paper needs to clarify what 'full-scale' means and explain how the truncated, top-100 evaluation supports the broad claim.","section":"Section 2.2 and Section 3.2"},{"comment":"The model architecture is described inconsistently: Section 2.3 states that the attention layers have 12 heads, while Section 2.4 states that the ProtBert module includes 16 attention heads. This is a concrete reproducibility-relevant contradiction that must be resolved, along with a precise specification of which layers are frozen and which are fine-tuned.","section":"Section 2.3 vs. Section 2.4"},{"comment":"No code, model checkpoints, evaluation scripts, or data-split definitions are provided. The paper also does not report the decision threshold used for evaluation or the source of the Proteinfer baseline numbers. For a paper whose central claim is empirical superiority over existing systems, these omissions prevent verification and make the state-of-the-art assertion unreproducible as written.","section":"Section 2.4 and Section 3"}],"minor_comments":[{"comment":"The phrase 'not unaffected by sequence length' is a double negative and appears to contradict the claim of negligible dependency on sequence length; it should be reworded.","section":"Abstract"},{"comment":"There is a typo: 'positional enbeddings' should be 'positional embeddings'.","section":"Section 2.3"},{"comment":"The text says 'Table 4 shows the performance of the proposed model on all three GO aspects ... on the random split dataset,' but Table 4 is labeled 'Clustered Split dataset.' The table cross-references need to be corrected.","section":"Section 3"},{"comment":"The ROC-curve figure is referenced but not actually included in the text, and the AUC values are reported without showing the corresponding curves or axis labels.","section":"Section 3.1"},{"comment":"The sentence 'An Negative Log Likelihood (NLL) loss function' contains a grammatical error and should be 'A Negative Log Likelihood.'","section":"Section 2.4"},{"comment":"The decision threshold is described as a hyperparameter 'set to the default value for all algorithms,' but the actual value is never given, nor is it stated what default threshold the baselines use.","section":"Section 2.4"}],"recommendation":"reject","confidential_remarks":"The decisive issue is the unstated baseline protocol and the undefined accuracy metric. These are not presentation details; they affect the core claim of state-of-the-art performance. The paper would need a substantial rewrite with a transparent, like-for-like evaluation and code/data release to be considered. I see no evidence of misconduct, but the current manuscript is not ready for publication in a serious journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my read. ProtGO is a simple but sensible idea: three ProtBert modules fine-tuned separately for BP, MF, and CC, then fused. The authors also report performance on a clustered split, which is harder than random and more realistic for distantly related proteins. That part is genuinely useful: if the numbers hold up, a single transformer that degrades gracefully on clustered splits is worth knowing about. The work is a direct follow-up to their ProtEC paper, and they say so.\n\nThe problem is the central claim: 'state-of-the-art accuracy' over Proteinfer and ProteinferEN. The paper never says whether the baselines were retrained on the same top-100 GO labels per aspect, the same SwissProt splits, the same accuracy metric, or the same decision threshold. Section 2.4 says a default threshold was used 'for all algorithms,' but no value is given. If the Proteinfer numbers come from its original full-ontology evaluation, then Tables 3 and 4 compare a 100-label classifier against a model evaluated on the full ontology. That's apples-to-oranges, and the claimed 3–7% gains could just be task narrowing. The stress-test note is right.\n\nThere are also smaller internal inconsistencies. The architecture section says 12 attention heads; the training section says ProtBert has 16. The abstract says 'not unaffected by sequence length' when the results show the model is stable to 1000 tokens and degrades after. The paper claims no development/validation set is needed, but Tables 1 and 2 have development splits. And there are no error bars or released code. For a paper whose main contribution is an empirical accuracy claim, that's a lot of missing context.\n\nWhat the paper does well: the clustered split evaluation is a real strength, and the focus on low memory (3.5GB on a single RTX2080Ti) is useful for practitioners. The fusion idea is not new in a deep conceptual sense, but the specific configuration with per-aspect fine-tuning is a reasonable thing to try.\n\nBottom line: the model may be fine, but the evidence as written does not support the SOTA claim. I'd encourage the authors to release code, rerun Proteinfer under the same protocol, and add confidence intervals. As is, I would not cite the SOTA claim. I'd still send it out for review, because the underlying idea is worth a careful look and the protocol gaps are fixable. If the baselines come out equivalent under a fair comparison, the paper becomes a modest but honest engineering contribution.\n\nBest.","headline":"Fusion of three ProtBert heads is a plausible engineering contribution, but the SOTA claim rests on an unstated baseline protocol and a reduced top-100 label set.","tokens_in":11168,"tokens_out":2823,"would_cite":false,"duration_ms":27975,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that ProtGO, a transformer-based fusion model fine-tuned from ProtBert, annotates full protein sequences with Gene Ontology terms more accurately than Proteinfer and ProteinferEN on both random and clustered dataset splits.","keywords":["gene ontology annotation","protein function prediction","transformer fusion model","protein language model","multi-label classification","selective fine-tuning","clustered data split"],"falsifier":"Re-run Proteinfer and ProteinferEN on the same top-100 per-aspect GO label sets and the same random and clustered splits used for ProtGO, with the threshold fixed at the same value; if their accuracy and F1 match or exceed the reported ProtGO numbers, the state-of-the-art claim collapses. A quicker check is to read the original Proteinfer evaluation and see whether its published numbers were computed over the full GO vocabulary rather than the 100-label subset.","tokens_in":10179,"feed_emoji":"🧬","tokens_out":8054,"duration_ms":72587,"temperature":0.7,"pith_summary":"ProtGO is a transformer-based fusion model that predicts Gene Ontology (GO) terms, covering biological process, molecular function, and cellular component, directly from full protein sequences. The paper claims it beats the Proteinfer and ProteinferEN systems by roughly 3 percentage points on a random split of a manually curated protein database and by around 7 points on a clustered split where training and test sequences share little similarity. The model is a single lightweight network rather than an ensemble, using three selectively fine-tuned transformers, one per GO aspect, whose predictions are combined. If the claim holds, newly sequenced proteins could receive accurate functional annotations from sequence alone at a fraction of the compute used by current benchmark methods.","feed_headline":"ProtGO tops Proteinfer by up to 7 percent in GO-term prediction","feed_subtitle":"Lightweight transformer fuses three GO-aspect heads to annotate full proteins, even on clustered splits.","key_machinery":"The architecture is a fusion of three transformer submodules, Prot_BP, Prot_MF, and Prot_CC, one for each GO aspect. Each submodule starts from the pretrained ProtBert protein language model (30 encoder layers, model dimension 1024, multihead attention) and is selectively fine-tuned on sequences labeled only with that aspect's top-100 GO terms: some layers are frozen and others updated, following the authors' earlier enzyme-number model. A protein sequence is tokenized into amino-acid tokens; positional, token, and segment embeddings are summed and fed through the encoder, mean-pooled, and passed to a classification layer. The three submodule outputs are concatenated to produce the final multi-label prediction, with cross-entropy loss and an Adam optimizer during fine-tuning.","core_discovery":"The paper's central claim is that one transformer fusion model can annotate full-length protein sequences with the 100 most frequent GO terms per aspect at higher accuracy than Proteinfer and ProteinferEN, while remaining robust to sequence length. On the random split it reports accuracies of 86.06% for biological process, 94.60% for molecular function, and 78.30% for cellular component; on the harder clustered split the corresponding numbers are 82.16%, 91.51%, and 73.28%. The margins over ProteinferEN are roughly 3 points on the random split and 6.6 to 7.6 points on the clustered split, with F1 and precision also higher. ROC AUC is reported above 99% on the random split and above 98% on the clustered split, and test accuracy stays approximately flat up to 1,000 amino acids, the truncation length used in the study.","pith_inferences":["Because the evaluation is restricted to the top-100 GO terms per aspect, the reported accuracy does not cover the long tail of rare GO terms; extending to that tail is a harder test and would likely reduce the margins.","The comparison assumes Proteinfer and ProteinferEN were evaluated under identical conditions, such as the same 100-label vocabulary, the same splits, the same accuracy definition, and the same threshold, and the paper does not document those details; a fair replication is needed before the state-of-the-art claim is taken at face value.","The same three-headed fusion design could be ported to other annotation tasks, such as Enzyme Commission numbers or subcellular localization, where whole-sequence context matters.","A natural stress test is to run the model on unreviewed, automatically generated protein sequences, where labels are noisier and the distribution differs from the curated training set."],"forward_implications":["A single ProtGO model could replace multi-model ensembles for GO annotation, using one GPU with about 3.5 GB of memory and roughly 80 hours of training while reporting higher accuracy.","The clustered-split results imply the model generalizes to protein families with little sequence similarity to its training data, which matters for uncharacterized or newly discovered sequences.","Stable accuracy up to the 1,000-residue truncation suggests the model can be applied to long proteins, and raising the truncation threshold would likely improve very-long-sequence performance.","Near-perfect ROC AUC indicates the model separates positive from negative GO assignments well enough to adjust precision and recall through a decision threshold, addressing the recall gap observed against ProteinferEN."],"supporting_citations":[{"why":"Provides the Proteinfer and ProteinferEN benchmarks that ProtGO compares against; without it the state-of-the-art comparison has no baseline.","marker":"[35]"},{"why":"Supplies the pretrained ProtBert transformer weights that each GO-aspect submodule is fine-tuned from.","marker":"[33]"},{"why":"The authors' earlier enzyme-number model, from which the selective fine-tuning and frozen-layer configuration are reused.","marker":"[36]"},{"why":"Supplies the protein sequences and their GO annotations used to build the random and clustered splits.","marker":"[1]"},{"why":"Provides the sequence-clustering method used to construct the clustered split with minimal similarity between train and test.","marker":"[37]"},{"why":"Provides the training framework and automated hyperparameter optimization used to fine-tune the model.","marker":"[38]"}],"fun_headline_variants":["ProtGO beats Proteinfer by up to 7.6% on hard splits","ProtGO: state-of-the-art GO prediction from full protein sequences","Lightweight transformer ProtGO wins on clustered splits","ProtGO predicts GO terms better, even on shifted protein sets","ProtGO: up to 7.6% better GO-term accuracy, robust to length"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The state-of-the-art claim rests on the assumption that Proteinfer and ProteinferEN numbers were produced under exactly the same evaluation protocol as ProtGO, meaning the same top-100 GO labels per aspect, the same data splits, the same accuracy metric, and the same decision threshold, which the paper does not explicitly confirm.","fun_headline_variants_meta":{"raw":{"variants":["ProtGO beats Proteinfer by up to 7.6% on hard splits","ProtGO: state-of-the-art GO prediction from full protein sequences","Lightweight transformer ProtGO wins on clustered splits","ProtGO predicts GO terms better, even on shifted protein sets","ProtGO: up to 7.6% better GO-term accuracy, robust to length"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00084,"raw_usage":{"total_tokens":3655,"prompt_tokens":937,"completion_tokens":2718,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":2624}},"tokens_in":553,"tokens_out":2718,"duration_ms":18548,"temperature":1.0,"reasoning_tokens":2624,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:22:06.739009+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run Proteinfer and ProteinferEN on the same top-100 per-aspect GO label sets and the same random and clustered splits used for ProtGO, with the threshold fixed at the same value; if their accuracy and F1 match or exceed the reported ProtGO numbers, the state-of-the-art claim collapses. A quicker check is to read the original Proteinfer evaluation and see whether its published numbers were computed over the full GO vocabulary rather than the 100-label subset.","supporting_citations":[{"cited_title":"Bileschi, David Belanger, and Lucy J","cited_arxiv_id":null,"evidence_quote":"Provides the Proteinfer and ProteinferEN benchmarks that ProtGO compares against; without it the state-of-the-art comparison has no baseline."},{"cited_title":"ProtTrans: Toward understanding the language of life through self-supervised learning","cited_arxiv_id":null,"evidence_quote":"Supplies the pretrained ProtBert transformer weights that each GO-aspect submodule is fine-tuned from."},{"cited_title":"Protec: A transformer based deep learning system for accurate annotation of enzyme commission numbers","cited_arxiv_id":null,"evidence_quote":"The authors' earlier enzyme-number model, from which the selective fine-tuning and frozen-layer configuration are reused."},{"cited_title":"UniProt: a worldwide hub of protein knowledge","cited_arxiv_id":null,"evidence_quote":"Supplies the protein sequences and their GO annotations used to build the random and clustered splits."},{"cited_title":"UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches","cited_arxiv_id":null,"evidence_quote":"Provides the sequence-clustering method used to construct the clustered split with minimal similarity between train and test."}],"review_version":1}