{"id":"e6904108-0170-40ab-947c-c84b491ba6a1","arxiv_id":"2501.07737","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Phenformer uses frozen Enformer embeddings of 512 gene windows to predict disease risk and cell type involvement, outperforming PRS baselines restricted to the same genomic regions.","lead":"This paper presents Phenformer, a deep learning model that reads DNA from 512 gene windows, totaling about 88 million base pairs, to predict risk for six common diseases and to point out which cell types and tissues may be involved. The results show small but significant improvements in risk prediction when combined with standard polygenic scores, and better agreement with published cell type-disease links than previous methods.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The mechanistic-superiority claim rests on an unvalidated LLM-generated literature gold standard and overlap-restricted baseline comparisons; a robust re-evaluation could overturn it.","rationale":"The reader's weakest assumption concerned the frozen Enformer embeddings being sufficient to capture variant effects. That is a real concern, acknowledged in the Discussion, but the paper's empirical results (cell-type F1, PRS ensemble gains) suggest the embeddings carry at least some usable signal. I find a more load-bearing problem in the evaluation supporting the paper's most novel claim: superiority over scRNA-based methods for mechanistic discovery. The gold standard is generated by an LLM without human validation, thresholds are arbitrary, and the comparison protocol (overlap-only, with a baseline augmented by data from a subset of diseases) can inflate Phenformer's relative performance. If this evaluation is unreliable, the paper loses its headline contribution, even if the risk prediction results stand. The concern is addressable by re-running the comparison with a validated gold standard and a fairer protocol, so the paper should remain CONDITIONAL rather than being rejected. I give credit for the large-scale end-to-end architecture, the honest limitation statements, and the modest but consistent PRS ensemble improvements, all of which are valuable independent of the mechanistic comparison.","tokens_in":50873,"tokens_out":3857,"duration_ms":44089,"concrete_test":"Recompute the Figure 2 comparison using a validated gold standard: (1) take a random sample of 200 disease–cell-type abstract pairs, have two independent human curators score them with the same -5..5 scale, measure inter-annotator agreement and the agreement of Claude Sonnet with the human consensus; (2) pre-specify the enrichment threshold before computing F1; (3) recompute Phenformer's F1 both on the full disease/cell-type matrix and on each baseline's own output set, rather than only the overlap. If Phenformer does not retain a significant F1 advantage under human-validated labels and full coverage, the mechanistic-superiority claim is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's most distinctive claim is that Phenformer's predicted disease–cell-type mechanisms match literature better than state-of-the-art methods that additionally use scRNA-seq (Figure 2). This claim depends entirely on the evaluation in the Section 'Cell type-disease associations supported by literature' and 'Baseline methods for cell type identification'. The gold standard is produced by Claude Sonnet scoring PubMed abstracts without any validation against human expert labels, using hand-set thresholds (at least 5% enrichment; at least 5 abstracts with score >= 4). The F1 scores in Figure 2a have no confidence intervals and are computed at a single arbitrary threshold. Baselines are compared on the overlap of diseases and cell types for which each method provides predictions, and for Jagadeesh et al. the authors generate associations from scRNA-seq datasets for only three diseases (T1D, psoriasis, COPD) specifically to increase overlap. This restricted overlap can systematically favor Phenformer, which outputs a fixed set of 21 cell types for all six diseases, while baselines may output fewer or different cell types. If the LLM gold standard is noisy or biased, or if the overlap restriction favors Phenformer, the central claim that sequence-only Phenformer beats methods requiring additional experimental data is not established. The risk prediction results (Figure 3) are more robust but less distinctive; the mechanistic comparison is what elevates the paper's significance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces Phenformer, a Transformer model that consumes 512 Enformer-derived embedding tokens centered on 196-kb transcription-start-site windows and predicts disease status for six diseases using UK Biobank whole-genome sequencing data from roughly 150,000 individuals. The authors report three main results: (i) Phenformer's cell-type and tissue attributions for disease match PubMed literature with higher F1 than established methods that additionally use single-cell RNA-seq; (ii) logistic-regression ensembles of Phenformer with five polygenic risk score methods improve held-out AUROC, with larger relative gains in non-European ancestry; and (iii) UMAP/HDBSCAN clustering of individual attributions produces comorbidity-associated disease subtypes. The abstract frames the contribution as multi-megabase-scale, sequence-only genome interpretation that both generates mechanistic hypotheses and improves disease risk prediction.","tokens_in":51084,"tokens_out":5200,"duration_ms":56789,"significance":"If the principal claims hold, this would be a meaningful advance: a sequence-only model with roughly 88 Mb of context that both predicts disease risk and proposes cell-type-level mechanisms would go beyond existing PRS and variant-effect methods. The study has clear strengths: a large WGS cohort, a fixed 60/20/20 split, six diseases, bootstrap uncertainty estimates, ancestry-stratified evaluation, and reliance on public data for the main training resource. The risk-prediction component is the more robust contribution: the reported gains are modest but consistent, and the ensemble comparison is in principle reproducible. However, the headline mechanistic claim currently depends on an unvalidated LLM-generated literature gold standard and on an overlap-restricted baseline design, and the risk comparison is labeled 'whole-genome' even though the PRS baselines are constructed from the same 3% of the genome as Phenformer. The manuscript is a strong candidate for publication after the load-bearing evaluation points are reworked.","major_comments":[{"comment":"The central claim that Phenformer outperforms state-of-the-art methods requiring scRNA-seq is not yet established. The pairwise overlap design and the decision to generate Jagadeesh et al. associations only for T1D, psoriasis, and COPD can systematically favor Phenformer, which emits a fixed set of 21 cell types for all six diseases while baselines may output fewer or different cell types. Please report the full disease-by-cell-type association matrices for every method, evaluate on a common predefined set of cell types rather than only the overlap, provide per-disease F1 with confidence intervals, and state explicitly how differences in the number and granularity of predicted cell types are handled.","section":"Baseline methods for cell type identification; Figure 2"},{"comment":"The literature gold standard is generated by Claude Sonnet with no reported validation against human expert labels, and the thresholds used to derive enrichment (at least 5% enrichment; at least 5 abstracts with score >= 4) are hand-set. Since the F1 values in Figure 2a are computed at these single thresholds, small changes in scoring behavior or threshold choice could change the ranking of methods. Please validate the LLM scores on a human-curated sample, report agreement statistics, and provide a threshold sweep or precision-recall analysis to show that the reported ranking is not an artifact of the chosen cutoffs.","section":"Cell type-disease associations supported by literature"},{"comment":"The labels 'whole-genome level' in Figure 3a-b and in the Results text are not supported by the Methods, which state that the GWAS and PRS baselines were trained using the 'exact same genomic information as provided to Phenformer' (i.e., SNPs inside the 512 selected windows). These panels therefore compare Phenformer to 3%-restricted PRS, not to genome-wide PRS. Please either add genuine whole-genome PRS baselines or reword the manuscript, including the abstract and figure captions, so that the comparison is described as restricted to the 3% of the genome used by Phenformer.","section":"Baselines; Figure 3"},{"comment":"All Phenformer inputs are frozen Enformer embeddings of 196-kb TSS-centered windows, and the paper itself acknowledges in the Discussion (citing Sasse et al., ref 64) that sequence-to-expression backbones 'perform not particularly well' at variant-induced effect prediction. This is load-bearing because any variant effects not captured by Enformer are invisible to Phenformer for both risk prediction and mechanism attribution. Please include a direct evaluation of the backbone's variant sensitivity in the relevant setting, such as an eQTL/MPRA benchmark on the selected windows, a comparison with an alternative sequence embedding, or an analysis of how many test-set variants actually alter the embeddings; absent such evidence, the mechanistic claims should be explicitly conditional on the backbone's variant-effect sensitivity.","section":"Section 4.1 Step 1; Discussion limitation"},{"comment":"The 512-gene window set is selected using Enformer-predicted CAGE changes in a 100-case/50-control psoriasis cohort and then used for all six diseases. This makes the input choice outcome-dependent for psoriasis and raises the possibility that the favorable psoriasis cell-type results in Figure 2, and any disease-general conclusions, are influenced by the gene selection step. Please report sensitivity of the main results to alternative gene sets (e.g., random gene sets of the same size or disease-specific sets) and clarify whether the 150 individuals used for gene selection overlap with the training set used for Phenformer.","section":"Gene set selection; Section 4.2 Data"}],"minor_comments":[{"comment":"The rendered draft contains uninterpretable '/uni...' glyph strings in Figures S1 and S2; these figure panels should be regenerated before resubmission.","section":"Supplementary Figures S1 and S2"},{"comment":"The cohort size is given as 150119 in one sentence and 150076 in the next; please clarify whether these are different inclusion criteria and harmonize the wording.","section":"Section 4.2 Data"},{"comment":"The saliency aggregation is restricted to 'true positive samples,' but the threshold that defines a true positive prediction is not specified; please state the operating point used.","section":"Model interpretation"},{"comment":"The code is promised 'upon publication'; for a computational manuscript of this type, please consider making the code, trained model weights, and exact hyperparameter settings available to reviewers, or at minimum specify all baseline software versions and random seeds in the Methods.","section":"Code availability"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially high-impact, but the mechanistic-superiority claim should not be accepted without the requested reanalysis: a human-validated literature gold standard, fair cell-type baseline comparisons, and a corrected whole-genome framing for the risk-prediction experiments. The risk-prediction results are more solid and could support a revised manuscript even if the mechanistic claim is substantially qualified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — worth reading, but the headline claim is currently oversold. The genuinely new contribution is the architecture: 512 TSS-centered 196 kb windows, frozen Enformer embeddings as tokens, a transformer with PMA pooling, trained on 150k UK Biobank WGS samples to predict six diseases. That is a real piece of engineering, and the risk prediction results show consistent, though modest, gains over PRS baselines built on the same subset — especially in non-European ancestry, where the transportability story is credible. The paper is also honest in the Discussion about the 3% genome coverage and the known weakness of Enformer at variant-effect prediction.\n\nThe soft spots are in the evaluation of the mechanistic claims. The cell-type comparison (Figure 2) depends on a literature gold standard generated by Claude Sonnet scoring PubMed abstracts, with hand-set thresholds and no human validation. The F1 scores have no confidence intervals. Baselines are compared only on the overlap of diseases and cell types, and for Jagadeesh et al. the authors ran scRNA-seq on just three diseases to increase overlap. Since Phenformer always outputs the same 21 cell types for all six diseases, the overlap restriction can systematically favor it. The claim that a sequence-only model beats methods using scRNA-seq is not established by this protocol.\n\nThe risk prediction comparison is cleaner but carries a labeling problem: the PRS baselines are built on SNPs within the same 512 windows — not genome-wide PRS — while Figure 3a-b is labeled 'whole genome.' That doesn't invalidate the head-to-head on the same input, but it does mean the paper does not show Phenformer beats genome-wide PRS. The gene set selection is also biased: genes were chosen using Enformer on a 100-case/50-control psoriasis cohort, then fixed for all diseases — a defensible heuristic, but a real limit for non-immune diseases that should be stated more directly.\n\nNet: the architecture and the risk prediction results deserve serious referee time. The mechanistic-superiority claim needs a better evaluation — a validated gold standard, error bars, and a fair overlap protocol — before it can be taken at face value. Send it to review with major revision.","headline":"A real engineering advance in sequence-to-phenotype modeling, but the mechanistic-superiority claim is oversold by an evaluation design that needs tightening.","tokens_in":51698,"tokens_out":3337,"would_cite":false,"duration_ms":33174,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Phenformer predicts disease risk and mechanistic hypotheses directly from up to 88 million base pairs of an individual's genome sequence.","keywords":["genetic language model","disease risk prediction","sequence-to-phenotype","Enformer","cell-type attribution","polygenic risk scores","genome interpretation","disease subtypes"],"falsifier":"Permute the disease labels and retrain Phenformer; if the cell-type and gene attributions still match literature at the reported F1 levels, the mechanistic signal is inherited from the frozen Enformer embeddings rather than learned from disease status.","tokens_in":50625,"feed_emoji":"🧬","tokens_out":8794,"duration_ms":79478,"temperature":0.7,"pith_summary":"Phenformer is a deep-learning model that takes individual genome sequence as input and predicts both disease risk and candidate mechanisms, following the information flow from DNA to cell-type-specific expression to phenotype. The paper's central claim is that a model using only sequence—through a frozen Enformer backbone that embeds 196-kilobase windows around 512 gene start sites—can produce disease-relevant cell- and tissue-type hypotheses that match published literature better than state-of-the-art methods that additionally require single-cell RNA sequencing data. The paper also claims that ensembling Phenformer with standard polygenic risk scores improves risk prediction accuracy in 86.7% of mixed-ancestry and 96.7% of non-European-ancestry disease/method combinations, with average AUROC gains up to 4.2% and 11.19%. If correct, this would make whole-genome sequence interpretation without experimental perturbation feasible at scale, and would extend risk prediction to more diverse populations.","feed_headline":"Sequence-only model beats polygenic scores on disease risk","feed_subtitle":"Phenformer reads 88 million base pairs per person and finds disease-relevant cell types without lab experiments.","key_machinery":"The central machinery is the frozen Enformer sequence-to-expression backbone used as a tokenizer, followed by a learned expression-to-phenotype transformer. Each individual is represented by 512 tokens; each token is the 3072-dimensional Enformer embedding of a 196 kb window centered on a gene's transcription start site, which encodes predicted expression and chromatin accessibility across many cell types. A shared projection maps tokens to 512 dimensions, Fourier position encodings supply genomic location, four transformer encoder layers model interactions across loci, and Pooling by Multihead Attention (PMA) collapses the set into a pooled representation for a two-layer disease-risk head. Attribution then works backward: saliency gradients on the input embeddings are used to perturb embeddings, and the same Enformer head translates the perturbation into CAGE-track changes, producing cell-type rankings. The Enformer head thus doubles as the interpretability interface that turns risk predictions into mechanistic hypotheses.","core_discovery":"The core discovery the paper argues for is that an end-to-end model can connect individual genomes to phenotypes by learning a mapping from sequence embeddings to disease risk, and that the intermediate attributions of this mapping recover known biology. Specifically, Phenformer extracts 3072-dimensional embeddings from a frozen Enformer model at 512 transcription-start-site-centered windows (about 88 million base pairs, roughly 3% of the genome), processes them through four transformer encoder layers with Fourier position encodings, pools them with multihead attention, and outputs a disease logit. Interpreting the model with saliency gradients and projecting back through Enformer's CAGE tracks yields per-window and per-cell-type importance rankings. The paper reports that these rankings achieve higher F1 against literature-curated cell-type–disease associations than five existing methods that require both genetic and single-cell RNA-seq data, and that the same model's risk predictions improve PRS ensembles and generalize better to non-European ancestries.","pith_inferences":["If the central claim holds, whole-genome risk prediction can bypass SNP lists and ancestry-specific LD panels: a frozen expression-model backbone acts as a universal tokenizer, and adding more gene windows should improve both risk and mechanism discovery.","A clean test of the mechanism claims would compare Phenformer's attributions against those from a variant-effect-tuned backbone; if literature enrichment persists, the signal is in the phenotype head, and if it disappears, it lives in the frozen embeddings.","Because only 3% of the genome (512 immune-associated windows) is used, applying the same architecture to the full set of about 21,725 transcription-start-site windows could reveal whether the non-European portability comes from the sequence-to-expression bottleneck or from the particular gene set chosen.","The literature baseline itself is produced by an LLM scoring PubMed abstracts, so part of Phenformer's apparent advantage may reflect shared vocabulary with the text-scoring process; a prospective validation on newly discovered cell-type–disease associations would be stronger evidence."],"forward_implications":["Ensembling Phenformer with standard PRS methods (Lassosum, LDpred2, PRS-CSx, Pthres, C+T) significantly improves AUROC in 86.7% of disease/method combinations in mixed-ancestry and 96.7% in non-European-ancestry test sets.","On the 512-gene windows Phenformer sees, it outperforms PRS methods built on the same windows by up to 5.49% AUROC in mixed ancestry and 14.59% in non-European ancestry.","Phenformer's cell- and tissue-type attributions match literature-reported disease associations with higher average F1 than five methods that require single-cell RNA sequencing data in addition to genetics.","The model surfaces molecular hypotheses for known but unexplained comorbidities, such as liver involvement in psoriasis and small-intestine/appendix involvement in type 1 diabetes.","Individual-level attribution embeddings cluster into subtypes with significantly different comorbidity rates, suggesting that sequence alone can stratify patients by disease mechanism."],"supporting_citations":[{"why":"Supplies the frozen Enformer sequence-to-expression backbone whose 3072-dimensional TSS-centered embeddings serve as the input tokens for Phenformer.","marker":"[2]"},{"why":"Supplies Pooling by Multihead Attention (PMA), the set-pooling mechanism that aggregates the 512 token embeddings before the risk head.","marker":"[4]"},{"why":"Describes the cohort resource from which the whole genome sequencing and phenotype data are drawn.","marker":"[7]"},{"why":"Provides the validated disease-phenotype definitions used to label the six diseases studied.","marker":"[69]"},{"why":"Supplies the sequences of 150,119 genomes that constitute the training data.","marker":"[70]"},{"why":"Cited by the paper as evidence that Enformer performs not particularly well at variant-induced effect prediction, acknowledged as a limitation of the frozen backbone.","marker":"[64]"},{"why":"Provides PRS-CSx, one of the five polygenic risk score baselines that Phenformer ensembles with and outperforms.","marker":"[28]"},{"why":"Provides LDpred2, another PRS baseline used in the ensemble and comparison.","marker":"[58]"},{"why":"One of the five cell-type identification baselines that Phenformer's literature-enrichment F1 is compared against.","marker":"[50]"},{"why":"Another cell-type identification baseline used in the pairwise comparison.","marker":"[51]"}],"fun_headline_variants":["Sequence-only Phenformer maps 88M base pairs to disease cell types","Phenformer reads 88M base pairs, finds disease cell types without lab","Phenformer: sequence-only model ties 88M base pairs to disease","Genome-scale risk: Phenformer uses only DNA to predict disease"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the frozen Enformer embeddings, computed only around 512 gene start sites, carry enough information about how an individual's genetic variants change expression and chromatin that disease risk and mechanisms can be read off from them.","fun_headline_variants_meta":{"raw":{"variants":["Sequence-only Phenformer maps 88M base pairs to disease cell types","Phenformer reads 88M base pairs, finds disease cell types without lab","Phenformer: sequence-only model ties 88M base pairs to disease","Genome-scale risk: Phenformer uses only DNA to predict disease"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001252,"raw_usage":{"total_tokens":5120,"prompt_tokens":920,"completion_tokens":4200,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":4119}},"tokens_in":536,"tokens_out":4200,"duration_ms":26643,"temperature":1.0,"reasoning_tokens":4119,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:37:04.686596+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Permute the disease labels and retrain Phenformer; if the cell-type and gene attributions still match literature at the reported F1 levels, the mechanistic signal is inherited from the frozen Enformer embeddings rather than learned from disease status.","supporting_citations":[{"cited_title":"Halldorsson, Hannes P","cited_arxiv_id":null,"evidence_quote":"Supplies the sequences of 150,119 genomes that constitute the training data."},{"cited_title":"How far are we from personalized gene expression prediction using sequence-to-expression deep neural networks? bioRxiv, pages 2023–03, 2023","cited_arxiv_id":null,"evidence_quote":"Cited by the paper as evidence that Enformer performs not particularly well at variant-induced effect prediction, acknowledged as a limitation of the frozen backbone."},{"cited_title":"Ldpred2: better, faster, stronger","cited_arxiv_id":null,"evidence_quote":"Provides LDpred2, another PRS baseline used in the ensemble and comparison."},{"cited_title":"Estimating the causal tissues for complex traits and diseases.Nature genetics, 49(12):1676–1683, 2017","cited_arxiv_id":null,"evidence_quote":"One of the five cell-type identification baselines that Phenformer's literature-enrichment F1 is compared against."},{"cited_title":"Heritability enrichment of specifically expressed genes identifies disease-relevant tissues and cell types.Nature genetics, 50(4):621–629, 2018","cited_arxiv_id":null,"evidence_quote":"Another cell-type identification baseline used in the pairwise comparison."}],"review_version":1}