{"id":"51e62cc3-0762-450c-8c46-9e8cf7340670","arxiv_id":"2412.13478","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A parameter-efficient adapter that injects drug embeddings into a frozen single-cell transformer achieves state-of-the-art perturbation response prediction, including zero-shot generalization to new cell lines.","lead":"This paper adds drug-conditioned adapter layers to a frozen single-cell foundation model to predict how cells respond to drug treatments. The approach reports state-of-the-art accuracy on a benchmark, including zero-shot prediction for cell lines never seen during fine-tuning.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The R2 metric is computed on an unspecified 'top 20 DEG' subset; if DEG selection uses test perturbation labels, every reported comparison is outcome-conditioned and the SOTA claim is unsupported.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the DEG-selection protocol is unspecified while every reported improvement is expressed on that subset. I agree. Section 4.1 only says 'In our experiments, we consider the top 20 DEGs'; it does not say whether the set is derived from training data, all data, or per test sample. Appendix A.7 confirms the metric is sensitive to DEG-set size, so the choice of 20 genes is not a harmless detail. Even if the selection is benign in the authors' implementation, the current text and absence of code make the state-of-the-art claim unverifiable. The method itself is well-motivated, parameter counts are believable, and the architecture is clearly described; the concern is about the evaluation protocol, not the construction. A precise, auditable DEG-selection rule applied within training splits, plus a full-gene robustness check, would settle it. Because the reader already conditioned acceptance on resolving this, I recommend no change to the CONDITIONAL verdict.","tokens_in":14788,"tokens_out":6430,"duration_ms":57256,"concrete_test":"Request the exact DEG-selection procedure (or code). Then re-run all four tasks with a pre-registered protocol: choose the top 20 genes using only training-split data (before any test labels are accessed; e.g., rank genes by F-statistic of drug effect across training drugs and training cell lines), hold that gene set fixed for every model, and also report R2 on the full 2,000-gene set. If scDCA's margins over ChemCPA, BioLORD, and SAMS-VAE shrink or reverse under either re-evaluation, the SOTA claim depends on the unspecified DEG filtering.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim--scDCA 'consistently outperforms the baseline models for all tasks' (Sec. 4.2)--is supported only by R2 values computed on the 'top 20 DEGs' (Sec. 4.1). Nowhere is it stated how these genes are selected, whether the selection is per cell line, per drug, or global, or whether the test perturbation labels contribute to the choice. If the 20 genes are chosen using the full dataset or per test drug-cell-line pair, the evaluation is conditioned on the outcome: the models are scored only on the genes with the largest, most predictable expression changes. That does not let a model see the test label, but it can inflate every model's R2 and, more importantly, can change the ranking if scDCA's pretrained representations are stronger on high-variance genes while baselines are better on the remaining 1,980 genes. Since all quantitative claims in Fig. 3 and Tables 1-3 and the paired t-tests in Appendix A.8 use this unspecified subset, the state-of-the-art conclusion is not currently supported. The fact that the zero-shot claim rests on only three cell lines, and that no code is released to resolve the ambiguity, makes the omission material rather than cosmetic.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces scDCA, a parameter-efficient fine-tuning strategy that couples a frozen single-cell foundation model (scGPT) with drug-conditional adapter layers. The molecule embedding from ChemBERTa is injected into each transformer layer via learned biases, while the original scGPT weights remain frozen, so that less than 1% of the backbone parameters are trainable. The method is evaluated on sci-Plex 3 in four settings: unseen drugs, unseen drug-cell-line combinations, few-shot unseen cell lines, and zero-shot unseen cell lines, with R2 computed on the top 20 differentially expressed genes. The authors report that scDCA outperforms ChemCPA, BioLORD, and SAMS-VAE across all tasks, with the largest gains in zero-shot and few-shot cell-line generalization, and support this with ablations, target-robustness analyses, and paired t-tests.","tokens_in":15051,"tokens_out":4815,"duration_ms":48081,"significance":"If the reported results are credible, the paper makes a useful contribution by showing that a cross-modal adapter, rather than full fine-tuning, can adapt a single-cell foundation model to chemical perturbation prediction and can generalize to cell lines not seen during fine-tuning. The study is reasonably broad: it compares several recent baselines, includes ablations of the adapter design, examines robustness across drug target clusters, and reports per-cell-line zero-shot results. The main barrier to accepting the central claim is that the evaluation protocol for selecting the genes on which all R2 metrics are computed is underspecified, and the zero-shot claim rests on only three cell lines. These issues are fixable with additional clarification and experiments, so the work warrants a major revision rather than rejection.","major_comments":[{"comment":"The protocol for selecting the top 20 differentially expressed genes (DEGs) is never specified. The text states only that the analysis 'targets those genes that exhibit noticeable changes in their expression levels compared to their initial (control) expression profiles' and that 'we consider the top 20 DEGs,' but it does not say whether the DEGs are chosen per drug, per cell line, or globally, nor whether the selection uses training-set labels only, the full dataset, or the test perturbation outcomes. Because every reported R2 value in Figure 3, Tables 1-3, and the paired t-tests in Appendix A.8 is computed on this subset, an outcome-dependent DEG choice could inflate and re-rank all models. The authors must specify the exact DEG-selection algorithm and report results on the full 2,000-gene set or on a pre-registered, training-only DEG criterion, so that the state-of-the-art claim can be verified.","section":"Section 4.1, Evaluation Metrics"},{"comment":"The zero-shot cell-line evaluation uses leave-one-out over only three cell lines (A549, K562, MCF7), but the paper reports means and standard errors over five runs with 'different random splits.' With three cell lines there are only three possible leave-one-out splits, so it is unclear what the five runs represent and what the degrees of freedom are for the paired t-tests in Tables 4-7. The authors should clarify the run structure and report per-split results for every method, including the fine-tuning and ablation baselines, so that the statistical significance claims are interpretable.","section":"Appendix A.9 and Section 4.1"},{"comment":"The claim of zero-shot generalization to 'unseen cell lines' is weakened by the fact that the scGPT pretraining corpus contains many human cell types and likely includes the A549, MCF7, and K562 cell lines. The model has therefore seen these biological contexts during pretraining, even if they are absent from the fine-tuning data. The manuscript should explicitly acknowledge this, quantify the overlap with pretraining data if possible, and ideally validate the method on a cell type or tissue that is genuinely absent from the pretraining corpus, to support the stronger novelty claim.","section":"Section 3.1 and Section 4.2, unseen cell line tasks"},{"comment":"The claim that 'scDCA predictions are within single-cell measurement uncertainty' is supported only by two example molecules (Quisinostat and Dacinostat) in Figure 4, with no quantitative definition of the uncertainty range and no aggregate statistics across the test set. The authors should define the measurement-noise threshold explicitly and report, for example, the fraction of test drugs and genes for which the prediction error lies within one or two standard deviations of the single-cell distribution, rather than relying on two selected examples.","section":"Section 4.2, Figure 4"}],"minor_comments":[{"comment":"In the first paragraph, the reference 'e.g., (Roohani et al., 2024)' appears mid-sentence without a closing period, and the quotation mark after '10^60' is unbalanced; these should be corrected.","section":"Section 2, Related Work"},{"comment":"The phrase 'as input for for all models considered here' contains a duplicated 'for'; please fix the typo.","section":"Section 3.1, Preliminaries"},{"comment":"The text says scDCA 'consistently outperforms the naive finetuning approach across all tasks,' but the unseen-drug row shows identical means (0.81 for both methods). Please qualify the claim or explain why the difference is not reflected in the reported means.","section":"Section 4.2, Table 1"},{"comment":"The notation X^(0)(d) for the control expression vector should likely be X^(0)(c), since the control profile depends on the cell line, not the drug; please correct the notation.","section":"Appendix A.2, Preprocessing"},{"comment":"The sentence 'ensuring that all annotated compounds are reserved for the test set' suggests that no annotated compound appears in the training set, but the target-cluster evaluation then has no in-training examples for those targets. Please clarify how the training set is formed for this experiment.","section":"Appendix A.6, Additional discussion"},{"comment":"The DEG-size sensitivity analysis in Figure 8 does not state whether the same unspecified DEG-selection criterion is used for each size; the figure is difficult to interpret without this information.","section":"Appendix A.7, Figure 8"},{"comment":"The summary sentence says scDCA improves over 'all baselines' for all tasks, but the zero-shot table (Table 7) contains no ChemCPA row because ChemCPA cannot run that task; please rephrase to reflect the actual comparisons made.","section":"Appendix A.8, Paired t-tests"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important problem and the proposed method is sensible, but the evaluation protocol is currently too underspecified to support the state-of-the-art claim. The DEG-selection ambiguity is the most serious issue because every quantitative result is computed on that subset. I recommend major revision with a request for the exact DEG-selection code or a precise algorithmic description, additional results on the full gene set, and a clearer description of the zero-shot run structure. Releasing code would substantially help reviewers and readers resolve the remaining ambiguity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: scDCA is a sensible, parameter-efficient way to adapt scGPT to chemical perturbations, and the unseen-cell-line zero-shot evaluation is a genuinely new piece of work. But the paper currently under-specifies how the top-20 DEGs are chosen for the R2 metric, and every headline number is computed on that subset. That is a load-bearing gap, not a cosmetic one.\n\nWhat is actually new: the adapter architecture — a small two-layer MLP that maps a ChemBERTa molecular embedding to biases for the down- and up-projections inside each adapter, with scGPT frozen — is a clean design. Training <1% of the parameters and still beating full fine-tuning, especially zero-shot, is a real result if the evaluation holds up. The four-task split (unseen drug, unseen drug-cell-line, few-shot cell line, zero-shot cell line) is well thought through, and the paper includes ablations, paired t-tests, per-cell-line results, target-cluster robustness, and training-time/memory comparisons. The writing is clear and the limitations section is honest.\n\nThe soft spot: Section 4.1 says 'we consider the top 20 DEGs' but never says how those genes are selected, whether the selection is global or per cell line/drug, or whether the test outcome labels contribute. Since all R2 values and the t-tests in A.8 use only this unspecified subset, the comparison could be conditioned on the outcome. I don't think the authors intended leakage — the approach reads as good-faith — but the omission makes the main claim unverifiable as written. The zero-shot cell line section is also thin: three cell lines via leave-one-out, and scGPT was pre-trained on public atlases that very likely include those lines, so 'unseen' is overstated. No code/data release compounds both problems. The 'predictions within measurement uncertainty' claim is supported by visual inspection of a couple of molecules, not a formal test; I'd call that minor.\n\nWho this is for: anyone working on perturbation prediction or fine-tuning single-cell FMs. The central idea is worth knowing about; the evaluation protocol is worth arguing about. I'd send it to review, but with a request that the authors specify the DEG selection, report results on a fixed gene set and on all genes, and release code so the splits can be checked.","headline":"Useful adapter method, but the DEG selection is underspecified and every headline R2 depends on it; the zero-shot claim is also thin at three cell lines.","tokens_in":15568,"tokens_out":2357,"would_cite":true,"duration_ms":22121,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T07","92C40"],"pacs":[],"model":"deepseek-v4-flash","headline":"Frozen single-cell foundation models can predict drug responses in unseen cell lines with a drug-conditional adapter.","keywords":["single-cell foundation models","drug-conditional adapter","perturbation prediction","parameter-efficient fine-tuning","zero-shot generalization","scGPT","ChemBERTa","sci-Plex3"],"falsifier":"Recompute R2 on a fixed gene set or on a DEG set chosen a priori from control versus treated data without using test labels, and check whether scDCA still beats the baselines on zero-shot cell lines; the current reported numbers cannot discriminate because the DEG-selection step is unspecified.","tokens_in":14609,"feed_emoji":"💊","tokens_out":4582,"duration_ms":38281,"temperature":0.7,"pith_summary":"The paper aims to show that a single-cell foundation model can be adapted to predict transcriptional responses to chemical perturbations by freezing the foundation model and training only small drug-conditioned adapter layers. The proposed method, scDCA, injects a molecule embedding into every transformer layer as biases for down- and up-projections, so the model stays in the gene-expression input space seen during pre-training. The authors report state-of-the-art R2 on sci-Plex3 across four generalization tasks, with the largest gains on unseen-cell-line zero- and few-shot settings, while training less than 1% of the original model parameters. If correct, this gives a practical way to extend pre-trained cell atlases to virtual screening and new biological contexts with very little paired perturbation data.","feed_headline":"Drug-conditional adapter predicts cell responses to unseen drugs","feed_subtitle":"A frozen scGPT with fewer than 1% trainable parameters beats full fine-tuning on zero-shot cell-line tasks.","key_machinery":"The drug-conditional adapter. In each of the 12 transformer layers of scGPT, a two-layer network $f^m_l$ maps a frozen ChemBERTa molecule embedding $\\mathrm{emb}_m(d)$ to a bias vector $b_l$. The adapter transforms the layer hidden state as $h^{\\mathrm{down}}_l = W^{\\mathrm{down}}_l h_l + \\Pi^{\\mathrm{down}} b_l$, passes it through a residual network, and outputs $h^{\\mathrm{up}}_l = W^{\\mathrm{up}}_l h^{\\mathrm{res\\_net}}_l + b_l$; the bias is what carries the drug signal into every layer. Original scGPT weights, gene tokens, expression embedding, and molecule encoder are frozen, so the input distribution stays close to pre-training and only the adapter parameters are learned, which the paper reports as less than 1% of the original model.","core_discovery":"The central claim is that molecular conditioning can be added to a frozen single-cell transformer through trainable adapter layers whose biases are generated from a molecular embedding, and that this preserves the pre-trained biological representation well enough to predict perturbation outcomes for drugs and cell lines never seen during fine-tuning. On the sci-Plex3 dataset (A549, MCF7, K562; 188 drugs), scDCA reports R2 of 0.81 for unseen drugs, 0.83 for unseen drug-cell-line combinations, 0.88 for few-shot unseen cell lines, and 0.82 for zero-shot unseen cell lines, outperforming full fine-tuning of scGPT and the ChemCPA, BioLORD, and SAMS-VAE baselines, especially in zero-shot cell-line generalization. The authors frame the result as evidence that careful parameter-efficient fine-tuning, not naive fine-tuning or feature extraction, is what unlocks foundation-model knowledge for molecular perturbation prediction.","pith_inferences":["If DEG selection is made label-free and fixed in advance, the zero-shot advantage may shrink or vanish; the paper's load-bearing comparison should be re-run under such a protocol.","The same bias-from-another-modality adapter could condition single-cell foundation models on other inputs, such as genetic perturbations or growth conditions, since the machinery does not depend on chemistry-specific details.","With only three cell lines, leave-one-out results are a weak test of unseen-cell-line generalization; validation on more diverse cell lines would show whether the effect holds beyond adenocarcinoma and leukemia splits.","The authors' own appendix shows performance varies by held-out line, so practical deployment would need a way to estimate which new cell lines are likely to be predictable."],"forward_implications":["A frozen single-cell foundation model plus small adapters can generalize to entirely unseen cell lines in zero-shot mode, provided control gene expression for the new cell line is available.","Because the adapter accepts any molecular embedding, the same recipe extends to other molecule encoders without retraining the single-cell foundation model.","The performance gap over baselines grows as the task becomes harder, suggesting the pre-trained representation is the main carrier of generalization to new biological contexts.","The approach is limited to transformer-based foundation models and requires control expression data, so it cannot be applied directly to datasets lacking controls or to non-transformer architectures."],"supporting_citations":[{"why":"Supplies scGPT, the frozen single-cell foundation model and its pre-training objective that scDCA adapts.","marker":"Cui et al., 2024"},{"why":"Supplies the sci-Plex3 dataset of 649,340 cells, 188 drugs, and 3 cell lines used in all evaluations.","marker":"Srivatsan et al., 2020"},{"why":"Supplies ChemBERTa, the frozen molecular encoder that produces the drug embeddings fed into the adapters.","marker":"Chithrananda et al., 2020"},{"why":"Supplies the ChemCPA baseline and the R2-on-DEG evaluation protocol that the paper adopts.","marker":"Hetzel et al., 2022"},{"why":"Supplies the GEARS baseline and the prior observation that focusing on a small set of differentially expressed genes is a robust evaluation target.","marker":"Roohani et al., 2024"},{"why":"Supplies the BioLORD baseline, a disentangled single-cell generative framework.","marker":"Piran et al., 2024"},{"why":"Supplies the SAMS-VAE baseline, an additive mechanism-shift variational autoencoder for perturbation modelling.","marker":"Bereket & Karaletsos, 2024"}],"fun_headline_variants":["Frozen scGPT with <1% adapter nails zero-shot drug responses","Tiny adapter turns single-cell AI into zero-shot perturbation predictor","Almost no fine-tuning needed: scDCA predicts unseen cell lines","Less than 1% trainable beats full fine-tuning for perturbation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All reported R2 values are computed on the top 20 differentially expressed genes, but the paper never specifies how those genes are selected; if the selection uses the test perturbation labels, the reported performance advantage could be an artifact of the chosen gene subset rather than a property of the model.","fun_headline_variants_meta":{"raw":{"variants":["Frozen scGPT with <1% adapter nails zero-shot drug responses","Tiny adapter turns single-cell AI into zero-shot perturbation predictor","Almost no fine-tuning needed: scDCA predicts unseen cell lines","Less than 1% trainable beats full fine-tuning for perturbation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000511,"raw_usage":{"total_tokens":2470,"prompt_tokens":912,"completion_tokens":1558,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":1483}},"tokens_in":528,"tokens_out":1558,"duration_ms":9776,"temperature":1.0,"reasoning_tokens":1483,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:05:03.682494+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute R2 on a fixed gene set or on a DEG set chosen a priori from control versus treated data without using test labels, and check whether scDCA still beats the baselines on zero-shot cell lines; the current reported numbers cannot discriminate because the DEG-selection step is unspecified.","supporting_citations":[{"cited_title":"scgpt: toward building a foundation model for single-cell multi-omics using generative ai","cited_arxiv_id":null,"evidence_quote":"Supplies scGPT, the frozen single-cell foundation model and its pre-training objective that scDCA adapts."},{"cited_title":"Chemberta: Large-scale self-supervised pretraining for molecular property prediction","cited_arxiv_id":null,"evidence_quote":"Supplies ChemBERTa, the frozen molecular encoder that produces the drug embeddings fed into the adapters."},{"cited_title":"Predicting cellular responses to novel drug perturbations at a single-cell resolution","cited_arxiv_id":null,"evidence_quote":"Supplies the ChemCPA baseline and the R2-on-DEG evaluation protocol that the paper adopts."},{"cited_title":"Disentanglement of single-cell data with biolord","cited_arxiv_id":null,"evidence_quote":"Supplies the BioLORD baseline, a disentangled single-cell generative framework."}],"review_version":1}