{"id":"1dbc3ea9-67de-4ae2-a294-e8a2c3c85892","arxiv_id":"1908.06197","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Applying PCA-based differential identifiability to functional connectomes improves the stability and, by a modest margin, the accuracy of predicting cognitive deficits in Alzheimer's disease.","lead":"Researchers combined a PCA-based denoising step with connectome predictive modeling to predict cognitive test scores from brain scans of Alzheimer's patients. They report that denoising makes feature selection more stable and modestly improves prediction on a separate validation set.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation re-estimates the 𝕀𝑓 PCA basis on the validation cohort, so the reported 'external' gains may reflect test-set adaptation rather than a deployable fixed model.","rationale":"The reader's conditional verdict already covers the main problem, but their weakest_assumption emphasizes the split-half proxy rather than the validation-set PCA refit. I agree the split-half proxy is an acknowledged idealization, but the more directly actionable issue for the external claim is that the validation pipeline re-estimates the denoising basis on the validation cohort. This is not an accusation of dishonest leakage: unsupervised PCA on unlabeled test features can be a legitimate transductive procedure, but it is not the same as evaluating a fixed model on external data, and the manuscript does not quantify how much of the 5/7 improvement depends on it. The paper has real internal support: the consistency of the improvements across outcomes, the restA/restB stability curves, and the RSN associations are suggestive that the method captures meaningful signal. Yet the headline claim invokes 'external FC data,' and that claim is only secure once validation is run with a training-fixed basis. The limitation section itself calls for ADNI3 as a truly external test, which reinforces that the current external-validation language is premature. A fixed-basis rerun on the existing data is cheap and would settle the concern; if it confirms the results, the paper's core contribution stands. I therefore keep the reader's conditional verdict rather than moving to reject or accept as-is.","tokens_in":16848,"tokens_out":6397,"duration_ms":72390,"concrete_test":"Re-run the Section 2.4 validation loop while fixing the PCA basis and the number of PCs from the training cohort only, then apply that fixed reconstruction operator to validation FCs instead of refitting 𝕀𝑓 on the validation cohort. If five of seven outcomes remain significantly improved over original FCs, the external claim survives; if the improvements shrink or lose significance, the headline depends on validation-set adaptation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.4 states that in each split-half fold '𝕀𝑓 was performed separately on training versus validation subjects,' and Section 3.1 reports Idiff peaking at 41 PCs in the validation cohort as well. This means the PCA basis and the PC count used to denoise validation FCs are estimated from the same validation FCs (restA/restB) whose cognition is then predicted. The claim in Section 4.2 that this 'fully maintain[s] the independence of training data and validation data' is too strong: outcome labels are not used, but the validation set does determine the preprocessing transform. Because the two halves come from the same session (Section 2.2), the Idiff used to select this transform is a within-session quantity and may capitalize on shared physiological noise. In a real deployment a single new subject cannot be denoised by re-estimating a group-level PCA basis; a fixed basis from the training cohort would be needed. The reported 5/7 significant validation improvements therefore measure a transductive pipeline rather than a fixed model's external generalizability, and may overstate the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a combined pipeline in which functional connectomes (FCs) from restA/restB halves of a single resting-state fMRI session are denoised by group-level PCA at the number of principal components that maximize differential identifiability (Idiff), and the denoised FCs are then used in connectome predictive modeling (CPM) to predict seven cognitive outcomes in a cohort of 82 ADNI2 participants spanning the Alzheimer's disease spectrum. The authors evaluate (i) stability of edge selection between restA and restB, (ii) specificity of edge selection across outcomes, (iii) restA-to-restB prediction within the training cohort, and (iv) generalization to a held-out validation cohort, comparing original FCs against Idiff-optimally reconstructed FCs. They report significantly improved validation correlations for five of seven outcomes and identify resting-state network interactions associated with each outcome.","tokens_in":17061,"tokens_out":8251,"duration_ms":79345,"significance":"The contribution is potentially valuable: it is one of the first to connect differential-identifiability denoising to predictive modeling of cognition in a clinical population, and the emphasis on feature-selection stability before model fitting is a useful methodological message. The split-half test-retest evaluation within the training cohort is a sensible minimum standard, and the RSN-overrepresentation analyses provide interpretable, hypothesis-generating maps. However, the headline validation result rests on a transductive preprocessing procedure and a statistical comparison that treats non-independent splits as independent. If the validation gains survive a training-locked preprocessing transform and a subject-clustered permutation test, the framework would be an important practical advance; the current evidence is suggestive rather than conclusive.","major_comments":[{"comment":"The validation pipeline re-estimates the PCA basis on the validation cohort, so the claim of 'external' generalization is not supported for a fixed model. Section 2.4 states that 𝕀𝑓 was 'performed separately on training versus validation subjects,' and Section 3.1 reports Idiff peaking at 41 PCs in the validation cohort as well. Thus the PCA basis and the PC count used to denoise validation FCs are estimated from the same validation subjects whose outcomes are later predicted. The statement in Section 4.2 that this 'fully maintain[s] the independence of training data and validation data' is too strong: outcome labels are not used, but the validation set does determine the preprocessing transform. In a real deployment a single new subject cannot be denoised by re-estimating a group-level PCA basis; a fixed basis from the training cohort would be needed. To support the central claim, the validation should be repeated with the PCA basis and the PC count locked from training data only, and the 'external' language should be reframed accordingly.","section":"Section 2.4, Section 4.2, Figure 6"},{"comment":"The paired t-test across the 200 random splits is not statistically valid because the same subjects appear in multiple folds. Each fold is a random permutation of the same 82 subjects into training and validation halves, so the 200 Pearson correlation values are not independent samples. This dependence likely makes the reported p-values anti-conservative and could affect the claim of significant improvement for 5 of 7 outcomes. Please use a permutation test that resamples subjects (e.g., permuting the outcome labels once and running the full pipeline) or a cluster-based bootstrap that accounts for subjects appearing in multiple folds, and report effect sizes with appropriate confidence intervals.","section":"Section 3.1.3, Figure 6"},{"comment":"The 'test-retest' evaluation is based on splitting a single fMRI session into restA and restB halves, which share physiological noise, head motion, and scanner drift. This is a best-case reliability scenario, and the denoising gains observed in validation may partly reflect removal of within-session artifacts that are not reproducible across true test-retest sessions. The paper acknowledges the need for external validation in Section 4.4, but the abstract and results use 'test-retest' and 'external data' without qualification; these terms should be replaced by 'split-half within-session' and 'held-out cohort from the same dataset' to avoid overstating generalizability.","section":"Section 2.2"}],"minor_comments":[{"comment":"Section 3.1 reports mean variance explained of 71.86% (training) and 71.80% (validation) at the optimal reconstruction, while Section 4.1 states that 'Optimal reconstruction retained 80% of the variance in the data (Fig. 3 purple line)'; these numbers are inconsistent and should be reconciled.","section":"Sections 3.1 and 4.1"},{"comment":"Section 3.1.1 reports restA-restB mask overlap in the optimal range as '68% Clock Score – 76% Trails B', while Section 4.2 states an 'average peak overlap of 65% across outcome measures'; since 65% is below the reported range, one of the statements is likely an error and should be corrected.","section":"Sections 3.1.1 and 4.2"},{"comment":"The supplementary figure captions refer to '1000 folds' while Figure 6 and the main text refer to '200 folds'; Section 2.4 does not specify the number of folds. Please clarify the number of random splits and use consistent terminology throughout.","section":"Supplementary Figures S1-S2 and Figure 6"},{"comment":"The terms 'training', 'testing', and 'validation' are used inconsistently (e.g., Figure 3 caption says 'Testing Cohort' while Section 2.4 uses 'validation cohort'); please adopt one consistent naming convention and define whether the training/validation splits are stratified by diagnostic group.","section":"Sections 2.4 and 3"},{"comment":"There are minor typographical issues, including 'rest/restest' and a missing space in '𝕀𝑓More importantly'; please proofread the manuscript.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper's core idea is interesting and the training-level stability analyses are solid, but the validation protocol as written cannot support the 'external' generalization claim because the preprocessing transform is re-estimated on the validation cohort. The authors should be asked to rerun the validation with a training-locked PCA basis and PC count, and to provide a statistically valid comparison across folds. If the revised analysis no longer shows significant improvements, the paper would still be publishable as a methodological demonstration of improved feature stability, but the abstract's stronger claim would need to be softened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a solid, workmanlike extension of the differential identifiability (Idiff) framework to connectome predictive modeling in Alzheimer's disease. What's genuinely new is the systematic evaluation of test-retest feature-selection stability, edge-selection specificity, and model generalizability across seven cognitive outcomes. The training-side results are convincing: rebuilding models on restA and testing on restB shows a clear peak in performance near the Idiff-optimal PC count, and the increase in mask overlap is substantial. The RSN association analyses are also a nice addition, tying the improved edge selection to interpretable networks. I'd give credit for the breadth of outcomes and the careful reporting of stability metrics.\n\nThe soft spot is the external validation step, and it's not minor. The PCA basis and the PC count are re-estimated on the validation cohort itself. That means the 'optimally reconstructed' validation FCs are not produced by a fixed, deployable transform; they're produced by a transform that has already 'seen' the validation subjects' resting-state data. The paper's claim that this 'fully maintains the independence of training data and validation data' is too strong. It doesn't use outcome labels, but it does use the validation FCs to determine the preprocessing. In practice, a new subject would need a basis estimated from training data alone, and we don't know how much of the reported 5/7 improvement survives under that protocol. The paired t-tests across folds also treat the 200 folds as independent, which they aren't. That inflates significance. The split-half proxy for test-retest is another caveat: within-session halves share physiological noise, so the gains could be somewhat optimistic.\n\nThat said, the core effect is probably real. The training-side test-retest improvement doesn't suffer from the validation leakage, and the consistency across all seven outcomes suggests the method does something useful. The absolute correlations are modest (0.20-0.33), so I wouldn't oversell clinical near-term impact, but as a preprocessing step it looks promising.\n\nThis paper deserves a serious peer review, but it needs a corrected validation protocol: fix the PCA basis from training only, and use proper statistical tests that account for fold non-independence. As it stands, the central claim is overstated, though the underlying approach is sound. I'd engage with it, but I'd want the revision before citing it as evidence for external generalizability.","headline":"Useful systematic demonstration that Idiff improves test-retest stability of CPM in AD, but the external-validation claim is weakened by re-estimating the PCA transform on the validation cohort.","tokens_in":17622,"tokens_out":1232,"would_cite":false,"duration_ms":13889,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a PCA-based denoising step applied to resting-state functional connectivity before predictive modeling makes connectome models more stable and improves prediction of cognitive deficits in Alzheimer's disease.","keywords":["Alzheimer's disease","functional connectivity","differential identifiability","connectome predictive modeling","resting-state fMRI","cognitive outcomes","test-retest reliability","PCA denoising"],"falsifier":"Run the same $\\mathbb{I}_f$-then-CPM pipeline on a dataset where each person has true retest fMRI from a different day or a different scanner: if the optimal PCA reconstruction fails to increase rest-to-retest edge overlap and validation correlations above those of raw connectomes, or if the model improvements disappear when training and validation connectomes come from different sessions, the central claim fails.","tokens_in":16643,"feed_emoji":"🧠","tokens_out":7823,"duration_ms":73389,"temperature":0.7,"pith_summary":"The paper tries to show that unreliable functional connectomes, rather than the predictive models themselves, are the main obstacle to subject-level prediction of cognition in Alzheimer's disease, and that a PCA-based denoising step fixes that obstacle. It combines the differential identifiability framework, which finds the connectome reconstruction where a subject's own scan halves match most strongly relative to other subjects, with connectome predictive modeling, which selects connectivity edges and fits linear models to cognitive scores. The authors report that the denoised connectomes make edge selection substantially more reproducible between split halves, make models trained on one half generalize better to the other half, and significantly improve prediction on a held-out validation cohort for five of seven cognitive outcomes. If true, this matters because the procedure can be applied retrospectively to existing clinical resting-state fMRI data, reducing the need for longer or repeated scans before connectivity-based biomarkers can be used.","feed_headline":"Denoising brain scans improves Alzheimer's cognition predictions","feed_subtitle":"A PCA cleanup step lifted validation accuracy in five of seven cognitive tests for Alzheimer's disease.","key_machinery":"$\\mathbb{I}_f$ (differential identifiability) is the central mechanism: a data-driven denoising procedure for functional connectomes. Each subject's single resting-state fMRI session is split into two halves, restA and restB, and an identifiability matrix records the pairwise correlations between every subject's restA and restB connectomes. Differential identifiability is defined as $I_\\mathrm{diff}=I_\\mathrm{self}-I_\\mathrm{others}$, the gap between same-subject and different-subject similarity. Group-level PCA is applied to the vectorized connectomes, and connectomes are reconstructed using an increasing number of principal components; the reconstruction maximizing $I_\\mathrm{diff}$ is selected. That reconstruction feeds into connectome predictive modeling, where edges are selected by their correlation with each cognitive outcome, positive and negative masks are formed, per-subject strengths in those masks are used as regressors, and a linear model is tested on a held-out half of the cohort. All of the paper's claims about improved stability, test-retest generalizability, and external prediction follow from the choice of this particular PCA reconstruction.","core_discovery":"The central discovery claimed is that optimizing differential identifiability in resting-state functional connectomes before connectome predictive modeling improves the stability of the entire prediction pipeline. In the authors' hands, the optimal PCA reconstruction, found at a number of principal components equal to the cohort size and retaining about 80% of the variance, raises test-retest overlap in edge selection by about 30 percentage points, improves generalization of models from one scan half to the other, and lifts validation-cohort correlations between predicted and observed cognition for five of the seven outcome measures examined. The stabilized edge selection then supports the paper's secondary claim: specific resting-state network interactions are reproducibly associated with distinct cognitive deficits, with executive control, default mode, and somatomotor networks involved across all outcomes and domain-specific motifs such as salience-executive control connectivity negatively associated with memory and attention tasks. The authors present these results as the first step toward clinically useful subject-level predictions of cognition from functional connectivity in Alzheimer's disease.","pith_inferences":["Editorial inference: the optimal number of principal components equaling the cohort size suggests a convenient stopping rule for new cohorts, but the authors do not establish that this rule is universal; it should be tested across datasets with different sizes and acquisition protocols.","Editorial inference: if the split-half proxy is optimistic because the two halves share within-session physiological noise, the real-world improvement on separate-day sessions could be smaller than the five-of-seven validation gain; the paper's own limitation section calls for external validation on a separate dataset.","Editorial inference: because the framework preserves the separation between outcome-specific maps while shrinking within-outcome session variability, it may be useful beyond Alzheimer's disease for dissociating cognitive domains generally, though the paper only claims the Alzheimer's application."],"forward_implications":["Denoised connectomes at the $I_\\mathrm{diff}$-optimal reconstruction can be used directly in connectome predictive modeling without changing the downstream regression pipeline.","A test-retest validation step from restA to restB becomes a practical minimum standard for judging whether a connectivity-based cognitive model is overfit before any external validation.","Functional network maps generated from the stable edge masks can be interpreted as candidate biomarkers of specific cognitive deficits in Alzheimer's disease, not just as correlates of diagnostic status.","Existing clinical datasets with short single-session resting-state scans can be re-analyzed with this preprocessing step, potentially improving subject-level prediction from already-collected data."],"supporting_citations":[{"why":"Introduces the differential identifiability framework and the PCA-based procedure that this paper applies before connectome predictive modeling.","marker":"[12]"},{"why":"Defines connectome-based predictive modeling, the model-building and feature-selection pipeline used throughout the paper.","marker":"[27]"},{"why":"Documents poor test-retest reliability of short resting-state scans and its unclear relation to behavioral utility, motivating the need for denoising.","marker":"[10]"},{"why":"Provides the test-retest reasoning and split-half validation approach used to evaluate whether models generalize to new data from the same subjects.","marker":"[22]"},{"why":"Shows that differential identifiability can be optimized in resting-state data from a clinical Alzheimer's cohort, extending the method beyond healthy young adults.","marker":"[13]"},{"why":"Details the fMRI processing pipeline and functional connectivity estimation approach applied to the imaging data.","marker":"[32]"},{"why":"Supplies the 278-region cortical parcellation used to define connectivity nodes in the functional connectomes.","marker":"[49]"}],"fun_headline_variants":["Optimizing connectome fingerprints boosts Alzheimer's cognition predictions","Enhanced connectome identifiability improves Alzheimer's prediction","Stabilized brain network edges sharpen Alzheimer's cognitive forecasts","PCA-tuned connectomes lift Alzheimer's cognition prediction accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that splitting a single fMRI session into two halves reproduces a true test-retest scenario; if shared in-session physiological noise inflates the apparent reliability of the denoised connectomes, the generalization gains reported here could shrink when the method is applied across separate sessions or scanners.","fun_headline_variants_meta":{"raw":{"variants":["Optimizing connectome fingerprints boosts Alzheimer's cognition predictions","Enhanced connectome identifiability improves Alzheimer's prediction","Stabilized brain network edges sharpen Alzheimer's cognitive forecasts","PCA-tuned connectomes lift Alzheimer's cognition prediction accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000431,"raw_usage":{"total_tokens":2170,"prompt_tokens":885,"completion_tokens":1285,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":1220}},"tokens_in":501,"tokens_out":1285,"duration_ms":9398,"temperature":1.0,"reasoning_tokens":1220,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:53:15.804721+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same $\\mathbb{I}_f$-then-CPM pipeline on a dataset where each person has true retest fMRI from a different day or a different scanner: if the optimal PCA reconstruction fails to increase rest-to-retest edge overlap and validation correlations above those of raw connectomes, or if the model improvements disappear when training and validation connectomes come from different sessions, the central claim fails.","supporting_citations":[{"cited_title":"Int J Neuropsychopharmacol, 2017","cited_arxiv_id":null,"evidence_quote":"Introduces the differential identifiability framework and the PCA-based procedure that this paper applies before connectome predictive modeling."},{"cited_title":"Neuroimage, 2013","cited_arxiv_id":null,"evidence_quote":"Defines connectome-based predictive modeling, the model-building and feature-selection pipeline used throughout the paper."},{"cited_title":"Misic, and O","cited_arxiv_id":null,"evidence_quote":"Documents poor test-retest reliability of short resting-state scans and its unclear relation to behavioral utility, motivating the need for denoising."},{"cited_title":"Xia, and D.S","cited_arxiv_id":null,"evidence_quote":"Provides the test-retest reasoning and split-half validation approach used to evaluate whether models generalize to new data from the same subjects."},{"cited_title":"Neuroimage, 2012","cited_arxiv_id":null,"evidence_quote":"Shows that differential identifiability can be optimized in resting-state data from a clinical Alzheimer's cohort, extending the method beyond healthy young adults."},{"cited_title":"Neuroimage, 2019","cited_arxiv_id":null,"evidence_quote":"Details the fMRI processing pipeline and functional connectivity estimation approach applied to the imaging data."},{"cited_title":"Front Aging Neurosci, 2018","cited_arxiv_id":null,"evidence_quote":"Supplies the 278-region cortical parcellation used to define connectivity nodes in the functional connectomes."}],"review_version":1}