{"id":"e06e7ade-c556-47d4-95ed-c10df2a58482","arxiv_id":"2506.15761","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A multi-task neural network trained on longitudinal multi-omics data reconstructs ME/CFS symptom scores and distinguishes patients from controls with a reported AUC of 0.91, though external validation is weak.","lead":"An AI framework called BioMapAI maps gut microbiome, blood metabolite, immune cell, and lab-test data to 12 ME/CFS symptom scores, and the thesis reports 91% accuracy in separating patients from healthy controls. The authors also build a healthy-versus-disease connectivity map of microbiome-immune-metabolome interactions and propose butyrate, tryptophan, and bile-acid related mechanisms.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 91% AUC claim likely rests on sample-level 5-fold CV that does not group repeated samples by participant; without patient-stratified CV, the headline classification performance is not established.","rationale":"The reader's verdict is CONDITIONAL, with the weakest assumption identified as participant-level leakage in cross-validation. I read the manuscript in good faith and reached the same conclusion: the Methods section describing BioMapAI's training and evaluation does not mention grouping by participant, and the dataset is explicitly longitudinal with 515 timepoints from 249 participants. Because both the input omics and the target symptom scores are patient-specific, sample-level folds allow the model to exploit repeated-measure identity, which would inflate all reported internal AUCs. This concern is load-bearing for the strongest claim that BioMapAI achieved state-of-the-art precision in disease classification. The external validation with zero-imputation and low feature overlap is a second serious concern, but the internal CV issue alone is sufficient to make the headline claim unverified. My recommendation is therefore UNCHANGED relative to the reader's CONDITIONAL verdict: the paper's dataset and framework retain independent value, but the central accuracy claim needs a patient-grouped reanalysis. The proposed test is a single, feasible re-run of the existing pipeline, and it would settle whether the 91% AUC is real or an artifact of leakage.","tokens_in":49210,"tokens_out":1719,"duration_ms":25166,"concrete_test":"Re-run BioMapAI's evaluation with GroupKFold stratified by participant ID, so all timepoints from the same individual are confined to a single fold, and otherwise keep the preprocessing, architecture, and hyperparameters identical. Compare the combined-omics 91% AUC (and each per-omics AUC) against the current sample-level 5-fold CV result. If the participant-grouped AUC drops materially (for example, below 80% or toward chance), the state-of-the-art classification claim is unsupported; if it remains around 91%, the concern is refuted. The same grouped CV should be applied to GDBT and DNN before comparing models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that BioMapAI's omics-derived scores achieve 91% AUC in separating ME/CFS patients from controls depends on the cross-validation procedure in Methods 'Cross-Validation and Model Training'. That section states only that 'we employed a robust 5-fold cross-validation' and does not state that folds were grouped by participant. The cohort contains 515 timepoints from 249 participants, so roughly two repeated samples per participant. With sample-level splits, the same individual can appear in both training and test folds. Since omics profiles and the 12 clinical scores are both strongly person-specific, the model can memorize a patient's identity rather than learning disease-general patterns. This inflates symptom reconstruction MSE and, in turn, the downstream binary classification AUC, because the optional ScoreLayer is fit to the predicted scores Y_hat. The Discussion itself acknowledges the model was 'trained on < 500 samples with fivefold cross-validation' without mentioning patient-level splitting, so the reader's concern is not contradicted elsewhere in the manuscript. This is the most load-bearing weakness because every state-of-the-art claim, including the reported superiority over GDBT and DNN, inherits it. Separately, the external validation protocol zero-imputes missing features with overlaps as low as 19%, which could also attenuate or distort performance, but the internal CV leakage is the more fundamental threat to the headline result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The thesis chapter presents a longitudinal multi-omics cohort study of ME/CFS (153 patients, 96 healthy controls, 515 timepoints) and introduces BioMapAI, an explainable multi-task deep neural network that maps omics matrices to 12 clinical symptom scores and then uses the predicted scores for binary disease classification. The paper claims that omics-derived predicted scores achieve a 91% AUC in distinguishing ME/CFS patients from healthy controls within the cohort, that BioMapAI outperforms LR, SVM, GDBT, and a standard DNN, and that it validates on four external cohorts with accuracies between 58% and 72%. The authors further use WGCNA and SHAP to identify disease- and symptom-specific biomarkers and to construct microbiome-immune-metabolome interaction networks, proposing mechanistic hypotheses involving butyrate, BCAA, tryptophan, benzoate, and mucosal immune cells.","tokens_in":49518,"tokens_out":2176,"duration_ms":27892,"significance":"If the central performance claims were established, the work would be valuable: the longitudinal multi-omics dataset is rich, the multi-output neural architecture is a sensible response to disease heterogeneity, the code and trained models are promised on GitHub, and the external cohort attempts go beyond most single-cohort studies. The confounder analyses (MaAsLin2, residual adjustment for immune and metabolome data) and the use of SHAP for interpretability are concrete strengths. However, the headline 91% AUC and the claimed superiority over baselines rest on a cross-validation procedure that is not shown to be participant-stratified in a repeated-measures design, and the external validation is quantitatively weak, with several accuracies near chance and feature overlaps as low as 19%. The microbial metabolite inference in Chapter 1 also has a circularity concern. These issues are load-bearing for the central claims, so the manuscript needs substantive revision before its main conclusions can be accepted.","major_comments":[{"comment":"The five-fold cross-validation used to produce the 91% AUC is described only as 'a robust 5-fold cross-validation' in the main text and in the Methods section. The cohort contains 515 timepoints from 249 participants, so repeated samples from the same individual are almost certainly present in multiple folds if splitting is performed on samples rather than on participants. Because both the omics features and the 12 clinical scores are strongly person-specific, sample-level folds allow the model to memorize patient identity, inflating the symptom reconstruction accuracy and hence the downstream binary classification AUC of the optional ScoreLayer. The Discussion acknowledges that the model was 'trained on < 500 samples with fivefold cross-validation' without stating that folds were grouped by participant. The authors must either demonstrate that cross-validation was stratified by participant (e.g., groupKFold) or re-run the entire evaluation with participant-level splitting and report the resulting AUCs and baseline comparisons. Without this, the central 'state-of-the-art precision' claim is not established.","section":"Chapter 2, Methods, 'Cross-Validation and Model Training'"},{"comment":"The external validation protocol zero-imputes all missing features to align external datasets to the BioMapAI feature set. With the Che metabolome cohort featuring only 19% feature overlap, and the Germain cohort 79%, the test effectively runs the model on matrices that are mostly zeros, which can attenuate performance in either direction and makes the reported accuracies (59% and 68%) difficult to interpret. Moreover, the claim that 'BioMapAI significantly surpassed GDBT and DNN in external cohort validation' is not supported by any statistical test, confidence interval, or repeated subsampling of the external data; the accuracies of 58-72% are also close to chance for a binary classification. The authors should report the external validation with only overlapping features (or a principled imputation), include uncertainty estimates, and test whether the difference from baselines is statistically significant.","section":"Chapter 2, Results, 'External Validation' and Methods, 'External Validation with Independent Dataset'"},{"comment":"The MCMC-based inference of gut isobutyrate (Figure 6B) accepts simulation steps only when the simulated growth rate correlates with the metagenomic species abundance profile (Pearson ρ > 0.6). This means the predicted gut metabolome is not independent of the species abundance differences used to define the patient groups. The subsequent claim that 'inferred concentrations of isobutyrate were significantly decreased in ME/CFS' is therefore not an independent confirmation of reduced gut butyrate; it is a re-statement of the species-level differences already used to calibrate the simulation. The authors should either present the species-abundance-based result as the primary evidence, or recalibrate the MCMC without the abundance-correlation acceptance criterion and show that the predicted butyrate difference still holds.","section":"Chapter 1, Methods, 'Gut metabolic status prediction' and Results, 'Gut and plasma butyrate is reduced in early-stage…"}],"minor_comments":[{"comment":"The learning rate is stated as 0.01 in the initial description ('The learning rate was set to 0.01') and then as 0.0005 in the 'Cross-Validation and Model Training' section ('a learning rate of 0.0005, optimized through grid search'). Please clarify which value was used for the reported models or describe the grid search that led to the final choice.","section":"Chapter 2, Methods, 'BioMapAI'"},{"comment":"The external validation bar chart reports accuracies without error bars or sample sizes per cohort in the main figure. Adding bootstrapped confidence intervals and the number of test samples would make the comparison with baselines interpretable.","section":"Chapter 2, Figure 2E"},{"comment":"There are several typos and formatting inconsistencies, for example 'intergrading' in Chapter 1 ('our 'omics workflows could be one of the guiding frameworks to intergrading'), 'in the in the TissueLyser' in Chapter 2 Methods, and the inconsistent use of 'GDBT' versus 'GBDT' across the thesis. A careful proofread is needed.","section":"General"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things worth knowing about this thesis. The real value is the longitudinal multi-omics resource: 515 timepoints from 249 people with paired gut metagenomics, plasma metabolome, immune profiling, blood labs, and 12 symptom scores. That is a genuinely strong dataset for ME/CFS. The headline claim that BioMapAI reaches a 91% AUC, however, is not supported as written because the five-fold cross-validation is sample-level rather than participant-level. With 515 timepoints from 249 people, repeated samples from the same person can appear in both training and test folds, allowing the model to memorize person-specific patterns and inflate the AUC. The paper never states that folds were grouped by participant, and the Discussion only says the model was 'trained on < 500 samples with fivefold cross-validation.'\n\nWhat is actually new in Chapter 2 is the multi-task DNN architecture with two shared hidden layers and task-specific sub-layers, plus SHAP-based attribution linking omics features to individual symptoms. That is a sensible response to disease heterogeneity, and the symptom-specific biomarker analysis is a useful extension. The WGCNA connectivity maps between microbiome, immune, and metabolome modules are exploratory but generate testable hypotheses. Code and data are on GitHub, and Chapter 1 is a peer-reviewed publication with its own external validation.\n\nThe soft spots line up with the reader's concerns. The cross-validation leak is load-bearing; every state-of-the-art claim inherits it. External validation is weak, with accuracies of 58-72%, several close to chance, and features zero-imputed with overlap as low as 19%. One external cohort, Raijmakers, is Q fever fatigue syndrome, not ME/CFS, so calling that a validation set misstates the target. The MCMC gut metabolome prediction also has a circular element because its acceptance criterion uses the species abundances being tested. That is a minor issue compared to the CV problem.\n\nIn short, the dataset and modeling framework are worth engaging, but the headline numbers must be re-derived with participant-stratified cross-validation, and external validation needs an honest treatment of feature overlap and cohort definitions. I would send this out for peer review rather than desk-reject because the resource is valuable and the identified flaws are fixable with reanalysis. A major revision with corrected numbers is required before the classification claims can be trusted.","headline":"A valuable longitudinal multi-omics resource and a sensible multi-task modeling idea, but the headline 91% AUC is likely inflated by sample-level cross-validation and needs a patient-level reanalysis before it can be trusted.","tokens_in":50052,"tokens_out":3928,"would_cite":false,"duration_ms":42432,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An explainable deep network reconstructs ME/CFS symptoms from multi-omics and classifies patients with 91% AUC.","keywords":["ME/CFS","multi-omics","deep learning","explainable AI","symptom reconstruction","biomarker discovery","gut microbiome","metabolomics"],"falsifier":"Re-run the five-fold cross-validation with all timepoints from each participant assigned to the same fold (grouped or leave-one-participant-out), and compare the AUC to the reported 91%. If the grouped AUC drops substantially toward chance, the in-cohort result is inflated by within-participant leakage rather than reflecting true disease classification.","tokens_in":48961,"feed_emoji":"🧬","tokens_out":7555,"duration_ms":73895,"temperature":0.7,"pith_summary":"This thesis reports a deep-learning framework, BioMapAI, that connects multiple biological omics layers—gut metagenomics, plasma metabolome, immune cell profiles, and blood labs—to a matrix of twelve clinical symptom scores in a longitudinal cohort of 249 participants followed for up to four years. The central claim is that the model's predicted symptom scores separate ME/CFS patients from healthy controls with 91% AUC, outperforming gradient boosting, support vector machines, and a standard deep network, and that it generalizes to independent external cohorts better than those baselines. The same model yields both disease-level and symptom-level biomarkers and produces the first connectivity map of microbiome-immune-metabolome interactions in health and disease. If correct, the work shows that a heterogeneous chronic illness can be characterized by linking many biological measurements to many clinical outcomes at once, rather than to a single diagnosis.","feed_headline":"AI maps multi-omics to symptoms, hits 91% AUC for ME/CFS","feed_subtitle":"A deep network trained on gut, blood, and immune data separates chronic fatigue patients from healthy controls.","key_machinery":"The central object is the BioMapAI architecture: a deep neural network whose hidden layers are split into two shared general-pattern layers and a third layer of parallel sub-networks, one per clinical outcome, so the model learns both disease-wide and symptom-specific representations. This design lets a single model map high-dimensional omics inputs to a mixed-type outcome matrix of 12 clinical scores, with loss functions chosen per outcome. Explainability comes from SHAP values computed on reconstructed per-symptom sub-models, and the connectivity map is built separately with WGCNA co-expression modules and Spearman correlations between modules across omics layers.","core_discovery":"BioMapAI is a fully connected neural network with two shared hidden layers (64 and 32 nodes) followed by a parallel layer of 12 outcome-specific sub-networks, one per clinical score, with each output assigned its own loss function. Trained with five-fold cross-validation on five omics inputs, the network reconstructs the distribution of the twelve symptom scores and, through an auxiliary classification layer, distinguishes ME/CFS from healthy controls with 91% AUC using integrated multi-omics. Immune profiling is the strongest single modality (80% AUC), followed by KEGG gene abundance (78%) and blood measures (71%). In external cohorts, BioMapAI outperforms GBDT and a standard DNN, supporting the paper's claim that connecting omics to symptoms improves generalizability. The paper further identifies shared disease-specific biomarkers (e.g., increased B cells and CD4 naive T cells, Dysosmobacter welbionis, bile acid changes) and symptom-specific biomarkers (e.g., Faecalibacterium prausnitzii showing a biphasic relationship with pain), and describes how healthy microbiome-immune-metabolome networks, including butyrate, BCAA, and tryptophan pathways, become dysbiotic in patients.","pith_inferences":["If the 91% AUC survives participant-grouped cross-validation, the strongest immediate use is not diagnosis but patient stratification: predicted symptom scores could guide personalized symptom management even without a biomarker-based diagnostic test.","The multi-task architecture suggests a general recipe for other omics studies of heterogeneous disease: predicting the full symptom vector rather than a single label may yield richer biomarkers and better external generalizability—an effect that could be tested on public single-cohort datasets.","The reported loss of butyrate/BCAA interactions with regulatory T cells implies a concrete experiment: short-term butyrate or BCAA supplementation in a small ME/CFS cohort should shift the SHAP-predicted pain, GI, and fatigue scores if the connectivity map reflects causal biology.","The absence of strong temporal signals over 3-4 years may indicate that baseline omics capture a stable trait-like disease state; testing this requires longer follow-up or event-based modeling of symptom flares rather than yearly averages."],"forward_implications":["Multi-omics profiles can reconstruct the twelve clinical symptom scores well enough that predicted scores serve as a disease classifier with 91% AUC.","Immune profiling is the most informative single modality for most symptoms, while the gut microbiome is the best predictor of gastrointestinal, emotional, and sleep scores.","The model's architecture and SHAP decoding separate disease-specific from symptom-specific biomarkers, enabling per-symptom biomarker discovery in heterogeneous chronic disease.","Healthy microbiome-immune-metabolome networks are established as a baseline, and their dysregulation in ME/CFS—loss of butyrate and BCAA interactions, gain of inflammatory gamma-delta T and MAIT cell connections—points to testable targets for intervention.","The same framework is portable to other heterogeneous chronic conditions, including long COVID, by swapping the input omics matrix and the output symptom matrix."],"supporting_citations":[{"why":"This is the prior multi-omics ME/CFS study whose cohort and design Chapter 2 extends into the BioMapAI model.","marker":"33"},{"why":"This external ME/CFS microbiome cohort is used to validate BioMapAI's species and KEGG gene models.","marker":"114"},{"why":"This external multi-omics cohort of Q fever fatigue syndrome is used to test generalizability of the model.","marker":"115"},{"why":"This external ME/CFS plasma metabolomics cohort is used to validate the metabolome model.","marker":"116"},{"why":"This external ME/CFS metabolomics cohort is used to validate the model on an independent patient group.","marker":"117"},{"why":"This reference supplies the SHAP method used to make BioMapAI's predictions explainable.","marker":"163"},{"why":"This reference supplies WGCNA, which is used to construct the co-expression modules underlying the healthy and disease connectivity maps.","marker":"166"}],"fun_headline_variants":["AI ties gut, blood, immune data to ME/CFS, 91% AUC","BioMapAI: explainable deep learning hits 91% AUC for ME/CFS","Multi-omics AI distinguishes ME/CFS from healthy with 91% AUC","AI maps omics to symptoms, hits 91% AUC in chronic fatigue","Explainable AI ties multi-omics to ME/CFS, 91% AUC"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline 91% AUC assumes that repeated samples from the same participant are effectively independent, so that five-fold cross-validation without participant-level grouping does not let the model memorize patient-specific signatures and inflate performance.","fun_headline_variants_meta":{"raw":{"variants":["AI ties gut, blood, immune data to ME/CFS, 91% AUC","BioMapAI: explainable deep learning hits 91% AUC for ME/CFS","Multi-omics AI distinguishes ME/CFS from healthy with 91% AUC","AI maps omics to symptoms, hits 91% AUC in chronic fatigue","Explainable AI ties multi-omics to ME/CFS, 91% AUC"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000884,"raw_usage":{"total_tokens":3818,"prompt_tokens":944,"completion_tokens":2874,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":2774}},"tokens_in":560,"tokens_out":2874,"duration_ms":21202,"temperature":1.0,"reasoning_tokens":2774,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:54:01.213082+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the five-fold cross-validation with all timepoints from each participant assigned to the same fold (grouped or leave-one-participant-out), and compare the AUC to the reported 91%. If the grouped AUC drops substantially toward chance, the in-cohort result is inflated by within-participant leakage rather than reflecting true disease classification.","supporting_citations":[{"cited_title":"& Horvath, S","cited_arxiv_id":null,"evidence_quote":"This reference supplies WGCNA, which is used to construct the co-expression modules underlying the healthy and disease connectivity maps."}],"review_version":1}