{"id":"cdd50163-2aad-4899-8276-04d9cc4bafda","arxiv_id":"1909.03029","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review that summarizes known deep learning applications in health care without contributing any new results.","lead":"This preprint surveys deep learning and big data methods in medical imaging, genomics, and electronic health records. It summarizes selected papers but presents no new experiments, data, or analysis.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The review's value depends on accurate summaries; Section 3 inverts He et al.'s conclusion on pre-training, so the paper misleads rather than informs.","rationale":"The reader's weakest assumption correctly identifies the load-bearing dependency: for a review, accuracy of literature summaries is the product. The He et al. misrepresentation is concrete, checkable, and directly contradicts the cited source. Because the paper offers no original contribution, this is not a salvageable error. I agree with the REJECT verdict; no adjustment is needed.","tokens_in":7381,"tokens_out":2392,"duration_ms":21862,"concrete_test":"Independently read the abstract and experiments of He et al., arXiv:1811.08883, and record whether the paper says pre-training helps or hurts on small target datasets. Then select five other cited papers (e.g., Esteva et al., Raghu et al., Miotto et al.) and compare their abstracts to the review's sentence. If He et al. is inverted and at least one more summary misrepresents its source, the review fails its purpose.","verdict_should_be":"UNCHANGED","load_bearing_attack":"This is a review with no original results; its central claim is that surveyed big-data methods are useful and promising for healthcare. That claim can only be as strong as the accuracy of the paper's one-sentence summaries. Section 3 states that Kaiming He et al. [11] 'show that... the accuracy of the network is no worse than training from scratch on even small datasets' and that pre-trained weights 'do not prevent over fitting except in a few cases.' The cited paper (arXiv:1811.08883) concludes the opposite: pre-training improves accuracy when the target dataset is small, and can hurt when the target dataset is large. This is a direct inversion of the source's main finding. Section 1 also claims 'Apple's latest smart watch can detect heart attacks,' which is unsupported and false (wearables detect arrhythmias, not heart attacks). These errors are material because a reader relying on the review would receive wrong guidance about transfer learning and consumer wearables. If other summaries are similarly unreliable, the paper's usefulness collapses.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a narrative survey of machine learning and big data methods in healthcare, organized into three application domains: medical imaging, genomics, and electronic health records. The authors describe a selection of techniques (CNNs, RNNs/LSTMs, SVMs, autoencoders, transfer learning) and summarize roughly a dozen representative papers in each domain. The central claim is that these methods are powerful, accurate, and have 'great scope' in healthcare, with future research likely focusing on small-data regimes and hybrid models.","tokens_in":7558,"tokens_out":6537,"duration_ms":57353,"significance":"Should the summaries be accurate, the paper would serve as a compact entry point for non-specialist readers, and Table 1 provides a convenient digest of methods and references. The paper does not claim original results, and it provides no machine-checked proofs, code, or empirical evaluations; its contribution is entirely the selection and synthesis of prior work. That contribution is currently compromised because at least one key summary inverts the cited paper's conclusion and another unsupported claim in the introduction is factually wrong. As a result, the paper in its current form cannot be relied upon as a guide to the literature.","major_comments":[{"comment":"The text states that pre-trained weights 'might help speed up convergence' but that 'the accuracy of the network is no worse than training from scratch on even small datasets with around 10,000 images,' and that pre-training does not prevent overfitting except in a few cases. This is the opposite of the finding in arXiv:1811.08883, whose abstract states that pre-training improves accuracy on small datasets and can hurt on large ones. Because this paragraph is the only discussion of transfer learning in the imaging section and the error appears in the summary table's implied guidance, it is a load-bearing misrepresentation. The paragraph should be rewritten to state the actual conclusion, and the implications for medical imaging should be reconciled with the paper's own conclusion in Section 6 that small-data learning is an important future direction.","section":"Section 3, paragraph on transfer learning (He et al. [11])"},{"comment":"The claim that 'Apple's latest smart watch can detect heart attacks' is false. The Apple Watch's FDA-cleared ECG and irregular-rhythm notification features detect atrial fibrillation; they do not diagnose myocardial infarction. Since this claim is made without citation and is used to motivate the pervasiveness of smart wearables in healthcare, it should be corrected or removed.","section":"Section 1, Introduction"}],"minor_comments":[{"comment":"The regularizer is written as 'λ||W 2||' and should read 'λ||W||²'; there is also a sign mismatch between the hyperplane definitions Wᵀx − b = ±1 and the hinge loss in Eq. (1).","section":"Section 2.3, Eq. (2)"},{"comment":"The text says 'ISIB EM segmentation challenge' but the correct acronym is ISBI (International Symposium on Biomedical Imaging).","section":"Section 3"},{"comment":"The name 'Swark et al.' appears in the text and table, but reference [27] is Skwark et al.; the spelling should match the reference list.","section":"Section 4 and Table 1"},{"comment":"The phrase 'Electric health records' should be 'Electronic health records'.","section":"Section 5"},{"comment":"The explanation of autoencoders with equal input and hidden sizes is unclear; the statement that the learned weights become 'essentially linear' should be rephrased or supported.","section":"Section 2.4"},{"comment":"The manuscript contains numerous language and typographical issues (e.g., 'activites,' 'an eternity in today's age') and informal statements that should be tightened.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a broad, non-systematic review whose value rests entirely on the accuracy of its one-sentence summaries. The He et al. misrepresentation is a serious factual inversion and must be corrected; the Apple Watch claim is also unsupported. These errors are localized and fixable, so I do not recommend rejection, but the revision should be checked carefully by the authors and the manuscript proofread. The authors may also want to add a brief statement of how the surveyed papers were selected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a survey paper, no new methods, no new data, no synthesis beyond 'neural networks work well in healthcare.' That by itself is not disqualifying for a review, but the paper's entire value rests on the accuracy of its one-paragraph summaries of other people's work. And at least one of those summaries is a direct inversion of the source.\n\nThe He et al. summary in Section 3 says pre-trained weights give accuracy 'no worse than training from scratch' on small datasets and don't prevent overfitting. The cited paper (arXiv:1811.08883) concludes the opposite: pre-training helps when the target dataset is small, and can hurt when the target dataset is large. A reader relying on this survey would come away with wrong guidance on transfer learning. That's a load-bearing error, not a typo.\n\nThe Apple Watch claim in the introduction—that it can 'detect heart attacks'—is also unsupported. The wearable detects atrial fibrillation, not heart attacks, and conflating the two is the kind of overstatement that a review should be correcting, not repeating. There are smaller issues too, like calling Skwark et al. 'Swark' in the running text.\n\nTo be fair, the paper does some things okay. The background sections on CNNs, RNNs, SVMs, and autoencoders are textbook-level and mostly correct. Several summaries—U-Net, Esteva et al., Choi et al.—are broadly faithful. The table of methods is a handy index. If you know nothing at all about medical deep learning, you could get a rough map from this paper, as long as you stayed away from the transfer-learning paragraph.\n\nBut that's a serious caveat. The paper is promoted as an 'analysis' but has no original evaluation or critical comparison. The conclusion is a generic statement that neural networks perform well. The authors did not seem to engage carefully with the papers they cite, and the one place where they try to make a critical point about transfer learning, they get it wrong.\n\nFor whom is this? Possibly an undergraduate looking for a starting bibliography, but even they should be told to go read the primary papers. As a scholarly contribution, it does not meet a publishable bar. I would desk-reject this without sending it to referees; the central purpose of the paper is undermined by the factual inversion, and there is no original contribution that would justify the reviewer time.","headline":"A survey with no new results whose only value is accurate summaries, and one of its key summaries (He et al. on transfer learning) is exactly backwards.","tokens_in":7996,"tokens_out":2554,"would_cite":false,"duration_ms":24948,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims that deep learning and other big-data technologies have great scope in health care, with neural networks performing well on imaging, genomics, and electronic health records.","keywords":["Medical Imaging","Machine Learning","Deep Learning","Convolutional Neural Networks","Recurrent Neural Networks","LSTM","Transfer Learning","Electronic Health Records"],"falsifier":"Compare the review's summary of any cited study against that study's abstract or conclusions. For example, the review says [11] showed pre-trained weights do not improve accuracy on small datasets, but the cited paper concludes the opposite; if many summaries are similarly off, the review's overall picture would not hold.","tokens_in":7218,"feed_emoji":"🩺","tokens_out":6123,"duration_ms":54022,"temperature":0.7,"pith_summary":"This paper is a review arguing that deep learning and other big-data techniques have become powerful and accurate enough to matter in health care. It surveys representative contributions in three data-rich areas — medical imaging, genomics, and electronic health records — and concludes that neural networks perform well across all three. The value the paper claims is practical: algorithms can support diagnosis, patient monitoring, and outcome prediction using the massive health data already being collected. A sympathetic reader should come away persuaded that these methods are useful and promising, and that future work should focus on learning from smaller datasets and combining models.","feed_headline":"Neural networks perform well across health care, review finds","feed_subtitle":"A survey of imaging, genomics, and EHR studies concludes deep learning is ready for practical medicine.","key_machinery":"The survey is organized around a small set of neural-network architectures that carry the evidence. Convolutional neural networks are presented as the main tool for image analysis, using shared weights and pooling to learn position-independent features. Recurrent networks — LSTM and GRU — are presented as the tool for time-dependent data such as ECG, EEG, and longitudinal electronic health records. Autoencoders, including stacked and denoising variants, are presented as the representation-learning module that compresses high-dimensional inputs before a classifier such as an SVM or shallow network makes the final prediction. Transfer learning is a recurring theme: pre-trained networks are used to overcome small medical datasets. These architectural categories, not any single formula, are what the review uses to connect each application to a method.","core_discovery":"The paper's central claim is that deep learning and related big-data technologies have 'great scope' in health care because health care generates massive data and modern algorithms can reach near-human accuracy. To support this, it reviews selected applications: convolutional neural networks for lung nodule detection, U-Net for segmentation, transfer-learning CNNs for skin disease and cancer metastasis detection, CNN+LSTM hybrids for cardiac sequences, stacked autoencoders for MRI denoising and gene-expression cancer detection, DEEP/PEDLA for enhancer prediction, and recurrent or neural models for predicting heart failure, diagnosis, medication, and readmission from electronic health records. The paper itself presents no new experiments; its contribution is a structured overview concluding that neural networks, especially CNNs for images, sequence models for temporal data, and autoencoders for representation learning, are the dominant and best-performing tools surveyed.","pith_inferences":["A natural testable extension the review does not perform is to benchmark transfer learning specifically on medical imaging datasets; its own comments suggest medical images differ enough from ImageNet that pre-trained features may help less than in natural-image tasks.","If the reviewed accuracy figures hold under clinical validation, deep learning could move from assistant roles toward triage in high-volume imaging and record review, but that step depends on prospective clinical studies the review does not cover.","Readers should treat the review's one-line descriptions of each cited paper as pointers rather than quotations; verifying a few original abstracts would be enough to tell whether the survey's overall optimism is well supported.","The apparent tension inside the review about pre-training — one source reportedly seeing no accuracy benefit on small datasets while another sees minimal gain — could be resolved by a direct comparison of pre-trained versus from-scratch training on a medical dataset."],"forward_implications":["If the reviewed results generalize, automated CNN-based screening tools could handle routine image-reading tasks such as lung nodule detection, skin disease classification, and cancer-metastasis detection with accuracy near that of trained professionals.","Sequence models trained on electronic health records could become practical for predicting heart failure, future diagnoses, medication needs, and hospital readmission, giving clinicians early-warning signals.","Autoencoders could provide a way to build compact patient representations from high-dimensional EHR data, making downstream prediction feasible even when labeled outcomes are scarce.","Hybrid architectures that combine CNNs with temporal models would let one system exploit both spatial image structure and time, as in cardiac video analysis.","If the paper's future-research forecast is right, progress will shift toward methods that learn from small medical datasets and toward multi-model or cross-domain combinations."],"supporting_citations":[{"why":"Supplies the U-Net segmentation architecture described as outperforming sliding-window methods on biomedical images.","marker":"[26]"},{"why":"Reports dermatologist-level skin disease classification via transfer-learned GoogLeNet, a key accuracy claim.","marker":"[9]"},{"why":"Reports 92.4% accuracy detecting cancer metastases with Inception v3, another accuracy benchmark.","marker":"[18]"},{"why":"Shows a CNN+LSTM temporal regression network for cardiac frame recognition with an average 0.4-frame difference.","marker":"[16]"},{"why":"Supports the use of stacked denoising autoencoders with an SVM for cancer detection from gene expression.","marker":"[8]"},{"why":"Provides the DEEP ensemble framework for predicting enhancers across multiple cell types.","marker":"[15]"},{"why":"Extends enhancer prediction to more cell types with PEDLA, showing how the framework generalizes.","marker":"[17]"},{"why":"Supports using deep denoising autoencoders on electronic health records to predict future disease and readmission.","marker":"[22]"},{"why":"Reports DoctorAI's recurrent GRU model for predicting diagnosis, medication, and visit time from EHRs.","marker":"[5]"},{"why":"Reports heart failure prediction from EHR text using word vectors and a neural network.","marker":"[6]"}],"fun_headline_variants":["Neural networks rule health care data, review finds","Deep learning review: AI excels in medical imaging and genomics","Neural nets lead in health care big data review","Health care AI sees neural networks dominate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's conclusions stand on its one-sentence summaries of each cited study being accurate, because it offers no independent experiments or further evidence.","fun_headline_variants_meta":{"raw":{"variants":["Neural networks rule health care data, review finds","Deep learning review: AI excels in medical imaging and genomics","Neural nets lead in health care big data review","Health care AI sees neural networks dominate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001037,"raw_usage":{"total_tokens":4269,"prompt_tokens":755,"completion_tokens":3514,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":371,"completion_tokens_details":{"reasoning_tokens":3453}},"tokens_in":371,"tokens_out":3514,"duration_ms":25869,"temperature":1.0,"reasoning_tokens":3453,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:52:05.782201+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the review's summary of any cited study against that study's abstract or conclusions. For example, the review says [11] showed pre-trained weights do not improve accuracy on small datasets, but the cited paper concludes the opposite; if many summaries are similarly off, the review's overall picture would not hold.","supporting_citations":[{"cited_title":"Scientiﬁc reports 6, 28517 (2016)","cited_arxiv_id":null,"evidence_quote":"Extends enhancer prediction to more cell types with PEDLA, showing how the framework generalizes."},{"cited_title":"Nature542(7639), 115 (2017)","cited_arxiv_id":null,"evidence_quote":"Reports dermatologist-level skin disease classification via transfer-learned GoogLeNet, a key accuracy claim."},{"cited_title":"In: Ourselin, S., Joskowicz, L., Sabuncu, M.R., Unal, G., Wells, W","cited_arxiv_id":null,"evidence_quote":"Shows a CNN+LSTM temporal regression network for cardiac frame recognition with an average 0.4-frame difference."},{"cited_title":"In: PACIFIC SYMPOSIUM ON BIOCOMPUTING 2017","cited_arxiv_id":null,"evidence_quote":"Supports the use of stacked denoising autoencoders with an SVM for cancer detection from gene expression."},{"cited_title":"Nucleic acids research 43(1), e6–e6 (2014)","cited_arxiv_id":null,"evidence_quote":"Provides the DEEP ensemble framework for predicting enhancers across multiple cell types."},{"cited_title":"Scientiﬁc reports 6, 26094 (2016)","cited_arxiv_id":null,"evidence_quote":"Supports using deep denoising autoencoders on electronic health records to predict future disease and readmission."},{"cited_title":"In: Machine Learning for Healthcare Conference","cited_arxiv_id":null,"evidence_quote":"Reports DoctorAI's recurrent GRU model for predicting diagnosis, medication, and visit time from EHRs."},{"cited_title":"Medical Concept Representation Learning from Electronic Health Records and its Application on Heart Failure Prediction","cited_arxiv_id":"1602.03686","evidence_quote":"Reports heart failure prediction from EHR text using word vectors and a neural network."}],"review_version":1}