{"id":"ac5b4011-0022-44b3-a06f-61c36298bd2f","arxiv_id":"1909.00384","paper_version":2,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review of AI health informatics spanning imaging, EHRs, genomics, sensing, and online health, cataloging methods and open challenges.","lead":"This paper is a survey of how artificial intelligence and deep learning are being applied across medical imaging, electronic health records, genomics, wearable sensors, and online health data. It catalogs common approaches and lists barriers such as missing labels, data bias, interpretability, and security.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'comprehensive' claim rests on an unreported and unreproducible paper-selection protocol in Section 3; without databases, dates, eligibility criteria, or screening counts, the reviewed corpus cannot be shown to be representative.","rationale":"The reader's verdict is UNVERDICTED, and the weakest assumption identified by the reader is the representativeness of the ad hoc paper selection in Section 3. My stress-test reaches the same point: the review's central claim is comprehensiveness, and that claim is only as strong as the selection protocol. The manuscript provides search terms but no database list, no search date range, no inclusion or exclusion criteria, no screening protocol, and no flow diagram or counts. Without these, the paper cannot be checked for coverage bias, and the claimed map of applications and challenges is not reproducible. The abstract's pre-processing claim is an additional accuracy issue, but it does not bear on the central comprehensiveness claim as directly. Because the reader already classified the paper as UNVERDICTED for essentially this reason, my analysis does not move the verdict. I recommend no change: the review may be useful as a narrative survey, but the 'comprehensive' framing remains unsupported by the reported methods.","tokens_in":46320,"tokens_out":3299,"duration_ms":32160,"concrete_test":"Reconstruct the Section 3 corpus by listing every cited application paper with its publication venue and year; then run the stated query in PubMed, Scopus, IEEE Xplore, and ACM DL for 2013-2020, and randomly sample 50 eligible papers from each of the five subfields. If the reviewed corpus mentions or discusses fewer than roughly 80% of these sampled papers, or if the authors cannot supply databases, dates, and screening counts sufficient to allow this reconstruction, the comprehensiveness claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the abstract's promise of a 'comprehensive review' of AI in health informatics over roughly 2013 to 2020. The load-bearing condition is that the papers selected in Section 3 are representative of the full literature in the five surveyed areas. The manuscript's own description of the method ('searched across several databases with the combination of search terms' plus 'significantly relevant papers ... briefly reviewed') does not state which databases, what date range was applied, how relevance was judged, or how many records were screened. Consequently, no check can distinguish a complete map from a selected sample. The risk is concrete: if the selection skews toward the authors' own interests, toward preprints or venues accessible to the authors, or toward positive results and prominent architectures, the application catalogue and the 'challenges' inferred from it are not comprehensive. A secondary internal inconsistency exists: the abstract's statement that recent AI approaches 'do not require domain-specific data pre-processing' is contradicted by Section 4.1.2, which argues that pre-processing is important and 'still a blind exploration process.' That contradiction weakens the review's accuracy, but the representativeness of the reviewed corpus is the load-bearing issue for the headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This preprint is a narrative review of artificial intelligence applications in health informatics. It covers five application areas: medical imaging, electronic health records, genomics, sensing, and online communication health. The paper gives an overview of common models, including SVM, tensor decomposition, word embeddings, CNNs, RNNs, autoencoders, deep belief networks, attention, transfer learning, and reinforcement learning. It then surveys representative studies in each application area and concludes with a taxonomy of data-side and model-side challenges, including high dimensionality, heterogeneity, missing labels, bias, interpretability, reliability, feasibility, security, and scalability. The abstract's headline claim is that the article is a comprehensive review of AI research in health informatics over roughly the last seven years.","tokens_in":46658,"tokens_out":2674,"duration_ms":27001,"significance":"If the review were comprehensive and accurate, it would provide a useful map of a fast-moving field, and its taxonomy of challenges would be a convenient reference for researchers entering health informatics. The paper has real strengths: it collects a large number of representative studies, organizes them by data modality and task, and accompanies them with a self-contained model overview including formal descriptions of SVM, tensor decomposition, and neural architectures. The challenges section (Section 4) is broad and reflects concerns that are widely discussed in the literature, such as data bias, missingness, interpretability, and scalability. However, the paper's value as a 'comprehensive review' is limited by the absence of a reproducible paper-selection protocol, and the abstract's claim that recent AI approaches 'do not require domain-specific data pre-processing' is internally inconsistent with Section 4.1.2. These issues affect the central characterization of the reviewed literature and the derived challenge list, though they are fixable in revision.","major_comments":[{"comment":"The load-bearing claim of 'comprehensive review' is not supported by the reported search method. The paper says the authors 'searched across several databases with the combination of search terms' and then 'significantly relevant papers ... were brieﬂy reviewed,' but it does not state which databases were searched, what date range was applied, how relevance was judged, or how many records were screened and excluded. Without this information, no reader can verify that the 300+ cited papers are representative of the literature in medical imaging, EHR, genomics, sensing, and online communication health rather than a selected sample. This is not a cosmetic omission: the application catalogue in Section 3 and the challenge taxonomy in Section 4 are both presented as field-wide conclusions, so the selection bias risk propagates to the paper's main claims. I recommend adding a reproducible search-and-screening protocol, for example a PRISMA-style flow diagram with database names, search dates, inclusion/exclusion criteria, and screening counts, or explicitly relabeling the article as a narrative review rather than a comprehensive one.","section":"Section 3, first paragraph"},{"comment":"The abstract states that 'recent artiﬁcial intelligence approaches do not require domain-speciﬁc data pre-processing,' but Section 4.1.2 argues that pre-processing is important for high-dimensional, sparse, irregular, and biased data and that 'pre-processing, normalization or change of input domain, class balancing and hyperparameters of models are still a blind exploration process.' Section 3.3 similarly notes that genomic datasets are 'incredibly high dimensional, heterogeneous, and unbalanced' and that domain experts' pre-processing 'was frequently required.' The abstract's claim is therefore contradicted by the paper's own detailed discussion. This inconsistency should be resolved by either removing or qualifying the abstract statement, since as written it misrepresents the review's substantive content.","section":"Abstract vs. Section 4.1.2"},{"comment":"The paper claims to focus on 'the last seven years,' but the text does not specify the search period, and the reference list includes foundational works from before 2013 (e.g., LeNet [64], LSTM [73], SVM [57], and the original word2vec paper [62]). It is possible that these older works are included as background for model descriptions, but the review does not distinguish background citations from papers that fall within the claimed seven-year window. Figure 1 also reports a distribution of published papers from PubMed without stating the query, date range, or deduplication method used. I ask the authors to clarify the temporal scope and to state explicitly which papers are part of the 'last seven years' corpus and which are cited only as background.","section":"Section 3 and Figure 1"}],"minor_comments":[{"comment":"The sentence following Equation (2), 'Having a regularization term ||w||2 and a small value parameter makes data can be ﬁnally linearly classiﬁable,' is grammatically unclear and should be rewritten to explain the role of the regularization parameter and the soft margin.","section":"Equation (2)"},{"comment":"In the paragraph discussing genetic variants, the text says 'Reference [212]' for the LPA SNP study; this same reference is already discussed in Section 3.2.2 as an EHR topic-modeling study. The cross-reference is acceptable but the sentence could more clearly indicate that this is a genomics example using EHR-derived phenotypes.","section":"Section 3.3.1"},{"comment":"The paper uses 'traditional models' and 'domain-speciﬁc data pre-processing' in the abstract but 'pre-processing' and 'feature extraction' elsewhere; the terminology should be harmonized so that the discussion of Section 4.1.2 is not confused with the more general claims in the introduction.","section":"Throughout"},{"comment":"The caption of Figure 1 says 'from PubMed' but gives no query details, date of search, or number of papers considered; please provide this information or remove the quantitative-sounding distribution claim.","section":"Figure 1"},{"comment":"Some references are incomplete or informal, such as [61] and [76], which cite Wikipedia pages without author names or version dates, and [187], a preprint under review. Please check all references for completeness and ensure published versions are cited where available.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful narrative review, but the 'comprehensive review' claim is currently overclaimed relative to the reported method. The central fix is methodological transparency: databases, search dates, inclusion/exclusion criteria, and screening counts. If the authors prefer to keep the current method, they should rename the paper to a narrative or scoping review and soften the abstract. The internal contradiction about pre-processing must also be resolved. I see no reason to reject the manuscript, as the issues are correctable within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent literature review with two caveats that matter — the \"comprehensive\" label is not backed by a reported search protocol, and the abstract contradicts the body on data pre-processing.\n\nWhat's actually new: not much as a scientific result, which is normal for a survey. It covers five application areas — imaging, EHR, genomics, sensing, and online communication health — and that last pair gives it slightly more breadth than earlier reviews. The model overview in Section 2 is clear, and the challenge taxonomy (data vs model; high-dimensionality, bias, interpretability, feasibility, scalability) is sensible. The reference list is broad and mostly fits the claims it supports. This is a fine orientation piece for a graduate student entering the field.\n\nSoft spots, in proportion. First, the central claim. The abstract promises a \"comprehensive review\" of the last seven years, but Section 3 says only that the authors \"searched across several databases\" and briefly reviewed \"significantly relevant papers.\" No databases, no search date range, no inclusion/exclusion criteria, no screening counts. That means the reviewed corpus cannot be checked for representativeness. I see no evidence the selection is actually skewed, but the method doesn't rule it out, and the claim is load-bearing. Second, an internal contradiction: the abstract says recent AI approaches \"do not require domain-specific data pre-processing,\" while Section 4.1.2 says pre-processing is important and \"still a blind exploration process.\" The body is right; the abstract should be fixed. Third, the online communication health section is noticeably thinner than the EHR and imaging sections, so the breadth claim is uneven. Minor point: the only self-citation, [187], appears as an example in the EHR review and is not load-bearing. Citation pattern looks fine.\n\nVerdict: this is not a new result; it is a synthesis that overlaps with earlier reviews such as [1-3,6,8,9]. It would help a newcomer, and with an honest title and a documented method it could be a useful published survey. I would send it to peer review — not because it is groundbreaking, but because the scope is broad and the references serious enough that a referee can push the authors to either justify or retract the comprehensiveness claim.","headline":"A serviceable survey that overclaims comprehensiveness: useful for newcomers, not a new result, and it needs a documented search protocol and a fixed abstract.","tokens_in":47044,"tokens_out":1969,"would_cite":false,"duration_ms":19429,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review maps roughly seven years of artificial-intelligence research across medical imaging, electronic health records, genomics, sensing, and online communication health, and consolidates the data and model barriers that still stand…","keywords":["artificial intelligence","health informatics","deep learning","medical imaging","electronic health records","genomics","wearable sensing","online health communities"],"falsifier":"A preregistered systematic search of the same five domains over 2013-2020 that applied explicit inclusion and exclusion criteria would settle whether the review's catalogue is comprehensive; if such a search surfaced major untouched research lines, or showed the cited papers to be concentrated in a few well-resourced settings, the paper's map and its challenge taxonomy would need revision.","tokens_in":46106,"feed_emoji":"🩺","tokens_out":5781,"duration_ms":50610,"temperature":0.7,"pith_summary":"This paper sets out to give a comprehensive map of artificial-intelligence research in health informatics over roughly the last seven years, organized around five application areas: medical imaging, electronic health records, genomics, sensing, and online communication health. It argues that deep learning in particular has shifted health-AI work away from hand-crafted feature extraction, enabling disease diagnosis, outcome prediction, genotype-phenotype analysis, treatment recommendations, and outbreak forecasting. The review then assembles the recurring obstacles into two buckets: data-side problems (high dimensionality, heterogeneity, time dependency, sparsity, irregularity, missing labels, bias) and model-side problems (reliability, interpretability, feasibility, security, scalability). A sympathetic reader would take the paper's contribution to be the consolidated landscape plus a usable challenge taxonomy, rather than any single new algorithm or result.","feed_headline":"Seven years of medical AI, mapped across five data domains","feed_subtitle":"Deep learning's uses in imaging, EHRs, genomics, sensors, and online health plus the barriers to the clinic.","key_machinery":"The organizing machinery is a three-part grid: data modality (imaging, EHR, genomics, sensing, online) times task type (classification, segmentation, prediction, extraction, generation, recommendation) times model family (CNN, RNN, autoencoders, DBN, attention, transfer learning, reinforcement learning). The grid carries the argument by letting the review show that the same model families recur across very different data, and that a single data/model challenge taxonomy explains why most systems stop short of the clinic.","core_discovery":"The authors claim that the past seven years have produced a recognizable set of deep-learning practices in each of five data modalities, and that the same short list of unsolved problems repeats across all of them. The review classifies the field's tasks—classification, segmentation, detection, outcome prediction, phenotyping, information extraction, representation learning, de-identification, intervention recommendation, gene and epigenome prediction, drug generation, biosignal monitoring, and social-media health surveillance—and matches each to the model families that dominate it: CNN and multi-stream, 2.5D, and 3D variants for imaging; word embeddings, RNN, LSTM, GRU, and attention for electronic health records; CNN, DBN, and autoencoders for genomics; CNN and deep deterministic learning for sensing; and natural-language models for online health data. The claim is descriptive: these applications have reached clinical alternative technology levels in some areas, but data scarcity, missing labels, bias, irreproducible pre-processing, black-box interpretability, feasibility, security, and scalability remain unsolved.","pith_inferences":["Extension: the 'comprehensive' label is best read as comprehensive within the five chosen domains, since adjacent areas such as nursing documentation and dental imaging are not covered.","Extension: if the data/model challenge taxonomy is stable, progress on interpretability in one domain should transfer to the others, which a cross-domain method-transfer experiment could test cheaply.","Extension: the emphasis on retrospective cohorts and shifting clinical protocols implies that prospective, multi-institution validation is the binding constraint on clinical adoption, not model capacity.","Extension: the credibility concerns raised for sensing and online data suggest privacy-preserving approaches will be a precondition for the social-media health surveillance strand."],"forward_implications":["Transfer learning and multi-stream architectures should be treated as the default scaffolding for medical-imaging models when annotated volumes are scarce, since the review finds them consistently outperforming single-stream training from scratch.","In EHR research, temporal modeling with RNN, LSTM, GRU, and attention, together with explicit missingness handling, is the prevailing route to outcome prediction and phenotyping.","Genomics and drug-design work increasingly relies on unsupervised representation learning, generative models, and reinforcement learning to compensate for high-dimensional, sparsely labeled molecular data.","Data pre-processing choices are themselves a source of bias, so the review implies that fully reporting them is a precondition for comparing models' true clinical performance.","Multi-modal and multi-task learning, plus on-device or privacy-preserving deployment, are the stated future directions for moving these systems from retrospective datasets toward prospective clinical use."],"supporting_citations":[{"why":"Provides the framing that deep learning in healthcare is a reviewable field with known opportunities and challenges.","marker":"[1]"},{"why":"Surveys deep learning across health informatics and defines the domain boundaries the paper extends.","marker":"[2]"},{"why":"Surveys deep learning in medical image analysis, grounding the imaging section's task taxonomy.","marker":"[3]"},{"why":"Surveys deep EHR analysis, supplying the structured and unstructured data modeling challenge baseline.","marker":"[6]"},{"why":"Reviews deep learning for computational biology, anchoring the genomics and epigenomics sections.","marker":"[7]"},{"why":"Presents a guide to deep learning in healthcare, supporting the claim that deep learning needs little domain-specific feature engineering.","marker":"[8]"},{"why":"Systematically reviews machine learning on online personal health data, grounding the sensing and online communication portion.","marker":"[9]"},{"why":"Reviews deep learning techniques for genomics, providing the molecular-data modeling context.","marker":"[16]"}],"fun_headline_variants":["AI in health: seven years, five data types, same hurdles","Deep learning's five-front health check","Medical AI review: what works, what's stuck","The state of AI in health: five domains, one set of barriers","AI's health report card: five data arenas, recurring limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The map's completeness assumes that the informally selected papers in Section 3 fairly represent the field over the past seven years, since no database list, search date range, inclusion criteria, or screening protocol is reported.","fun_headline_variants_meta":{"raw":{"variants":["AI in health: seven years, five data types, same hurdles","Deep learning's five-front health check","Medical AI review: what works, what's stuck","The state of AI in health: five domains, one set of barriers","AI's health report card: five data arenas, recurring limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000633,"raw_usage":{"total_tokens":2941,"prompt_tokens":983,"completion_tokens":1958,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":1875}},"tokens_in":599,"tokens_out":1958,"duration_ms":13521,"temperature":1.0,"reasoning_tokens":1875,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:53:48.549827+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A preregistered systematic search of the same five domains over 2013-2020 that applied explicit inclusion and exclusion criteria would settle whether the review's catalogue is comprehensive; if such a search surfaced major untouched research lines, or showed the cited papers to be concentrated in a few well-resourced settings, the paper's map and its challenge taxonomy would need revision.","supporting_citations":[],"review_version":1}