{"id":"9504d05f-8af2-470f-b84f-e854bcc4b061","arxiv_id":"2505.10261","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A topic-modeling survey of 19,123 medical NLP papers shows generative LLM research is concentrated in open-ended tasks, while traditional NLP remains dominant in information extraction and named-entity recognition.","lead":"This paper analyzed 19,123 medical AI studies to compare where generative large language models and older natural language processing methods are used. It found that LLM research clusters around open-ended tasks such as medical education and summarization, while traditional NLP papers focus more on structured tasks like extracting information from health records.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The keyword-regex split into generative LLM vs traditional NLP (Supplementary D) is unvalidated and both over- and under-inclusive, so the 4,295/14,828 grouping and all topic proportions that support the central claim are not trustworthy.","rationale":"Good-faith reading: this is a useful descriptive bibliometric study with a transparent workflow (BERTopic, MedCPT embeddings, expert topic merging) and a large corpus. The problem is not the pipeline per se but the unvalidated keyword rule that defines the two comparison groups. The specific regex patterns in Supplementary D create obvious false positives (RAG gene, llama nanobodies, palmar) and obvious false negatives (GPT-family model names), and neither the search strategy nor the classifier includes GPT-family terms, suggesting systematic under-coverage of the generative LLM literature. Since the paper's conclusion is phrased as 'demonstrate advantages,' it goes beyond what a publication-count analysis can support even with perfect grouping; but the descriptive landscape claim is the load-bearing part, and it fails if the grouping is noisy. The reader's weakest assumption identifies the same issue, so I agree. Because the concern is concrete and testable - a small annotation study would settle it - the appropriate verdict remains conditional pending validation and rewording, not outright rejection.","tokens_in":10462,"tokens_out":7361,"duration_ms":69683,"concrete_test":"Stratified by the 26 merged topics, randomly sample 300 abstracts from the 19,123 corpus. Two independent annotators, blind to the keyword labels, classify each abstract as generative LLM, traditional NLP, both, or neither, with adjudication by a third. Compute precision and recall of the Supplementary D regex against this gold standard, then re-estimate all topic proportions using corrected labels. If the generative-LLM share in 'Medical Education' or 'Text Summarization' shifts by more than 5 percentage points, or if any topic's rank ordering by generative-LLM share changes, the reported split is an artifact of the keyword rule.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that generative LLMs concentrate in open-ended tasks while traditional NLP dominates extraction/analysis rests on the binary split made in Data Processing using Supplementary D. That split is not validated and is both over- and under-inclusive. Over-inclusive examples: the regex \\b(rag)\\b matches 'rag' as a word, which in biomedical text is commonly the recombination-activating gene; \\b(palm)\\b matches 'palmar' and other anatomical uses; \\b(llama)\\b matches llama-derived nanobodies (single-domain antibodies). Any abstract containing one of these plus any NLP/language-model phrase is counted as a generative LLM study, inflating LLM shares in immunology/oncology and similar topics. Under-inclusive examples: the list contains no pattern for GPT, GPT-3, GPT-4, OpenAI, o1, DeepSeek, Claude, Mistral, or 'generative pre-trained transformer' as a standalone phrase; only the generic 'large language model' catches such papers, and the search strategy in Supplementary C also omits GPT-family terms, so a substantial body of model-named generative LLM work may be missing from the corpus entirely. No precision/recall, manual audit, or inter-annotator agreement is reported. Because every downstream number - 4,295 vs 14,828, the topic percentages, and the qualitative 'advantages' conclusion - is computed from this rule, the headline finding is not established unless the rule is shown to be accurate.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a bibliometric and topic-modeling analysis of 19,123 medical NLP studies retrieved from PubMed, Embase, Scopus, and Web of Science. The studies are classified into generative LLM and traditional NLP groups by keyword regex matching, embedded with the MedCPT Article Encoder, reduced with UMAP, clustered with BERTopic, and merged into final topic labels by medical experts. The authors report that generative LLM studies concentrate in topics such as Medical Education and Text Summarization, while traditional NLP studies concentrate in Electronic Health Records and Named Entity Recognition, and they conclude from these publication proportions that generative LLMs demonstrate advantages in open-ended tasks while traditional NLP dominates information extraction and analysis tasks.","tokens_in":10745,"tokens_out":5042,"duration_ms":47822,"significance":"If the descriptive claims were established, the paper would offer a useful map of where medical NLP research effort has been concentrated across two technological paradigms. The study has strengths: a large multi-database corpus, a domain-specific embedding model, an explicit keyword list in Supplementary D, and topic keyword tables in Supplementary B. However, the headline 'advantages' claim goes beyond the descriptive evidence because no performance or comparative evaluation data are reported, and the validity of the keyword-based binary split is not demonstrated. Reproducibility is also limited: no code repository is provided and the data availability statement is ambiguous. The paper's contribution is therefore better characterized as a descriptive landscape analysis than as an evaluation of task-specific advantages.","major_comments":[{"comment":"The abstract claims that 'generative LLMs demonstrate advantages in open-ended tasks, while traditional NLP dominates in information extraction and analysis tasks,' but the study reports no task-performance data or head-to-head evaluations. The cited evidence (e.g., Medical Education 72.23%, Text Summarization 19.95%, EHR 23.62%) consists of proportions of published studies per topic, and a higher publication share does not establish an advantage in task quality or accuracy. This inference is load-bearing for the central claim; the claim should be rephrased as a statement about research focus and publication activity, or supplemented with comparative performance evidence.","section":"Abstract"},{"comment":"The binary split into generative LLM versus traditional NLP is made by keyword regex matching in Supplementary D and is neither validated nor audited. The regex \\b(palm)\\b will match 'palmar' and other anatomical uses; \\b(llama)\\b will match llama-derived nanobodies; and \\b(rag)\\b will match the recombination-activating gene in biomedical text, so abstracts containing such terms plus a generic NLP phrase are counted as generative LLM studies. Conversely, the list omits patterns for GPT-3, GPT-4, OpenAI, o1, DeepSeek, Claude, and Mistral, and the search strategy in Supplementary C omits these terms except 'chatgpt', so many model-named generative LLM papers are likely absent from the corpus entirely. No precision, recall, manual audit, or inter-annotator agreement is reported. Because the counts 4,295 versus 14,828 and all topic proportions derive from this rule, the central comparison is not trustworthy until the classification rule is validated.","section":"Methods: Data Processing / Supplementary D"},{"comment":"The manuscript states that 40 initial topics were merged by medical experts into 26 topics, but Supplementary B lists 40 keyword sets with several duplicate topic labels (e.g., Medical Image Analysis appears as Topics 5, 9, and 34, and Mental Health & Psychology appears multiple times), so the mapping from initial to final topics is not fully documented. No inter-annotator agreement or reliability measure is reported for the expert merging. Since topic-level proportions are the basis for the main conclusions, the merging procedure needs to be transparent and reproducible.","section":"Methods: Topic Modeling / Supplementary B"},{"comment":"The near-complete separation in Dimension 1 of the UMAP embedding is presented as evidence of disparity between the two groups, but this separation is expected because the grouping labels were assigned using keywords that appear in the same titles and abstracts used to compute the embeddings. This visualization should be presented as descriptive, with the caveat that it is partly a consequence of the classification rule, rather than as independent evidence for two distinct research paradigms.","section":"Main Text, Figure 2 paragraph"}],"minor_comments":[{"comment":"The Code Availability section states that 'The data used in this study is available upon request,' which conflicts with the preceding Data Availability statement that the full data cannot be publicly shared; please clarify what artifacts (code, topic assignments, classification labels) are actually available and provide a code repository.","section":"Code Availability"},{"comment":"Supplementary A is described as a PRISMA flow diagram, but the text provided contains only a narrative summary; please confirm that the actual flow diagram is included in the supplement.","section":"Supplementary Information A"},{"comment":"Reference 19 is incomplete, consisting only of 'Website.' followed by a URL, and references 18 and 19 are non-archival web sources; please provide full citation details and access dates.","section":"References"},{"comment":"The sentence 'Figure 2 shows the distribution ratio ...' appears to refer to Figure 3, which displays topic distributions; please correct the cross-reference.","section":"Main Text, Figure 3 cross-reference"},{"comment":"Footnote 1 contains the phrase 'a continuously period'; this should read 'a continuous period.'","section":"Footnote 1"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the unvalidated keyword classification and the unsupported 'advantages' language in the abstract. If the authors provide a manual audit or precision/recall evaluation of the classification rule and reframe the central claim as a descriptive landscape analysis, the paper could be publishable. Please consider whether the journal's scope requires the comparative performance claim to be supported by data rather than by publication counts."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful descriptive map of 19,123 medical NLP papers, but the binary split that drives every percentage is built from an unvalidated keyword regex, and the abstract's 'advantages' language is not supported by activity counts. I'd send it to serious peer review only with major revision.\n\nWhat's actually new: the scale. Four databases searched, deduped to 20,228, cleaned to 19,123, embedded with MedCPT, reduced with UMAP, clustered with BERTopic, and 40 raw topics merged by medical experts into 26. As far as I can tell, this is the first time this particular generative-vs-traditional comparison has been done at this size. The PRISMA flow is standard and the methods are described plainly. If you want a first-pass map of where LLM and traditional NLP work in medicine is concentrated, the topic proportions are a reasonable starting point. Citation pattern is clean; the self-citations are relevant to the method.\n\nThe soft spot is load-bearing. Supplementary D lists the classification regexes: llama, palm, rag, gemini, llm, chatgpt. The stress-test note is correct. In biomedical text, '\\b(rag)\\b' matches recombination-activating gene, '\\b(palm)\\b' matches palmar, '\\b(llama)\\b' matches llama-derived nanobodies. Any abstract with one of those plus any NLP phrase gets counted as generative LLM, which inflates LLM shares in immunology and other topics. Worse, the list omits GPT, GPT-4, OpenAI, o1, DeepSeek, Claude, Mistral, and 'generative pre-trained transformer'. The search strategy in Supplementary C also omits GPT-family terms. So a substantial body of model-named generative LLM work may be missing from the corpus entirely. No precision/recall, no manual audit, no inter-annotator agreement is reported. Every downstream number—4,295 vs 14,828, the topic percentages, the complementary-advantages conclusion—is computed from this rule. Until the rule is validated, the headline finding is not established.\n\nThere's also an interpretive overreach: the abstract says generative LLMs 'demonstrate advantages' in open-ended tasks. The data are research-activity counts, not performance comparisons. 'Advantages' is the wrong word; 'research focus' would be honest.\n\nMinor issues: no confidence intervals on the percentages, no code or data release (the statement says data available upon request but gives no mechanism), and the expert merging has no reliability check. These are minor relative to the split problem.\n\nWho it's for: researchers, funders, educators wanting a rough landscape. It is not a methods contribution. It deserves a serious referee because the corpus is large and the map is useful, but the classification needs external validation, the missing GPT-family terms need to be added, and the claims need to be scaled back before the numbers can be trusted.","headline":"A useful 19k-paper map of medical NLP, undercut by an unvalidated keyword regex split and an overreaching 'advantages' claim.","tokens_in":11321,"tokens_out":4017,"would_cite":false,"duration_ms":34594,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that across 19,123 medical NLP studies, generative LLMs and traditional NLP occupy complementary niches: LLMs concentrate in open-ended tasks like education and summarization, while traditional NLP still leads structured…","keywords":["generative large language models","traditional NLP","medical natural language processing","topic modeling","BERTopic","research landscape","information extraction","medical education"],"falsifier":"Manually label a random sample of 300 abstracts from each group, compute the topic proportions from the corrected labels, and compare them to the paper's Figure 3; if the LLM group's share in open-ended topics drops sharply toward the traditional group's once passing mentions are removed, the central split is an artifact of the keyword classifier.","tokens_in":10268,"feed_emoji":"📊","tokens_out":4713,"duration_ms":43573,"temperature":0.7,"pith_summary":"This paper maps the current research landscape of medical natural language processing by analyzing 19,123 studies retrieved from four bibliographic databases. It argues that generative large language models and traditional NLP have complementary strengths, with LLM research concentrated in open-ended tasks such as medical education and text summarization, and traditional NLP still dominating structured information-extraction tasks such as named entity recognition and electronic health record mining. The authors combine keyword-based classification with neural topic modeling to measure how much of each topic's literature belongs to each technological paradigm. If the mapping is right, it tells researchers and clinicians which tool class to reach for, and where hybrid systems may be needed.","feed_headline":"19,123 studies: LLMs and traditional NLP split medical work by task","feed_subtitle":"Generative models concentrate in education and summarization; traditional NLP still leads EHR and entity recognition.","key_machinery":"The argument is carried by a topic-modeling pipeline: each study is embedded with a biomedical article encoder trained on PubMed query-article pairs, embeddings are reduced to four dimensions, clusters are found by density-based clustering, and topics are labeled and merged with expert review (with assistance from a large language model) to yield 26 topics. Keyword-based classification into generative LLM versus traditional NLP groups is applied to titles and abstracts first, then topic proportions are computed per group. The mechanism that produces the conclusion is the joint distribution of those two groupings: the same topic space, partitioned by paradigm, reveals where each paradigm's research attention concentrates.","core_discovery":"The central discovery is an empirical division of labor: in a corpus of 19,123 studies, generative LLM papers (4,295) cluster in open-ended and generative use cases, while traditional NLP papers (14,828) keep the lead in structured extraction and analysis. Concretely, the topic 'Medical Education' has 72.23% of its studies in the LLM group, and 'Text Summarization' and 'Medical Image Analysis' are also LLM-heavy, whereas 'Electronic Health Records' (23.62%), 'Named Entity Recognition' (13.70%), and 'Semantic and Lexical Processing' (9.95%) dominate the traditional NLP share. The semantic-embedding visualization shows near-complete separation between the two groups along one dimension, with LLM studies more concentrated and traditional NLP more dispersed. The paper reads this as evidence that the two paradigms are complementary rather than mutually replacing: LLMs bring flexibility and generation, traditional methods bring controllability and precision.","pith_inferences":["The paper measures research activity, not measured performance; the word 'advantages' is an inference from where studies concentrate, not from benchmark results.","The near-perfect separation in the embedding space may be partly a mechanical effect of the keyword classification, since keywords such as 'LLM' appear in the abstracts that are embedded.","A cheap test of the mapping is to rerun the same pipeline on the same corpus with a manually labeled random sample of a few hundred abstracts; stable topic proportions would strengthen the claim, large shifts would call it into question.","Extending the search window past March 2025 could show whether the 72% share in medical education persists or migrates toward extraction tasks as LLMs gain instruction-following precision."],"forward_implications":["Medical NLP research appears to be splitting by task type, so future LLM work should target open-ended and generative functions while traditional NLP remains the default for structured extraction.","Hybrid systems that combine LLM generation with the precision of traditional extraction methods are a natural next step for electronic health record applications.","The near-complete semantic separation between the two groups suggests the field has not yet converged; as reasoning-capable LLMs mature, their share may extend into clinical reasoning and decision support.","For clinicians and educators, the concentration of LLM studies in medical education points to scalable simulation and self-assessment tools, provided ethical and privacy safeguards are in place."],"supporting_citations":[{"why":"Supplies the biomedical article encoder that embeds all 19,123 studies into the semantic space used for topic modeling.","marker":"11"},{"why":"Provides the UMAP dimension-reduction algorithm that maps embeddings to four dimensions before clustering.","marker":"12"},{"why":"Supplies the BERTopic package used to generate the initial 40 topics via class-based TF-IDF.","marker":"13"},{"why":"Supports the interpretation that conventional NLP methods retain an advantage in tasks requiring precision and controllability, such as clinical named entity recognition.","marker":"17"},{"why":"Provides the HDBSCAN density-based clustering algorithm that automatically determines the number of semantic clusters.","marker":"22"},{"why":"Describes the GPT-4o model used to optimize topic representations in the pipeline.","marker":"23"}],"fun_headline_variants":["LLMs for open-ended, traditional NLP for extraction—19k studies","19,123 studies: LLMs for generation, NLP for analysis","LLMs lead open-ended tasks, traditional NLP leads extraction","Medical AI: LLMs plus NLP—division of labor across tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison rests on classifying each paper by whether its title or abstract mentions generative-LLM keywords, so papers that merely compare against an LLM or mention one in passing are counted as LLM studies, and that noise could change the topic proportions.","fun_headline_variants_meta":{"raw":{"variants":["LLMs for open-ended, traditional NLP for extraction—19k studies","19,123 studies: LLMs for generation, NLP for analysis","LLMs lead open-ended tasks, traditional NLP leads extraction","Medical AI: LLMs plus NLP—division of labor across tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000695,"raw_usage":{"total_tokens":3078,"prompt_tokens":812,"completion_tokens":2266,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":2192}},"tokens_in":428,"tokens_out":2266,"duration_ms":17973,"temperature":1.0,"reasoning_tokens":2192,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:12:49.669222+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Manually label a random sample of 300 abstracts from each group, compute the topic proportions from the corrected labels, and compare them to the paper's Figure 3; if the LLM group's share in open-ended topics drops sharply toward the traditional group's once passing mentions are removed, the central split is an artifact of the keyword classifier.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the biomedical article encoder that embeds all 19,123 studies into the semantic space used for topic modeling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the interpretation that conventional NLP methods retain an advantage in tasks requiring precision and controllability, such as clinical named entity recognition."},{"cited_title":"& Astels, S","cited_arxiv_id":null,"evidence_quote":"Provides the HDBSCAN density-based clustering algorithm that automatically determines the number of semantic clusters."}],"review_version":1}