{"id":"6179c2a1-00f3-4949-9905-d5363153d908","arxiv_id":"2505.12220","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A scoping review of 32 studies finds that ML models for suicide risk in military and veteran groups achieve moderate accuracy, identify consistent risk factors, but lack PPV/NPV, survival metrics, and clinical integration.","lead":"This paper reviews 32 studies that use machine learning to predict suicide risk in military and veteran populations. It summarizes common risk factors and highlights gaps in how predictions are tested and used.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's claim of consistent risk factors is undercut by the review's own reported heterogeneity, especially for PTSD and other mental health conditions, so the central synthesis overstates cross-study agreement.","rationale":"The reader's weakest assumption concerned corpus representativeness (single-author screening and title-only Google Scholar search). That is a legitimate methodological limitation, but it affects completeness rather than the truth of the synthesis within the included corpus. The more directly load-bearing weakness is internal: the abstract's 'consistently identified' claim conflicts with the paper's own reported heterogeneity. If the synthesis is meant to show common factors despite variability, the review needs to demonstrate that agreement, not merely assert it. The paper gives credit where due: it follows PRISMA, provides detailed tables, and includes explicit caveats about mixed PTSD findings and missing PPV/NPV metrics; those caveats make the overstatement fixable without changing the overall scope. Therefore, I would keep the reader's CONDITIONAL verdict, but for a slightly different reason: revise the consistency language and add a frequency-based support check rather than only address search strategy. Agreement with the reader is partial because the reader's weakest assumption (search bias) differs from this concern, although both point to the reliability of the synthesis.","tokens_in":16949,"tokens_out":7598,"duration_ms":76732,"concrete_test":"Re-extract from each of the 32 included studies the reported leading risk factors, grouped by the paper's outcome categories (SI, SA, SM), and compute for each abstract-level factor class (depression, PTSD, prior suicidal thoughts/attempts, physical health, demographics) the fraction of studies in each outcome group that list it as a leading factor. If, for example, PTSD appears as a leading factor in fewer than half of the SA and SM studies, the consistency claim should be revised to a weaker formulation such as 'recurring across some outcomes but with notable heterogeneity.' For each row of Table 3, also verify that the cited papers actually report the listed factors as leading, and flag any row supported by only one study in the final stated claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the 32 studies 'consistently identified' a common set of risk factors, including PTSD and other mental health conditions, is not fully supported by the evidence the review itself presents. Section 3.4.2 states that 'PTSD was not a major factor in predicting suicide attempts and mortality, despite many in the sample having a PTSD diagnosis' and that 'there were mixed findings regarding the importance of mental health disorders in predicting suicide attempts for veterans.' The Discussion similarly concedes that 'the role of PTSD and other mental health conditions varies considerably across different study samples.' Table 3 is organized by outcome and population, and several rows rest on a single citation (e.g., suicide attempt in general active service members cites ref. 28; suicide mortality after psychiatric outpatient visits cites ref. 13). The abstract's phrase 'consistently identified' implies a level of cross-study agreement that a scoping narrative does not establish; no vote-count, frequency analysis, or any other systematic comparison of factor importance is provided. The review's own caveats about mixed findings are in tension with the headline synthesis, and that tension is load-bearing because the paper's main contribution is that synthesis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a PRISMA-based scoping review of machine learning (ML) studies that assess or predict suicide-related outcomes (suicidal ideation, attempts, mortality) in active-duty service members and veterans. The authors searched PubMed, IEEE Xplore, ACM Digital Library, and Google Scholar, and retained 32 studies published between 2014 and 2024. The review summarizes study characteristics (datasets, sample sizes, ML techniques, outcome definitions, performance metrics) in Table 1, and it organizes reported risk factors into categories in Tables 2–4. The stated central findings are that, despite large variability across studies, a consistent set of risk factors emerges (depression, PTSD, prior suicidal ideation/attempts, physical health problems, demographic characteristics), and that ML models achieve reasonable predictive accuracy, with AUCs typically around 0.80–0.85. The Discussion also identifies research gaps: underuse of PPV/NPV and survival-specific metrics, limited longitudinal modeling, insufficient attention to clinical rationale, and narrow demographic diversity. The conclusion emphasizes that ML has verified known risk factors on a large scale and that prevention strategies must be comprehensive and flexible.","tokens_in":17127,"tokens_out":5314,"duration_ms":51542,"significance":"If the synthesis is accurate, this review provides a useful map of a rapidly growing but fragmented literature. Its strengths include a reproducible search protocol, detailed extraction tables, a 2014–2024 coverage window, and explicit recognition of heterogeneity in the Discussion. The paper also flags practical gaps (e.g., underreporting of PPV/NPV, absence of c-index in survival-frame studies) that are actionable for future work. However, the significance is diminished because the headline claim of 'consistently identified' risk factors is not backed by a systematic comparison of factor importance across studies, and the review's own text reports contrary evidence for PTSD and other mental-health diagnoses. The paper would be strengthened by either softening that claim or adding a formal cross-study synthesis (e.g., frequency counts of factor categories, direction of association, number of studies supporting each row in Table 3). The topic is timely for ML suicide research, and the paper has the potential to serve as a reference for interdisciplinary teams, but the current framing overstates the certainty of the evidence.","major_comments":[{"comment":"The abstract's claim that the 32 studies 'consistently identified' a common set of risk factors (depression, PTSD, suicidal ideation, prior attempts, physical health, demographics) is not reconciled with the review's own findings. Section 3.4.2 states that 'PTSD was not a major factor in predicting suicide attempts and mortality, despite many in the sample having a PTSD diagnosis' and that 'there were mixed findings regarding the importance of mental health disorders in predicting suicide attempts for veterans.' The Discussion likewise concedes that 'the role of PTSD and other mental health conditions varies considerably across different study samples.' Because the synthesis is the paper's main contribution, this internal tension is load-bearing; the authors should either rephrase the conclusion to acknowledge heterogeneity explicitly or provide a systematic, reproducible comparison (e.g., counting how many studies report each factor as leading, with direction of effect) to substantiate 'consistently identified.'","section":"Abstract; §3.4.2; §4"},{"comment":"The corpus underlies all subsequent claims, but the screening process has a material risk of selection bias: title/abstract screening appears to have been conducted by a single author (YZ, per Figure 1), and the Google Scholar search was restricted to a title-only query. Single-reviewer screening is known to miss potentially relevant records, and a title-only search will systematically exclude studies where 'machine learning' or 'suicide' appear only in the abstract or full text. The authors should either re-screen a sample independently and report inter-rater agreement, or at least discuss these as limitations in the Discussion; the current manuscript does neither.","section":"§2.3; §2.4; Figure 1"},{"comment":"The inclusion criteria do not require completed data analysis, and accordingly two of the 32 included items are study protocols without results: ref. 40 (Brown et al., a study protocol) and ref. 41 (Meerwijk et al., a protocol for a mixed-method study). Including protocols in the count and in Table 1 inflates the corpus and does not provide evidence for the review's claims about predictive performance or risk factors. These items should be excluded from the synthesis or analyzed separately and clearly labeled as protocols, with the study total adjusted accordingly.","section":"Table 1; §2.1"},{"comment":"Table 3 presents 'Leading factors' for each outcome–population subgroup, but many rows rest on a single citation (e.g., 'Suicide attempt — General active service members population' cites ref. 28; 'Suicide mortality — Current service members after psychiatric outpatient visits' cites ref. 13). Since the table is the principal evidence for cross-study consistency, the authors should indicate how many studies support each row and, where a row reflects only one study, say so explicitly. Without this, the visual layout of the table implies a degree of replication that the underlying studies may not provide.","section":"Table 3"}],"minor_comments":[{"comment":"The Discussion states that 'Thompson et al. represented a pioneering effort, utilizing social media data for assessing suicide risk' (with ref. 10), but Table 1 lists Thompson et al. as using EHR data. The social-media study appears to be Zuromski et al. (ref. 23). This attribution needs correction.","section":"§4"},{"comment":"The Introduction cites the '2024 National Suicide Prevention Annual Report,' but reference 7 is titled '2023 national veteran suicide prevention annual report.' The year in the text and the reference year should be aligned.","section":"§1"},{"comment":"The PRISMA flow diagram shows 'Records identified from databases (n = 1,110),' but subsection 2.3 reports two separate queries, with the second (deep-learning query) run on December 15, 2024. It is unclear whether the 1,110 includes both queries; the figure and text should be made consistent.","section":"Figure 1"},{"comment":"The contributorship statement mentions experimental studies by 'G.H., M.L.' and data interpretation by 'J.B.'; these initials do not correspond to any listed author (the author list includes G.L.H., J.M.B., and others with different initials). This appears to be a typographical error and should be corrected.","section":"Contributorship statement"},{"comment":"The sentence 'The most statistically significant predictive factors ... were identified from articles and included in the tables as leading factors' does not specify the criterion used to designate a factor as 'leading' (e.g., rank order, effect size, author-reported importance). A brief operational definition would improve transparency.","section":"§2.4"},{"comment":"The Discussion notes that one study reported an AUC of 1.00 'which needs further clarification and explanation,' but no resolution is proposed. Since such a result is a red flag for overfitting or leakage, the authors should recommend a concrete reporting standard or explain how to interpret such values in future reviews.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a competent scoping review with useful tables, but the central synthesis claim is stronger than the evidence presented, and the corpus has a few inclusion/screening issues (single-reviewer screening, title-only Google Scholar query, protocols counted as studies). These are fixable within the manuscript's scope: adjust the abstract and conclusion to reflect heterogeneity, add a systematic factor-frequency analysis, and revisit the inclusion of protocols. I recommend major revision rather than rejection. I also suggest the editor ask the authors to double-check the internal consistency of references 7 and 10, as those affect reader trust."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nQuick take: this is the first scoping review specifically on ML for suicide outcomes in military and veteran populations, and it is a genuinely useful structured map of 32 studies. The extraction tables and the sex-specific breakdown are the most valuable parts. The abstract, though, says the studies 'consistently identified' factors like PTSD when the body of the review reports mixed findings. That tension needs fixing before the synthesis is cited as solid.\n\nWhat's good: the PRISMA flow is documented, the inclusion criteria are reasonable, and the tables give detailed study-level information—population, outcome, metric, and ML technique. The authors also flag real gaps: missing PPV/NPV reporting, few survival-specific metrics, limited deep learning, underrepresentation of non-U.S. and minority samples, and the odd AUC of 1.00 that needs explanation. The discussion of sex differences is careful and grounded in the included studies. For a review, the descriptive synthesis of common risk factors (depression, prior attempts, etc.) is a fair reading of most of the included work.\n\nSoft spots: (1) The claim of consistency is stronger than the evidence. Section 3.4.2 explicitly says PTSD was not a major factor in some samples and that mental health findings were mixed; the Discussion later says the role of PTSD varies considerably across samples. The abstract and conclusion should mirror that nuance. (2) Title/abstract screening was done by a single author, and Google Scholar used a title-only search; that's a legitimate limitation for a scoping review but should be acknowledged in the methods. (3) Two included items are protocols without results (refs 40–41); fine as descriptions of future work, but counting them as evidence for the synthesis is questionable. (4) The Discussion claims ML achieved 'higher prediction accuracy' and 'noticeable improvements' over traditional models without a systematic comparison; most included studies don't compare against a well-specified traditional model, so that is overreach.\n\nNone of these are fatal for a scoping review. The paper maps the area and identifies gaps, and it does that well. The fix is mostly about calibrating claims and being transparent about the screening process.\n\nBottom line: worth a serious referee. If you work on suicide prediction or military/veteran mental health, this is a useful reference. Read it with the abstract in mind, but trust the tables more than the headline.","headline":"A useful scoping review of 32 ML studies on suicide in military/veteran populations, but the abstract overstates the consistency of risk factors that the body itself shows are mixed.","tokens_in":17715,"tokens_out":2342,"would_cite":false,"duration_ms":22543,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This scoping review of 32 studies argues that machine learning studies of suicide in military and veteran populations, despite their differences, converge on a common set of risk factors—depression, PTSD, prior suicidal thoughts or…","keywords":["suicide prevention","machine learning","military veterans","risk factors","suicidal ideation","suicide attempt","suicide mortality","scoping review"],"falsifier":"Run the same review with dual independent screening, full-text keyword searches in all databases, and no country restriction; if the study set changes the core risk-factor list or moves the typical AUC range outside roughly 0.80 to 0.85, the convergence claim is weakened. A complementary test is to prospectively apply the common factor set to a new military or veteran cohort and check whether predictive accuracy reproduces.","tokens_in":16748,"feed_emoji":"🧠","tokens_out":10189,"duration_ms":95353,"temperature":0.7,"pith_summary":"This paper tries to establish that machine learning studies of suicide in military and veteran populations, despite broad differences in samples, data sources, outcomes, and algorithms, converge on a common set of risk factors: depression and other mental-health problems, post-traumatic stress disorder (PTSD), prior suicidal thoughts or attempts, physical health problems, and demographic characteristics. The authors screened 1,110 records, retained 32 studies, and synthesized their reported predictors and performance. They argue that this convergence matters because machine learning can be applied at scale in large health systems to flag at-risk individuals, and because the breadth of factors shows that effective prevention must be multi-component. They also claim that current studies are missing the metrics and modeling choices needed for prevention policy: most omit positive and negative predictive values (the rates of false alarms and missed cases), few treat time-to-event properly, and most do not connect model outputs to clinical reasoning. A sympathetic reader would care because the synthesis points to which risk signals are robust enough to build screening programs on and where the evidence is still thin.","feed_headline":"Machine learning finds shared suicide risks across 32 military studies","feed_subtitle":"Typical accuracy scores reach 0.80–0.85, yet most studies skip the error metrics prevention policies need.","key_machinery":"The argument is carried by a structured literature-review and synthesis process. Four literature databases were searched, records were screened against five eligibility criteria, full texts were reviewed by two reviewers, and the retained studies were compared on study population, data modality, outcome, metrics, and leading risk factors. The load-bearing analytic device is the risk-factor categorization scheme, which groups predictors reported across studies into common domains such as mental-health problems, substance use, physical health, military experience, demographics, and trauma or interpersonal violence. This categorization is what allows the authors to demonstrate convergence despite methodological variation, while the tables of leading factors per outcome and the reported AUC values supply the evidence for that convergence.","core_discovery":"The central discovery is a convergence result across heterogeneous studies: no single dataset or algorithm dominates, yet the same risk clusters recur. Mental-health problems, PTSD, prior suicidal thoughts and attempts, physical health problems, traumatic brain injury, relationship and financial stressors, military experience variables, and demographic factors appear again and again as leading predictors. The paper reports that most models achieve AUC values around 0.80 to 0.85, where AUC is a standard ranking-accuracy score, and notes one reported AUC of 1.00 that it flags as needing further explanation. It also identifies sex-specific patterns, with alcohol misuse and sexual abuse weighing more heavily for female veterans and traumatic brain injury more prominently for males. From this, the paper concludes that machine learning has verified on a large scale the risk factors previously found by manual analytic methods, and that prevention strategies must be comprehensive and flexible.","pith_inferences":["An extension the paper leaves implicit: if the factor convergence is real, the same core set should perform reasonably in non-U.S. military cohorts, and a prospective external-validation study would test this transferability.","Not stated in the paper but implied by its metric critique: choosing a decision threshold changes both which factors look important and how many false alarms a program must absorb; reporting threshold-dependent metrics would likely alter the priority ordering of risk factors.","The single AUC of 1.00 that the review flags as needing clarification is, in editorial inference, most plausibly a sign of data leakage or overfitting; routine sharing of code and data would let readers distinguish that from a genuine result.","Because most studies were conducted in the U.S. with predominantly White samples, the paper's factor list is a claim about that population; applying the same pipelines to underrepresented racial and ethnic groups would reveal whether the common factors are universal or context-bound."],"forward_implications":["If the convergence claim holds, suicide-screening systems for military and veteran populations can be built around a stable core of variables—mental-health diagnoses, prior suicidal thoughts or attempts, physical health problems, and demographic indicators—without depending on any single algorithm.","The typical AUC range of 0.80 to 0.85 implies models are informative but far from deterministic, so real deployments must plan for both false alarms and missed cases.","Because most studies omit positive and negative predictive values, current evidence cannot tell prevention programs how many flagged individuals would be false positives; better error reporting is needed before cost-based policy decisions.","The scarcity of survival-specific metrics such as the c-index means the evidence is weak on when risk peaks; studies that treat time-to-event explicitly would strengthen intervention timing.","The diversity of leading factors across studies implies that prevention strategies must be multi-component and tailored to sex and subpopulation rather than a single risk score."],"supporting_citations":[{"why":"Supplies a veteran suicidal-ideation model with AUC 0.91 and a sex-specific analysis of trauma and mental-health predictors.","marker":"[18]"},{"why":"Provides a veteran suicidal-ideation prediction model based on psychosocial well-being indicators with AUC 0.83.","marker":"[12]"},{"why":"Identifies leading suicidal-ideation predictors in the first post-service year, including sex differences.","marker":"[21]"},{"why":"Shows a survival and ensemble model predicting suicide attempts among soldiers who deny suicidal ideation, with AUC 0.83.","marker":"[26]"},{"why":"Contributes pre-deployment predictors of suicide attempt from large military cohort data, with AUC 0.78.","marker":"[28]"},{"why":"Predicts post-separation suicide attempts with AUC 0.74 and identifies lifetime suicide plans and trauma as top factors.","marker":"[29]"},{"why":"Provides a model of suicide after outpatient mental-health visits with AUC 0.75, a key mortality-outcome example.","marker":"[13]"},{"why":"Provides an administrative-data model of suicide after psychiatric hospitalization with AUC 0.79, a key mortality example.","marker":"[35]"},{"why":"Demonstrates a large-scale deep neural network for suicide attempt risk reporting PPV 0.55, illustrating the metric gap.","marker":"[32]"},{"why":"Shows an NLP-based clinical-note model with AUC 0.69, evidence for textual data and modest accuracy.","marker":"[37]"}],"fun_headline_variants":["ML review: 32 studies, same suicide risk factors","Suicide risk factors recur across 32 military ML studies","ML confirms shared suicide risks, but metrics gaps remain","ML review highlights suicide risk factors, yet misses key metrics","Common suicide risks surface across 32 military ML studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's conclusions rest on the assumption that the 32 retained studies fairly represent the full body of relevant research; because one reviewer performed title and abstract screening and one large database was searched by title only, a different or broader search could change the common risk-factor list or the reported accuracy range.","fun_headline_variants_meta":{"raw":{"variants":["ML review: 32 studies, same suicide risk factors","Suicide risk factors recur across 32 military ML studies","ML confirms shared suicide risks, but metrics gaps remain","ML review highlights suicide risk factors, yet misses key metrics","Common suicide risks surface across 32 military ML studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000546,"raw_usage":{"total_tokens":2631,"prompt_tokens":984,"completion_tokens":1647,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":1568}},"tokens_in":600,"tokens_out":1647,"duration_ms":13292,"temperature":1.0,"reasoning_tokens":1568,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:37:32.766653+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same review with dual independent screening, full-text keyword searches in all databases, and no country restriction; if the study set changes the core risk-factor list or moves the typical AUC range outside roughly 0.80 to 0.85, the convergence claim is weakened. A complementary test is to prospectively apply the common factor set to a new military or veteran cohort and check whether predictive accuracy reproduces.","supporting_citations":[{"cited_title":"Gender differences in machine learning models of trauma and suicidal ideation in veterans of the iraq and afghanistan wars","cited_arxiv_id":null,"evidence_quote":"Supplies a veteran suicidal-ideation model with AUC 0.91 and a sex-specific analysis of trauma and mental-health predictors."},{"cited_title":"How well can U.S","cited_arxiv_id":null,"evidence_quote":"Provides a veteran suicidal-ideation prediction model based on psychosocial well-being indicators with AUC 0.83."},{"cited_title":"Predicting suicide attempts among soldiers who deny suicidal ideation in the army study to assess risk and resilience in servicemembers (army STARRS)","cited_arxiv_id":null,"evidence_quote":"Shows a survival and ensemble model predicting suicide attempts among soldiers who deny suicidal ideation, with AUC 0.83."},{"cited_title":"Pre-deployment predictors of suicide attempt during and after combat deployment: Results from the army study to assess risk and resilience in servicemembers","cited_arxiv_id":null,"evidence_quote":"Contributes pre-deployment predictors of suicide attempt from large military cohort data, with AUC 0.78."},{"cited_title":"Predicting suicide at- tempts among U.S","cited_arxiv_id":null,"evidence_quote":"Predicts post-separation suicide attempts with AUC 0.74 and identifies lifetime suicide plans and trauma as top factors."},{"cited_title":"Predicting suicides after outpatient mental health visits in the army study to assess risk and resilience in servicemembers (army STARRS)","cited_arxiv_id":null,"evidence_quote":"Provides a model of suicide after outpatient mental-health visits with AUC 0.75, a key mortality-outcome example."},{"cited_title":"Using administrative data to predict suicide after psychiatric hospitalization in the veterans health administration system","cited_arxiv_id":null,"evidence_quote":"Provides an administrative-data model of suicide after psychiatric hospitalization with AUC 0.79, a key mortality example."},{"cited_title":"Deep sequential neural network models improve stratification of suicide attempt risk among US veterans","cited_arxiv_id":null,"evidence_quote":"Demonstrates a large-scale deep neural network for suicide attempt risk reporting PPV 0.55, illustrating the metric gap."},{"cited_title":"Leveraging natural language processing to improve electronic health record suicide risk prediction for veterans health administration users","cited_arxiv_id":null,"evidence_quote":"Shows an NLP-based clinical-note model with AUC 0.69, evidence for textual data and modest accuracy."}],"review_version":1}