{"id":"33e24389-7c72-46b7-8321-070a14fb925a","arxiv_id":"2501.06859","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A benchmark of 8 LLMs on Arabic mental health datasets shows structured prompts and model choice drive performance more than language, while few-shot prompting yields large gains for GPT-4o Mini.","lead":"This paper tests eight large language models on Arabic mental health datasets, comparing prompt styles, Arabic versus English text, and few-shot examples. It finds that prompt structure and model choice matter more than language, and that few-shot examples help one model substantially on multi-class tasks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth label quality is the load-bearing assumption: several included datasets have unverified or weak annotation protocols, and the AMI exclusion (Appendix 6.4) proves such errors can be severe; all model rankings and prompt-effect estimates rest on these labels.","rationale":"The paper's primary contributions are empirical rankings and effect-size estimates; every one of those numbers is a function of the ground-truth labels. The strongest claim—prompt structure matters by 14.5 points on multi-class, model selection dominates, few-shot helps—would be meaningless if the labels are systematically wrong. The manuscript itself provides prima facie evidence that label risk is real: AMI was excluded only after all eight models scored below chance and manual inspection revealed widespread labeling faults. The same section that describes dataset collection states that quality could be assessed through the collective judgment of LLMs, a circular standard when those same LLMs are under evaluation. Several retained datasets (DCAT, MCD, ARADEPSU, CAIRODEP, MDE) have unknown or non-expert annotation provenance, yet they contribute to every model-rank and prompt-effect average in Tables 6, 7, and Figures 5, 7, 9. Other concerns—lack of significance tests, few-shot evidence limited to two models, MEDMCQA leakage suspicion—are real but secondary; they affect the strength of inferences, not the validity of the measurement substrate. The label-reliability concern, by contrast, threatens the substrate itself. I therefore agree with the reader's weakest_assumption. The conditional verdict is appropriate, but the 'condition' should explicitly include an independent label audit of the retained datasets; if such an audit substantially changes model ordering or the multi-class prompt gap, the paper's conclusions would need to be revised.","tokens_in":29458,"tokens_out":7886,"duration_ms":76618,"concrete_test":"Have two qualified clinical annotators independently re-label a random stratified sample (200–300 posts per dataset, blind to original labels) from DCAT, MCD, MDE, and CAIRODEP, the datasets with unknown annotator expertise; compute Cohen's kappa and per-class disagreement against the original labels. Then re-run the main model comparisons on the subset with high-confidence corrected labels (e.g., where both annotators agree) and compare rankings and the multi-class prompt gap (ZS-2 vs ZS-1). If the top models change order or the 14.5-point gap collapses below the 5% random-fluctuation threshold used in Section 4.2, the paper's central claims require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claims—Phi-3.5 MoE best balanced accuracy, Mistral NeMo best severity MAE, and the 14.5-point structured-prompt advantage on multi-class tasks—are all computed against dataset labels of unverified reliability. Sections 2.1.1, 2.1.2, 2.1.3, 2.1.5, and 2.2 document that several retained datasets have no stated annotator expertise: DCAT is a minimal Dataverse deposit; MCD was initially keyword-labeled and only later reviewed by unspecified persons; ARADEPSU's annotators are unnamed with authors' revisions; MDE is a student project; CAIRODEP used keyword-based labeling. Section 2's note that 'dataset quality can later be assessed through the collective judgment of LLMs' is circular when the same LLMs are being scored. The AMI episode (Appendix 6.4) is the concrete proof of risk: every model scored below 50% BA on every disorder split, and manual inspection found widespread faulty labeling; the dataset was then excluded. If any retained dataset contains similar errors, model rankings and the prompt/few-shot effect sizes could change materially. The paper itself flags SDCNL Depression as near-random-guessing, which it attributes to difficulty but could equally reflect label noise. Because label quality is the measurement foundation, this assumption is the most load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates eight large language models on Arabic mental-health diagnostic tasks, combining native Arabic datasets with Google-translated versions of English datasets. It compares two zero-shot prompt templates and a few-shot variant across binary, multi-class, severity, and MCQ tasks, reporting balanced accuracy and normalized mean absolute error. The main findings are that the structured prompt ZS-2 outperforms the less structured ZS-1 largely by improving instruction following, with a 14.5-point balanced-accuracy advantage on multi-class datasets; that model choice is the dominant factor, with Phi-3.5 MoE best on balanced accuracy and Mistral NeMo best on severity MAE; that language/translation effects are modest; and that few-shot prompting helps, especially for GPT-4o Mini on multi-class tasks. The authors transparently report invalid-response rates, perform an intersection-of-valid-responses analysis, and exclude the AMI dataset after discovering widespread labeling faults.","tokens_in":29737,"tokens_out":7178,"duration_ms":72069,"significance":"If the findings hold, the paper offers useful practical guidance for LLM-based mental-health screening in Arabic: prompt structure matters mainly through instruction following, smaller open models can be competitive, and cross-lingual evaluations need to be interpreted with translation-quality caveats. The study is valuable for its breadth—eight models, multiple dataset types, and both native and translated Arabic data—and for its unusually candid treatment of parsing and label-quality confounds, including the explicit analysis of the 14.5-point prompt advantage when only parsable responses are considered. However, the headline quantitative claims currently rest on unverified dataset labels and on point estimates without uncertainty quantification, so the conclusions should be treated as provisional until those load-bearing assumptions are addressed.","major_comments":[{"comment":"Ground-truth label reliability is a load-bearing assumption that is not established. Several retained datasets have no documented annotator expertise (DCAT, MCD, ARADEPSU, CAIRODEP, MDE) or use automated labels (SDCNL), and Appendix 6.4 shows that the AMI dataset was excluded only after all models scored below 50% balanced accuracy and manual inspection revealed widespread faulty labeling. Because every model ranking and prompt-effect estimate in Tables 5–9 is computed against these labels, hidden errors of the same kind in any retained dataset could materially change the conclusions. The paper should audit labels on retained datasets, report any available annotator-agreement statistics, and/or run a sensitivity analysis excluding datasets with weak annotation provenance. In addition, the statement in Section 2.1 that 'dataset quality can later be assessed through the collective judgment of LLMs' is circular when the same LLMs are being scored.","section":"§2.1, §2.1.5, §2.2, §2.3.2, Appendix 6.4"},{"comment":"The few-shot claim is overgeneralized. The few-shot experiment covers only two models, Phi-3.5 MoE and GPT-4o Mini, yet the abstract states that 'few-shot prompting consistently improved performance.' Table 9 shows that Phi-3.5 MoE actually lost performance on binary tasks in both the ALL (-0.64) and AR (-1.27) groups, so 'consistently' is not supported even for the tested model. The claim should be restricted to the tested models and task types, or the experiment should be extended to more models before making a general statement.","section":"§4.5, Table 9, Abstract"},{"comment":"The headline quantities—the 14.5 BA multi-class prompt difference, the model ranking, and the language-effect averages—are reported as point estimates without confidence intervals, significance tests, or repeated runs. The 'random fluctuation threshold' of 5% in Section 4.2 is ad hoc, and the evaluation is itself a sample (Section 3.4). Since API outputs can be stochastic and the samples are finite, the paper should provide uncertainty measures (e.g., bootstrap confidence intervals or per-seed variances) for at least the abstract-level claims, or explicitly label them as exploratory. Without this, statements such as 'significantly influences' and 'crucial' go beyond what the data demonstrate.","section":"§4.2, §4.3, §3.4"},{"comment":"The conclusion that language influence is modest is confounded by machine translation quality. All translated corpora were produced by Google Translate with no human evaluation or translation-quality metric, and Section 4.4.2 shows that different translations of the same source (Google Translate vs. BiMediX) differ by up to 16.7 BA. The final section acknowledges this limitation, but the language-effect claims in Sections 4.4.1 and 4.4.3 are still presented as substantive results. A translation-quality check on a sample (e.g., human adequacy/fluency scores or back-translation agreement) would be needed to separate language effects from translation artifacts.","section":"§4.4.1, §4.4.3, §5"}],"minor_comments":[{"comment":"The text says the final DREADDIT dataset contains 3,553 labeled segments, while Table 2 reports a sample size of 1,000 posts; please reconcile these numbers.","section":"§2.3.1, Table 2"},{"comment":"The MDE class counts in Table 10 (600 + 597 + 600 = 1,797) do not match the stated total of 1,800 records; please verify.","section":"§2.2, Table 10"},{"comment":"The term 'fair random sampling' should be described as stratified sampling, and the paper should list which datasets were sampled and the actual sample sizes used, since several datasets in Table 10 have fewer than 1,000 instances.","section":"§3.4"},{"comment":"The confusion-matrix subtraction display is difficult to interpret; please clarify whether the entries are ZS-1 minus ZS-2 or the reverse, and explain how the reported 2.8% increase in negative predictions is derived from the table.","section":"§4.2, Table 4"},{"comment":"The 'Best' prompt row is described as the optimal prompt for each model-dataset pair; please clarify that this is an oracle choice and note that it can overstate the performance achievable without access to ground-truth labels.","section":"Table 7"},{"comment":"Several dataset citations are incomplete or lack venue and publication details (e.g., references [22], [26], and [27]); please complete the bibliographic information.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The authors are unusually transparent about the parsing confound and the AMI label-quality episode, which is a real strength. However, the paper relies heavily on datasets whose annotation provenance is unverified, and the few-shot and statistical-support issues affect the abstract's central claims. I would encourage the editor to request that the authors either provide label audits and uncertainty estimates or substantially soften the corresponding claims. The authors should also consider releasing the evaluation scripts and sampled data to make the study reproducible, since no artifact is currently mentioned."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful, honest benchmark paper for LLM-based mental health diagnosis on Arabic text. The central claims — structured prompts help mainly through instruction following, model choice matters more than language, few-shot helps — are supported by the reported experiments, with the usual benchmark caveats. It extends the authors' English study rather than introducing new methods, but the Arabic/translation axis is genuinely new and practically motivated.\n\nWhat it does well: eight models are evaluated across native Arabic and translated English datasets, with explicit parsing-invalidity analysis and a controlled translation-direction experiment. The prompt analysis is particularly candid: the 14.5-point multi-class BA gain for ZS-2 shrinks to 1.6 in favor of ZS-1 when only parsable responses are compared, so the authors correctly attribute the effect to instruction following rather than diagnostic skill. The MEDMCQA anomaly is investigated rather than waved away, and the AMI dataset is excluded with clear evidence that all models scored below chance and the labels were faulty.\n\nSoft spots, in proportion. The biggest is label quality. Several retained datasets have unverified annotator expertise (DCAT, MCD, ARADEPSU, MDE, CAIRODEP), and the AMI episode proves such noise can be severe. The paper itself flags SDCNL Depression as near-random guessing, which could be difficulty or label noise. Since every ranking and effect size is computed against these labels, this is a real limitation. I would not call it fatal — the rankings are plausible and consistent across datasets — but a sensitivity analysis or a more prominent caveat is needed. Second, the few-shot claim is overgeneralized: only Phi-3.5 MoE and GPT-4o Mini were tested, with no significance tests or confidence intervals, and the \"consistent improvement\" is driven mostly by GPT-4o Mini on multi-class tasks; Phi-3.5 MoE actually lost slightly on binary tasks. Third, the model ranking uses the best prompt per model-dataset pair in places, which can favor unstable models; the AVG columns are better, and fortunately the top model is robust either way.\n\nWho this is for: anyone building or evaluating Arabic mental-health NLP tools, and researchers interested in prompt parsing and label noise in LLM evaluation. It deserves a serious referee. I would send it to peer review with a request for label-quality sensitivity analysis and a narrower few-shot conclusion.","headline":"A solid, transparent benchmark for LLMs on Arabic mental-health text, with real label-quality caveats and an overbroad few-shot claim.","tokens_in":30247,"tokens_out":2051,"would_cite":true,"duration_ms":21988,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Structured prompts outperform loose prompts by 14.5 points on multi-class Arabic mental-health classification, mainly through better instruction following.","keywords":["large language models","Arabic mental health NLP","prompt engineering","few-shot prompting","cross-lingual evaluation","psychiatric diagnostics","instruction following","balanced accuracy"],"falsifier":"Have expert clinicians re-annotate a random sample of ARADEPSU and MCD posts, then recompute the ZS-1 vs ZS-2 balanced-accuracy gap on the corrected labels; if the 14.5-point advantage disappears or reverses, the paper's headline prompt-engineering result is an artifact of label noise rather than instruction following.","tokens_in":29280,"feed_emoji":"🧠","tokens_out":9740,"duration_ms":84218,"temperature":0.7,"pith_summary":"The paper evaluates eight LLMs on Arabic mental-health datasets, both native Arabic corpora and English datasets machine-translated into Arabic, to find what drives diagnostic accuracy. Its central claim is that prompt structure is a major lever: a structured, explicitly formatted prompt beats a looser 'act as a psychologist' prompt by an average of 14.5 balanced-accuracy points on multi-class tasks, and this gap almost entirely disappears when only responses that both prompts can parse are compared. Model choice matters more than language: Phi-3.5 MoE leads in balanced accuracy, Mistral NeMo leads in severity error, and translating datasets from English to Arabic costs only about 6% accuracy on average. Few-shot prompting with one example per class consistently helps, boosting GPT-4o Mini by roughly 20% overall and 57.6% on multi-class tasks. The authors also show that label quality is a live threat: one public dataset was excluded after every model scored at or below chance, underscoring that benchmark labels must be audited before rankings are trusted.","feed_headline":"Structured prompts lift Arabic mental-health LLM scores by 14.5 points","feed_subtitle":"Structured prompts win on multiclass Arabic tasks; few-shot examples add up to 58 percent, and label quality is the hidden risk.","key_machinery":"The central object is the controlled two-template prompt design with parse-validity filtering: for each task type, two semantically equivalent prompt templates differ only in structure and role framing, and a parsing pipeline classifies each model response as a valid label or invalid. By comparing performance on all responses versus only the subset where both prompts yield parsable output, the design isolates instruction-following failures from genuine diagnostic differences. A companion few-shot variant appends one example per class, and the same parsing and balanced-accuracy/MAE metrics are applied across all datasets and models.","core_discovery":"This study demonstrates that, for LLM-based psychiatric diagnosis in Arabic, how you prompt the model can shift performance as much as which model you choose. On multi-class datasets, a structured prompt with explicit formatting instructions outperformed a semantically identical but loosely structured 'act as a psychologist' prompt by 14.5 balanced-accuracy points; when only responses that both prompts could parse are compared, the gap collapses to 1.6 points, showing the difference is driven by instruction following, not clinical judgment. Model identity is the largest factor in diagnostic accuracy: Phi-3.5 MoE achieves the highest balanced accuracy, especially on binary tasks, while Mistral NeMo has the lowest mean absolute error on severity ratings. Language effects are modest, with English-native datasets beating their translated Arabic versions by about 6% on average, but translation quality and possible training-data leakage can create large artifacts, as seen in the 69% native-English advantage on MedMCQA. Few-shot prompting with one example per class improves performance consistently, with the largest gains on multi-class tasks.","pith_inferences":["The parsing-validity lens suggests that current Arabic mental-health LLM rankings partly measure output-formatting compliance; benchmarks that treat unparseable answers as errors rather than excluding them would narrow reported model gaps.","The modest language effect implies that well-translated English curricula could bootstrap Arabic psychiatric diagnostics, but the translation-quality bottleneck makes investment in native Arabic clinical corpora a higher-leverage intervention than further model tuning.","The MedMCQA native-English advantage is consistent with training-data memorization; a newly written Arabic clinical MCQ set that has never appeared online would separate true knowledge from leakage.","The AMI experience could be recycled as a cheap data-quality screen: if every tested model scores far below chance on a new dataset, suspect the labels before suspecting the models."],"forward_implications":["Structured, explicitly formatted prompts should be the default for Arabic mental-health LLM tasks; loosely worded 'act as a psychologist' prompts lose about 14.5 balanced-accuracy points on multi-class datasets, mostly by producing unparseable responses.","Model choice outweighs language choice: Phi-3.5 MoE is the strongest option for balanced accuracy, and Mistral NeMo is the best pick when severity rating error is the priority.","One-example-per-class few-shot prompting is a cheap, consistent improvement, about 20% for GPT-4o Mini overall and 57.6% on multi-class tasks, and should be adopted before considering fine-tuning.","English-to-Arabic machine translation costs only about 6% balanced accuracy on average, but translation quality and possible training-data leakage can create artifacts larger than the language effect itself.","Arabic mental-health benchmark labels need independent auditing before results are trusted; the excluded AMI dataset shows that faulty labels can push every model to chance performance."],"supporting_citations":[{"why":"Supplies the earlier English evaluation whose prompt designs and model shortlist this study adapts, and the observation that minor prompt tweaks cause unpredictable swings.","marker":"[18]"},{"why":"Provides the only other bilingual Arabic-English medical LLM evaluation and the expert-reviewed Arabic MedMCQA translation used to separate translation quality from language effects.","marker":"[20]"},{"why":"The DCAT dataset, the easiest Arabic depression binary corpus and an anchor for dataset-difficulty ranking.","marker":"[22]"},{"why":"ARADEPSU, the native Arabic multi-class dataset where the ZS-1 vs ZS-2 prompt gap is most visible.","marker":"[24]"},{"why":"AMI, the dataset excluded after all models scored at or below chance, demonstrating that label faults can invalidate evaluation sets.","marker":"[26]"},{"why":"DREADDIT, a translated English stress dataset used to measure cross-lingual performance.","marker":"[33]"},{"why":"SDCNL, a translated depression/suicide dataset whose depression split approached random-guessing accuracy, marking the difficult end of binary tasks.","marker":"[34]"},{"why":"DEPTWEET, a large severity dataset used in the binarized balanced-accuracy and MAE comparisons.","marker":"[36]"},{"why":"RED SAM, a severity dataset central to the MAE analyses in which Mistral NeMo led.","marker":"[38]"}],"fun_headline_variants":["Structured prompts beat loose ones by 14.5 points on Arabic mental-health","Prompt style shifts Arabic mental-health LLM performance as much as model choice","Few-shot examples boost Arabic mental-health LLM accuracy by up to 58%","Phi-3.5 tops balanced accuracy; Mistral NeMo best for severity scoring"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes the ground-truth labels of the Arabic and translated datasets are accurate; if hidden labeling faults like those found in the excluded AMI dataset exist elsewhere, the model rankings and prompt-effect sizes could shift.","fun_headline_variants_meta":{"raw":{"variants":["Structured prompts beat loose ones by 14.5 points on Arabic mental-health","Prompt style shifts Arabic mental-health LLM performance as much as model choice","Few-shot examples boost Arabic mental-health LLM accuracy by up to 58%","Phi-3.5 tops balanced accuracy; Mistral NeMo best for severity scoring"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000656,"raw_usage":{"total_tokens":3043,"prompt_tokens":1027,"completion_tokens":2016,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":1930}},"tokens_in":643,"tokens_out":2016,"duration_ms":15569,"temperature":1.0,"reasoning_tokens":1930,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:49:39.033338+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have expert clinicians re-annotate a random sample of ARADEPSU and MCD posts, then recompute the ZS-1 vs ZS-2 balanced-accuracy gap on the corrected labels; if the 14.5-point advantage disappears or reverses, the paper's headline prompt-engineering result is an artifact of label noise rather than instruction following.","supporting_citations":[{"cited_title":"A comprehen- sive evaluation of large language models on mental illnesses.arXiv preprint, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the earlier English evaluation whose prompt designs and model shortlist this study adapts, and the observation that minor prompt tweaks cause unpredictable swings."},{"cited_title":"Depression corpus of arabic tweets, 2022","cited_arxiv_id":null,"evidence_quote":"The DCAT dataset, the easiest Arabic depression binary corpus and an anchor for dataset-difficulty ranking."},{"cited_title":"Aradepsu: Detecting depression and suicidal ideation in arabic tweets using transformers","cited_arxiv_id":null,"evidence_quote":"ARADEPSU, the native Arabic multi-class dataset where the ZS-1 vs ZS-2 prompt gap is most visible."},{"cited_title":"Transfer learning-based automatic sentiment annotation of a twitter-based arabic mental illness (ami) dataset, 2023","cited_arxiv_id":null,"evidence_quote":"AMI, the dataset excluded after all models scored at or below chance, demonstrating that label faults can invalidate evaluation sets."},{"cited_title":"Deep learning for suicide and depression identification with unsupervised label correction","cited_arxiv_id":null,"evidence_quote":"SDCNL, a translated depression/suicide dataset whose depression split approached random-guessing accuracy, marking the difficult end of binary tasks."},{"cited_title":"Deptweet: A typology for social media texts to detect depression severities","cited_arxiv_id":null,"evidence_quote":"DEPTWEET, a large severity dataset used in the binarized balanced-accuracy and MAE comparisons."},{"cited_title":"No Disorder","cited_arxiv_id":null,"evidence_quote":"RED SAM, a severity dataset central to the MAE analyses in which Mistral NeMo led."}],"review_version":1}