Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Positioning AI Tools to Support Online Harm Reduction Practice: Applications and Design Directions

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that large language models, when guided by a harm-reduction prompt, can broaden access to drug-safety information, but only if future systems address ethical alignment, context, and clear limits.

desk verdict A genuinely under-explored topic with a thoughtful workshop study, but the Appendix's January 2025 're-engineered' example undercuts the August 2023 evidence base and needs fixing before the findings are treated as established. read the letter →

arxiv 2506.22941 v3 pith:VXF6OSST submitted 2025-06-28 cs.HC cs.AI

classification cs.HCcs.AI
keywords harmreductionlargelanguagemodelspeoplewhousedrugsstigmaco-designresponsibleAIonlinehealthinformationqualitativeworkshop
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that large language models could fill real gaps in online harm reduction information for people who use drugs, particularly by being instantly available, multilingual, and less judgmental than existing channels. It also argues that these benefits only hold if the systems are deliberately aligned with harm reduction ethics, can read unstated risk context, communicate briefly and clearly, and know when to stop and refer. To test this, the authors ran a qualitative workshop with practitioners, researchers, a forum moderator, and computer scientists, using a single LLM with a custom harm reduction prompt. If the findings hold, LLM tools are best positioned as complements to human and peer support rather than replacements.

What carries the argument

The load-bearing mechanism is a custom system-prompt template that encodes the core tenets of harm reduction, that drug use is a reality, that safety outranks abstinence, that responses must be non-judgmental, and that social contexts shape risk, and injects them into ChatGPT through custom instructions. The prompt is what turns generic, often abstinence-oriented or refusal-heavy outputs into actionable harm-reduction advice; the paper's findings about capability are findings about this specific prompt-model pairing. A second mechanism is the four-activity qualitative workshop, which uses real practitioner-derived queries with manipulated contexts of health, geography, and ethical boundary to elicit and evaluate model behaviour.

What would settle it

Run the same workshop queries, such as safe MDMA use in English and Chinese, fentanyl-testing advice in Scotland, and refusal of speedball instructions, on several current LLMs using the published prompt; if most models refuse or produce abstinence-only, judgmental, or inaccurate answers, or fail to match the multilingual and geographic adaptivity reported here, the central claim of LLM potential collapses.

Watch

Extended reading notes

Core claim

The central discovery is a conditional one: a single LLM, when given a carefully engineered prompt embedding harm reduction principles, can produce responsive, non-judgmental, multilingual, and context-sensitive answers to PWUD's safety questions, yet it cannot be trusted to probe unstated risk factors, stay current, cite sources, or manage crises. The paper claims that this mix of capability and limitation means LLMs should be designed as complementary front-line information tools, co-designed with experts and PWUD, with explicit operational boundaries and referral pathways.

Load-bearing premise

The findings rest on one chatbot model (ChatGPT/GPT-4, August 2023) paired with one custom prompt; if other LLMs do not behave like that combination, the claimed capabilities may not transfer to real systems or later models.

Editorial extensions

If this is right

  • LLM-based tools could serve as first-response information sources when human moderators or services are unavailable, covering common safety questions around dosing, adulterants, and drug interactions.
  • Multilingual responses from LLMs could extend harm reduction information to non-English-speaking PWUD communities currently underserved by existing resources.
  • Systems must be evaluated on harm-reduction-specific criteria, such as non-judgmental framing, handling of incomplete contexts, refusal quality, and source attribution, not just general language quality.
  • Future systems need live, curated knowledge bases and expert plus PWUD governance to avoid giving outdated or jurisdictionally wrong advice.
  • LLMs cannot replace peer support and should be positioned as complements with clear referral pathways to human services.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The prompt-engineering result suggests a testable extension: a benchmark of harm-reduction queries could grade LLMs on refusal rate, actionability, tone, and context sensitivity, giving programmes a way to compare systems over time.
  • The paper's finding that the model did not proactively ask about unstated risk factors points to a specific design target: conversational agents that ask one or two safety questions before answering, similar to human triage.
  • The same design logic likely extends to other stigmatised health information domains, such as sexual health, self-managed abortion, or mental health, where low-stigma access and non-judgmental tone are critical.
  • Because the evidence comes from a single model-prompt pair, the reported capabilities are a floor, not a ceiling; newer models with better instruction-following may exceed or fail these observed behaviours.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a qualitative workshop study exploring how large language models (LLMs) might be designed to support harm reduction information provision for people who use drugs (PWUD). Eleven participants—harm reduction practitioners, academics, an online community moderator, and computer scientists—took part in four activities: establishing shared technical understanding, identifying use cases through live demonstrations, probing contextual factors, and evaluating responses to derive design considerations. The paper's central claim is that LLMs can address some information barriers (responsiveness, multilingual access, reduced perceived stigma) but that their effectiveness depends on resolving challenges in ethics alignment, contextual understanding, communication, and operational boundaries. It contributes participant-derived design pathways, including co-design with experts and PWUD, transparent source attribution, and retrieval-augmented knowledge grounding.

Significance. The topic is timely and socially important, and the paper is honest about its exploratory scope. Its strongest contribution is the set of participant-derived design considerations for a high-stakes, underserved domain, informed by a stakeholder group that includes a moderator of a large PWUD community and front-line practitioners. The paper also ships useful artifacts: a detailed harm-reduction-aligned prompt template in the appendix and several response examples that can be inspected. If the evidence provenance and methodological reporting issues are resolved, the design recommendations would be a useful empirical starting point for future work on LLM-based harm reduction tools. However, the current manuscript does not yet fully substantiate the empirical basis for its central capability claims.

major comments (3)
  1. [§4.3 and Appendix A] The contextual-adaptation finding—that adding a heart-disease history to an MDMA dosage query made ChatGPT shift to cardiac risk information—is presented in §4.3 as workshop evidence, but the appendix version is headed 'Example of ChatGPT’s responses adapted to specific health conditions (re-engineered in January 2025)'. Since §3.1 states that 'ChatGPT GPT-4, August 2023' was used for all workshop activities, this indicates the example was produced after the workshop, presumably on a later model version, and the main text does not disclose the mismatch. The general caveat in §6 that model versions limit generalisability does not address an example presented as an in-workshop observation. The authors should state which examples in §4.2 and §4.3 are authentic August 2023 outputs, and either remove or explicitly label and relegate the January 2025 reconstruction to supplementary status. This is required before the contextual-adaptation claim can be treated as empirically established.
  2. [§3.2 and §4] The paper does not report a formal qualitative analysis procedure. The methods describe a 'qualitative descriptive design' and a structured workshop, but there is no account of how the workshop documentation was coded, how themes were extracted, or how the quotes and paraphrases selected in §4 relate to the full corpus. Without such an analysis protocol, the reader cannot distinguish systematic thematic findings from illustrative anecdotes. A subsection describing the analytic procedure—including data sources, coding steps, and any triangulation or reliability measures—should be added.
  3. [§3.1, Appendix A, and §4.2] Because the demonstrated outputs were generated with a custom prompt that explicitly encodes harm reduction principles, the observation that the LLM produced harm-reduction-aligned, multilingual, and non-judgmental responses is partly built into the test setup. This does not invalidate the design considerations, which came from participants, but the paper should frame these demonstrations as evidence about a designed system (model plus custom instructions) rather than about unmodified LLMs. The contrast in Fig. 1 already makes this point, but the text of §4.2 should consistently attribute observed capabilities to the prompted system and clarify that the findings are not evidence about out-of-the-box LLM behaviour.
minor comments (5)
  1. [§3.1] The 'Discrepancy in System Capability' material appears under the same numbered subsection heading as 'Prompt Engineering', which is confusing; these should be separate subsections or clearly distinguished in the heading hierarchy.
  2. [§4.1 and §4.4] There are grammatical slips that should be corrected, including 'a imbalance' (§4.1) and 'a evaluation of LLM-generated responses' (§4.4).
  3. [Title page and running footer] The ACM reference block includes 'Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009', and the article is dated August 2023 despite the arXiv version being July 2025; these template artifacts should be corrected.
  4. [Fig. 3] The figure caption and layout should clarify which panel is the user query and which is the LLM response, since the current display shows the question text and then a long answer with no clear separation between the two.
  5. [Appendix A] The appendix examples should each carry a label stating the model version and date of generation, matching the provenance disclosure requested for §4.3, so that readers can tell which examples are workshop outputs from August 2023 and which are later reconstructions.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; prompt-engineering effects are disclosed and the design considerations are participant-derived.

full rationale

No significant circularity. The paper's empirical chain is a qualitative workshop with diverse stakeholders, participant-proposed harm reduction queries, ChatGPT responses generated with a disclosed harm-reduction prompt, and participant evaluation leading to design considerations. The only step that resembles an input-output equivalence is the prompt template: Section 3.1 states the template "integrated the key harm reduction tenets, such as acknowledging drug use as a reality, respecting individual autonomy, prioritising users' safety over abstinence, using non-judgmental language," and Section 4.2 then observes the LLM answered "without explicit moral judgment." This is a built-in effect, but the paper does not disguise it: it explicitly frames the prompt as an engineering intervention, contrasts it with un-prompted abstinence-oriented responses, and repeatedly qualifies the observed behavior as occurring "when appropriately prompted." The core contribution -- participant-derived design considerations about ethical alignment, contextual reasoning, communication, operational boundaries, and governance -- comes from independent stakeholder discussion in Activities 3-4 and Section 5, not from the prompt construction. There is no load-bearing self-citation, no imported uniqueness theorem, and no ansatz smuggled in via citation; the references are external literature. One correctness and evidence-provenance concern should be noted separately from circularity: the Section 4.3 MDMA/heart-disease example is labeled in the Appendix as "re-engineered in January 2025," while the method states ChatGPT GPT-4 August 2023 was used for all workshop activities. This is an internal inconsistency that the Limitations section (Sec. 6) does not address, and it may weaken the empirical basis of the contextual-adaptation claim. However, that issue concerns evidence quality and reporting, not circularity, because the claim does not reduce to the constructed example by definition.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities apply to this qualitative study. The core input assumptions are the choice of harm reduction principles as the evaluative framework, the representativeness of the expert sample, and the representativeness of a single LLM.

assumptions (3)
  • domain assumption Harm reduction principles as published by the National Harm Reduction Coalition are the correct normative framework for evaluating LLM outputs.
    The prompt template in Section 3.1 is built on these principles, and participant evaluation in Activity 4 uses them as the yardstick.
  • domain assumption A single qualitative workshop with 11 expert stakeholders yields transferable design considerations.
    The paper generalizes from a small, UK-centric sample and acknowledges this in the Limitations section.
  • domain assumption ChatGPT GPT-4 (August 2023) is a representative LLM for probing LLM capabilities in this domain.
    The study standardizes on one model to avoid inter-system confounds, then draws conclusions about LLMs broadly.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Positioning AI Tools to Support Online Harm Reduction Practice: Applications and Design Directions." pith.science (2026). https://pith.science/paper/VXF6OSST

@misc{pith2026250622941,
  author       = {Pith},
  title        = {Pith review of: Positioning AI Tools to Support Online Harm Reduction Practice: Applications and Design Directions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VXF6OSST}},
  note         = {Machine review of arXiv:2506.22941}
}
read the original abstract

Access to accurate and actionable harm reduction information can directly impact the health outcomes of People Who Use Drugs (PWUD), yet existing online channels often fail to meet their diverse and dynamic needs due to limitations in adaptability, accessibility, and the pervasive impact of stigma. Large Language Models (LLMs) present a novel opportunity to enhance information provision, but their application in such a high-stakes domain is under-explored and presents socio-technical challenges. This paper investigates how LLMs can be responsibly designed to support the information needs of PWUD. Through a qualitative workshop involving diverse stakeholder groups (academics, harm reduction practitioners, and an online community moderator), we explored LLM capabilities, identified potential use cases, and delineated core design considerations. Our findings reveal that while LLMs can address some existing information barriers (e.g., by offering responsive, multilingual, and potentially less stigmatising interactions), their effectiveness is contingent upon overcoming challenges related to ethical alignment with harm reduction principles, nuanced contextual understanding, effective communication, and clearly defined operational boundaries. We articulate design pathways emphasising collaborative co-design with experts and PWUD to develop LLM systems that are helpful, safe, and responsibly governed. This work contributes empirically grounded insights and actionable design considerations for the responsible development of LLMs as supportive tools within the harm reduction ecosystem.

Figures

Figures reproduced from arXiv: 2506.22941 by the authors.

Figure 1
Figure 1. A comparison of ChatGPT’s responses to an inquiry about MDMA use. The generic response, from a simple query, [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. A comparison of responses to a query about marijuana use and uncharacteristic aggression. The response on the right, [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. The LLM’s response to a query about experiencing atypical fatigue after cocaine use. The model provides potential [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: The LLM’s response to a query about fentanyl testing, localised to Scotland. The model correctly identifies and [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

    cs.CL 2025-07 conditional novelty 6.0 of 10

    State-of-the-art LLMs are frequently inaccurate, and sometimes dangerous, when answering harm reduction questions about drug use, even when given retrieved source material.

Reference graph

Works this paper leans on

64 extracted references · 51 canonical work pages · cited by 1 Pith paper

  1. [1]

    Khalid AJ Al Khaja, Alwaleed K AlKhaja, and Reginald P Sequeira. 2018. Drug information, misinformation, and disinformation on social media: a content analysis study. Journal of public health policy 39 (2018), 343–357

  2. [2]

    Hilal Al Shamsi, Abdullah G Almutairi, Sulaiman Al Mashrafi, and Talib Al Kalbani. 2020. Implications of language barriers for healthcare: a systematic review. Oman medical journal 35, 2 (2020), e122

  3. [3]

    Ahmed Shihab Albahri, Ali M Duhaim, Mohammed A Fadhel, Alhamzah Alnoor, Noor S Baqer, Laith Alzubaidi, Osamah Shihab Albahri, Abdullah Hussein Alamoodi, Jinshuai Bai, Asma Salhi, et al. 2023. A systematic review of trustworthy and explainable artificial intelligence in healthcare: Assessment of quality, bias risk, and data fusion. Information Fusion 96 (2...

  4. [4]

    Katherine A Allen, Victoria Charpentier, Marissa A Hendrickson, Molly Kessler, Rachael Gotlieb, Jordan Marmet, Emily Hause, Corinne Praska, Scott Lunos, and Michael B Pitt. 2023. Jargon be gone–patient preference in doctor communication. Journal of patient experience 10 (2023), 23743735231158942

  5. [5]

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency . 610–623

  6. [6]

    Carmel Bradshaw, Sandra Atkinson, and Owen Doody. 2017. Employing a qualitative description approach in health care research. Global qualitative nursing research 4 (2017), 2333393617742282

  7. [7]

    Janet N Chu, Urmimala Sarkar, Natalie A Rivadeneira, Robert A Hiatt, and Elaine C Khoong. 2022. Impact of language preference and health literacy on health information-seeking experiences among a low-income, multilingual cohort. Patient Education and Counseling 105, 5 (2022), 1268–1275

  8. [8]

    Kasey R Claborn, Suzannah Creech, Quanisha Whittfield, Ruben Parra-Cardona, Andrea Daugherty, and Justin Benzer. 2022. Ethical by design: engaging the community to co-design a digital health ecosystem to improve overdose prevention efforts among highly vulnerable people who use drugs. Frontiers in Digital Health 4 (2022), 880849

Show all 64 references
  1. [9]

    Karen Jiggins Colorafi and Bronwynne Evans. 2016. Qualitative descriptive methods in health science research. HERD: Health Environments Research & Design Journal 9, 4 (2016), 16–25

  2. [10]

    Arsen Davitadze, Peter Meylakhs, Aleksey Lakhov, and Elizabeth J King. 2020. Harm reduction via online platforms for people who use drugs in Russia: a qualitative analysis of web outreach work. Harm Reduction Journal 17 (2020), 1–9

  3. [11]

    Luigi De Angelis, Francesco Baglivo, Guglielmo Arzilli, Gaetano Pierpaolo Privitera, Paolo Ferragina, Alberto Eugenio Tozzi, and Caterina Rizzo. 2023. ChatGPT and the rise of large language models: the new AI-driven infodemic threat in public health. Frontiers in public health...

  4. [12]

    Louise Doyle, Catherine McCabe, Brian Keogh, Annemarie Brady, and Margaret McCann. 2020. An overview of the qualitative descriptive design within nursing research. Journal of research in nursing 25, 5 (2020), 443–455

  5. [13]

    Samer El Hayek, Wael Foad, Renato de Filippis, Abhishek Ghosh, Nadine Koukach, Aala Mahgoub Mohammed Khier, Sagun Ballav Pant, Vanessa Padilla, Rodrigo Ramalho, Hossameldin Tolba, et al. 2024. Stigma toward substance use disorders: a multinational perspective and call for acti...

  6. [14]

    Batya Friedman, Peter H Kahn, Alan Borning, and Alina Huldtgren. 2013. Value sensitive design and information systems. Early engagement and new technologies: Opening up the laboratory (2013), 55–95

  7. [15]

    Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024. Bias and fairness in large language models: A survey. Computational Linguistics 50, 3 (2024), 1097–1179

  8. [16]

    Ariana Genovese, Sahar Borna, Cesar A Gomez-Cabello, Syed Ali Haider, Srinivasagam Prabha, Antonio J Forte, and Benjamin R Veenstra. 2024. Artificial intelligence in clinical settings: a systematic review of its role in language translation and interpretation. Annals of Transl...

  9. [17]

    Tarleton Gillespie. 2018. Custodians of the Internet: Platforms, content moderation, and the hidden decisions that shape social media . Yale University Press

  10. [18]

    André Belchior Gomes and Aysel Sultan. 2024. Problematizing content moderation by social media platforms and its impact on digital harm reduction. Harm Reduction Journal 21, 1 (2024), 194. Proc. ACM Meas. Anal. Comput. Syst., Vol. 37, No. 4, Article 111. Publication date: Augu...

  11. [19]

    Chloe Grace Rose, Victoria Kulbokas, Emir Carkovic, Todd A Lee, and A Simon Pickard. 2023. Contextual factors affecting the implementation of drug checking for harm reduction: a scoping literature review from a North American perspective. Harm Reduction Journal 20, 1 (2023), 124

  12. [20]

    Chloe Grace Rose, A Simon Pickard, Victoria Kulbokas, Stacey Hoferka, Kaitlyn Friedman, Jennifer Epstein, and Todd A Lee. 2023. A qualitative assessment of key considerations for drug checking service implementation. Harm Reduction Journal 20, 1 (2023), 151

  13. [21]

    Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L Isbell, and Andrea L Thomaz. 2013. Policy shaping: Integrating human feedback with reinforcement learning. Advances in neural information processing systems 26 (2013)

  14. [22]

    Mary Hawk, Robert WS Coulter, James E Egan, Stuart Fisk, M Reuel Friedman, Monique Tula, and Suzanne Kinsky. 2017. Harm reduction principles for healthcare settings. Harm reduction journal 14 (2017), 1–9

  15. [23]

    Jiayu He, Ying Wang, Zhicheng Du, Jing Liao, Na He, and Yuantao Hao. 2020. Peer education for HIV prevention among high-risk groups: a systematic review and meta-analysis. BMC infectious diseases 20 (2020), 1–20

  16. [24]

    Dagmar Hedrich and Richard Lionel Hartnoll. 2021. Harm-reduction interventions. Textbook of addiction treatment: international perspectives (2021), 757–775

  17. [25]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2025. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Info...

  18. [26]

    Bogeum Kim, Hyejin Sim, Jiwon Yun, Jaehan Cho, and Howon Kim. 2024. LLM Guardrail Framework: A Novel Approach for Implementing Zero Trust Architecture. In International Conference on Information Security Applications . Springer, 138–148

  19. [27]

    Margaret E Kruk, Anna D Gage, Catherine Arsenault, Keely Jordan, Hannah H Leslie, Sanam Roder-DeWan, Olusoji Adeyi, Pierre Barker, Bernadette Daelmans, Svetlana V Doubova, et al. 2018. High-quality health systems in the Sustainable Development Goals era: time for a revolution....

  20. [28]

    Yi-Chieh Lee, Yichao Cui, Jack Jamieson, Wayne Fu, and Naomi Yamashita. 2023. Exploring effects of chatbot-based social contact on reducing mental illness stigma. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–16

  21. [29]

    Simon Lenton and Eric Single. 1998. The definition of harm reduction. Drug and alcohol review 17, 2 (1998), 213–219

  22. [30]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing s...

  23. [31]

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM computing surveys 55, 9 (2023), 1–35

  24. [32]

    G Alan Marlatt. 1996. Harm reduction: Come as you are. Addictive behaviors 21, 6 (1996), 779–788

  25. [33]

    David N Milne, Kathryn L McCabe, and Rafael A Calvo. 2019. Improving moderator responsiveness in online peer support through automated triage. Journal of medical Internet research 21, 4 (2019), e11410

  26. [34]

    Levente Móró and József Rácz. 2013. Online drug user-led harm reduction in Hungary: a review of “Daath”. Harm Reduction Journal 10 (2013), 1–11

  27. [35]

    Sarah Myers West. 2018. Censored, suspended, shadowbanned: User interpretations of content moderation on social media platforms. New Media & Society 20, 11 (2018), 4366–4383

  28. [36]

    Mette Asbjoern Neergaard, Frede Olesen, Rikke Sand Andersen, and Jens Sondergaard. 2009. Qualitative description–the poor cousin of health research? BMC medical research methodology 9 (2009), 1–5

  29. [37]

    Karen Ka Yan Ng, Izuki Matsuba, and Peter Chengming Zhang. 2025. RAG in Health Care: A Novel Framework for Improving Communication and Decision-Making by Addressing LLM Limitations. NEJM AI 2, 1 (2025), AIra2400380

  30. [38]

    Lena Katharina Oeltjen, Maike Schulz, Imke Heuer, Georg Knigge, Rebecca Nixdorf, Denis Briel, Patricia Hamer, Werner Brannath, Jörg Utschakowski, Candelaria Mahlke, et al. 2024. Effectiveness of a peer-supported crisis intervention to reduce the proportion of compulsory admiss...

  31. [39]

    Jianing Qiu, Kyle Lam, Guohao Li, Amish Acharya, Tien Yin Wong, Ara Darzi, Wu Yuan, and Eric J Topol. 2024. LLM-based agentic systems in medicine and healthcare. Nature Machine Intelligence 6, 12 (2024), 1418–1420

  32. [40]

    Sandeep Reddy. 2023. Evaluating large language models for use in healthcare: A framework for translational value assessment.Informatics in Medicine Unlocked 41 (2023), 101304

  33. [41]

    Alison Ritter and Jacqui Cameron. 2006. A review of the efficacy and effectiveness of harm reduction strategies for alcohol, tobacco and illicit drugs. Drug and alcohol review 25, 6 (2006), 611–624

  34. [42]

    Sara Rolando, Giulia Arrighetti, Elisa Fornero, Ombretta Farucci, and Franca Beccaria. 2023. Telegram as a space for peer-led harm reduction communities and Netreach interventions. Contemporary Drug Problems 50, 2 (2023), 190–201

  35. [43]

    Rikard Rosenbacke, Åsa Melhus, Martin McKee, and David Stuckler. 2024. How Explainable Artificial Intelligence Can Increase or Decrease Clinicians’ Trust in AI Applications in Health Care: Systematic Review. JMIR AI 3 (2024), e53207

  36. [44]

    Margarete Sandelowski. 2000. Whatever happened to qualitative description? Research in nursing & health 23, 4 (2000), 334–340. Proc. ACM Meas. Anal. Comput. Syst., Vol. 37, No. 4, Article 111. Publication date: August 2023. Positioning AI Tools to Support Online Harm Reduction...

  37. [45]

    Thomas Savage, Ashwin Nayak, Robert Gallo, Ekanath Rangan, and Jonathan H Chen. 2024. Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine. NPJ Digital Medicine 7, 1 (2024), 20

  38. [46]

    Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. 2024. Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems 36 (2024)

  39. [47]

    Frank J Schwebel and Daniel G Orban. 2023. Online support for all: Examining participant characteristics, engagement, and perceived benefits of an online harm reduction, abstinence, and moderation focused support group for alcohol and other drugs. Psychology of Addictive Behav...

  40. [48]

    The human body is a black box

    Mark Sendak, Madeleine Clare Elish, Michael Gao, Joseph Futoma, William Ratliff, Marshall Nichols, Armando Bedoya, Suresh Balu, and Cara O’Brien. 2020. " The human body is a black box" supporting clinical decision-making with deep learning. In Proceedings of the 2020 conferenc...

  41. [49]

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023. Large language models encode clinical knowledge. Nature 620, 7972 (2023), 172–180

  42. [50]

    Mingyang Song and Mao Zheng. 2024. A Survey of Query Optimization in Large Language Models. arXiv preprint arXiv:2412.17558 (2024)

  43. [51]

    Alex Stevens. 2020. Critical realism and the ‘ontological politics of drug policy’. International Journal of Drug Policy 84 (2020), 102723

  44. [52]

    Waleed M Sweileh. 2024. Analysis and mapping of harm reduction research in the context of injectable drug use: identifying research hotspots, gaps and future directions. Harm Reduction Journal 21, 1 (2024), 131

  45. [53]

    Boden Tighe, Matthew Dunn, Fiona H McKay, and Timothy Piatkowski. 2017. Information sought, information shared: exploring performance and image enhancing drug user-facilitated harm reduction information in online forums. Harm reduction journal 14 (2017), 1–9

  46. [54]

    Babak Tofighi, Chemi Chemi, Jose Ruiz-Valcarcel, Paul Hein, Lu Hu, et al. 2019. Smartphone apps targeting alcohol and illicit substance use: systematic search in in commercial app stores and critical content analysis. JMIR mHealth and uHealth 7, 4 (2019), e11831

  47. [55]

    Tao Tu, Mike Schaekermann, Anil Palepu, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Yong Cheng, et al. 2025. Towards conversational diagnostic artificial intelligence. Nature (2025), 1–9

  48. [56]

    Roxanne Turuba, Christina Katan, Kirsten Marchand, Chantal Brasset, Alayna Ewert, Corinne Tallon, Jill Fairbank, Steve Mathias, and Skye Barbic. 2024. Weaving community-based participatory research and co-design to improve opioid use treatments and services for youth, caregive...

  49. [57]

    United Nations Office on Drugs and Crime. 2023. World Drug Report 2023: Executive Summary. https://www.unodc.org/res/WDR- 2023/WDR23_Exsum_fin_SP.pdf Accessed: 2025-01-25

  50. [58]

    Karen Urbanoski, Bernadette Pauly, Dakota Inglis, Fred Cameron, Troy Haddad, Jack Phillips, Paige Phillips, Conor Rosen, Grant Schlotter, Elizabeth Hartney, et al. 2020. Defining culturally safe primary care for people who use substances: a participatory concept mapping study....

  51. [59]

    Ibo Van de Poel. 2020. Embedding values in artificial intelligence (AI) systems. Minds and machines 30, 3 (2020), 385–409

  52. [60]

    A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017)

  53. [61]

    Li Wang, Xi Chen, XiangWen Deng, Hao Wen, MingKe You, WeiZhi Liu, Qi Li, and Jian Li. 2024. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. NPJ digital medicine 7, 1 (2024), 41

  54. [62]

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682 (2022)

  55. [63]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  56. [64]

    indigenous knowledge

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023). Proc. ACM Meas. Anal. Comput. Syst., Vol. 37, No. 4, Articl...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.