{"id":"5efd80c0-338c-4daa-be01-8281b43efa52","arxiv_id":"2506.22941","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A stakeholder workshop identifies opportunities and design requirements for using LLMs to deliver harm reduction information to people who use drugs.","lead":"This paper reports a workshop where harm reduction experts tested ChatGPT's responses to drug safety questions and sketched design guidelines for safer AI support. It maps where LLMs could help people who use drugs find non-judgmental, multilingual information, and where human professionals remain irreplaceable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Key Appendix example is labeled 're-engineered in January 2025' while the method says ChatGPT GPT-4 August 2023; the Section 4.3 contextual-adaptation finding may rest on a different model, leaving the empirical basis of the central claim unverified.","rationale":"The reader's weakest assumption is that a single model—ChatGPT GPT-4, August 2023—cannot bear the weight of general claims about LLM capabilities. My concern is more specific and more concrete: an appendix example that directly supports the contextual-adaptation finding is labeled as re-engineered in January 2025, seventeen months after the stated model date. That creates an internal inconsistency in the empirical evidence, not merely a generalizability caveat. The paper already acknowledges model-version limitations in Section 6, which is a point in its favor, but it does not disclose that a displayed output was generated outside the workshop. If the provenance issue is resolved by raw logs, the paper's design directions remain plausible and the conditional verdict stands. If not, the specific claim in Sec. 4.3 loses support. I am not alleging fraud; post-hoc reconstruction for clearer figures is understandable. The problem is the absence of disclosure and the centrality of that example to the claim that LLMs adapt to explicitly stated health contexts. Other aspects of the workshop—diverse stakeholders, iterative prompt engineering, and explicit design considerations—are genuine strengths, but they do not substitute for verifiable evidence provenance. I therefore keep the reader's conditional verdict unchanged, with the added condition that the authors either release the workshop record or clearly mark and revalidate the re-engineered examples.","tokens_in":22617,"tokens_out":6011,"duration_ms":72866,"concrete_test":"Ask the authors to release the Activity 3 raw data: the participant-generated queries, the verbatim ChatGPT outputs captured during the workshop, and timestamps. Then verify whether the MDMA/heart-disease exchange in Appendix A appears in that raw record. If it does not, re-run that query on a ChatGPT GPT-4 August 2023 snapshot with the exact Appendix A prompt. If the reported contextual shift to cardiac-risk advice does not reproduce, Section 4.3's adaptation finding is unsupported for the stated testbed, and the central claim should be narrowed to later model versions or reworded to rely only on workshop-generated examples.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that LLMs can help overcome information barriers while remaining contingent on design challenges rests on workshop observations of a single system: ChatGPT GPT-4, August 2023 (Sec. 3.1). One of the most consequential demonstrations is in Sec. 4.3, where modifying an MDMA-dosage query to include a history of heart disease reportedly made 'ChatGPT shift from dosage information to explaining cardiac risks and recommending against use.' That exchange is presented as workshop evidence, but the Appendix version is headed 'Example of ChatGPT’s responses adapted to specific health conditions (re-engineered in January 2025).' This label reveals that at least this example was produced after the workshop, presumably on a later model version. The main text does not disclose this temporal mismatch. The paper's own Section 6 acknowledges that findings may not generalize across model versions, but it does not say that a displayed example was generated later. If the contextual-adaptation finding depends on a post-hoc reconstruction, then the central claim's empirical support is weaker than presented, and readers cannot tell which other examples in Sec. 4.2 and Sec. 4.3 are authentic workshop outputs from the stated system. This is not a claim of misconduct; it is an internal inconsistency in evidence provenance that should be resolved before the findings are treated as established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative workshop study exploring how large language models (LLMs) might be designed to support harm reduction information provision for people who use drugs (PWUD). Eleven participants—harm reduction practitioners, academics, an online community moderator, and computer scientists—took part in four activities: establishing shared technical understanding, identifying use cases through live demonstrations, probing contextual factors, and evaluating responses to derive design considerations. The paper's central claim is that LLMs can address some information barriers (responsiveness, multilingual access, reduced perceived stigma) but that their effectiveness depends on resolving challenges in ethics alignment, contextual understanding, communication, and operational boundaries. It contributes participant-derived design pathways, including co-design with experts and PWUD, transparent source attribution, and retrieval-augmented knowledge grounding.","tokens_in":22874,"tokens_out":3567,"duration_ms":40202,"significance":"The topic is timely and socially important, and the paper is honest about its exploratory scope. Its strongest contribution is the set of participant-derived design considerations for a high-stakes, underserved domain, informed by a stakeholder group that includes a moderator of a large PWUD community and front-line practitioners. The paper also ships useful artifacts: a detailed harm-reduction-aligned prompt template in the appendix and several response examples that can be inspected. If the evidence provenance and methodological reporting issues are resolved, the design recommendations would be a useful empirical starting point for future work on LLM-based harm reduction tools. However, the current manuscript does not yet fully substantiate the empirical basis for its central capability claims.","major_comments":[{"comment":"The contextual-adaptation finding—that adding a heart-disease history to an MDMA dosage query made ChatGPT shift to cardiac risk information—is presented in §4.3 as workshop evidence, but the appendix version is headed 'Example of ChatGPT’s responses adapted to specific health conditions (re-engineered in January 2025)'. Since §3.1 states that 'ChatGPT GPT-4, August 2023' was used for all workshop activities, this indicates the example was produced after the workshop, presumably on a later model version, and the main text does not disclose the mismatch. The general caveat in §6 that model versions limit generalisability does not address an example presented as an in-workshop observation. The authors should state which examples in §4.2 and §4.3 are authentic August 2023 outputs, and either remove or explicitly label and relegate the January 2025 reconstruction to supplementary status. This is required before the contextual-adaptation claim can be treated as empirically established.","section":"§4.3 and Appendix A"},{"comment":"The paper does not report a formal qualitative analysis procedure. The methods describe a 'qualitative descriptive design' and a structured workshop, but there is no account of how the workshop documentation was coded, how themes were extracted, or how the quotes and paraphrases selected in §4 relate to the full corpus. Without such an analysis protocol, the reader cannot distinguish systematic thematic findings from illustrative anecdotes. A subsection describing the analytic procedure—including data sources, coding steps, and any triangulation or reliability measures—should be added.","section":"§3.2 and §4"},{"comment":"Because the demonstrated outputs were generated with a custom prompt that explicitly encodes harm reduction principles, the observation that the LLM produced harm-reduction-aligned, multilingual, and non-judgmental responses is partly built into the test setup. This does not invalidate the design considerations, which came from participants, but the paper should frame these demonstrations as evidence about a designed system (model plus custom instructions) rather than about unmodified LLMs. The contrast in Fig. 1 already makes this point, but the text of §4.2 should consistently attribute observed capabilities to the prompted system and clarify that the findings are not evidence about out-of-the-box LLM behaviour.","section":"§3.1, Appendix A, and §4.2"}],"minor_comments":[{"comment":"The 'Discrepancy in System Capability' material appears under the same numbered subsection heading as 'Prompt Engineering', which is confusing; these should be separate subsections or clearly distinguished in the heading hierarchy.","section":"§3.1"},{"comment":"There are grammatical slips that should be corrected, including 'a imbalance' (§4.1) and 'a evaluation of LLM-generated responses' (§4.4).","section":"§4.1 and §4.4"},{"comment":"The ACM reference block includes 'Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009', and the article is dated August 2023 despite the arXiv version being July 2025; these template artifacts should be corrected.","section":"Title page and running footer"},{"comment":"The figure caption and layout should clarify which panel is the user query and which is the LLM response, since the current display shows the question text and then a long answer with no clear separation between the two.","section":"Fig. 3"},{"comment":"The appendix examples should each carry a label stating the model version and date of generation, matching the provenance disclosure requested for §4.3, so that readers can tell which examples are workshop outputs from August 2023 and which are later reconstructions.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper's empirical core is a small workshop, and the strongest value is in the design directions rather than in the specific capability demonstrations. The provenance mismatch for the January 2025 example is the most serious issue, but it appears fixable by honest labeling and by presenting that example as supplementary rather than as workshop data. The absence of an analysis procedure is a standard qualitative-methods concern that should be surfaced prominently to the authors. I would also gently check whether the ACM template artifacts (received dates from 2007 and 2009) indicate a formatting oversight rather than a substantive problem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, honest qualitative HCI paper on a genuinely under-studied topic — using LLMs to provide harm reduction information to people who use drugs. The contribution is real: eleven stakeholders (practitioners, drug-policy researchers, an r/Drugs moderator, and LLM folks) in a structured workshop, with real-time demos, producing domain-specific design principles. The design considerations — align with harm reduction ethics, handle implicit risk contexts, define operational limits, keep information current and attributed — are sensible and grounded in the participants' observations, not just generic responsible-AI talk. The paper is well-written and appropriately modest about scope.\n\nThe main soft spot is a real one: the Appendix labels the MDMA/heart-disease example \"re-engineered in January 2025\" while the methods say all workshop outputs came from ChatGPT GPT-4 in August 2023. That example is the key evidence in Section 4.3 for contextual adaptation — the claim that the model shifted from dosage information to cardiac risk. If it was generated later on a different model, the central empirical support is weaker than presented, and the reader cannot tell which other examples are authentic. The paper's own limitation section acknowledges single-model generalizability limits, but it does not disclose this provenance gap. This is an internal inconsistency, not misconduct, and it should be fixable — re-run the example on the stated model version or clearly report the version and date.\n\nOther soft spots are milder. There is no described analysis protocol — no coding or thematic extraction procedure — so the findings appear as themes without showing how they were derived. The circularity issue is real but partial: the custom prompt encodes harm reduction principles, so observing aligned responses is partly built into the test setup; the design considerations come from participants and do not depend on that. Single-model use is acknowledged and defensible as a way to avoid inter-system confounds.\n\nWho this is for: HCI researchers working on responsible AI in public health, harm reduction practitioners interested in AI tools, and anyone studying qualitative evidence in LLM-era studies. With the provenance issue fixed, it deserves publication; even as-is, it merits serious review rather than desk rejection.","headline":"A genuinely under-explored topic with a thoughtful workshop study, but the Appendix's January 2025 're-engineered' example undercuts the August 2023 evidence base and needs fixing before the findings are treated as established.","tokens_in":23334,"tokens_out":1800,"would_cite":false,"duration_ms":18680,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that large language models, when guided by a harm-reduction prompt, can broaden access to drug-safety information, but only if future systems address ethical alignment, context, and clear limits.","keywords":["harm reduction","large language models","people who use drugs","stigma","co-design","responsible AI","online health information","qualitative workshop"],"falsifier":"Run the same workshop queries, such as safe MDMA use in English and Chinese, fentanyl-testing advice in Scotland, and refusal of speedball instructions, on several current LLMs using the published prompt; if most models refuse or produce abstinence-only, judgmental, or inaccurate answers, or fail to match the multilingual and geographic adaptivity reported here, the central claim of LLM potential collapses.","tokens_in":22442,"feed_emoji":"💬","tokens_out":4235,"duration_ms":44834,"temperature":0.7,"pith_summary":"The paper argues that large language models could fill real gaps in online harm reduction information for people who use drugs, particularly by being instantly available, multilingual, and less judgmental than existing channels. It also argues that these benefits only hold if the systems are deliberately aligned with harm reduction ethics, can read unstated risk context, communicate briefly and clearly, and know when to stop and refer. To test this, the authors ran a qualitative workshop with practitioners, researchers, a forum moderator, and computer scientists, using a single LLM with a custom harm reduction prompt. If the findings hold, LLM tools are best positioned as complements to human and peer support rather than replacements.","feed_headline":"AI chatbots can expand drug-safety access if designed for it","feed_subtitle":"Workshop evidence shows LLMs answer multilingual, stigma-free queries, but need ethical alignment and clear limits to be safe.","key_machinery":"The load-bearing mechanism is a custom system-prompt template that encodes the core tenets of harm reduction, that drug use is a reality, that safety outranks abstinence, that responses must be non-judgmental, and that social contexts shape risk, and injects them into ChatGPT through custom instructions. The prompt is what turns generic, often abstinence-oriented or refusal-heavy outputs into actionable harm-reduction advice; the paper's findings about capability are findings about this specific prompt-model pairing. A second mechanism is the four-activity qualitative workshop, which uses real practitioner-derived queries with manipulated contexts of health, geography, and ethical boundary to elicit and evaluate model behaviour.","core_discovery":"The central discovery is a conditional one: a single LLM, when given a carefully engineered prompt embedding harm reduction principles, can produce responsive, non-judgmental, multilingual, and context-sensitive answers to PWUD's safety questions, yet it cannot be trusted to probe unstated risk factors, stay current, cite sources, or manage crises. The paper claims that this mix of capability and limitation means LLMs should be designed as complementary front-line information tools, co-designed with experts and PWUD, with explicit operational boundaries and referral pathways.","pith_inferences":["The prompt-engineering result suggests a testable extension: a benchmark of harm-reduction queries could grade LLMs on refusal rate, actionability, tone, and context sensitivity, giving programmes a way to compare systems over time.","The paper's finding that the model did not proactively ask about unstated risk factors points to a specific design target: conversational agents that ask one or two safety questions before answering, similar to human triage.","The same design logic likely extends to other stigmatised health information domains, such as sexual health, self-managed abortion, or mental health, where low-stigma access and non-judgmental tone are critical.","Because the evidence comes from a single model-prompt pair, the reported capabilities are a floor, not a ceiling; newer models with better instruction-following may exceed or fail these observed behaviours."],"forward_implications":["LLM-based tools could serve as first-response information sources when human moderators or services are unavailable, covering common safety questions around dosing, adulterants, and drug interactions.","Multilingual responses from LLMs could extend harm reduction information to non-English-speaking PWUD communities currently underserved by existing resources.","Systems must be evaluated on harm-reduction-specific criteria, such as non-judgmental framing, handling of incomplete contexts, refusal quality, and source attribution, not just general language quality.","Future systems need live, curated knowledge bases and expert plus PWUD governance to avoid giving outdated or jurisdictionally wrong advice.","LLMs cannot replace peer support and should be positioned as complements with clear referral pathways to human services."],"supporting_citations":[{"why":"Supplies the definition and philosophy of harm reduction that the custom prompt encodes and that workshop participants use as their evaluative baseline.","marker":"[32]"},{"why":"Provides the harm-reduction principles for healthcare settings that shape the evaluative lens in Activity 4 (accuracy, non-judgment, user autonomy).","marker":"[22]"},{"why":"Establishes the 'emergent abilities' of large language models that motivate exploring them for this new application.","marker":"[62]"},{"why":"Underpins the prompt-engineering methodology and the rationale for the custom harm-reduction template.","marker":"[31]"},{"why":"Grounds the hallucination and stochastic-output risks that motivate the paper's caution and its call for guardrails and verification.","marker":"[5]"},{"why":"Proposed as the technical route for keeping LLM responses current and grounded through retrieval-augmented generation, addressing the 'information currency' challenge.","marker":"[30]"},{"why":"Supports the need for expert stakeholder involvement when designing high-stakes AI systems with potential health implications.","marker":"[48]"}],"fun_headline_variants":["LLMs need design guardrails to safely aid drug users","AI can answer drug safety questions if co-designed with experts","Harm reduction AI needs ethical alignment and clear limits","Chatbots can de-stigmatize drug info, but boundaries needed","Drug-safety AI needs careful design to be effective"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings rest on one chatbot model (ChatGPT/GPT-4, August 2023) paired with one custom prompt; if other LLMs do not behave like that combination, the claimed capabilities may not transfer to real systems or later models.","fun_headline_variants_meta":{"raw":{"variants":["LLMs need design guardrails to safely aid drug users","AI can answer drug safety questions if co-designed with experts","Harm reduction AI needs ethical alignment and clear limits","Chatbots can de-stigmatize drug info, but boundaries needed","Drug-safety AI needs careful design to be effective"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000997,"raw_usage":{"total_tokens":4191,"prompt_tokens":884,"completion_tokens":3307,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":3227}},"tokens_in":500,"tokens_out":3307,"duration_ms":27144,"temperature":1.0,"reasoning_tokens":3227,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:54:03.191794+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same workshop queries, such as safe MDMA use in English and Chinese, fentanyl-testing advice in Scotland, and refusal of speedball instructions, on several current LLMs using the published prompt; if most models refuse or produce abstinence-only, judgmental, or inaccurate answers, or fail to match the multilingual and geographic adaptivity reported here, the central claim of LLM potential collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the definition and philosophy of harm reduction that the custom prompt encodes and that workshop participants use as their evaluative baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the harm-reduction principles for healthcare settings that shape the evaluative lens in Activity 4 (accuracy, non-judgment, user autonomy)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the hallucination and stochastic-output risks that motivate the paper's caution and its call for guardrails and verification."},{"cited_title":"The human body is a black box","cited_arxiv_id":null,"evidence_quote":"Supports the need for expert stakeholder involvement when designing high-stakes AI systems with potential health implications."}],"review_version":1}