{"id":"2cbc59e2-a650-47c7-8604-429a7ad19f80","arxiv_id":"2608.11830","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 47 therapeutic LLM configurations, the safest models had estimated energy use up to 60 times higher than slightly less safe efficient models, and extra reasoning did not reliably improve safety.","lead":"This paper measures how much extra electricity and water is used by AI chatbots for mental health as their safety scores get higher. It finds the safest models can use about 60 times more energy for only 2.6 points of extra safety, and more thinking time does not always help.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 60x ratio rests on EcoLogits' model-level, configuration-invariant estimates with unquantified hardware and electricity-mix assumptions; a sensitivity envelope is required before the disproportionality claim is accepted.","rationale":"The reader's weakest assumption already identifies EcoLogits modeled estimates as the load-bearing premise. My read agrees and sharpens it: the mismatch is not only statistical uncertainty but also a configuration-to-model pairing problem, because safety is per configuration while footprints are per base model. The proposed sensitivity test would settle whether the 60x number is a robust feature of the data or an artifact of default hardware and electricity-mix assumptions. I do not see a reason to move the verdict, because the paper explicitly flags the LCA modeling limitation and frames results as estimates; however, the conditional status is correct until such a sensitivity bound is reported, and no independent evidence such as direct measurements or provider-reported data is supplied.","tokens_in":8985,"tokens_out":6642,"duration_ms":71258,"concrete_test":"Recompute Table 1's headline pair in EcoLogits v0.11.0 for openai/gpt-5.5 and anthropic/claude-haiku-4.5 across all supported hardware profiles, PUE values (e.g., 1.1-1.6), and electricity-mix zones, reporting the energy ratio for each combination. If the minimum plausible energy ratio across this envelope falls below about 10x, the disproportionate-trade-off claim is not robust; if the ratio remains above about 20x under all supported assumptions, the qualitative conclusion stands despite unverifiable provider-side hardware.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central result is a single-pair comparison: openai/gpt-5.5 (Best Risk 96.02, 6.4390 kWh/1M tokens) versus anthropic/claude-haiku-4.5 (93.41, 0.1095 kWh/1M tokens). As Section 4 states, 'EcoLogits returned identical impact estimates within each base model because its fixed-token estimates do not distinguish configuration-specific costs.' Thus the safety score is configuration-specific (prompt and reasoning setting), while the environmental cost is base-model-level and determined by EcoLogits' assumptions about hardware (H100/A100), PUE, and electricity mix. The paper provides no uncertainty intervals and no provider-side validation for closed models. If gpt-5.5 is served on newer, more efficient accelerators, different caching, or different routing than the default profile, and claude-haiku-4.5 on less efficient hardware, the 58.8x ratio could shrink substantially; the same sensitivity applies to GWP, water, and ADPe because EcoLogits derives these indicators from shared assumptions. The qualitative conclusion is plausible, but the quantitative '~60-fold for 2.61 points' claim cannot be evaluated without a sensitivity bound. The secondary reasoning finding is independent and appropriately hedged.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper combines K-Bench clinical safety scores with EcoLogits life-cycle assessment estimates for 47 supported model configurations across 13 base models, and reports a non-linear safety--sustainability trade-off. The headline result is that the highest-scoring configuration, openai/gpt-5.5 (Best Risk 96.02), is estimated to use 6.4390 kWh per million output tokens, whereas anthropic/claude-haiku-4.5 achieves 93.41 at 0.1095 kWh, an approximately 60-fold energy increase for a 2.61-point safety gain. The paper also reports row-level analyses suggesting that additional test-time reasoning does not consistently improve safety, and recommends dynamic model selection and model cascading as a way to reduce environmental impact while preserving clinical performance in high-risk cases.","tokens_in":9200,"tokens_out":3350,"duration_ms":36726,"significance":"The paper addresses a genuine gap: clinical safety and environmental impact are rarely evaluated together for therapeutic LLMs, and the qualitative direction of the trade-off is visible in the data. Strengths include reliance on public benchmark data, use of a reproducible estimation package, absence of fitted parameters, and candid discussion of limitations. The secondary finding on test-time reasoning is appropriately hedged and does not depend on the environmental estimates. However, the central quantitative claim is a single-pair comparison built on modeled, configuration-invariant EcoLogits estimates that carry no uncertainty bounds; without a sensitivity analysis, the specific 60-fold disproportionality claim cannot be fully evaluated. This is a load-bearing issue for the paper's main conclusion, so the manuscript needs revision before the claim is accepted.","major_comments":[{"comment":"The headline ratio of approximately 60x energy per million output tokens (gpt-5.5 at 6.4390 kWh vs. claude-haiku-4.5 at 0.1095 kWh) rests entirely on EcoLogits model-level estimates. The paper states that these estimates depend on assumed hardware, PUE, and electricity mix, and Section 4 notes that EcoLogits returns identical impacts for every configuration of a base model. For closed models, provider-side serving details are not observable, so the true ratio could be materially different if gpt-5.5 is served on newer accelerators, different caching, or different routing than EcoLogits assumes. Please add a sensitivity analysis that varies hardware profile, PUE, electricity-mix zone, and token-normalization assumptions, and report how the headline ratio and the Pareto frontier change. Without such bounds, the quantitative disproportionality claim cannot be evaluated.","section":"Table 1 and Section 4"},{"comment":"Because EcoLogits returns identical impact estimates within each base model, the comparison pairs configuration-specific safety scores with base-model-level environmental cost. The '2.61 points for ~60x' comparison uses the Best Risk score for each model but does not know the energy use of that particular prompt--reasoning configuration; the 6.4390 kWh figure is a base-model-level estimate. The paper should explicitly acknowledge this unit mismatch and, if possible, bound it by estimating configuration-specific inference cost, for example by accounting for the actual number of output tokens generated by each configuration in the K-Bench evaluation.","section":"Section 4, Configuration-Level Variation"},{"comment":"Of the 28 base architectures in the initial K-Bench dataset, 15 were excluded because EcoLogits lacked the required metadata, leaving 13 base models. The paper reports this transparently, but it does not discuss how this selection could bias the Pareto frontier. Excluded models such as llama-4-maverick, deepseek-v4-flash, claude-fable-5, and grok-4.20 may be systematically newer, less documented, or at different points on the safety--efficiency curve. Please characterize the K-Bench safety scores of the excluded models and state whether their inclusion or exclusion could change the location of the frontier or the qualitative conclusion.","section":"Section 3, Data Synthesis and Filtering"}],"minor_comments":[{"comment":"The abstract says '47 supported model configurations' while Table 1 and Figure 1 present 13 base models; the relationship between configurations and base models could be stated more explicitly in the methodology.","section":"Abstract"},{"comment":"The example comparing gpt-5.5 with 'none' reasoning (96.02) and 'low' reasoning (95.50) would be more reproducible if the full configuration details, including the prompt variant, were given.","section":"Section 5, Test-Time Compute"},{"comment":"Some model labels in Figure 1 appear crowded or overlapping (e.g., the gemini labels in panels a and b); consider using a legend with distinct markers to improve readability.","section":"Figure 1"},{"comment":"There is a minor typo in the ACM reference line: 'InACM/IEEE' should read 'In ACM/IEEE'.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is appropriate for a workshop-length venue and the authors are candid about limitations. The main concern is that the central quantitative claim needs a sensitivity analysis before it can be accepted; this is fixable within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a small, honest paper that quantifies something people have been hand-waving about—the safety-vs-environment trade-off in therapeutic LLMs—and it's worth a look despite a fragile headline number. The 60x energy ratio for a 2.61-point safety gain is the kind of stat that gets quoted, so it matters that it rests on modeled, configuration-invariant EcoLogits estimates.\n\nWhat's new: they merge the K-Bench leaderboard with EcoLogits LCA data across 47 configurations, draw Pareto frontiers for energy, GWP, water, and ADPe, and report a secondary finding that extra test-time reasoning didn't consistently improve safety (4/12 high-reasoning, 5/9 low-reasoning). That secondary finding is independent of the environmental modeling and is appropriately hedged.\n\nCredit where due: the paper is transparent about its weak spots. It says EcoLogits returns identical estimates within each base model, so configuration-level safety variation is paired with base-model-level environmental cost. It excludes 15 architectures it couldn't estimate rather than guessing. It warns that the indicators are not fully independent. That's good practice.\n\nThe soft spots: the headline '~60-fold for 2.61 points' is a single-pair comparison (gpt-5.5 vs claude-haiku-4.5) and the environmental numbers are model-level estimates with no uncertainty intervals and no provider-side validation. If gpt-5.5 is served on newer accelerators or different caching than EcoLogits assumes, the ratio could shrink a lot. The same sensitivity applies to the other three indicators because they share assumptions. So the precise magnitude should be treated as a rough illustration, not a measurement. The qualitative conclusion—top-end safety gains come with disproportionate environmental cost—is plausible and visible in the data, but the quantitative claim is only as good as EcoLogits' default hardware profile.\n\nMinor: the abstract and text say 'combined risk score' while Table 1 and Figure 1 use 'Best Risk'—worth aligning.\n\nWho benefits: deployment engineers choosing models for mental-health chatbots, Green AI people, and anyone building model routers. The paper is a decent data point for a workshop/companion venue. I'd send it to a serious referee; the limitations are stated clearly enough that the referee can focus on whether the sensitivity analysis is adequate.\n\nMy recommendation: engage, but require an explicit sensitivity bound or at least a sentence saying the 60x is an upper-bound-style estimate under EcoLogits assumptions. Without that, the headline will be over-interpreted.","headline":"A small, honest paper that quantifies the safety-environment trade-off in therapeutic LLMs with public data; the qualitative pattern is real, but the headline 60x figure rests on modeled, configuration-invariant estimates and needs a sensitivity bound before it can be quoted as a measurement.","tokens_in":9735,"tokens_out":2277,"would_cite":false,"duration_ms":22518,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"For therapeutic LLMs, near-maximal clinical safety scores come with disproportionately large estimated environmental footprints: roughly a 60-fold energy increase for a 2.61-point safety gain, while added test-time reasoning does not…","keywords":["clinical AI safety","therapeutic LLMs","K-Bench","EcoLogits","life-cycle assessment","sustainable AI","test-time compute","Pareto frontier"],"falsifier":"Provider-side metered energy use per million output tokens for gpt-5.5 and claude-haiku-4.5 on comparable hardware, made public, would settle it: if the measured ratio is far below the modeled 60-fold, the central trade-off is an artifact of EcoLogits assumptions; likewise, a K-Bench evaluation showing high-reasoning configurations consistently outperform no-reasoning configurations across all matched comparisons would falsify the secondary claim.","tokens_in":8795,"feed_emoji":"⚡","tokens_out":6726,"duration_ms":60220,"temperature":0.7,"pith_summary":"This paper tries to establish that, among large language models proposed for therapeutic use, the last few points of clinical safety come at a disproportionate environmental price. By pairing K-Bench clinical safety scores with EcoLogits life-cycle estimates for 47 model configurations, the authors find that the highest-scoring model (gpt-5.5, 96.02) is estimated to use about 60 times more energy per million output tokens than a model scoring 2.61 points lower (claude-haiku-4.5, 93.41). They also find that adding test-time reasoning does not consistently improve safety scores across configurations. If true, this matters because it gives deployers a concrete reason to treat model choice as a multi-objective decision, reserving the largest models for high-risk cases rather than defaulting to them everywhere.","feed_headline":"The last 2.6 safety points cost 60 times the energy","feed_subtitle":"Top-rated therapeutic models are estimated to use ~60 times more energy per token than models just 2.6 points behind.","key_machinery":"The argument runs on two instruments and one analytic device. K-Bench is a transcript-based benchmark that evaluates therapist-style AI in simulated multi-turn mental health conversations: a factorial vignette generator creates risk scenarios across suicide, self-harm, domestic violence, and substance misuse, and a clinician-calibrated automated judge (Cohen's kappa of 0.850 on the calibration set) scores risk recognition and exploration on a 0-100 scale. EcoLogits is a Python package that converts model metadata, assumed hardware, Power Usage Effectiveness, and regional electricity mix into life-cycle impact estimates per million output tokens for energy, global warming potential, water consumption, and abiotic depletion. The Pareto frontier then sorts the 47 supported configurations so that each point on a frontier is not outperformed by another model on both safety and impact; this device turns raw pairs into the claim that the highest safety scores sit on a steep-impact tail.","core_discovery":"The central claim is that clinical safety and environmental footprint are in a non-linear trade-off for therapeutic LLMs, with sharply rising impact at the top of the safety distribution. On K-Bench's Best Risk score, gpt-5.5 reaches 96.02 at an estimated 6.4390 kWh per million output tokens, while claude-haiku-4.5 reaches 93.41 at 0.1095 kWh, a 98.3 percent reduction in energy for a 2.61-point safety gap, equivalently a roughly 60-fold energy increase for the higher score. The pattern repeats across global warming potential, water consumption, and abiotic depletion. A secondary claim is that additional test-time reasoning does not reliably buy safety: across 12 high-reasoning comparisons, only 4 improved over no reasoning, and across 9 low-reasoning comparisons, only 5 improved. The authors conclude that selecting solely by safety score, or assuming larger models and more compute are the route to safety, is inefficient, and they recommend dynamic model selection and cascading.","pith_inferences":["A natural but implicit extension: the Pareto 'knee' where added safety per unit of environmental cost drops sharply could be formalized as a selection rule, such as choosing the model with the best safety score per kilowatt-hour above a clinically acceptable threshold; this is my inference, not stated in the paper.","If the EcoLogits hardware assumptions are right, provider-published energy meters should show the same ordering across models, though not necessarily the same ratios; a testable prediction is that the gpt-5.5 versus claude-haiku-4.5 energy ratio lies well above 10 in real deployments.","The reasoning findings suggest a mechanism worth testing: K-Bench's rubric may reward direct, well-bounded responses, while longer reasoning chains introduce more opportunities for unsafe or off-topic statements, a hypothesis the authors did not explore.","The same paired-analysis approach could be applied to other clinical tasks or to fine-tuned smaller models; if domain-specific training closes the safety gap, the case for frontier models in therapy weakens further."],"forward_implications":["A deployer who wants the very highest K-Bench risk score should expect an estimated environmental footprint roughly 60 times larger than a model only 2.61 points behind, across energy, carbon, water, and resource depletion.","Efficient models such as claude-haiku-4.5 can sit on the Pareto frontier, meaning no other evaluated model achieves both a higher safety score and a lower estimated impact.","Test-time reasoning should not be treated as a reliable safety lever: in the evaluated configurations, added reasoning often lowered K-Bench scores, so compute budgets may be spent with no safety return.","Model selection in therapeutic AI should be framed as multi-objective optimization, with dynamic routing or cascading used to send high-risk cases to larger models and routine interactions to smaller ones.","Because EcoLogits returns identical environmental estimates for every configuration of a given base model, prompt- and reasoning-level safety differences are evaluated against a fixed per-model footprint rather than a per-configuration one."],"supporting_citations":[{"why":"Supplies the EcoLogits package that produces the life-cycle environmental impact estimates for each model endpoint.","marker":"[23]"},{"why":"Provides evidence that inference, especially autoregressive decoding, dominates energy use, motivating the per-token footprint framing.","marker":"[17]"},{"why":"Analyzes the energy cost of test-time reasoning, grounding the secondary claim that reasoning compute may not pay off.","marker":"[14]"},{"why":"Offers recent estimates of AI inference energy and efficiency pathways, supporting the discussion of test-time scaling.","marker":"[21]"},{"why":"Shows that larger parameter counts are not always the most efficient route to better performance, a premise for preferring smaller models.","marker":"[11]"},{"why":"Argues that small models can suffice for many tasks, which the paper cites to motivate model cascading.","marker":"[3]"},{"why":"Frames green AI as evaluating efficiency alongside task performance, grounding the multi-objective recommendation.","marker":"[24]"},{"why":"Provides evidence that safety and accuracy follow different scaling laws in clinical LLMs, relevant to interpreting marginal safety gains.","marker":"[27]"},{"why":"Measures the environmental impact of delivering AI at scale, justifying concern with inference footprint.","marker":"[6]"}],"fun_headline_variants":["2.6 safety points cost 60x energy in therapeutic LLMs","Top safety scores come at 60x energy cost","Safety vs. sustainability: 2.6 points = 60x energy","More compute doesn't guarantee safer therapeutic AI","Therapeutic LLMs: 2.6 safety points cost 60x energy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that EcoLogits' modeled life-cycle estimates, with their assumed hardware, data-center efficiency, and electricity mix, faithfully represent the relative environmental cost of serving each configuration; if providers serve flagship models on different hardware or with caching that the model does not capture, the reported 60-fold energy ratio could change materially.","fun_headline_variants_meta":{"raw":{"variants":["2.6 safety points cost 60x energy in therapeutic LLMs","Top safety scores come at 60x energy cost","Safety vs. sustainability: 2.6 points = 60x energy","More compute doesn't guarantee safer therapeutic AI","Therapeutic LLMs: 2.6 safety points cost 60x energy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000713,"raw_usage":{"total_tokens":3214,"prompt_tokens":957,"completion_tokens":2257,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":2168}},"tokens_in":573,"tokens_out":2257,"duration_ms":15940,"temperature":1.0,"reasoning_tokens":2168,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:25:09.354984+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Provider-side metered energy use per million output tokens for gpt-5.5 and claude-haiku-4.5 on comparable hardware, made public, would settle it: if the measured ratio is far below the modeled 60-fold, the central trade-off is an artifact of EcoLogits assumptions; likewise, a K-Bench evaluation showing high-reasoning configurations consistently outperform no-reasoning configurations across all matched comparisons would falsify the secondary claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the EcoLogits package that produces the life-cycle environmental impact estimates for each model endpoint."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Analyzes the energy cost of test-time reasoning, grounding the secondary claim that reasoning compute may not pay off."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues that small models can suffice for many tasks, which the paper cites to motivate model cascading."},{"cited_title":"Linear Motility Maps in Nonlinear Viscous Fluids","cited_arxiv_id":"2606.00063","evidence_quote":"Provides evidence that safety and accuracy follow different scaling laws in clinical LLMs, relevant to interpreting marginal safety gains."}],"review_version":1}