{"id":"5ee7ecfa-0b4d-4e96-bf06-7e8fff1c8fc3","arxiv_id":"2501.14334","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"A corporate AI portfolio LCA model projects that high generative-AI adoption could increase AI electricity use about 24-fold by 2030.","lead":"This paper proposes a step-by-step method for companies to estimate the electricity, water, and resource footprint of their AI systems, then uses it to sketch 2030 scenarios. It reports that large generative models can use thousands of times more energy per task than traditional models, and that unchecked adoption could multiply AI electricity use by roughly 24 by 2030.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 24.4x scenario index rests on an unsupported 2024 model-size distribution (85.8% of GenAI use cases in the high bucket), derived from a 6x OpenAI-vs-Ollama proxy; a direct sensitivity test is needed before the headline is trusted.","rationale":"The reader's weakest_assumption correctly identifies the 2024 company portfolio and the 85.8% high-model-size share as the anchor for all projections. I agree that this is the most load-bearing spot: the 24.4x headline cannot be separated from that distribution. My partial disagreement is on the claim that the headline would 'change roughly proportionally' if the distribution differs. Because the 2024 denominator is itself inflated by the High share, correcting the distribution can either raise or lower the ratio depending on what the 2030 mix becomes; the direction is not predetermined. That makes a direct recomputation with alternative mixes the right test. The paper's own external calibration against Luccioni et al. supports the per-inference energy levels for small and medium open models, but it does not validate the corporate portfolio weights or the scenario CAGRs. The Table 5 row swap is a real editorial defect and supports the need for the simulator to be released, but it is not the central mathematical vulnerability. The overall verdict should remain CONDITIONAL: the method is coherent and useful as a scenario framework, but the headline 24.4x should be treated as an illustration pending a robustness check of the 2024 model-size mix and publication of the simulator.","tokens_in":1087,"tokens_out":1013,"duration_ms":74815,"concrete_test":"Recompute Table 2 'High adoption without boundaries' with two alternative 2024 model-size mixes: (a) allocate OpenAI usage proportionally to the observed HuggingFace download ratios of Llama-3.1 8B/70B/405B instead of assigning all OpenAI traffic to High; (b) use an actual enterprise telemetry mix, e.g., OpenAI API token volumes by model family or a corporate LLM gateway log. If the energy index moves by more than ±25% from 2440, the headline is not robust to the 2024 size-mix assumption; if it stays within roughly 18x-30x, the qualitative conclusion survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The 2030 projection inherits the 2024 portfolio baseline, and the baseline is anchored by the model-size mix in Table 25. In the supplementary 'Model Size' section, the ratio of model sizes is computed from Llama-3.1 HuggingFace downloads plus the assumption that OpenAI's GPT-4 has 6x the usage of Ollama, with all OpenAI usage assigned to the 'High' bucket. This yields 31.4M estimated GPT-4 'downloads' against 30k Llama-405B downloads, pushing the High share to 85.8%. That is not an observed enterprise distribution: it ignores GPT-4o-mini and GPT-3.5 families, conflates Ollama local usage with HuggingFace downloads, and assumes all OpenAI traffic is the largest model. Since high-size inference is about 11x medium and 186x low for chat (and similar for RAG and agents), the 2024 denominator and the 2030 numerator both scale with this allocation. The direction of the resulting error is not obvious: if the 2024 High share is overestimated, the 2024 baseline is inflated, so the 24.4x ratio may understate growth; if the 2030 shift toward high-size models is also overestimated, the ratio may overstate it. The paper's sensitivity analysis varies the 2030 model-size evolution, but never perturbs the 2024 size-mix baseline, so the headline is not stress-tested where it is most assumption-heavy. Without the Excel simulator or the underlying data, an independent reader cannot determine whether the 24.4x factor is robust. The per-inference calibration against Luccioni et al. is a genuine external check, but it does not validate the portfolio composition or the scenario CAGR inputs.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes a four-layer methodology for estimating the environmental impacts of a company's AI portfolio: a life-cycle assessment model for hardware, an AI use-case clustering model, a representative corporate portfolio model, and a set of 2030 projection scenarios. The authors calibrate per-inference energy figures against external measurements from Luccioni et al. and Artificial Analysis, report per-inference impacts across GHG, water, primary energy, and resource depletion, and present scenario indices relative to a fictional 2024 portfolio. Headline results are that large generative AI models consume up to about 4,600x more energy per inference than traditional NLP models and that the \"High adoption without boundaries\" 2030 scenario yields a 24.4x increase in AI electricity use relative to 2024. The paper also contains a mitigation thought experiment suggesting that 175x-565x hardware efficiency improvements would be needed to offset 90% GHG reductions under high adoption scenarios.","tokens_in":33700,"tokens_out":6278,"duration_ms":55631,"significance":"The paper's main value is practical: it offers a parameterized, multi-criteria framework that companies could use without deep LCA expertise, and it is anchored by a genuine external calibration to Luccioni et al. The scenario logic is transparent, and the attention to water, resource depletion, and embodied impacts goes beyond the usual carbon-only analyses. If the assumptions are made fully testable, the framework could be a useful decision-support tool for sustainability practitioners. However, the headline 24.4x factor is not an empirical forecast but an index built on hand-set growth assumptions and a fragile 2024 portfolio baseline; the current sensitivity analysis does not stress-test the most assumption-heavy input. With corrections, released data, and clearly conditional framing, the contribution would be sound; as it stands, several central quantitative claims need further support.","major_comments":[{"comment":"Table 5 contains internally inconsistent row labels. The row labeled \"High adoption without boundaries\" reports 755/576/650 for energy/GHG/water, which are exactly the Intermediate scenario values in Table 2, while the row labeled \"Intermediate scenario\" reports 2440/1862/2102, the high-adoption values. The two \"Offset scenario\" rows then associate the 565x and 175x hardware-efficiency factors in an order that contradicts the Discussion's statement that the high-adoption scenario requires 565x efficiency. Please relabel the rows and reassign the 565x/175x factors so the thought experiment is internally consistent.","section":"Results, Table 5"},{"comment":"The 2024 baseline places 85.8% of GenAI usage in the High model-size bucket. This share is derived by adding an estimated GPT-4 \"downloads\" count of about 31.4M on the assumption that OpenAI has 6x Ollama's usage, assigning all OpenAI traffic to the High bucket, and comparing with Llama-3.1 Hugging Face downloads. Because per-inference energy for High chat is about 186x Low chat (Table 1), this single assumption dominates both the 2024 denominator and the 2030 numerator of every scenario index, including the 24.4x headline. The sensitivity analysis in Table 3 perturbs only the 2030 model-size evolution and never varies the 2024 size-mix baseline. A direct sensitivity test over plausible 2024 High/Low/Medium shares is needed before the headline factor can be considered robust.","section":"Supplementary, Company Portfolio Model, Model Size (Table 25)"},{"comment":"The abstract says the model \"forecasts AI electricity use up to 2030\" and that AI electricity use \"is projected to rise by a factor of 24.4,\" but the 24.4x is the output of a deliberately extreme scenario with hand-set CAGRs of 47% for GenAI and 55% for agentic use cases, a threefold model-size increase, and a threefold output-token increase (Tables 2, 28-31). No uncertainty intervals or alternative-parameter ranges are reported for this scenario index. Please state consistently that these are conditional \"if-then\" projections, and report a sensitivity of the 24.4x factor to the main CAGR and token assumptions, as is already done for the Intermediate scenario.","section":"Abstract and Results, Table 2"},{"comment":"The statistical portfolio that anchors all results is constructed from non-public Capgemini data: a list of 350+ client use cases labeled with an LLM, and \"typical company\" usage frequencies calibrated to Capgemini experience. The \"excel simulator\" mentioned in Methods is not provided, so an independent reader cannot recompute the 2024 portfolio or the scenario indices. Please release the simulator or a parameterized, anonymized version with all distributions explicitly tabulated, or clearly mark the relevant sections as illustrative rather than reproducible.","section":"Methods, Model 3 and Supplementary, Use Case usage"}],"minor_comments":[{"comment":"The text contains the placeholder \"Error! Reference source not found.\" in place of a citation to Table 2; this must be fixed.","section":"Results, before Table 2"},{"comment":"The text refers to \"Paccou et al. study6,\" but reference 6 is Wijnhoven and Paccou; please correct the citation style.","section":"Supplementary, 2030 Systemic projections, Model Efficiency"},{"comment":"The caption states twice that \"CAGR, by definition, represents exponential growth over time\"; please remove the redundant clause.","section":"Figure 2 caption"},{"comment":"The \"Return on Environment\" metric is introduced as a recommendation but is not defined quantitatively; a short definition or reference would help readers understand what is being proposed.","section":"Discussion, Return on Environment"}],"recommendation":"major_revision","confidential_remarks":"The paper is close to a consulting-style framework paper: the methodology is reasonable, the external calibration is a genuine strength, but the central 24.4x result rests on an unverified 2024 model-size distribution and on several hand-set CAGRs. The Table 5 row swap is a concrete error that must be corrected. The missing Excel simulator and reliance on internal Capgemini use-case data also limit reproducibility. These are fixable, so I see no grounds for rejection, but the authors should be required to address the baseline sensitivity and to clarify the conditional nature of the scenario results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my read on Desroches et al. The useful core is the integrated accounting framework: per-inference energy from measured sources, LCA for embodied impacts, a 192-cluster company portfolio, and 2030 scenarios. That modular structure is genuinely practical, and the small-model calibration against Luccioni et al. (0.093 Wh vs 0.082/0.104 for OPT/BLOOMz) is a real external check. The multi-criteria breakdown (GHG, water, resources, embodied vs operational) is a step beyond most AI-energy papers and would serve a corporate sustainability team well.\n\nThe soft spots are also real. The headline 24.4x scenario index rests on a 2024 model-size distribution that says 85.8% of GenAI usage is in the high bucket, derived from HuggingFace downloads for Llama plus an assumption that OpenAI has 6x Ollama usage. That conflates local usage with HF downloads, ignores small OpenAI models, and is not an observed enterprise distribution. The paper's own sensitivity analysis perturbs the 2030 model-size evolution but never the 2024 baseline, so the most assumption-heavy input is not stress-tested. The direction of error is ambiguous, as your note says; that makes the problem worse, not better, because a reader cannot tell whether 24.4x under- or overstates. Table 5 has a clear label swap (the 'High adoption' and 'Intermediate' rows are inverted relative to the text), and the 'Error! Reference source not found' placeholder should have been caught. The abstract says 'forecasts' when the paper itself describes the scenarios as conditionals; that mismatch should be fixed.\n\nI do not think this is a fatal flaw. The per-inference numbers are anchored to external measurements, the scenario logic is transparent, and the limitations section is unusually candid about cluster oversimplification, missing pre-training, and uncertainty. The missing piece is the Excel simulator and data. The paper says the models are 'merged into an excel simulator for analysis' but no artifact is provided. Without it, an independent reader cannot decompose the portfolio ratios or rerun the sensitivity analysis. That is the difference between a useful framework and a claimed one.\n\nWho is this for? Practitioners who need a structured way to estimate corporate AI footprints and researchers working on AI LCA. It deserves a serious referee, but it needs the simulator released and the 2024 model-size baseline either replaced with real enterprise data or subjected to a direct sensitivity test. I would not desk-reject; I would send to review with a request for the artifact and a careful look at Table 5 and the baseline.","headline":"A genuinely integrative corporate AI footprint framework, but the headline 2030 numbers rest on an unvalidated 2024 model-size baseline and no released simulator.","tokens_in":34234,"tokens_out":1652,"would_cite":false,"duration_ms":16110,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a portfolio-level method to estimate AI environmental impacts, showing that large generative models use up to 4,600 times more energy per inference and that unconstrained adoption could raise AI electricity use…","keywords":["AI environmental impact","generative AI","life cycle assessment","energy consumption","2030 scenario","corporate AI portfolio","sustainability","LLM inference"],"falsifier":"Run the three reference Llama models on a p4de.24xlarge instance and measure energy per chat inference; the model predicts roughly 0.093 Wh (8B), 1.55 Wh (70B), and 17.3 Wh (405B), and measured values more than about twice those numbers would break the 4,600x ratio and the portfolio projections built on it.","tokens_in":1821,"feed_emoji":"⚡","tokens_out":2266,"duration_ms":108612,"temperature":0.7,"pith_summary":"This paper builds a practical, top-down method for estimating the environmental footprint of a company's entire AI portfolio without requiring deep LCA or AI expertise. It anchors the method on a typical large-company portfolio of 100 use cases and shows that generative AI already dominates the energy bill: large-model chat inference uses up to 4,600 times more electricity than a conventional NLP inference, and RAG/agent workflows push that to 25,000 times. It then projects the same portfolio to 2030 under four boundary scenarios and an intermediate scenario. The headline result is that an unconstrained high-adoption path raises AI electricity use by a factor of 24.4, while a frugal-efficiency path could cut it by 70%, and that no single efficiency lever reaches a 90% GHG reduction alone. The authors argue for standardized assessment, provider transparency, and a Return on Environment metric as the practical consequences.","feed_headline":"Generative AI could drive a 24.4x rise in AI electricity by 2030","feed_subtitle":"A portfolio-level model quantifies corporate AI footprints and finds isolated efficiency fixes can't reach net-zero.","key_machinery":"The carrying mechanism is a four-sub-model projection chain. First, an LCA model assigns embodied and operational impacts to three capacity types—compute, storage, and network—using a bill-of-materials for a cloud server with 8 A100-class GPUs. Second, a use-case clustering scheme splits AI into 192 clusters by AI type, task (chat, RAG, agents, tabular, computer vision, NLP), model-size bucket, user count, and usage frequency. Third, a representative company portfolio distributes use cases as 71% traditional AI and 29% GenAI, with the GenAI model-size mix anchored on open download counts scaled by a 6x closed-provider usage factor. Fourth, 2030 scenarios combine usage-growth CAGRs with systemic-efficiency factors for hardware FLOPs/W, PUE, grid decarbonization, quantization, model size, and output tokens. The per-inference impact formula sums contributions over the two lifecycle steps (fine-tuning and inference), the four component categories, and the two embodied/operational stages, which lets the same machinery compute GHG, water, primary energy, resource depletion, and final electricity use for any portfolio.","core_discovery":"The paper's central claim is that a company-level AI environmental footprint can be approximated from a small set of publicly available parameters—model size, use-case type, user count, usage frequency, and geographic distribution—and that doing so reveals a scale problem. Per-inference calculations show a high-size generative chat inference at about 17 Wh versus 0.0037 Wh for a traditional NLP inference, a 4,600-fold gap; high-size agentic workflows reach roughly 96 Wh, about 25,000 times the traditional baseline. Aggregated over a representative 100-use-case portfolio, generative AI accounts for 99.9% of inference energy despite being only 29% of use cases. Projecting to 2030, the high-adoption scenario lifts portfolio electricity use by factor 24.4 and GHG emissions by 18.6, whereas a scenario combining moderate adoption with ambitious hardware and grid improvements cuts energy by 70%. The authors further show that embodied impacts dominate water use (around 30%) and resource depletion (89%) while remaining small for GHG (5%), meaning carbon-only accounting misses the other environmental dimensions.","pith_inferences":["Beyond the paper's own proposals, the same 192-cluster taxonomy could be turned into a public task-level eco-score for any model, making energy per task the unit of comparison rather than a model-level rating.","The projection is conditional on the 85.8% high-model-size share, itself derived from open download counts times a 6x closed-provider usage factor; direct telemetry from closed providers would tighten this number more than any other single measurement.","The Return on Environment metric is only outlined; a natural completion is to pair it with consequential LCA so that AI's indirect energy savings are weighed against its direct footprint, a step the paper explicitly leaves out."],"forward_implications":["If the methodology is adopted, companies can estimate AI footprint from public parameters without waiting for providers to disclose internal data, lowering the barrier to net-zero accounting.","The per-inference ratios imply that shifting a workload from traditional NLP to a large generative or agentic system changes the energy profile by orders of magnitude, making use-case selection a first-order sustainability lever.","If the high-adoption 2030 scenario holds, corporate AI electricity scales by 24.4x; even the intermediate path (about 7.6x) strains net-zero targets without additional efficiency gains.","The 175x to 565x hardware efficiency improvement required for a 90% GHG reduction under high adoption indicates that hardware progress alone cannot offset usage growth, so adoption limits, frugal model design, grid decarbonization, and transparency are structural necessities.","Because embodied impacts dominate resource depletion and a substantial share of water use, focusing only on operational GHG reductions can leave water and minerals problems unaddressed, supporting the paper's call for multi-criteria assessment and an eco-score."],"supporting_citations":[{"why":"Supplies empirical per-inference energy measurements for standard CV/NLP models and for OPT-6.7B/BLOOM used to calibrate the traditional-AI and small-LLM baselines.","marker":"9"},{"why":"Supplies time-to-first-token and output throughput for Llama 3.1 models across cloud providers, the basis for generative inference energy calculations.","marker":"31"},{"why":"Supplies global data-center electricity growth projections that validate the 2024 portfolio aggregate and frame the 2030 scenarios.","marker":"4"},{"why":"Defines the Llama 3.1 model family used as reference architectures for the low, medium, and high GenAI model-size buckets.","marker":"18"},{"why":"Supplies the 6x usage ratio between OpenAI and local Llama serving that sets the 85.8% high-model-size share in the 2024 portfolio.","marker":"62"},{"why":"Supplies the historical 1.28x/year hardware efficiency trend used to set the x4.4 efficiency factor by 2030 in low-efficiency scenarios.","marker":"69"},{"why":"Establishes the training-energy accounting and transparency agenda that motivates the method's boundary choices and disclosure recommendations.","marker":"20"},{"why":"Supplies the regional datacenter capacity distribution used to weight electricity-grid impacts across the portfolio.","marker":"63"}],"fun_headline_variants":["Generative AI could spike electricity 24.4x by 2030","AI's 2030 energy use: 24.4x higher under high adoption","Isolated efficiency gains can't offset AI's 24.4x energy rise","Generative AI uses 4600x more energy per inference"],"cache_read_input_tokens":36224,"weakest_assumption_plain":"The load-bearing premise is that the 2024 representative portfolio—71% traditional AI, 29% GenAI, with 85.8% of GenAI calls in the largest model-size bucket—and the hand-selected 2030 adoption rates (up to 47% GenAI CAGR and 55% agentic CAGR) reflect reality; if the true mix is smaller models or slower adoption, the 24.4x headline shrinks roughly proportionally.","fun_headline_variants_meta":{"raw":{"variants":["Generative AI could spike electricity 24.4x by 2030","AI's 2030 energy use: 24.4x higher under high adoption","Isolated efficiency gains can't offset AI's 24.4x energy rise","Generative AI uses 4600x more energy per inference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1569,"prompt_tokens":1030,"completion_tokens":539,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":456}},"tokens_in":646,"tokens_out":539,"duration_ms":5319,"temperature":1.0,"reasoning_tokens":456,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:15:11.653494+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the three reference Llama models on a p4de.24xlarge instance and measure energy per chat inference; the model predicts roughly 0.093 Wh (8B), 1.55 Wh (70B), and 17.3 Wh (405B), and measured values more than about twice those numbers would break the 4,600x ratio and the portfolio projections built on it.","supporting_citations":[{"cited_title":"Artificial Analysis https://artificialanalysis.ai","cited_arxiv_id":null,"evidence_quote":"Supplies time-to-first-token and output throughput for Llama 3.1 models across cloud providers, the basis for generative inference energy calculations."},{"cited_title":"Electricity 2024 – Analysis","cited_arxiv_id":null,"evidence_quote":"Supplies global data-center electricity growth projections that validate the 2024 portfolio aggregate and frame the 2030 scenarios."},{"cited_title":"https://blog.langchain.dev/langchain-state-of-ai-2024/ (2024)","cited_arxiv_id":null,"evidence_quote":"Supplies the 6x usage ratio between OpenAI and local Llama serving that sets the 85.8% high-model-size share in the 2024 portfolio."},{"cited_title":"Can AI Scaling Continue Through 2030? Epoch AI https://epoch.ai/blog/can-ai-scaling- continue-through-2030 (2024)","cited_arxiv_id":null,"evidence_quote":"Supplies the historical 1.28x/year hardware efficiency trend used to set the x4.4 efficiency factor by 2030 in low-efficiency scenarios."},{"cited_title":"https://www.cbre.com/insights/reports/global-data-center-trends- 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the regional datacenter capacity distribution used to weight electricity-grid impacts across the portfolio."}],"review_version":1}