Pith. sign in

REVIEW 4 major objections 6 minor 51 references

Green AI patenting is shifting from combustion engines to data processing, microgrids, and agricultural water management.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 18:05 UTC pith:M6UJDEEJ

load-bearing objection A solid new Green AI patent dataset and taxonomy, but the headline 'shift' claim and market-value rankings need robustness work before policy use. the 4 major comments →

arxiv 2509.10109 v1 pith:M6UJDEEJ submitted 2025-09-12 econ.GN q-fin.EC

The anatomy of Green AI technologies: structure, evolution, and impact

classification econ.GN q-fin.EC
keywords green AIclimate patentspatent analysistopic modelingtechnological changeforward citationsmarket valueinnovation policy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper builds a new dataset of about 63,000 U.S. patents that are both AI-related and climate-related, then uses an unsupervised machine-learning topic model on their abstracts to map sixteen technological domains grouped into four broader macro-domains. It argues that Green AI patenting is undergoing a structural shift: legacy fields such as combustion-engine control and exhaust treatment are stagnating, while data processing and memory management, microgrids and distributed energy, and irrigation and agricultural water management are growing quickly. It also claims that technological impact and market value do not always align across these domains, so some climate-relevant fields with high scientific influence but low private returns may need public support. The value is in offering a systematic, data-driven taxonomy of where AI is actually being applied to climate problems, and in identifying where private incentives may be too weak to sustain innovation.

Core claim

The paper claims that Green AI patenting is not a monolithic field but a set of sixteen distinct technological domains with different life cycles and value profiles. Using a machine-learning topic model on roughly 63,000 U.S. patents that are both AI-related and climate-related, it finds a clear temporal shift away from combustion-engine control, exhaust and emission treatment, and mature aerodynamics, toward data processing, microgrids, battery management, photovoltaic devices, and agricultural water management. On impact, the paper finds that some domains, such as clinical microbiome research and agricultural water management, combine high market value with strong or moderate citation impa

What carries the argument

The central instrument is a purpose-built corpus of roughly 63,000 U.S. patents that are simultaneously classified as AI-related and climate-related, combined with a machine-learning topic model that extracts latent themes from patent abstracts. The resulting sixteen-domain taxonomy, grouped into four macro-domains and validated by coherence and diversity metrics, carries the paper's temporal-shift and impact-divergence arguments; forward citations and stock-market reactions to patent grants provide the two complementary impact measures.

Load-bearing premise

The impact rankings assume that the 42.2% of patents with stock-market value data are representative of all Green AI patents; if patents held by unlisted firms are valued differently by domain, the market-value findings are biased.

What would settle it

Compute alternative market-value estimates for the 57.8% of patents lacking stock-market data (e.g., by tracking renewal payments or later acquisitions) and re-rank the sixteen domains; the paper's claim that meteorological radar and emission treatment have weak private incentives would be falsified if those domains rise above combustion and fuel technologies in the imputed rankings.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If the shift continues, R&D targeting and policy incentives for climate AI should focus on data processing, microgrids, and agricultural water management rather than legacy combustion and emission systems.
  • Domains with high citations but low market value, such as meteorological radar and weather forecasting, are likely underfunded by private markets and are candidates for public subsidy or procurement.
  • Rising Gini concentration alongside a widening long tail means a small number of firms will hold large Green AI patent portfolios even as overall entry broadens, making monitoring for monopolistic behavior relevant.
  • The release of the dataset allows other researchers to track these sixteen domains over time and test whether the observed shift accelerates, reverses, or spreads to other patent systems.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • One extension the paper does not pursue: applying the same topic-modeling pipeline to European or Chinese patent data would show whether the U.S.-observed shift from combustion to data-centric climate AI is global or an artifact of U.S. patent-eligibility rules.
  • If the value gap for meteorological radar and emission treatment reflects a genuine market failure, a natural policy experiment is to track whether government procurement or data-sharing mandates increase patenting and citations in those domains.
  • The reported growth of the data-processing domain could partly reflect a patent-drafting fashion rather than new technical activity; checking citation-weighted domain shares would separate count inflation from technological influence.
  • The concentration finding suggests that open-licensing or patent-pool arrangements may become important for keeping Green AI accessible, though the paper itself does not assess this.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper constructs a novel dataset of 63,326 U.S. 'Green AI' patents by intersecting the CPC Y02/Y04S climate-related classes with the USPTO Artificial Intelligence Patent Dataset (AIPD). It describes the temporal dynamics, assignee concentration, geographic distribution, and technological composition of this corpus. Using BERTopic, the authors identify 16 technological domains and group them into 4 macro-domains; they then track domain-level patenting trends and combine forward citations with Kogan et al. market-value estimates to assess technological and economic impact. The headline claim is a 'clear shift' from legacy combustion-engine domains toward data processing, microgrids, and agricultural water management, with implications for climate policy and R&D priorities.

Significance. The manuscript offers a valuable descriptive contribution: the data construction is transparent, the dataset and code are deposited on Zenodo, and the systematic hyperparameter search for BERTopic is a reproducible step beyond much of the applied topic-modeling literature. If the headline shift claim and the domain-level impact rankings survive closer scrutiny, the paper provides a useful map of an important and understudied patent population. The combination of green-technology CPC tags with AIPD AI labels is a sensible and potentially reusable design. The main weaknesses are not in data construction but in the inference drawn from the topic model and in the unquantified impact comparisons; these are the load-bearing points for the paper's conclusions.

major comments (4)
  1. [Results: Technological (Macro-)Domains and Their Temporal Dynamics; Methods: Hyperparameter Tuning and Model Selection] The headline 'clear shift' claim rests on comparing topic-specific patent counts across 1976–2023 from a single BERTopic model fit to the full corpus. Patent language and drafting conventions change substantially over five decades, so a static 16-topic partition does not guarantee that 'Combustion Engine Control' in 1985 measures the same construct as in 2020. The Methods only say that 'consistency checks observed high stability' without reporting them, and UMAP is stochastic. If modern combustion patents use computing-oriented vocabulary, they could be absorbed into Data Processing & Memory Management, manufacturing the apparent shift. Please provide (i) multiple stochastic runs with topic-alignment statistics, (ii) period-wise refits (e.g., 1976–1995, 1996–2010, 2011–2023) with cross-period topic mapping, and (iii) validation of topic labels against CPC subclasses over time. This is ne
  2. [Results: Impact of Green AI Technological Domains on Knowledge Flows and Market Value; Methods: Impact analysis with for] Figure 5 compares raw average forward citations across domains whose grant-year distributions differ sharply. Older patents have had more time to accumulate citations, and no truncation correction, citation-lag adjustment, or grant-year fixed effects are used. The statement that Clinical Microbiome & Therapeutics is a 'clear outlier' and that Meteorological Radar has 'high citations but low economic return' is therefore not established. At a minimum, report citation counts by grant-year cohort and re-estimate domain differences with year fixed effects or a citation-lag model. Adding bootstrap confidence intervals would also show whether the domain-level rankings in Figure 5 are distinguishable from noise.
  3. [Results: Impact of Green AI Technological Domains on Knowledge Flows and Market Value; Methods: Impact analysis with for] The market-value analysis covers only 42.2% of the patents. The assertion that 'there is no reason to believe that the value of patents held by unlisted firms should behave any differently at the micro- and macro-domains level' is an untested missing-at-random assumption. Listing propensity is likely correlated with domain (e.g., automotive and electronics incumbents vs. agricultural water startups), which would bias the domain-level market-value rankings in Figure 5. Please compare matched and unmatched patents on observable characteristics (grant year, CPC subclasses, forward citations, assignee type) and provide a sensitivity analysis, such as Lee bounds or an inverse-probability weighting exercise, or explicitly relabel these rankings as exploratory.
  4. [Results: Technological (Macro-)Domains and Their Temporal Dynamics, Figure 4] The 'shift' from legacy to emerging domains is inferred from absolute patent counts, not relative shares. Total Green AI patenting grew by orders of magnitude over the sample period; a flat or declining absolute count for Combustion Engine Control is not the same as a structural shift in the composition of innovation. To support the headline claim, the paper should plot annual topic shares (or normalized counts) in addition to the 3-year rolling averages, and report the relative growth rates across domains.
minor comments (6)
  1. [Table 3] The column header 'prob' suggests these are probabilities, but the values are BERTopic c-TF-IDF scores. Please relabel and explain the metric.
  2. [Figure 5] The x-axis tick labels (5.0, 7.5, 10.0, 17.5, 20.0, 22.5) show an irregular gap; this is likely a plotting artifact or a typo. Please correct.
  3. [Discussion] There is a typo: 'hotsposts' should be 'hotspots'.
  4. [Conclusions] The domain name 'Meteorological Radar & Weather Management' in the Conclusions differs from 'Meteorological Radar & Weather Forecasting' used elsewhere. Please standardize.
  5. [Methods: Descriptive Analysis, Eq. (1)] The Gini coefficient formula contains LaTeX formatting artifacts in the displayed equation; please typeset it cleanly.
  6. [Methods: Hyperparameter Tuning and Model Selection] The 12.1% of documents treated as BERTopic outliers are excluded from all subsequent domain-level analyses. Please state whether the outlier rate varies systematically by grant year and whether the temporal trends in Figure 4 are robust to reassigning outliers or to alternative outlier thresholds.

Circularity Check

0 steps flagged

No significant circularity: descriptive empirical study with external data inputs; the only self-citation is non-load-bearing.

full rationale

The paper's derivation chain is a descriptive empirical analysis. The Green AI dataset is constructed by intersecting external, independently published classification systems (CPC Y02/Y04S and USPTO's AIPD AI labels). BERTopic topics are fitted to patent abstracts, and the temporal 'shift' claim is a direct summary of topic membership counts over time, not a prediction derived from fitted parameters. No equation in the paper defines a target variable in terms of the model's outputs; impact measures use external forward citations and precomputed Kogan et al. values. The only self-citation is ref. 20 (Biggi et al. 2025, co-authored by A. Mina) used to support the literature-gap statement that prior Green AI research relies on CPC rather than topic modeling; this is a literature claim and does not feed into any quantitative result, so it is not load-bearing. The assumption that the 42.2% of patents with market-value data are representative is a stated limitation, and the BERTopic temporal stability concern (single run, no cross-validation against CPC subclasses over time) is a correctness/robustness risk, not a circularity: the trends are not true by construction or by definition. The paper is self-contained against external benchmarks, so score is 0.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 0 invented entities

The paper is descriptive; its central claims rest on external taxonomies (CPC, AIPD), a topic-modeling pipeline with hyperparameters tuned on the same corpus, and proxy assumptions for impact and value. The most fragile input is the missing-at-random assumption for market values.

free parameters (3)
  • BERTopic hyperparameters (n_neighbors=30, min_dist=0.1, min_cluster_size=200, n_components=2) = n_neighbors=30, min_dist=0.1, min_cluster_size=200, n_components=2
    Selected by grid search over 54 combinations on the same corpus, optimizing UMass coherence and topic diversity. These choices determine the 16-topic structure and the domain taxonomy.
  • AIPD AI classification threshold = predict50_any_ai >= 0.5
    Imported from the external AIPD dataset, not fitted here, but the dataset boundary and all downstream results depend on this cutoff.
  • Number of topics (16) = 16
    Not directly chosen, but bounded by criterion C2 (5-30) and determined by the selected UMAP/HDBSCAN hyperparameters; the 16-topic solution was selected as best among grid results.
axioms (6)
  • domain assumption AIPD predict50_any_ai=1 correctly identifies AI inventions.
    Dataset construction section; the Green AI universe is defined by this classifier output.
  • domain assumption CPC Y02/Y04S tags correctly identify climate mitigation and adaptation technologies.
    Dataset construction section; relies on the EPO tagging scheme.
  • domain assumption BERTopic topics over patent abstracts correspond to meaningful technological domains.
    Topic modeling section; labels are manually assigned via keyword inspection.
  • domain assumption Forward citations measure technological impact.
    Impact analysis section; standard in the innovation literature.
  • domain assumption Kogan et al. market value estimates measure private economic value of patents.
    Impact analysis section; uses precomputed values from the Kogan et al. model.
  • ad hoc to paper Missing market values are missing at random across domains.
    Impact analysis section: 'there is no reason to believe that the value of patents held by unlisted firms should behave any differently', an explicitly stated and untested assumption.

pith-pipeline@v1.3.0-alltime-deepseek · 17217 in / 10881 out tokens · 97240 ms · 2026-08-04T18:05:10.879444+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of The anatomy of Green AI technologies: structure, evolution, and impact." pith.science (2026). https://pith.science/paper/M6UJDEEJ

@misc{pith2026250910109,
  author       = {Pith},
  title        = {Pith review of: The anatomy of Green AI technologies: structure, evolution, and impact},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/M6UJDEEJ}},
  note         = {Machine review of arXiv:2509.10109}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Artificial intelligence (AI) is a key enabler of innovation against climate change. In this study, we investigate the intersection of AI and climate adaptation and mitigation technologies through patent analyses of a novel dataset of approximately 63 000 Green AI patents. We analyze patenting trends, corporate ownership of the technology, the geographical distributions of patents, their impact on follow-on inventions and their market value. We use topic modeling (BERTopic) to identify 16 major technological domains, track their evolution over time, and identify their relative impact. We uncover a clear shift from legacy domains such as combustion engines technology to emerging areas like data processing, microgrids, and agricultural water management. We find evidence of growing concentration in corporate patenting against a rapidly increasing number of patenting firms. Looking at the technological and economic impact of patents, while some Green AI domains combine technological impact and market value, others reflect weaker private incentives for innovation, despite their relevance for climate adaptation and mitigation strategies. This is where policy intervention might be required to foster the generation and use of new Green AI applications.

Figures

Figures reproduced from arXiv: 2509.10109 by Andrea Mina, Andrea Vandin, Lorenzo Emer.

Figure 1
Figure 1. Figure 1: Development over time of the number of Green AI patents, both yearly and cumulatively, 1976-2023. semantic representations are then grouped using clustering algorithms, which organize similar documents into coherent topical clusters. To interpret each cluster, BERTopic applies Term Frequency–Inverse Document Frequency (TF-IDF), a technique that highlights the most informative and distinctive words within e… view at source ↗
Figure 2
Figure 2. Figure 2: (left) Country-level distribution of Green AI patents, using assignee location as a proxy. The United States is obscured in grey to better visualize activity in other countries. The heatmap reflects the relative share of non-US patents across countries. (right) US-level distribution of Green AI patents, using assignee location as a proxy. CPC Green Subclasses and AI Functional Categories We investigate the… view at source ↗
Figure 3
Figure 3. Figure 3: Projection in 2D-space of technological domains derived from BERTopic through dimensionality reduction (UMAP). Each grey bubble represents a single technological domain. The size of the bubbles reflects the domain importance in terms of number of patents. Each colored circle constitutes a manually labeled macro-domain, aggregating domains characterized by proximity in the Euclidean space. We display the fo… view at source ↗
Figure 4
Figure 4. Figure 4: Temporal evolution of patenting activity for each technological domain, displayed as individual line charts with a 3-year rolling average. Each panel illustrates the number of granted patents per year for a specific domain, capturing trends in innovation intensity from 1976 to 2023. The color of each line corresponds to the macro-domain to which the technological domain belongs as per [PITH_FULL_IMAGE:fig… view at source ↗
Figure 5
Figure 5. Figure 5: Scatter plot of average forward citations versus average market-value (deflated to 1982 million USD) per patent for sixteen Green AI technological domains, sized by the number of patents in each domain and colored by macro-domain. Marker size scales with topic patent volume (largest for Data Processing & Memory Management), while position reflects technological influence (horizontal) and private economic r… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

51 extracted references · 28 canonical work pages

  1. [2]

    F., Mendlik, T

    Maule, C. F., Mendlik, T. & Christensen, O. B. The effect of the pathway to a two degrees warmer world on the regional temperature change of europe.Clim. Serv.7, 3–11, DOI: 10.1016/j.cliser.2016.07.002 (2017)

  2. [3]

    S., Anjum, E., Francis, C.et al.Climate change, environmental disasters, and health inequities: The underlying role of structural inequalities.Curr

    Smith, G. S., Anjum, E., Francis, C.et al.Climate change, environmental disasters, and health inequities: The underlying role of structural inequalities.Curr. Environ. Heal. Reports9, 80–89, DOI: 10.1007/s40572-022-00336-w (2022)

  3. [4]

    Rogelj, J., Luderer, G., Pietzcker, R.et al.Energy system transformations for limiting end-of-century warming to below 1.5 °c.Nat. Clim. Chang.5, 519–527, DOI: 10.1038/nclimate2572 (2015)

  4. [5]

    & Haenlein, M

    Kaplan, A. & Haenlein, M. Siri, siri, in my hand: Who’s the fairest in the land? on the interpretations, illustrations, and implications of artificial intelligence.Bus. Horizons62, 15–25, DOI: 10.1016/j.bushor.2018.08.004 (2019)

  5. [6]

    & Floridi, L

    Cowls, J., Tsamados, A., Taddeo, M. & Floridi, L. The ai gambit: leveraging artificial intelligence to combat climate change—opportunities, challenges, and recommendations.AI & Soc.38, 283–307, DOI: 10.1007/s00146-021-01294-x (2023)

  6. [7]

    H., Jaffe, A

    Hall, B. H., Jaffe, A. B. & Trajtenberg, M. The nber patent citation data file: Lessons, insights and methodological tools (2001)

  7. [8]

    Benson, C. L. & Magee, C. L. Quantitative determination of technological improvement from patent data.PloS one10, e0121635, DOI: 10.1371/journal.pone.0121635 (2015)

  8. [9]

    USPTO Economic Working Paper 2024-4, USPTO (2024)

    Pairolero, N.et al.The artificial intelligence patent dataset (aipd) 2023 update. USPTO Economic Working Paper 2024-4, USPTO (2024). Available at https://www.uspto.gov/sites/default/files/documents/oce-aipd-2023.pdf

  9. [10]

    & Hassan, A

    Abdelrazek, A., Eid, Y ., Gawish, E., Medhat, W. & Hassan, A. Topic modeling algorithms and applications: A survey.Inf. Syst.112, 102131, DOI: 10.1016/j.is.2022.102131 (2023)

  10. [11]

    & Rost, K

    Momeni, A. & Rost, K. Identification and monitoring of possible disruptive technologies by patent-development paths and topic modeling.Technol. Forecast. Soc. Chang.104, 16–29, DOI: 10.1016/j.techfore.2015.12.003 (2016)

  11. [12]

    & Zellou, A

    Najmani, K., Ajallouda, L., Benlahmar, E., Sael, N. & Zellou, A. BERTopic and LDA, which topic modeling technique to extract relevant topics from videos in the context of massive open online courses (MOOCs)? In Motahhir, S. & Bossoufi, B. (eds.)Digital Technologies and Applications. ICDTA 2023, vol. 669 ofLecture Notes in Networks and Systems, DOI: 10.100...

  12. [13]

    In Lu, H

    Gan, L.et al.Experimental comparison of three topic modeling methods with lda, top2vec and bertopic. In Lu, H. & Cai, J. (eds.)Artificial Intelligence and Robotics. ISAIR 2023, vol. 1998 ofCommunications in Computer and Information Science, DOI: 10.1007/978-981-99-9109-9_37 (Springer, Singapore, 2024). 14/17

  13. [14]

    BERTopic: Neural topic modeling with a class-based tf-idf procedure.arXiv preprint arXiv:2203.05794 (2022)

    Grootendorst, M. BERTopic: Neural topic modeling with a class-based tf-idf procedure.arXiv preprint arXiv:2203.05794 (2022)

  14. [15]

    & Toutanova, K

    Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C. & Solorio, T. (eds.)Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), ...

  15. [16]

    T., Mirandola, P

    Guizzardi, S., Colangelo, M. T., Mirandola, P. & Galli, C. Modeling new trends in bone regeneration, using the BERTopic approach.Regen. Medicine18, 719–734, DOI: 10.2217/rme-2023-0096 (2023)

  16. [17]

    & Yun, S

    Yun, B., Yoon, J.-I., Bae, J. & Yun, S. Analysis of autonomous vehicles patent trends between korea and overseas using BERTopic. In2024 4th International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME), 1–6, DOI: 10.1109/ICECCME62383.2024.10796068 (IEEE, 2024)

  17. [18]

    Kim, K., Kogler, D. F. & Maliphol, S. Identifying interdisciplinary emergence in the science of science: combination of network analysis and bertopic.Humanit. Soc. Sci. Commun.11, 1–15, DOI: 10.1057/s41599-024-03044-y (2024)

  18. [19]

    & Huang, L

    Hou, K., Wu, M., Wu, W. & Huang, L. Topic mining and forecasting on patent map for gpu technology.The J. Supercomput. 81, 710, DOI: 10.1007/s11227-025-07215-9 (2025)

  19. [20]

    & Mina, A

    Biggi, G., Iori, M., Mazzei, J. & Mina, A. Green intelligence: The AI content of green technologies.Eurasian Bus. Rev. 1–38, DOI: 10.1007/s40821-024-00288-1 (2025)

  20. [21]

    A penny for your quotes: patent citations and the value of innovations.The Rand journal economics 172–187 (1990)

    Trajtenberg, M. A penny for your quotes: patent citations and the value of innovations.The Rand journal economics 172–187 (1990). 22.Hall, B. H., Jaffe, A. B. & Trajtenberg, M. Market value and patent citations: A first look (2000)

  21. [23]

    & Stoffman, N

    Kogan, L., Papanikolaou, D., Seru, A. & Stoffman, N. Technological innovation, resource allocation, and growth.The Q. J. Econ.132, 665–712, DOI: 10.1093/qje/qjw040 (2017)

  22. [24]

    Morris, J. J. & Alam, P. Value relevance and the dot-com bubble of the 1990s.The Q. Rev. Econ. Finance52, 243–255, DOI: 10.1016/j.qref.2012.04.001 (2012)

  23. [25]

    Teplykh, G. V . Innovations and productivity: the shift during the 2008 crisis.Ind. Innov.25, 53–83, DOI: 10.1080/ 13662716.2017.1286461 (2017)

  24. [26]

    & Park, W

    Hingley, P. & Park, W. G. Do business cycles affect patenting? evidence from european patent office filings.Technol. Forecast. Soc. Chang.116, 76–86, DOI: 10.1016/j.techfore.2016.11.003 (2017)

  25. [27]

    & Zingg, R

    Whalen, R. & Zingg, R. Innovating under uncertainty: The patent-eligibility of artificial intelligence after alice corp. v. cls bank international. In Langenfeld, J., Fagan, F. & Clark, S. (eds.)The Law and Economics of Privacy, Personal Data, Artificial Intelligence, and Incomplete Monitoring, vol. 30 ofResearch in Law and Economics, 59–81, DOI: 10.1108/...

  26. [28]

    & Xie, Y

    Hu, S., Xia, Q. & Xie, Y . Technological innovation under trade disputes: how does product market competition matter? Eur. J. Innov. Manag.28, 1202–1223, DOI: 10.1108/EJIM-04-2023-0318 (2025)

  27. [29]

    Farris, F. A. The gini index and measures of inequality.The Am. Math. Mon.117, 851–864, DOI: 10.4169/ 000298910X523344 (2010)

  28. [30]

    The origin and growth of industry clusters: The making of silicon valley and detroit.J

    Klepper, S. The origin and growth of industry clusters: The making of silicon valley and detroit.J. Urban Econ.67, 15–32, DOI: 10.1016/j.jue.2009.09.004 (2010)

  29. [31]

    & Wang, N

    Mattoon, R. & Wang, N. Industry clusters and economic development in the seventh district’s largest cities.Econ. Perspectives Q2, 52–66 (2014)

  30. [32]

    y02-y04s

    Angelucci, S., Hurtado-Albir, F. J. & V olpe, A. Supporting global initiatives on climate change: The epo’s “y02-y04s” tagging scheme.World Pat. Inf.54, S85–S92, DOI: 10.1016/j.wpi.2017.04.006 (2018)

  31. [33]

    Heal.4, e271–e279, DOI: 10.1016/S2542-5196(20)30121-2 (2020)

    Lenzen, M.et al.The environmental footprint of health care: a global assessment.The Lancet Planet. Heal.4, e271–e279, DOI: 10.1016/S2542-5196(20)30121-2 (2020)

  32. [34]

    Aus der Beek, T.et al.Pharmaceuticals in the environment—global occurrences and perspectives.Environ. Toxicol. Chem. 35, 823–835, DOI: 10.1002/etc.3339 (2016)

  33. [35]

    & Tietze, F

    Aristodemou, L. & Tietze, F. Citations as a measure of technological impact: A review of forward citation-based measures. World Pat. Inf.53, 39–44, DOI: 10.1016/j.wpi.2018.05.001 (2018). 15/17

  34. [36]

    Graham, S. J. H. & Mowery, D. C. Intellectual property protection in the us software industry. In Cohen, W. M. & Merrill, S. A. (eds.)Patents in the Knowledge-Based Economy, 219–231 (National Academies Press, Washington, DC, 2003)

  35. [37]

    A., Jones, C

    Toole, A. A., Jones, C. & Madhavan, S. Patentsview: An open data platform to advance science and technology policy. Tech. Rep., USPTO Economic Working Paper No. 2021-1 (2021). DOI: 10.2139/ssrn.3874213. Available at SSRN: https://ssrn.com/abstract=3874213

  36. [38]

    & Lerner, J

    Webb, M., Short, N., Bloom, N. & Lerner, J. Some facts of high-tech patenting. Working Paper 24793, National Bureau of Economic Research (2018). DOI: 10.3386/w24793

  37. [39]

    & Gurevych, I

    Reimers, N. & Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks (2019). https://www.sbert. net, 1908.10084

  38. [40]

    & Hsu, C.-C

    Wang, J. & Hsu, C.-C. A topic-based patent analytics approach for exploring technological trends in smart manufacturing. J. Manuf. Technol. Manag.32, 110–135, DOI: 10.1108/JMTM-03-2020-0106 (2021)

  39. [41]

    & Lee, J

    Kim, M. & Lee, J. What are the future trends in natural gas technology to address climate change? patent analysis through large language model.Energy312, 133644, DOI: 10.1016/j.energy.2024.133644 (2024)

  40. [42]

    & Großberger, L

    McInnes, L., Healy, J., Saul, N. & Großberger, L. Umap: Uniform manifold approximation and projection.J. Open Source Softw.3, 861, DOI: 10.21105/joss.00861 (2018)

  41. [43]

    & Healy, J

    McInnes, L. & Healy, J. Accelerated hierarchical density based clustering. In2017 IEEE International Conference on Data Mining Workshops (ICDMW), 33–42, DOI: 10.1109/ICDMW.2017.12 (2017)

  42. [44]

    & Allen, C

    Murdock, J. & Allen, C. Visualization techniques for topic model checking. InProceedings of the AAAI Conference on Artificial Intelligence, vol. 29, DOI: 10.1609/aaai.v29i1.9268 (2015)

  43. [45]

    & Singh, S

    Kumar, A., Karamchandani, A. & Singh, S. Topic modeling of neuropsychiatric diseases related to gut microbiota and gut brain axis using artificial intelligence based bertopic model on pubmed abstracts.Neurosci. Informatics4, 100175, DOI: 10.1016/j.neuri.2024.100175 (2024)

  44. [46]

    & Rodríguez-Rodríguez, L

    Madrid-García, A., Freites-Núñez, D., Merino-Barbancho, B., Pérez Sancristobal, I. & Rodríguez-Rodríguez, L. Map- ping two decades of research in rheumatology-specific journals: a topic modeling analysis with bertopic.Ther. Adv. Musculoskelet. Dis.16, 1759720X241308037, DOI: 10.1177/1759720X241308037 (2024)

  45. [47]

    & Jose, J

    Borˇcin, M. & Jose, J. M. Optimizing bertopic: Analysis and reproducibility study of parameter influences on topic modeling. In Goharian, N.et al.(eds.)Advances in Information Retrieval. ECIR 2024, vol. 14611 ofLecture Notes in Computer Science, 179–193, DOI: 10.1007/978-3-031-56066-8_14 (Springer, Cham, 2024)

  46. [48]

    & Emmert-Streib, F

    Farea, A., Tripathi, S., Glazko, G. & Emmert-Streib, F. Investigating the optimal number of topics by advanced text-mining techniques: Sustainable energy research.Eng. Appl. Artif. Intell.136, 108877, DOI: 10.1016/j.engappai.2024.108877 (2024)

  47. [49]

    & McCallum, A

    Mimno, D., Wallach, H., Talley, E., Leenders, M. & McCallum, A. Optimizing semantic coherence in topic models. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, 262–272 (Association for Computational Linguistics, 2011)

  48. [50]

    Topic modeling with gensim

    One-OffCoder. Topic modeling with gensim. https://datascience.oneoffcoder.com/topic-modeling-gensim.html (2020). Accessed: 2025-05-16

  49. [51]

    & Lee, J

    Ko, J. & Lee, J. Discovering research areas from patents: A case study in autonomous vehicles industry. In2021 IEEE International Conference on Big Data and Smart Computing (BigComp), 203–209, DOI: 10.1109/BigComp51126.2021. 00046 (IEEE, 2021)

  50. [52]

    M., Anggai, S., Tukiyat, Musyafa, A

    Zain, R. M., Anggai, S., Tukiyat, Musyafa, A. & Waskita, A. A. Revealing a country’s government discourse through bert- based topic modeling in the us presidential speeches. In2024 International Conference on Computer, Control, Informatics and its Applications (IC3INA), 191–196, DOI: 10.1109/IC3INA64086.2024.10732578 (IEEE, Bandung, Indonesia, 2024)

  51. [53]

    & Vandin, A

    Emer, L., Mina, A. & Vandin, A. The anatomy of green ai technologies: structure, evolution, and impact - dataset and replicability material, DOI: https://doi.org/10.5281/zenodo.15545360 (2025). Acknowledgements The work has been partially supported by project SMaRT COnSTRUCT (CUP J53C24001460006), in the context of FAIR (PE0000013, CUP B53C22003630006) un...