Pith. sign in

REVIEW 4 major objections 6 minor 85 references

MOFA: Discovering Materials for Carbon Capture with a GenAI- and Simulation-Based Workflow

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A generative-AI plus simulation workflow produced MOFs with CO2 adsorption capacities ranking among the best in a 137,652-structure dataset, in a single 450-node, 3-hour run.

desk verdict The systems-scale workflow results are credible, but the headline hMOF ranking claim is not established because it compares MOFA's GCMC protocol to published hMOF values without recomputing them under a common protocol. read the letter →

arxiv 2501.10651 v1 pith:I6CBQTKW submitted 2025-01-18 cs.DC cond-mat.mtrl-scics.LG

classification cs.DCcond-mat.mtrl-scics.LG
keywords generativeAImetal-organicframeworkscarboncapturehigh-performancecomputingonlinelearningdiffusionmodelCO2adsorptionheterogeneousworkflow
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MOFA is an open-source workflow that combines a generative diffusion model with atomistic simulation screening to discover metal-organic frameworks (MOFs) for carbon capture. In a single 450-node, 3-hour run it produced 114 novel MOFs per hour, including one with a CO2 uptake of 4.05 mol/kg at 0.1 bar, a value that ranks among the top five in the hypothetical-MOF dataset, and ten more in the top 10% of the structurally similar subset. The paper also reports that the rate of generating stable, high-performing MOFs grows approximately linearly with the number of compute nodes, and that periodically retraining the generative model on the best MOFs found so far substantially increases discovery rates. If these results hold, MOFA offers a way to navigate the enormous space of possible MOFs with far fewer guesses than brute-force enumeration.

What carries the argument

The mechanism that carries the argument is the online-learning generation-screening loop. A diffusion model fine-tuned on existing high-performing MOF linkers produces candidate linkers; the candidates are assembled with metal nodes, then passed through stability, density-functional, and Monte Carlo screens; the survivors are used to retrain the model. The retraining step is what makes the search adaptive: the paper reports that retraining increased the number of stable MOFs found at 90 minutes from 133 to 313 on 32 nodes and from 393 to 641 on 64 nodes. The workflow's scheduling layer keeps all stages running concurrently with low latencies, which is why throughput scales approximately linearly with node count.

What would settle it

Recompute the CO2 adsorption capacity at 0.1 bar and 300 K for the generated MOFs and for a matched random sample of hMOF structures under identical force-field, charge, and grand-canonical Monte Carlo settings; if the 4.05 mol/kg MOF no longer ranks in the top five of the full or similar-subset comparison, the headline ranking claim is not supported.

Watch

Extended reading notes

Core claim

The central claim is that a tight loop between generative AI and simulation can rapidly produce rare, high-quality MOFs for carbon capture. MOFA fine-tunes an E(3)-equivariant diffusion model for molecular linkers on data from the hypothetical-MOF (hMOF) dataset, then runs the generated linkers through a funnel of increasingly expensive screens: molecular dynamics for stability, density-functional-theory optimization, charge assignment, and grand-canonical Monte Carlo for CO2 adsorption. MOFs that pass these screens are fed back to retrain the generator, so later iterations produce linkers biased toward stable, high-capacity structures. The paper's headline result is that one 450-node, 3-hour run generated 114 MOFs per hour, with one MOF reaching 4.05 mol/kg CO2 capacity at 0.1 bar (in the top five of the hMOF dataset) and ten others in the top 10% of a 4,547-MOF structurally similar subset, while worker utilization stayed above 99%.

Load-bearing premise

The ranking claims depend on the assumption that the CO2 capacities computed in this work are directly comparable to the published hMOF dataset values, but the paper does not report recomputing the hMOF capacities with the same force-field and partial-charge settings.

Editorial extensions

If this is right

  • If the results generalize, a single large HPC run can search MOF chemical space fast enough to find top-ranking carbon-capture candidates in hours rather than via exhaustive enumeration.
  • Linear scaling with node count implies that adding compute nodes directly increases the rate of stable, high-adsorption MOF production, so larger machines accelerate discovery without changing the algorithm.
  • The measured benefit of retraining on intermediate results suggests that online learning is an effective strategy for inverse material design, not just for MOFs but for any property computable by simulation.
  • The modular architecture allows swapping the target property and screening criteria, so the same workflow can be pointed at catalysis, hydrogen storage, or other MOF applications.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A decisive validation the paper does not report is a same-protocol recomputation of hMOF capacities; until that is done, the top-5 and top-10% ranks are conditional on force-field and charge-method equivalence.
  • The 0.1 bar, 300 K condition is relevant to post-combustion capture, but real flue gas contains water and other gases; whether these MOFs retain capacity and stability under humid mixed-gas conditions is untested, so the practical carbon-capture claim is not yet established.
  • The linear-scaling result was measured on one machine up to 450 nodes; a stronger claim would require testing beyond that size or on different interconnect topologies, where communication or scheduling bottlenecks could appear.
  • The workflow's output is a set of computationally promising structures; the paper lists robotic synthesis as future work, so the near-term contribution is prescreening candidates for synthesis, not a guarantee that they can be made.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents MOFA, an open-source workflow that couples a diffusion-based generative model (MOFLinker) for MOF linker generation with a multi-stage screening pipeline (RDKit/OpenBabel assembly checks, LAMMPS stability simulations, CP2K/DDEC6 charge assignment, and RASPA GCMC adsorption estimates) in an online learning loop, orchestrated by Colmena and Parsl. The main reported results are: (i) on a 450-node, 3-hour Polaris run, MOFA generated about 114 MOFs per hour and produced one MOF with a CO2 capacity claimed to be in the top five and ten MOFs in the top 10% of a 4,547-MOF structurally similar subset of the hMOF dataset; (ii) worker utilization exceeds 99%; and (iii) throughput and stable-MOF discovery scale approximately linearly with node count.

Significance. The systems contributions are real and well presented: the integration of generator tasks into Colmena, the use of ProxyStore to decouple control and data transfer, the resource-allocation strategies, and the careful latency measurements provide a useful blueprint for GenAI-plus-simulation workflows on HPC systems. The scaling data in Figures 5-7 and the utilization analysis are internally consistent and appear reliable. The scientific discovery claim, if it survives a same-protocol comparison to hMOF, would be interesting, but it is currently not established because the ranking rests on an undocumented cross-protocol comparison and an undefined comparison subset. The paper would still be a solid systems contribution even if the discovery claim is softened.

major comments (4)
  1. [Section V-D; contribution 2 in Section I] The 'top five' and 'top 10%' assertions compare MOFA's GCMC results (RASPA with UFF4MOF Lennard-Jones parameters, the RASPA default CO2 model, DDEC6 partial charges, rigid framework, 300 K) against published hMOF dataset values without establishing that the two protocols are equivalent. The paper does not state that the hMOF reference capacities were recomputed with the MOFA protocol, nor does it report the force field, charge model, and CO2 model used in the original hMOF screening. Because CO2 uptake at 0.1 bar is strongly influenced by electrostatics, differences in partial charges alone could reorder the ranks substantially. Please recompute the hMOF subset under the identical GCMC protocol or provide a quantitative validation that the published values are directly comparable.
  2. [Section V-D; contribution 2 in Section I] The '4,547-MOF structurally similar subset' is not defined in the paper. The paper does not explain how this subset was selected, what 'structurally similar' means, or which features were used. Without this description, the top-five and top-10% statements cannot be reproduced, and the choice of subset could materially bias the percentile claims. In addition, the abstract's 'ranking among the top 10 in the hypothetical MOF (hMOF) dataset' is not what the body shows: the body reports one top-five result within the subset and ten results in the top 10% of the subset.
  3. [Sections V-C and V-D; Figures 5 and 7; retraining comparison] All scientific-output numbers come from single runs at each node count, and the retraining/no-retraining comparison reports no replication or error bars. For example, the statement that retraining increases the number of stable MOFs at 90 minutes from 133 to 313 on 32 nodes and from 393 to 641 on 64 nodes is presented without any measure of run-to-run variability, even though the generative and screening processes are stochastic. The linear-scaling conclusion would also be more robust with repeated runs or an uncertainty analysis. Please add replication or quantify the expected variability.
  4. [Section III-B, 'Estimate adsorption'] The GCMC simulations are described only by the pressure and temperature (0.1 bar and 300 K). The number of cycles, number of equilibration steps, number of independent GCMC runs, and the statistical uncertainty of the reported capacities (including the 4.05 mol/kg value) are missing. A single GCMC value without an error estimate cannot support a rank claim, even after protocol comparability is established.
minor comments (6)
  1. [Abstract] The statement 'CO2 adsorption capacities ranking among the top 10 in the hypothetical MOF (hMOF) dataset' is not supported by the body text; please revise to match the actual 'top five of a 4,547-MOF subset' and 'top 10% of the subset' results.
  2. [Table I] The 'Remain (%)' column reports average values but no variability or run source; please state whether these are from the 450-node run and, if so, over which time window.
  3. [Section V-C vs. Section III-C] The strain thresholds are inconsistent across the text: stable MOFs are defined as having '<10% chemical strain' in Section V-C, while Section III-C uses a 25% lattice-strain threshold for retraining triggers; please clarify the distinction or reconcile the definitions.
  4. [Section V-C] The sentence 'We attribute the modest increase over time in the rate at which stable MOFs are generated to repeated retraining' is an attribution; the no-retraining runs described later do control for this, but the attribution should be explicitly tied to that comparison.
  5. [Section III-B] The relationship between the LLST eigenvalue metric and the later term 'chemical strain' is never defined; please specify what '<10% chemical strain' means operationally.
  6. [Section VII and footnotes] There are minor presentation issues: 'in-silica' should be 'in silico' in Section VII, and the repository URL is redacted and should be restored in the camera-ready version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the central discovery claim is an external benchmark, and the active-learning loop is a search strategy, not a derivation from its own outputs.

full rationale

MOFA's workflow is an online learning loop: linkers generated by a diffusion model are screened with RDKit, LAMMPS, CP2K, and RASPA, and the model is retrained on screened stable/high-capacity MOFs. The 'improvement from retraining' is measured by the same stability metric used to select training data, but that is a self-consistent search objective, not a logical derivation of the conclusion from itself; the comparison runs with retraining disabled provide an independent control. The headline material-science claim is a benchmark of MOFA-generated MOFs against the external hMOF dataset, not a fit to hMOF values. The only notable caveat is that MOFA's GCMC protocol (UFF4MOF, RASPA default CO2 model, DDEC6 charges) may differ from the protocol used to generate the published hMOF capacities, which would undermine the 'top 5' rank; that is a correctness/comparability risk, not a circularity. Self-citations (e.g., [11], [17]) are to prior software and design papers and are not load-bearing for the paper's central derivation. No equation or claimed prediction reduces by construction to its inputs.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The workflow depends on several domain assumptions about force fields, simulation protocols, and the comparability of GCMC results to the hMOF benchmark. The retraining and stability thresholds are hand-chosen policy parameters. No new physical entities, forces, or conserved quantities are introduced.

free parameters (3)
  • Lattice strain threshold for 'stable MOF' = 10% (Fig. 7); 25% for retraining eligibility
    MOFs are counted as stable if the maximum absolute eigenvalue of the Linear Lagrangian Strain Tensor is below 10% after LAMMPS; retraining uses <25%. These thresholds are hand-chosen and directly determine the reported stable MOF counts.
  • Retraining schedule thresholds = 64 stability calculations; 32-8192 linkers; switch to adsorption-based selection after 64 adsorption calculations
    The policy that triggers retraining and the training set selection rules are ad hoc choices that influence the measured benefit of learning.
  • hMOF comparison subset size = 4,547 MOFs ('structurally similar subset')
    The comparison set is a subset of the full 137,652-MOF hMOF dataset; the selection criteria for this subset are not specified, and the abstract's 'top 10' language does not match the 'top 10%' claim.
assumptions (4)
  • domain assumption UFF4MOF plus the RASPA default CO2 model yields accurate CO2 adsorption capacities for MOFs.
    The GCMC estimate of CO2 uptake uses UFF4MOF Lennard-Jones parameters for the framework and the default RASPA CO2 model, with rigid MOF structures. The accuracy of this protocol for ranking against hMOF is assumed, not benchmarked here.
  • domain assumption The hMOF dataset's published CO2 capacities are computed with a protocol comparable to MOFA's RASPA setup.
    Used in Section V-D to rank MOFA outputs against hMOF; no recomputation or protocol comparison is reported.
  • domain assumption LAMMPS with UFF4MOF can identify chemically stable MOFs via lattice strain.
    The workflow defines stability by the maximum eigenvalue of the LLST from a 2x2x2 supercell NpT simulation; the force field's reliability for this screening is assumed.
  • domain assumption DiffLinker's E(3)-equivariant diffusion model can be fine-tuned to generate linkers that assemble into valid MOFs.
    Section III-B relies on fine-tuning a drug-discovery linker model on hMOF fragments; the representational fit to MOF linkers is assumed and only indirectly validated by the assembly and screening filters.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MOFA: Discovering Materials for Carbon Capture with a GenAI- and Simulation-Based Workflow." pith.science (2026). https://pith.science/paper/I6CBQTKW

@misc{pith2026250110651,
  author       = {Pith},
  title        = {Pith review of: MOFA: Discovering Materials for Carbon Capture with a GenAI- and Simulation-Based Workflow},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/I6CBQTKW}},
  note         = {Machine review of arXiv:2501.10651}
}
abstract

We present MOFA, an open-source generative AI (GenAI) plus simulation workflow for high-throughput generation of metal-organic frameworks (MOFs) on large-scale high-performance computing (HPC) systems. MOFA addresses key challenges in integrating GPU-accelerated computing for GPU-intensive GenAI tasks, including distributed training and inference, alongside CPU- and GPU-optimized tasks for screening and filtering AI-generated MOFs using molecular dynamics, density functional theory, and Monte Carlo simulations. These heterogeneous tasks are unified within an online learning framework that optimizes the utilization of available CPU and GPU resources across HPC systems. Performance metrics from a 450-node (14,400 AMD Zen 3 CPUs + 1800 NVIDIA A100 GPUs) supercomputer run demonstrate that MOFA achieves high-throughput generation of novel MOF structures, with CO$_2$ adsorption capacities ranking among the top 10 in the hypothetical MOF (hMOF) dataset. Furthermore, the production of high-quality MOFs exhibits a linear relationship with the number of nodes utilized. The modular architecture of MOFA will facilitate its integration into other scientific applications that dynamically combine GenAI with large-scale simulations.

Figures

Figures reproduced from arXiv: 2501.10651 by the authors.

Figure 1
Figure 1. MOFA implements an online learning loop that refines a generative AI model, MOFLinker, using the MOFs it has generated. The initial steps in the workflow validate linker molecules produced by the generative model before using those that pass validation to assemble MOFs. New MOFs are placed in a LIFO queue, from which they are retrieved to be evaluated for stability, and the gas capacity of the most stable are furthe… view at source ↗
Figure 2
Figure 2. Task and resource allocation in the MOFA workflow. The top section shows the Colmena Thinker, containing seven agents (rounded-corner boxes), each corresponding to one of the seven tasks. The bottom section depicts five types of MOFA workers, each with a 32-core CPU and four GPUs, with distinct resource allocation schemata for different MOFA tasks. 128 256 450 # Nodes 98.5 99.0 99.5 100.0 Worker Util. 1-HourAverage … view at source ↗
Figure 3
Figure 3. Active time of compute nodes on Polaris, as measured by the average [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: Sustained throughput in tasks per hour for the four main workflow [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: Number of stable MOFs found over time for [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: The generated MOF with highest CO2 capacity (4.05 mol/kg at 0.1 bar) produced by a 450-node, 3-hour MOFA run on Polaris. Brown: carbon; red: oxygen; white: hydrogen; yellow: sulfur; grey (big): zinc; blue white (small): nitrogen [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: UMAP plot of the diversity of MOFA-generated linkers compared to linkers from the hMOF database (represented with an RDKit embedding). While some regions of chemical space overlap between hMOF and MOFA￾generated linkers, the latter explores structures and moieties that…
Figure 10
Figure 10. Figure 10: The empirical cumulative distribution of the stability of MOFs [PITH_FULL_IMAGE:figures/full_fig_p010_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 78 canonical work pages

  1. [1]

    Climate impact of increasing atmospheric carbon dioxide,

    J. Hansen, D. Johnson, A. Lacis, S. Lebedeff, P. Lee, D. Rind, and G. Russell, “Climate impact of increasing atmospheric carbon dioxide,” Science, vol. 213, pp. 957–966, Aug. 1981

  2. [2]

    Ultrahigh metal– organic framework loading and flexible nanofibrous membranes for efficient CO2 capture with long-term, ultrastable recyclability,

    Y . Zhang, Y . Zhang, X. Wang, J. Yu, and B. Ding, “Ultrahigh metal– organic framework loading and flexible nanofibrous membranes for efficient CO2 capture with long-term, ultrastable recyclability,” ACS Applied Materials & Interfaces, vol. 10, no. 40, pp. 34802–34810, 2018

  3. [3]

    Rapid and accurate machine learning recognition of high performing metal organic frameworks for CO2 capture,

    M. Fernandez, P. G. Boyd, T. D. Daff, M. Z. Aghaji, and T. K. Woo, “Rapid and accurate machine learning recognition of high performing metal organic frameworks for CO2 capture,” The Journal of Physical Chemistry Letters, vol. 5, no. 17, pp. 3056–3060, 2014

  4. [4]

    Recent ad- vances in gas storage and separation using metal–organic frameworks,

    H. Li, K. Wang, Y . Sun, C. T. Lollar, J. Li, and H.-C. Zhou, “Recent ad- vances in gas storage and separation using metal–organic frameworks,” Materials Today, vol. 21, no. 2, pp. 108–121, 2018

  5. [5]

    Recent advances on preparation and environmental applications of MOF-derived carbons in catalysis,

    M. Hao, M. Qiu, H. Yang, B. Hu, and X. Wang, “Recent advances on preparation and environmental applications of MOF-derived carbons in catalysis,” Science of the Total Environment, vol. 760, p. 143333, 2021

  6. [6]

    Metal–organic frameworks for drug delivery: A design perspective,

    H. D. Lawson, S. P. Walton, and C. Chan, “Metal–organic frameworks for drug delivery: A design perspective,” ACS Applied Materials & Interfaces, vol. 13, no. 6, pp. 7004–7020, 2021

  7. [7]

    Lu- minescent sensors based on metal-organic frameworks,

    Y . Zhang, S. Yuan, G. Day, X. Wang, X. Yang, and H.-C. Zhou, “Lu- minescent sensors based on metal-organic frameworks,” Coordination Chemistry Reviews, vol. 354, 08 2017

  8. [8]

    High- resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High- resolution image synthesis with latent diffusion models,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 10684– 10695, 2022

Show all 85 references
  1. [9]

    Big-data science in porous materials: Materials genomics and machine learning,

    K. M. Jablonka, D. Ongari, S. M. Moosavi, and B. Smit, “Big-data science in porous materials: Materials genomics and machine learning,” Chemical Reviews, vol. 120, no. 16, pp. 8066–8129, 2020

  2. [10]

    ChatMOF: An artificial intelligence system for predicting and generating metal-organic frameworks using large language models,

    Y . Kang and J. Kim, “ChatMOF: An artificial intelligence system for predicting and generating metal-organic frameworks using large language models,” Nature Communications, vol. 15, no. 1, p. 4705, 2024

  3. [11]

    A generative artificial intelligence framework based on a molecular diffusion model for the design of metal-organic frameworks for carbon capture,

    H. Park, X. Yan, R. Zhu, E. A. Huerta, S. Chaudhuri, D. Cooper, I. Fos- ter, and E. Tajkhorshid, “A generative artificial intelligence framework based on a molecular diffusion model for the design of metal-organic frameworks for carbon capture,” Communications Chemistry , vol....

  4. [12]

    Understanding the diversity of the metal-organic framework ecosystem,

    S. M. Moosavi, A. Nandy, K. M. Jablonka, D. Ongari, J. P. Janet, P. G. Boyd, Y . Lee, B. Smit, and H. J. Kulik, “Understanding the diversity of the metal-organic framework ecosystem,” Nature Communications , vol. 11, no. 1, pp. 1–10, 2020

  5. [13]

    CP2K: An electronic structure and molecular dynamics software package - Quickstep: Efficient and accurate electronic structure calculations,

    T. D. K ¨uhne, M. Iannuzzi, M. D. Ben, V . V . Rybkin, P. Seewald, F. Stein, T. Laino, R. Z. Khaliullin, O. Sch ¨utt, F. Schiffmann, D. Golze, J. Wilhelm, S. Chulkov, M. H. Bani-Hashemian, V . Weber, U. Borˇstnik, M. Taillefumier, A. S. Jakobovits, A. Lazzaro, H. Pabst, T. M ¨...

  6. [14]

    LAMMPS - A flexible simulation tool for particle- based materials modeling at the atomic, meso, and continuum scales,

    A. P. Thompson, H. M. Aktulga, R. Berger, D. S. Bolintineanu, W. M. Brown, P. S. Crozier, P. J. in ’t Veld, A. Kohlmeyer, S. G. Moore, T. D. Nguyen, R. Shan, M. J. Stevens, J. Tranchida, C. Trott, and S. J. Plimpton, “LAMMPS - A flexible simulation tool for particle- based mat...

  7. [15]

    RASPA: Molecular simulation software for adsorption and diffusion in flexible nanoporous materials,

    D. Dubbeldam, S. Calero, D. E. Ellis, and R. Q. Snurr, “RASPA: Molecular simulation software for adsorption and diffusion in flexible nanoporous materials,” Molecular Simulation, vol. 42, no. 2, pp. 81–101, 2016

  8. [16]

    Parsl: Pervasive Parallel Programming in Python,

    Y . Babuji, A. Woodard, Z. Li, D. S. Katz, B. Clifford, R. Kumar, L. Lacinski, R. Chard, J. Wozniak, I. Foster, M. Wilde, and K. Chard, “Parsl: Pervasive Parallel Programming in Python,” in 28th ACM In- ternational Symposium on High-Performance Parallel and Distributed Computi...

  9. [17]

    Colmena: Scalable machine-learning-based steering of ensemble sim- ulations for high performance computing,

    L. Ward, G. Sivaraman, J. Pauloski, Y . Babuji, R. Chard, N. Dandu, P. C. Redfern, R. S. Assary, K. Chard, L. A. Curtiss, R. Thakur, and I. Foster, “Colmena: Scalable machine-learning-based steering of ensemble sim- ulations for high performance computing,” in IEEE/ACM Worksho...

  10. [18]

    Structure–property relationships of porous materials for carbon dioxide separation and capture,

    C. E. Wilmer, O. K. Farha, Y .-S. Bae, J. T. Hupp, and R. Q. Snurr, “Structure–property relationships of porous materials for carbon dioxide separation and capture,” Energy & Environmental Science, vol. 5, no. 12, p. 9849, 2012

  11. [19]

    State of the art and prospects in metal–organic framework (MOF)-based and MOF-derived nanocatalysis,

    Q. Wang and D. Astruc, “State of the art and prospects in metal–organic framework (MOF)-based and MOF-derived nanocatalysis,” Chemical reviews, vol. 120, no. 2, pp. 1438–1511, 2019

  12. [20]

    Stability of metal-organic frameworks: Recent advances and future trends,

    L. P. L. Mosca, A. B. Gapan, R. A. Angeles, and E. C. R. Lopez, “Stability of metal-organic frameworks: Recent advances and future trends,” in 4th International Electronic Conference on Applied Sciences , p. 146, MDPI, Nov. 2023

  13. [21]

    Preparation, clathration ability, and catalysis of a two-dimensional square network material composed of cadmium (II) and 4, 4’-bipyridine,

    M. Fujita, Y . J. Kwon, S. Washizu, and K. Ogura, “Preparation, clathration ability, and catalysis of a two-dimensional square network material composed of cadmium (II) and 4, 4’-bipyridine,” Journal of the American Chemical Society , vol. 116, no. 3, pp. 1151–1152, 1994

  14. [22]

    Engineering metal organic frameworks for heterogeneous catalysis,

    A. Corma, H. Garcia, and F. Llabr ´es i Xamena, “Engineering metal organic frameworks for heterogeneous catalysis,” Chemical Reviews , vol. 110, no. 8, pp. 4606–4655, 2010

  15. [23]

    Metal–organic frameworks meet metal nanoparticles: Synergistic effect for enhanced catalysis,

    Q. Yang, Q. Xu, and H.-L. Jiang, “Metal–organic frameworks meet metal nanoparticles: Synergistic effect for enhanced catalysis,” Chemical Society Reviews, vol. 46, no. 15, pp. 4774–4808, 2017

  16. [24]

    Metal–organic framework-derived porous materials for catalysis,

    Y .-Z. Chen, R. Zhang, L. Jiao, and H.-L. Jiang, “Metal–organic framework-derived porous materials for catalysis,” Coordination Chem- istry Reviews, vol. 362, pp. 1–23, 2018

  17. [25]

    Large-scale screening of hypothetical metal–organic frameworks,

    C. E. Wilmer, M. Leaf, C. Y . Lee, O. K. Farha, B. G. Hauser, J. T. Hupp, and R. Q. Snurr, “Large-scale screening of hypothetical metal–organic frameworks,” Nature Chemistry, vol. 4, no. 2, pp. 83–89, 2012

  18. [26]

    Computational screening of metal–organic frameworks for membrane-based CO2/N2/H2O separations: Best ma- terials for flue gas separation,

    H. Daglar and S. Keskin, “Computational screening of metal–organic frameworks for membrane-based CO2/N2/H2O separations: Best ma- terials for flue gas separation,” The Journal of Physical Chemistry C , vol. 122, no. 30, pp. 17347–17357, 2018

  19. [27]

    Geometrical properties can predict CO2 and N2 adsorption performance of metal–organic frameworks (MOFs) at low pressure,

    M. Fernandez and A. S. Barnard, “Geometrical properties can predict CO2 and N2 adsorption performance of metal–organic frameworks (MOFs) at low pressure,” ACS Combinatorial Science , vol. 18, no. 5, pp. 243–252, 2016

  20. [28]

    Sequential design of adsorption simulations in metal–organic frameworks,

    K. Mukherjee, A. W. Dowling, and Y . J. Col ´on, “Sequential design of adsorption simulations in metal–organic frameworks,” Molecular Systems Design & Engineering , vol. 7, no. 3, pp. 248–259, 2022

  21. [29]

    Generative AI for designing and validating easily synthesizable and structurally novel antibiotics,

    K. Swanson, G. Liu, D. B. Catacutan, A. Arnold, J. Zou, and J. M. Stokes, “Generative AI for designing and validating easily synthesizable and structurally novel antibiotics,” Nature Machine Intelligence , vol. 6, no. 3, pp. 338–353, 2024

  22. [30]

    Efficient aerodynamic shape optimization with deep-learning-based geometric filtering,

    J. Li, M. Zhang, J. R. Martins, and C. Shu, “Efficient aerodynamic shape optimization with deep-learning-based geometric filtering,” AIAA Journal, vol. 58, no. 10, pp. 4243–4259, 2020

  23. [31]

    Virtual screening of inorganic materials synthesis parameters with deep learning,

    E. Kim, K. Huang, S. Jegelka, and E. Olivetti, “Virtual screening of inorganic materials synthesis parameters with deep learning,” npj Computational Materials, vol. 3, no. 1, p. 53, 2017

  24. [32]

    Equivariant 3D-conditional diffusion models for molecular linker design,

    I. Igashov, H. St ¨ark, C. Vignac, A. Schneuing, V . G. Satorras, P. Frossard, M. Welling, M. Bronstein, and B. Correia, “Equivariant 3D-conditional diffusion models for molecular linker design,” Nature Machine Intelli- gence, pp. 1–11, 2024

  25. [33]

    MOFDiff: Coarse- grained diffusion for metal-organic framework design,

    X. Fu, T. Xie, A. S. Rosen, T. Jaakkola, and J. Smith, “MOFDiff: Coarse- grained diffusion for metal-organic framework design,” arXiv preprint arXiv:2310.10732, 2023

  26. [34]

    Inverse design of nanoporous crystalline reticular materials with deep generative models,

    Z. Yao, B. S ´anchez-Lengeling, N. S. Bobbitt, B. J. Bucior, S. G. H. Kumar, S. P. Collins, T. Burns, T. K. Woo, O. K. Farha, R. Q. Snurr, and A. Aspuru-Guzik, “Inverse design of nanoporous crystalline reticular materials with deep generative models,” Nature Machine Intelligen...

  27. [35]

    Learning everywhere: A taxonomy for the in- tegration of machine learning and simulations,

    G. Fox and S. Jha, “Learning everywhere: A taxonomy for the in- tegration of machine learning and simulations,” in 15th International Conference on eScience , pp. 439–448, IEEE, 2019

  28. [36]

    Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster,

    N. Dey, G. Gosal, Zhiming, Chen, H. Khachane, W. Marshall, R. Pathria, M. Tom, and J. Hestness, “Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster,” 2023

  29. [37]

    Dask: Parallel computation with blocked algorithms and task scheduling,

    M. Rocklin, “Dask: Parallel computation with blocked algorithms and task scheduling,” in 14th Python in Science Conference, vol. 130, p. 136, 2015

  30. [38]

    FireWorks: A dynamic workflow system designed for high-throughput applications,

    A. Jain, S. P. Ong, W. Chen, B. Medasani, X. Qu, M. Kocher, M. Brafman, G. Petretto, G.-M. Rignanese, G. Hautier, D. Gunter, and K. A. Persson, “FireWorks: A dynamic workflow system designed for high-throughput applications,” Concurrency and Computation: Practice and Experienc...

  31. [39]

    Pegasus, a workflow management system for science automation,

    E. Deelman, K. Vahi, G. Juve, M. Rynge, S. Callaghan, P. J. Maechling, R. Mayani, W. Chen, R. Ferreira da Silva, M. Livny, and K. Wenger, “Pegasus, a workflow management system for science automation,” Future Generation Computer Systems , vol. 46, pp. 17–35, 2015

  32. [40]

    Swift: A language for distributed parallel scripting,

    M. Wilde, M. Hategan, J. M. Wozniak, B. Clifford, D. S. Katz, and I. Foster, “Swift: A language for distributed parallel scripting,” Parallel Computing, vol. 37, no. 9, pp. 633–652, 2011

  33. [41]

    Ray: A dis- tributed framework for emerging AI applications,

    P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, and I. Stoica, “Ray: A dis- tributed framework for emerging AI applications,” in 13th USENIX Con- ference on Operating Systems Design and Implementation , OSDI’18, (USA)...

  34. [42]

    TaskVine: Managing in-cluster storage for high-throughput data intensive workflows,

    B. Sly-Delgado, T. S. Phung, C. Thomas, D. Simonetti, A. Hennessee, B. Tovar, and D. Thain, “TaskVine: Managing in-cluster storage for high-throughput data intensive workflows,” in SC ’23 Workshops of The International Conference on High Performance Computing, Network, Storage...

  35. [43]

    AWS Lambda

    “AWS Lambda.” https://aws.amazon.com/lambda. Accessed Jan 2023

  36. [44]

    Serverless execution of scientific workflows: Experiments with Hyperflow, AWS Lambda and Google Cloud Functions,

    M. Malawski, A. Gajek, A. Zima, B. Balis, and K. Figiela, “Serverless execution of scientific workflows: Experiments with Hyperflow, AWS Lambda and Google Cloud Functions,” Future Generation Computer Systems, vol. 110, pp. 502–514, 2020

  37. [45]

    FuncX: A federated function serving fabric for science,

    R. Chard, Y . Babuji, Z. Li, T. Skluzacek, A. Woodard, B. Blaiszik, I. Foster, and K. Chard, “FuncX: A federated function serving fabric for science,” in 29th International Symposium on High-performance Parallel and Distributed Computing , pp. 65–76, 2020

  38. [46]

    Exaworks: Workflows for exascale,

    A. Al-Saadi, D. H. Ahn, Y . Babuji, K. Chard, J. Corbett, M. Hategan, S. Herbein, S. Jha, D. Laney, A. Merzky, et al., “Exaworks: Workflows for exascale,” in IEEE Workshop on Workflows in Support of Large-Scale Science, pp. 50–57, IEEE, 2021

  39. [47]

    High throughput training of deep surrogates from large ensemble runs,

    L. T. Meyer, M. Schouler, R. A. Caulk, A. Rib ´es, and B. Raffin, “High throughput training of deep surrogates from large ensemble runs,” in International Conference for High Performance Computing, Networking, Storage and Analysis , pp. 1–16, 2023

  40. [48]

    GenSLMs: Genome- scale language models reveal SARS-CoV-2 evolutionary dynamics,

    M. Zvyagin, A. Brace, K. Hippe, Y . Deng, B. Zhang, C. O. Bohorquez, A. Clyde, B. Kale, D. Perez-Rivera, H. Ma, C. M. Mann, M. Irvin, J. G. Pauloski, L. Ward, V . Hayot, M. Emani, S. Foreman, Z. Xie, D. Lin, M. Shukla, W. Nie, J. Romero, C. Dallago, A. Vahdat, C. Xiao, T. Gibb...

  41. [49]

    Composition-transferable machine learning potential for LiCl-KCl molten salts validated by high-energy X- ray diffraction,

    J. Guo, L. Ward, Y . Babuji, N. Hoyt, M. Williamson, I. Foster, N. Jack- son, C. Benmore, and G. Sivaraman, “Composition-transferable machine learning potential for LiCl-KCl molten salts validated by high-energy X- ray diffraction,” Physical Review B , vol. 106, no. 1, p. 014209, 2022

  42. [50]

    Extreme scale survey simulation with Python workflows,

    A. S. Villarreal, Y . Babuji, T. Uram, D. S. Katz, K. Chard, and K. Heitmann, “Extreme scale survey simulation with Python workflows,” in IEEE 17th International Conference on eScience, pp. 206–214, IEEE, 2021

  43. [51]

    Inverse design of materials by multi-objective differential evolution,

    Y .-Y . Zhang, W. Gao, S. Chen, H. Xiang, and X.-G. Gong, “Inverse design of materials by multi-objective differential evolution,” Computa- tional Materials Science , vol. 98, pp. 51–55, 2015

  44. [52]

    Generative adversarial networks (GAN) based efficient sampling of chemical composition space for inverse design of inorganic materials,

    Y . Dan, Y . Zhao, X. Li, S. Li, M. Hu, and J. Hu, “Generative adversarial networks (GAN) based efficient sampling of chemical composition space for inverse design of inorganic materials,” npj Computational Materials, vol. 6, no. 84, 2020

  45. [53]

    Constrained crystals deep convolutional generative adversarial network for the inverse design of crystal struc- tures,

    T. Long, N. M. Fortunato, I. Opahle, Y . Zhang, I. Samathrakis, C. Shen, O. Gutfleisch, and H. Zhang, “Constrained crystals deep convolutional generative adversarial network for the inverse design of crystal struc- tures,” npj Computational Materials , vol. 7, no. 66, 2021

  46. [54]

    Inverse design of porous materials using artificial neural networks,

    B. Kim, S. Lee, and J. Kim, “Inverse design of porous materials using artificial neural networks,” Science Advances, vol. 6, no. 1, 2020

  47. [55]

    Optimal experimental design: Formulations and computations,

    X. Huan, J. Jagalur, and Y . Marzouk, “Optimal experimental design: Formulations and computations,” Acta Numerica, vol. 33, pp. 715–840, 2024

  48. [56]

    E(n) equivariant graph neural networks,

    V . G. Satorras, E. Hoogeboom, and M. Welling, “E(n) equivariant graph neural networks,” in International Conference on Machine Learning , pp. 9323–9332, 2021

  49. [57]

    E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials,

    S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, “E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials,” Nature communications, vol. 13, no. 1, p. 2453, 2022

  50. [58]

    GEOM, energy-annotated molecular conformations for property prediction and molecular genera- tion,

    S. Axelrod and R. G ´omez-Bombarelli, “GEOM, energy-annotated molecular conformations for property prediction and molecular genera- tion,” Scientific Data, vol. 9, no. 1, p. 185, 2022

  51. [59]

    Open Babel: An open chemical toolbox,

    N. M. O’Boyle, M. Banck, C. A. James, C. Morley, T. Vandermeersch, and G. R. Hutchison, “Open Babel: An open chemical toolbox,” Journal of Cheminformatics, vol. 3, p. 33, Dec. 2011

  52. [60]

    Merck molecular force field. II. MMFF94 van der Waals and electrostatic parameters for intermolecular interactions,

    T. Halgren, “Merck molecular force field. II. MMFF94 van der Waals and electrostatic parameters for intermolecular interactions,” J Comput Chem, vol. 17, pp. 520–552, 1996

  53. [61]

    Rdkit documentation,

    G. Landrum, “Rdkit documentation,” Release, vol. 1, no. 1-79, p. 4, 2013

  54. [62]

    The Reticular Chemistry Structure Resource (RCSR) Database of, and Symbols for, Crystal Nets,

    M. O’Keeffe, M. A. Peskov, S. J. Ramsden, and O. M. Yaghi, “The Reticular Chemistry Structure Resource (RCSR) Database of, and Symbols for, Crystal Nets,” Accounts of Chemical Research , vol. 41, pp. 1782–1789, Dec. 2008

  55. [63]

    OChemDb: The free on-line Open Chemistry Database portal for searching and analysing crystal structure information,

    A. Altomare, N. Corriero, C. Cuocci, A. Falcicchio, A. Moliterni, and R. Rizzi, “OChemDb: The free on-line Open Chemistry Database portal for searching and analysing crystal structure information,” Journal of Applied Crystallography, vol. 51, pp. 1229–1236, 2018

  56. [64]

    cif2lammps

    R. Anderson, “cif2lammps.” https://github.com/rytheranderson/ cif2lammps

  57. [65]

    Extension of the universal force field to metal–organic frameworks,

    M. A. Addicoat, N. Vankova, I. F. Akter, and T. Heine, “Extension of the universal force field to metal–organic frameworks,” Journal of Chemical Theory and Computation , vol. 10, no. 2, pp. 880–891, 2014

  58. [66]

    Extension of the universal force field for metal–organic frameworks,

    D. E. Coupry, M. A. Addicoat, and T. Heine, “Extension of the universal force field for metal–organic frameworks,” Journal of Chemical Theory and Computation, vol. 12, no. 10, pp. 5215–5225, 2016

  59. [67]

    Quickstep: Fast and accurate density functional calcu- lations using a mixed Gaussian and plane waves approach,

    J. VandeV ondele, M. Krack, F. Mohamed, M. Parrinello, T. Chassaing, and J. Hutter, “Quickstep: Fast and accurate density functional calcu- lations using a mixed Gaussian and plane waves approach,” Computer Physics Communications, vol. 167, no. 2, pp. 103–128, 2005

  60. [68]

    On the limited memory BFGS method for large scale optimization,

    D. C. Liu and J. Nocedal, “On the limited memory BFGS method for large scale optimization,” Mathematical programming, vol. 45, no. 1, pp. 503–528, 1989

  61. [69]

    Generalized gradient approximation made simple,

    J. P. Perdew, K. Burke, and M. Ernzerhof, “Generalized gradient approximation made simple,” Physical Review Letters, vol. 77, p. 3865, May 1996

  62. [70]

    Separable dual-space Gaussian pseudopotentials,

    S. Goedecker, M. Teter, and J. Hutter, “Separable dual-space Gaussian pseudopotentials,” Physical Review B, vol. 54, pp. 1703–1710, Jul 1996

  63. [71]

    Gaussian basis sets for accurate calcu- lations on molecular systems in gas and condensed phases,

    J. VandeV ondele and J. Hutter, “Gaussian basis sets for accurate calcu- lations on molecular systems in gas and condensed phases,” The Journal of Chemical Physics , vol. 127, p. 114105, Sept. 2007

  64. [72]

    A consistent and accu- rateab initioparametrization of density functional dispersion correction (dft-d) for the 94 elements h-pu,

    S. Grimme, J. Antony, S. Ehrlich, and H. Krieg, “A consistent and accu- rateab initioparametrization of density functional dispersion correction (dft-d) for the 94 elements h-pu,” The Journal of Chemical Physics , vol. 132, Apr. 2010

  65. [73]

    Introducing DDEC6 atomic population analysis: Part 1. Charge partitioning theory and methodology,

    T. A. Manz and N. G. Limasa, “Introducing DDEC6 atomic population analysis: Part 1. Charge partitioning theory and methodology,” RSC Advances, vol. 6, pp. 47771––47801, 2016

  66. [74]

    Introducing DDEC6 atomic population analysis: Part 2. Computed results for a wide range of periodic and nonperiodic materials,

    T. A. Manz and N. G. Limasa, “Introducing DDEC6 atomic population analysis: Part 2. Computed results for a wide range of periodic and nonperiodic materials,” RSC Advances, vol. 6, pp. 45727––45747, 2016

  67. [75]

    Cloud services enable efficient AI-guided simulation workflows across heterogeneous resources,

    L. Ward, J. G. Pauloski, V . Hayot-Sasson, R. Chard, Y . Babuji, G. Sivara- man, S. Choudhury, K. Chard, R. Thakur, and I. Foster, “Cloud services enable efficient AI-guided simulation workflows across heterogeneous resources,” in Heterogeneity in Computing Workshop , New York...

  68. [76]

    Employing artificial intelli- gence to steer exascale workflows with Colmena,

    L. Ward, J. G. Pauloski, V . Hayot-Sasson, Y . Babuji, A. Brace, R. Chard, K. Chard, R. Thakur, and I. Foster, “Employing artificial intelli- gence to steer exascale workflows with Colmena,” The International Journal of High Performance Computing Applications , vol. 0, no. 0, ...

  69. [77]

    Accelerating communications in federated applications with transparent object proxies,

    J. G. Pauloski, V . Hayot-Sasson, L. Ward, N. Hudson, C. Sabino, M. Baughman, K. Chard, and I. Foster, “Accelerating communications in federated applications with transparent object proxies,” in International Conference for High Performance Computing, Networking, Storage and A...

  70. [78]

    Object proxy patterns for accelerating distributed appli- cations,

    J. G. Pauloski, V . Hayot-Sasson, L. Ward, A. Brace, A. Bauer, K. Chard, and I. Foster, “Object proxy patterns for accelerating distributed appli- cations,” 2024. Preprint ArXiv:2407.01764

  71. [79]

    NVIDIA Multi Process Service

    “NVIDIA Multi Process Service.” https://docs.nvidia.com/deploy/mps/ index.html

  72. [80]

    CD-MOFs for CO2 capture and sep- aration: Current research and future outlook,

    E. C. R. Lopez and J. V . D. Perez, “CD-MOFs for CO2 capture and sep- aration: Current research and future outlook,” Engineering Proceedings, vol. 56, no. 1, 2023

  73. [81]

    Rational design of a low-cost, high-performance metal-organic framework for hydrogen storage and carbon capture,

    M. Witman, S. Ling, A. Gładysiak, K. Stylianou, B. Smit, B. Slater, and M. Haranczyk, “Rational design of a low-cost, high-performance metal-organic framework for hydrogen storage and carbon capture,” The Journal of Physical Chemistry C , vol. 121, 12 2016

  74. [82]

    Chapter 5 - Removal of toxic/radioactive metal ions by metal-organic framework-based materials,

    J. Li, W. Ye, and C. Chen, “Chapter 5 - Removal of toxic/radioactive metal ions by metal-organic framework-based materials,” in Emerging Natural and Tailored Nanomaterials for Radioactive Waste Treatment and Environmental Remediation (C. Chen, ed.), vol. 29 of Interface Scienc...

  75. [83]

    A review on metal-organic frameworks: Synthesis and applications,

    M. Safaei, M. M. Foroughi, N. Ebrahimpoor, S. Jahani, A. Omidi, and M. Khatami, “A review on metal-organic frameworks: Synthesis and applications,” TrAC Trends in Analytical Chemistry , vol. 118, pp. 401– 425, 2019

  76. [84]

    Advances and applications of metal-organic frameworks (MOFs) in emerging technologies: A comprehensive review,

    D. Li, A. Yadav, H. Zhou, K. Roy, P. Thanasekaran, and C. Lee, “Advances and applications of metal-organic frameworks (MOFs) in emerging technologies: A comprehensive review,” Global Challenges , vol. 8, no. 2, p. 2300244, 2024

  77. [85]

    The present state and challenges of active learning in drug discovery,

    L. Wang, Z. Zhou, X. Yang, S. Shi, X. Zeng, and D. Cao, “The present state and challenges of active learning in drug discovery,” Drug Discovery Today, p. 103985, 2024

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.