{"id":"b6d9adbf-f9ad-4b08-b9b5-8664fc45e424","arxiv_id":"2501.10651","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MOFA couples a fine-tuned diffusion model with LAMMPS, CP2K, and RASPA simulations in an online learning loop to generate stable MOFs with high CO2 adsorption, demonstrating near-linear scaling on up to 450 nodes.","lead":"MOFA is an open-source workflow that pairs a generative AI model with atomistic simulations to create and screen new metal-organic frameworks (MOFs) for carbon capture on supercomputers. In a three-hour run on 450 nodes it generated stable MOFs with CO2 uptakes that rank among the best in the hypothetical MOF dataset, at throughput that scales with system size.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The top-5 carbon-capture claim rests on an unverified comparison between MOFA's UFF4MOF/RASPA/DDEC6 GCMC protocol and the published hMOF reference values; if those protocols differ, the ranking is not meaningful.","rationale":"The reader's weakest assumption is exactly this protocol-comparability issue, and I agree. The central discovery claim—one MOFA-generated MOF ranking in the top five of hMOF—is an inter-dataset comparison; it is only valid if MOFA's GCMC protocol and the hMOF reference protocol produce commensurate capacities. The paper provides no evidence of that commensurability, and the strong electrostatic sensitivity of low-pressure CO2 uptake makes the concern concrete rather than hypothetical. This does not move the verdict because the reader already assigned CONDITIONAL; the condition should explicitly include recomputing reference values with the MOFA protocol. A secondary issue is that the abstract says 'top 10 in the hMOF dataset' while the body reports 'top 10% of a 4,547-MOF structurally similar subset'; this is an overclaim that should be corrected, but it is distinct from the protocol-comparability concern. I credit the paper for its careful throughput and latency measurements at multiple node counts, which make the linear-scaling systems claim credible independent of the chemistry ranking. The redacted repository URLs hinder reproducibility but are not a logical flaw in the argument. If the same-protocol recomputation preserves the reported ranks, the chemistry claim would be substantially stronger.","tokens_in":19556,"tokens_out":6011,"duration_ms":60269,"concrete_test":"Recompute CO2 uptake at 0.1 bar and 300 K for the 4,547 MOFs in the structurally similar hMOF subset using the exact MOFA pipeline (UFF4MOF, RASPA default CO2, DDEC6 charges, rigid framework), then re-rank the MOFA-generated MOFs against these recomputed values. If the 4.05 mol/kg structure remains in the top 5 (or the ten reported structures remain in the top 10%) of the recomputed subset, the ranking claim holds; if ranks drop materially, the comparison to the published hMOF values is invalid. A lighter first check is to compare the original hMOF dataset's documented GCMC parameters (force field, charges, CO2 model, fugacity/pressure treatment) with MOFA's and determine whether any differ.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section V-D reports one MOFA-generated MOF with 4.05 mol/kg CO2 at 0.1 bar as 'in the top five of the hMOF dataset,' and the contribution list narrows this to the 4,547-MOF structurally similar subset. Either way, the rank is a cross-dataset comparison: MOFA capacities are computed with RASPA using UFF4MOF Lennard-Jones parameters, the RASPA default CO2 model, DDEC6 partial charges from CP2K/Chargemol, and a rigid-framework GCMC at 300 K, while the published hMOF capacities were produced by the original hMOF screening protocol. The paper does not state that the hMOF reference values were recomputed with the MOFA protocol, nor does it document the force fields, charge assignment, and CO2 model used in the hMOF dataset. CO2 adsorption at 0.1 bar is strongly influenced by electrostatics, so differences in partial charges alone can shift ranks substantially. Without a same-protocol comparison, the 'top 5' and 'top 10%' claims are not established; the reported numbers are only as good as the hidden assumption that the two protocols are interchangeable. This is the load-bearing concern because it directly controls whether the central discovery claim is valid. The systems scaling result is independent and appears well supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents MOFA, an open-source workflow that couples a diffusion-based generative model (MOFLinker) for MOF linker generation with a multi-stage screening pipeline (RDKit/OpenBabel assembly checks, LAMMPS stability simulations, CP2K/DDEC6 charge assignment, and RASPA GCMC adsorption estimates) in an online learning loop, orchestrated by Colmena and Parsl. The main reported results are: (i) on a 450-node, 3-hour Polaris run, MOFA generated about 114 MOFs per hour and produced one MOF with a CO2 capacity claimed to be in the top five and ten MOFs in the top 10% of a 4,547-MOF structurally similar subset of the hMOF dataset; (ii) worker utilization exceeds 99%; and (iii) throughput and stable-MOF discovery scale approximately linearly with node count.","tokens_in":19784,"tokens_out":6682,"duration_ms":63508,"significance":"The systems contributions are real and well presented: the integration of generator tasks into Colmena, the use of ProxyStore to decouple control and data transfer, the resource-allocation strategies, and the careful latency measurements provide a useful blueprint for GenAI-plus-simulation workflows on HPC systems. The scaling data in Figures 5-7 and the utilization analysis are internally consistent and appear reliable. The scientific discovery claim, if it survives a same-protocol comparison to hMOF, would be interesting, but it is currently not established because the ranking rests on an undocumented cross-protocol comparison and an undefined comparison subset. The paper would still be a solid systems contribution even if the discovery claim is softened.","major_comments":[{"comment":"The 'top five' and 'top 10%' assertions compare MOFA's GCMC results (RASPA with UFF4MOF Lennard-Jones parameters, the RASPA default CO2 model, DDEC6 partial charges, rigid framework, 300 K) against published hMOF dataset values without establishing that the two protocols are equivalent. The paper does not state that the hMOF reference capacities were recomputed with the MOFA protocol, nor does it report the force field, charge model, and CO2 model used in the original hMOF screening. Because CO2 uptake at 0.1 bar is strongly influenced by electrostatics, differences in partial charges alone could reorder the ranks substantially. Please recompute the hMOF subset under the identical GCMC protocol or provide a quantitative validation that the published values are directly comparable.","section":"Section V-D; contribution 2 in Section I"},{"comment":"The '4,547-MOF structurally similar subset' is not defined in the paper. The paper does not explain how this subset was selected, what 'structurally similar' means, or which features were used. Without this description, the top-five and top-10% statements cannot be reproduced, and the choice of subset could materially bias the percentile claims. In addition, the abstract's 'ranking among the top 10 in the hypothetical MOF (hMOF) dataset' is not what the body shows: the body reports one top-five result within the subset and ten results in the top 10% of the subset.","section":"Section V-D; contribution 2 in Section I"},{"comment":"All scientific-output numbers come from single runs at each node count, and the retraining/no-retraining comparison reports no replication or error bars. For example, the statement that retraining increases the number of stable MOFs at 90 minutes from 133 to 313 on 32 nodes and from 393 to 641 on 64 nodes is presented without any measure of run-to-run variability, even though the generative and screening processes are stochastic. The linear-scaling conclusion would also be more robust with repeated runs or an uncertainty analysis. Please add replication or quantify the expected variability.","section":"Sections V-C and V-D; Figures 5 and 7; retraining comparison"},{"comment":"The GCMC simulations are described only by the pressure and temperature (0.1 bar and 300 K). The number of cycles, number of equilibration steps, number of independent GCMC runs, and the statistical uncertainty of the reported capacities (including the 4.05 mol/kg value) are missing. A single GCMC value without an error estimate cannot support a rank claim, even after protocol comparability is established.","section":"Section III-B, 'Estimate adsorption'"}],"minor_comments":[{"comment":"The statement 'CO2 adsorption capacities ranking among the top 10 in the hypothetical MOF (hMOF) dataset' is not supported by the body text; please revise to match the actual 'top five of a 4,547-MOF subset' and 'top 10% of the subset' results.","section":"Abstract"},{"comment":"The 'Remain (%)' column reports average values but no variability or run source; please state whether these are from the 450-node run and, if so, over which time window.","section":"Table I"},{"comment":"The strain thresholds are inconsistent across the text: stable MOFs are defined as having '<10% chemical strain' in Section V-C, while Section III-C uses a 25% lattice-strain threshold for retraining triggers; please clarify the distinction or reconcile the definitions.","section":"Section V-C vs. Section III-C"},{"comment":"The sentence 'We attribute the modest increase over time in the rate at which stable MOFs are generated to repeated retraining' is an attribution; the no-retraining runs described later do control for this, but the attribution should be explicitly tied to that comparison.","section":"Section V-C"},{"comment":"The relationship between the LLST eigenvalue metric and the later term 'chemical strain' is never defined; please specify what '<10% chemical strain' means operationally.","section":"Section III-B"},{"comment":"There are minor presentation issues: 'in-silica' should be 'in silico' in Section VII, and the repository URL is redacted and should be restored in the camera-ready version.","section":"Section VII and footnotes"}],"recommendation":"major_revision","confidential_remarks":"The systems claims in this paper are solid and appropriately evaluated, but the discovery claim (top-five/top-10% in hMOF) rests on a cross-protocol comparison that is not documented. I recommend requiring a same-protocol recomputation of the hMOF reference subset before acceptance, along with a correction of the abstract's overstatement. If the authors cannot provide the recomputation, the discovery claims should be substantially weakened and reframed as preliminary. The undefined 'structurally similar subset' should also be clarified, as it directly affects the reported ranks."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dave — quick take. This is a solid systems paper with one load-bearing chemistry claim that doesn't hold up as written. The genuinely new pieces are the online retraining loop around MOFLinker (a fine-tuned DiffLinker), the Colmena generator-task extension, and the 32-to-450-node scaling study. The scaling results look internally consistent: throughput for assembly, validation, optimization, and retraining rises about linearly with node count; worker utilization is above 99%; and the inter-stage latencies stay flat up to 450 nodes. The retraining ablation is also sensible, showing roughly 2–3x more stable MOFs with retraining than without. On the systems side, this deserves to be taken seriously.\n\nThe soft spot is the scientific-output claim. The abstract says MOFA-generated MOFs rank 'among the top 10 in the hMOF dataset.' The contributions say one MOF is in the top 5 and ten in the top 10% of a 4,547-MOF structurally similar subset of hMOF; the body repeats 'top five of the hMOF dataset' without the subset qualifier. Those are different claims, and the abstract's is the least supported. More importantly, the ranking compares MOFA's own GCMC results—UFF4MOF, RASPA default CO2 model, DDEC6 charges, rigid framework at 300 K—to published hMOF capacities from a different screening protocol. The paper never says the hMOF reference values were recomputed with the same protocol. At 0.1 bar, CO2 uptake is electrostatic-sensitive, so partial-charge differences alone could shift ranks. Without a same-protocol comparison, the 'top 5' and 'top 10%' results are not established. This is not a nitpick; it controls the paper's central discovery claim.\n\nOther issues are minor by comparison: single runs at each scale, no error bars on the retraining comparison, and the code/data URLs are redacted, so reproducibility is currently limited. The heavy self-citation is understandable given the direct extension of GHP-MOFassemble and Colmena, but more external baselines would help. These are all fixable.\n\nBottom line: the systems contribution is real and the scaling study is worth a referee's time. The chemistry ranking needs to be either recomputed against hMOF with a common protocol or substantially softened. I'd send it to peer review with a request for those fixes, and I'd be cautious about citing the top-5 result until the protocol comparison is added.","headline":"The systems-scale workflow results are credible, but the headline hMOF ranking claim is not established because it compares MOFA's GCMC protocol to published hMOF values without recomputing them under a common protocol.","tokens_in":20445,"tokens_out":3592,"would_cite":true,"duration_ms":35547,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A generative-AI plus simulation workflow produced MOFs with CO2 adsorption capacities ranking among the best in a 137,652-structure dataset, in a single 450-node, 3-hour run.","keywords":["generative AI","metal-organic frameworks","carbon capture","high-performance computing","online learning","diffusion model","CO2 adsorption","heterogeneous workflow"],"falsifier":"Recompute the CO2 adsorption capacity at 0.1 bar and 300 K for the generated MOFs and for a matched random sample of hMOF structures under identical force-field, charge, and grand-canonical Monte Carlo settings; if the 4.05 mol/kg MOF no longer ranks in the top five of the full or similar-subset comparison, the headline ranking claim is not supported.","tokens_in":19314,"feed_emoji":"🧪","tokens_out":6642,"duration_ms":61305,"temperature":0.7,"pith_summary":"MOFA is an open-source workflow that combines a generative diffusion model with atomistic simulation screening to discover metal-organic frameworks (MOFs) for carbon capture. In a single 450-node, 3-hour run it produced 114 novel MOFs per hour, including one with a CO2 uptake of 4.05 mol/kg at 0.1 bar, a value that ranks among the top five in the hypothetical-MOF dataset, and ten more in the top 10% of the structurally similar subset. The paper also reports that the rate of generating stable, high-performing MOFs grows approximately linearly with the number of compute nodes, and that periodically retraining the generative model on the best MOFs found so far substantially increases discovery rates. If these results hold, MOFA offers a way to navigate the enormous space of possible MOFs with far fewer guesses than brute-force enumeration.","feed_headline":"AI workflow found a top-5 CO2-capture MOF in a 3-hour run","feed_subtitle":"Generative AI plus simulation produced 114 new MOFs per hour, scaling linearly to 450 nodes.","key_machinery":"The mechanism that carries the argument is the online-learning generation-screening loop. A diffusion model fine-tuned on existing high-performing MOF linkers produces candidate linkers; the candidates are assembled with metal nodes, then passed through stability, density-functional, and Monte Carlo screens; the survivors are used to retrain the model. The retraining step is what makes the search adaptive: the paper reports that retraining increased the number of stable MOFs found at 90 minutes from 133 to 313 on 32 nodes and from 393 to 641 on 64 nodes. The workflow's scheduling layer keeps all stages running concurrently with low latencies, which is why throughput scales approximately linearly with node count.","core_discovery":"The central claim is that a tight loop between generative AI and simulation can rapidly produce rare, high-quality MOFs for carbon capture. MOFA fine-tunes an E(3)-equivariant diffusion model for molecular linkers on data from the hypothetical-MOF (hMOF) dataset, then runs the generated linkers through a funnel of increasingly expensive screens: molecular dynamics for stability, density-functional-theory optimization, charge assignment, and grand-canonical Monte Carlo for CO2 adsorption. MOFs that pass these screens are fed back to retrain the generator, so later iterations produce linkers biased toward stable, high-capacity structures. The paper's headline result is that one 450-node, 3-hour run generated 114 MOFs per hour, with one MOF reaching 4.05 mol/kg CO2 capacity at 0.1 bar (in the top five of the hMOF dataset) and ten others in the top 10% of a 4,547-MOF structurally similar subset, while worker utilization stayed above 99%.","pith_inferences":["A decisive validation the paper does not report is a same-protocol recomputation of hMOF capacities; until that is done, the top-5 and top-10% ranks are conditional on force-field and charge-method equivalence.","The 0.1 bar, 300 K condition is relevant to post-combustion capture, but real flue gas contains water and other gases; whether these MOFs retain capacity and stability under humid mixed-gas conditions is untested, so the practical carbon-capture claim is not yet established.","The linear-scaling result was measured on one machine up to 450 nodes; a stronger claim would require testing beyond that size or on different interconnect topologies, where communication or scheduling bottlenecks could appear.","The workflow's output is a set of computationally promising structures; the paper lists robotic synthesis as future work, so the near-term contribution is prescreening candidates for synthesis, not a guarantee that they can be made."],"forward_implications":["If the results generalize, a single large HPC run can search MOF chemical space fast enough to find top-ranking carbon-capture candidates in hours rather than via exhaustive enumeration.","Linear scaling with node count implies that adding compute nodes directly increases the rate of stable, high-adsorption MOF production, so larger machines accelerate discovery without changing the algorithm.","The measured benefit of retraining on intermediate results suggests that online learning is an effective strategy for inverse material design, not just for MOFs but for any property computable by simulation.","The modular architecture allows swapping the target property and screening criteria, so the same workflow can be pointed at catalysis, hydrogen storage, or other MOF applications."],"supporting_citations":[{"why":"Provides the hypothetical-MOF dataset that supplies fine-tuning data and the ranking baseline.","marker":"[18]"},{"why":"Supplies the diffusion-model architecture that MOFLinker fine-tunes for linker generation.","marker":"[32]"},{"why":"Provides the steering layer that expresses the workflow policies in the online-learning loop.","marker":"[17]"},{"why":"Provides task scheduling across heterogeneous CPU/GPU resources.","marker":"[16]"},{"why":"Supplies the Monte Carlo simulation code used to estimate CO2 adsorption capacities.","marker":"[15]"},{"why":"Supplies the density-functional-theory code used for cell optimization and electronic-density calculations.","marker":"[13]"},{"why":"Supplies the molecular-dynamics code used to screen generated MOFs for structural stability.","marker":"[14]"},{"why":"Defines the MOF-specific force field used in both stability and adsorption simulations.","marker":"[65]"},{"why":"Extends the MOF force field parameters used in the screening calculations.","marker":"[66]"}],"fun_headline_variants":["GenAI-sim loop yields top-5 CO2 MOF at scale","AI plus simulation discovers top CO2-capture MOF","MOFA workflow: 114 MOFs per hour, one top-5 for CO2","Scalable GenAI-sim pipeline tops CO2 MOF rankings","GenAI + simulation: top-5 CO2 MOF at 114 per hour"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking claims depend on the assumption that the CO2 capacities computed in this work are directly comparable to the published hMOF dataset values, but the paper does not report recomputing the hMOF capacities with the same force-field and partial-charge settings.","fun_headline_variants_meta":{"raw":{"variants":["GenAI-sim loop yields top-5 CO2 MOF at scale","AI plus simulation discovers top CO2-capture MOF","MOFA workflow: 114 MOFs per hour, one top-5 for CO2","Scalable GenAI-sim pipeline tops CO2 MOF rankings","GenAI + simulation: top-5 CO2 MOF at 114 per hour"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000705,"raw_usage":{"total_tokens":3191,"prompt_tokens":970,"completion_tokens":2221,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":2121}},"tokens_in":586,"tokens_out":2221,"duration_ms":14657,"temperature":1.0,"reasoning_tokens":2121,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:01:42.895965+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the CO2 adsorption capacity at 0.1 bar and 300 K for the generated MOFs and for a matched random sample of hMOF structures under identical force-field, charge, and grand-canonical Monte Carlo settings; if the 4.05 mol/kg MOF no longer ranks in the top five of the full or similar-subset comparison, the headline ranking claim is not supported.","supporting_citations":[{"cited_title":"Structure–property relationships of porous materials for carbon dioxide separation and capture,","cited_arxiv_id":null,"evidence_quote":"Provides the hypothetical-MOF dataset that supplies fine-tuning data and the ranking baseline."},{"cited_title":"Equivariant 3D-conditional diffusion models for molecular linker design,","cited_arxiv_id":null,"evidence_quote":"Supplies the diffusion-model architecture that MOFLinker fine-tunes for linker generation."},{"cited_title":"Colmena: Scalable machine-learning-based steering of ensemble sim- ulations for high performance computing,","cited_arxiv_id":null,"evidence_quote":"Provides the steering layer that expresses the workflow policies in the online-learning loop."},{"cited_title":"Parsl: Pervasive Parallel Programming in Python,","cited_arxiv_id":null,"evidence_quote":"Provides task scheduling across heterogeneous CPU/GPU resources."},{"cited_title":"RASPA: Molecular simulation software for adsorption and diffusion in flexible nanoporous materials,","cited_arxiv_id":null,"evidence_quote":"Supplies the Monte Carlo simulation code used to estimate CO2 adsorption capacities."},{"cited_title":"CP2K: An electronic structure and molecular dynamics software package - Quickstep: Efficient and accurate electronic structure calculations,","cited_arxiv_id":null,"evidence_quote":"Supplies the density-functional-theory code used for cell optimization and electronic-density calculations."},{"cited_title":"LAMMPS - A flexible simulation tool for particle- based materials modeling at the atomic, meso, and continuum scales,","cited_arxiv_id":null,"evidence_quote":"Supplies the molecular-dynamics code used to screen generated MOFs for structural stability."},{"cited_title":"Extension of the universal force field to metal–organic frameworks,","cited_arxiv_id":null,"evidence_quote":"Defines the MOF-specific force field used in both stability and adsorption simulations."},{"cited_title":"Extension of the universal force field for metal–organic frameworks,","cited_arxiv_id":null,"evidence_quote":"Extends the MOF force field parameters used in the screening calculations."}],"review_version":1}