{"id":"22ef375c-b551-4e1e-a0bd-4113ac05a746","arxiv_id":"2411.11783","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"OCx24 is a new experimental catalyst dataset plus a computational adsorption-energy screen whose HER model generalizes to unseen compositions and recovers platinum as a top candidate.","lead":"This paper introduces OCx24, an open dataset of 572 synthesized catalyst samples with X-ray characterization and electrochemical testing for CO2 reduction and hydrogen evolution. A companion computational screen of 19,406 materials produces a model that, without any platinum in its training data, ranks platinum as a top HER catalyst.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"HER volcano claim rests on an unvalidated linear extrapolation to Pt, far outside the training domain; the dome in Fig. 4b may be an artifact of the inference-set distribution rather than the learned model.","rationale":"The paper is a valuable resource: the dataset, the transparent CO2RR analysis, and the open code are real contributions. The HER volcano is the headline scientific claim, though, and it is the part that needs the most scrutiny. My concern differs from the reader's weakest assumption (surface averaging, Eq. 2). I agree that the mean aggregation is a potential weakness, and the paper honestly reports that Wulff and Boltzmann weighting did not improve results, which is some evidence against that concern. The more load-bearing issue is that the volcano/Pt result is an extrapolation outside the training domain, with no uncertainty quantification, and the volcano shape cannot be produced by the fitted linear model itself. The apparent dome must come from the distribution of the inference set or the projection. That makes the \"independent identification of Pt\" fragile in a way that LOCO CV does not address. The concrete test of ranking a benchmark set of known HER catalysts would directly validate (or refute) the generalization claim. If the model ranks the benchmark correctly, the concern is largely resolved; if not, the Pt result should be treated as an unvalidated extrapolation. This supports the reader's CONDITIONAL verdict rather than changing it.","tokens_in":30808,"tokens_out":11785,"duration_ms":123494,"concrete_test":"Compile a benchmark of about 20 well-studied HER catalysts (Pt, Pd, Ni, Cu, Ag, Au, MoS2, MoSe2, Ni-Mo, etc.) with literature exchange-current densities or overpotentials. Apply the exact linear model (mean E_H/OH, same standardization) to these materials and compute Spearman rank correlation against the experimental benchmark; report whether Pt is significantly separated from the rest. Separately, re-run inference on a uniform random sample drawn from the same 2D feature bounding box as the 19,406 materials and check whether a volcano still appears and Pt remains at the apex. If the volcano disappears or the rank correlation is poor, the headline result is an artifact of the inference-set distribution or an unvalidated extrapolation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central HER claim is a single extrapolation. The model is a linear regression on mean E_H and E_OH (Eq. 5, Section 5.1) trained on 179 experimental targets, overwhelmingly Cu-based alloys, occupying a compact feature region. LOCO CV (R2 = 0.59) tests interpolation among compositions similar to training; it does not test extrapolation to Pt, whose descriptor values lie far outside the training cloud and which was deliberately excluded. The paper provides no error bars, confidence intervals, or applicability-domain analysis for the 19,406 inference predictions, so \"Pt at the apex\" is a point extrapolation with unknown uncertainty. Moreover, a linear model in two features cannot itself produce a volcano: it has no interior maximum in the descriptor plane. The dome in Fig. 4b is imposed by the joint distribution of the 19,406 database-filtered materials (e.g., OH-H scaling) and the PC1 projection. The peak position is therefore sensitive to which materials are in the inference set and to the projection, not solely to the experimentally learned mapping. This undermines the claim that the model \"recovered a Sabatier volcano\" and independently identified Pt; the volcano may be a visualization artifact, and the Pt ranking has no demonstrated out-of-domain validity.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces OCx24, a high-throughput experimental dataset of 572 synthesized samples (441 gas diffusion electrode tests) for CO2RR and HER, with XRF/XRD characterization, combined with a computational screening of 19,406 materials using ML-accelerated adsorption energy calculations (AdsorbML). The authors train linear regression models to predict HER cell voltage from mean H and OH adsorption energies, reporting LOO R2=0.61 and LOCO R2=0.59, and apply the model to the 19,406 materials, claiming to recover a Sabatier volcano with Pt at its apex despite no Pt in the training data. For CO2RR, random forest models achieve LOO R2~0.4-0.5 but LOCO R2 near zero, which the authors report transparently. The paper provides code and data under open licenses.","tokens_in":31066,"tokens_out":7964,"duration_ms":74274,"significance":"The experimental dataset is a substantial community resource: it is large, characterized, and measured under industrially relevant conditions, and the code/data are openly released. The LOCO HER correlation (R2=0.59) is a credible benchmark for composition-out-of-distribution prediction. The computational screening effort is the largest of its kind and will be useful for future descriptor studies. However, the headline claim that the model 'recovers a Sabatier volcano' with Pt at its apex is not supported by the analysis as presented, because the HER model is linear and the volcano is a density artifact. The dataset and screening remain valuable independent of that claim.","major_comments":[{"comment":"The claim that the model 'recovers a Sabatier volcano' is not supported by the analysis. The HER model is linear in the mean H and OH adsorption energies (Eq. 5 with f linear); a linear function has no interior maximum in the descriptor plane. The dome-shaped envelope in Figure 4b arises from the joint distribution of the 19,406 inference materials and the PC1 projection (e.g., the density of materials in the E_H–E_OH plane), not from a non-monotonic learned response. To support the volcano claim, the authors should fit a model with an explicit nonlinearity (e.g., quadratic in E_H and E_OH) and show that the data support an interior optimum, or reframe the volcano as a property of the inference-set distribution rather than the learned mapping.","section":"Section 5.1, Eq. (5), Figure 4b"},{"comment":"The identification of Pt as a top HER catalyst is a single extrapolation. The training set (179 targets) is overwhelmingly Cu-based and occupies a compact feature region, while Pt's H and OH adsorption energies lie far outside that region. The LOCO R2=0.59 tests interpolation among compositions similar to the training set, not extrapolation to Pt. The paper reports no confidence intervals, bootstrap errors, or applicability-domain analysis for the 19,406 inference predictions, so the Pt ranking has unknown uncertainty. Please provide uncertainty quantification for the Pt prediction (e.g., conformal intervals or a bootstrap over the training data) and an assessment of feature-space overlap between training and inference sets, or soften the claim accordingly.","section":"Section 5.1, Pt extrapolation"},{"comment":"The peak position in Figure 4b is sensitive to two choices that are not tested: the Pourbaix stability filter that determines the 19,406-material inference set (Section 4.1) and the PC1 projection used for visualization. A linear model's predicted voltage for a material is a linear function of the descriptors, so the material that attains the maximum depends on the convex hull of the inference-set feature distribution. The claim that Pt is 'at the peak' should be robust to reasonable changes in the inference-set construction (e.g., different Pourbaix thresholds, or removing chalcogenides) and to the projection (e.g., plotting against E_H or E_OH directly). The authors should demonstrate this robustness or restrict the claim to the specific screen performed.","section":"Section 5.1, Figure 4b, inference-set sensitivity"}],"minor_comments":[{"comment":"The abstract and title use 'OCX24' and 'OCx24' inconsistently; please standardize the dataset name throughout.","section":"Abstract/Title"},{"comment":"The abbreviation 's.c.c.m.' should be 'sccm' (standard cubic centimeters per minute).","section":"Section 2.2"},{"comment":"The phrase 'within twice the mean absolute error (MAE) of Pt or better' is ambiguous; please specify the exact inequality used to define the 3,869 materials.","section":"Section 5.1"},{"comment":"Citation markers such as 'GemNet-OC109' and 'AdsorbML13' need a space before the reference numeral in running text.","section":"Section C.2.2"},{"comment":"The sentence 'Vienna Ab initio Simulation Package (VASP) with projector augmented wave (PAW) pseudopotentials and the RPBE functional were used' has a subject-verb agreement error; use 'was used.'","section":"Section C.2.3"},{"comment":"The caption does not define PC1; add a definition or refer to the text in Section 5.1 where it is introduced.","section":"Figure 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to attract attention because of the Pt/volcano claim, but that claim as currently analyzed is not supported; the dataset itself is the strongest contribution. The authors should be encouraged to either add a nonlinear model and uncertainty quantification or substantially temper the abstract. The manuscript is within scope for an applied ML/data journal, but the central claim needs revision before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe OCx24 dataset is the real contribution here. 572 synthesized samples, 441 tested electrodes, consistent high-throughput protocol, XRD/XRF characterization, and a 19,406-material computational screen with six adsorbates—that's a serious reusable resource. The code and data are open, and the authors are transparent about their CO2RR failures, which earns them credit. The HER LOCO R2 of 0.59 is credible given the experimental noise (σ ≈ 43 mV), and the scaling analysis with dataset size is a useful guide for future efforts.\n\nNow the soft spots. The headline claim—\"data-driven Sabatier volcano independently identified Pt\"—does not hold up. The model is linear in two features (E_H, E_OH). A linear function has no interior maximum, so the dome in Fig. 4b cannot come from the learned mapping. It is imposed by the joint distribution of the 19,406-material inference set and the PC1 projection used for visualization. The Pt point is a single extrapolation far outside the training cloud (training is mostly Cu-based alloys), and the paper gives no error bars, confidence intervals, or applicability-domain analysis. LOCO CV tests interpolation among compositions similar to training; it does not validate out-of-domain extrapolation. So \"independently identified Pt\" is an overstatement. What the model does show is a plausible linear trend between adsorption energies and HER voltage within the training domain—that is useful, but it is not a volcano.\n\nOther soft spots are more minor. The aggregation from 230 to 179 targets is manual, though the paper documents the rules. The computational energies are ML-relaxed structures with DFT single points rather than full relaxations; that is a known approximation in the AdsorbML pipeline, but the paper could state the uncertainty this introduces. The mean adsorption energy over all terminations is a crude aggregate; the authors tried Wulff and Boltzmann weighting and found they didn't help, which is honest but also a reminder that the descriptor link is not yet well understood.\n\nNet: the dataset and the computational screen deserve to be a standard benchmark. The volcano claim needs to be rewritten as a rank-order observation with uncertainty bounds, not a recovery of a Sabatier relationship. Who is this for? Computational catalysis and ML-for-materials folks who need experimental benchmarks with consistent protocols; also experimentalists who want a well-characterized dataset. I'd send this to a serious referee. The paper needs revision, but the resource is real.","headline":"The dataset is a real gift to the community; the Sabatier volcano is a visualization artifact that overstates what a linear model can do.","tokens_in":31711,"tokens_out":2382,"would_cite":true,"duration_ms":24623,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A model that never saw platinum still ranks it best for hydrogen evolution.","keywords":["OCx24 dataset","hydrogen evolution reaction","Sabatier volcano","adsorption energies","high-throughput experimentation","machine learning","CO2 reduction","catalyst screening"],"falsifier":"Measure HER cell voltage for a set of the 436 predicted low-cost non-noble-metal candidates (e.g., Mo-S, Mo-Se alloys) under the same MEA conditions; if these materials do not outperform or match the copper baseline or platinum, the ranking is refuted. Alternatively, computing adsorption energies with Wulff or Boltzmann weighting for a subset of materials and checking whether platinum stays at the apex would test the mean-aggregation assumption directly.","tokens_in":30614,"feed_emoji":"⚡","tokens_out":4547,"duration_ms":42254,"temperature":0.7,"pith_summary":"The paper builds a bridge between computational catalysis and high-throughput experiments by releasing a dataset (OCx24) of 572 synthesized samples, 441 tested electrodes, and adsorption energies for six adsorbates on 19,406 materials computed with an AI-accelerated pipeline. Its central claim is that a linear model trained on measured HER voltages using only the mean adsorption energies of hydrogen and hydroxyl can generalize to the full computational space and recover a Sabatier volcano with platinum at the peak, even though no platinum-containing sample was in the training data. If true, this supports adsorption energies as generalizable descriptors for HER and suggests that modest experimental sets can anchor computational screens of large material spaces.","feed_headline":"179 experiments rebuild the HER volcano, platinum on top","feed_subtitle":"A model that never saw platinum ranks it best across 19,406 materials, and flags low-cost alternatives.","key_machinery":"The load-bearing object is the mean adsorption energy E_ads,mean (Eq. 2), the arithmetic average of adsorption energies across all enumerated surface terminations up to Miller index 2. This single surface-level aggregation, combined with a linear regression trained on 179 experimental targets, allows inference over the full 19,406-material space; the Sabatier volcano emerges as the model's predicted activity landscape, and the paper's key comparison is the failure of Matminer-only features to reproduce the platinum apex.","core_discovery":"On the paper's own terms: using OCx24's experimental HER voltages at 50 mA/cm² as targets and the mean adsorption energy across all surface terminations (Eq. 2) of H and OH as features, a linear model trained under leave-one-composition-out cross-validation reaches R² = 0.59. Applied to the 19,406 stable and metastable materials from Materials Project, OQMD, and Alexandria, the model's predictions form a Sabatier volcano, with platinum predicted near the apex despite platinum being absent from the training set even as an alloy. The same model trained on Matminer bulk-element features reaches similar in-sample correlation but fails to place platinum at the top, indicating that adsorption energies carry the transferable signal. The paper also reports that CO2RR product rates are far harder to predict, with near-zero LOCO correlation, and attributes this to the reaction's network complexity and the limited expressiveness of single-adsorbate descriptors.","pith_inferences":["The volcano's recovery suggests that the linear map learned from 179 samples may be capturing a genuine physical trend rather than a compositional artifact, but the claim would be strengthened by testing a set of the predicted low-cost candidates in the same MEA setup.","The paper's weakest link—the unweighted mean over surfaces—could be probed directly by comparing predicted rankings against measurements on well-faceted nanoparticles where the Wulff shape is known; if weighted energies change Pt's ranking, the mean aggregation is a lucky accident.","A natural extension is to use the same pipeline for other reactions whose selectivity is controlled by a small number of adsorption-energy descriptors, with the expectation that the CO2RR failure is due to descriptor insufficiency rather than the experimental bridge itself."],"forward_implications":["The 3,869 materials predicted within twice the mean absolute error of platinum, including 436 that contain no noble metals, become concrete candidates for experimental HER testing.","Adding more experimental targets should sharpen the mapping; the paper projects from its scaling curve that 10⁴–10⁵ samples would yield substantially more predictive models.","Adsorption energies of H and OH, aggregated as a simple mean, are shown to be more transferable descriptors than bulk elemental features for HER activity ranking.","The OCx24 dataset, with its matched XRD/XRF characterization, provides a common testbed for future models that map computational descriptors to experimental performance."],"supporting_citations":[{"why":"Supplies the AdsorbML hybrid ML+DFT pipeline that produced the adsorption energies for 19,406 materials.","marker":"[13]"},{"why":"Materials Project is one of the three bulk databases whose stable/metastable materials form the computational search space.","marker":"[21]"},{"why":"Open Catalyst 2020 dataset trained the GemNet-OC ML potential used for the AI-accelerated relaxations.","marker":"[22]"},{"why":"Provides the theoretical Sabatier volcano trend for HER that the data-driven volcano reproduces with platinum at the apex.","marker":"[51]"},{"why":"Matminer supplies the bulk elemental features used as the comparison model that fails to crown platinum.","marker":"[67]"},{"why":"Demonstrates prior computational screening of HER catalysts that this work extends to a much larger experimental-anchored space.","marker":"[59]"}],"fun_headline_variants":["AI picks platinum as top HER catalyst without ever seeing it","Volcano model ranks Pt #1 for hydrogen evolution, no Pt in training","OCx24: 572 samples, AI predicts Pt best, despite no Pt data","Data-driven volcano: Pt tops HER list, model never trained on it","From 20k materials, AI flags Pt for HER, absent from training set"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The single mean adsorption energy over all computed surface terminations represents the catalytically relevant behavior of the real experimental nanoparticles; if that average is not representative, the predicted rankings of all 19,406 materials, including platinum's top position, lose their basis.","fun_headline_variants_meta":{"raw":{"variants":["AI picks platinum as top HER catalyst without ever seeing it","Volcano model ranks Pt #1 for hydrogen evolution, no Pt in training","OCx24: 572 samples, AI predicts Pt best, despite no Pt data","Data-driven volcano: Pt tops HER list, model never trained on it","From 20k materials, AI flags Pt for HER, absent from training set"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000679,"raw_usage":{"total_tokens":3120,"prompt_tokens":1013,"completion_tokens":2107,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":2008}},"tokens_in":629,"tokens_out":2107,"duration_ms":14779,"temperature":1.0,"reasoning_tokens":2008,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:08:40.202685+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure HER cell voltage for a set of the 436 predicted low-cost non-noble-metal candidates (e.g., Mo-S, Mo-Se alloys) under the same MEA conditions; if these materials do not outperform or match the copper baseline or platinum, the ranking is refuted. Alternatively, computing adsorption energies with Wulff or Boltzmann weighting for a subset of materials and checking whether platinum stays at the apex would test the mean-aggregation assumption directly.","supporting_citations":[{"cited_title":"Trends in the exchange current for hydrogen evolution.Journal of The Electrochemical Society, 152(3):J23, 2005","cited_arxiv_id":null,"evidence_quote":"Provides the theoretical Sabatier volcano trend for HER that the data-driven volcano reproduces with platinum at the apex."},{"cited_title":"Computational high-throughput screening of electrocatalytic materials for hydrogen evolution","cited_arxiv_id":null,"evidence_quote":"Demonstrates prior computational screening of HER catalysts that this work extends to a much larger experimental-anchored space."}],"review_version":1}