{"id":"3dc67af5-912c-4fee-b0f2-f717ac5aae4c","arxiv_id":"2506.21788","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A method that distributes per-dataset decoding heads across GPUs enables pre-training of a multi-task graph neural network on 24 million atomistic structures from five datasets.","lead":"The authors train graph neural networks on 24 million atomistic structures from five different datasets, giving each dataset its own output head inside a shared model. They add a way to spread those heads across GPUs so the training can scale to large supercomputers, and they report scaling results on three DOE machines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's GFM-Baseline-All force row is identical to Model-MPTrj; this internal inconsistency undermines the force-accuracy evidence for MTL and should be resolved by reproducing the evaluation.","rationale":"The paper's central claim is that MTL-GFM pre-training significantly improves accuracy and transferability over single-task and single-head baselines. The main evidence is Tables 1 and 2. Table 1 is plausible; Table 2 has a structural red flag: the baseline-all and Model-MPTrj rows are identical. No training procedure or physical mechanism would produce identical test-set force MAEs across all five datasets for two different models unless the evaluation is broken or the wrong row was pasted. This is not a matter of disagreement with a baseline; it is an internal contradiction in the presented evidence. The reader's energy-alignment concern is a legitimate external threat to the energy results: Sec. 4.1 asserts alignment without describing offsets or procedures, and if alignment values depend on dataset statistics, the per-dataset heads can trivially fit the bias. However, the Table 2 duplication is more load-bearing because it directly undermines the numerical basis of the force half of the claim. A conditional verdict remains appropriate: once the table is corrected and the evaluation reproduced, the force comparison either survives or collapses. If the corrected baseline row is different, the authors need to re-run the comparison; until then the force-accuracy claim should not be accepted as stated. Both concerns point at the same claim, but the table issue is internal and more concrete, so agreement with the reader is only partial.","tokens_in":12170,"tokens_out":3616,"duration_ms":32836,"concrete_test":"Obtain the exact GFM-Baseline-All checkpoint and the exact test split used for Table 2; run the HydraGNN force evaluation and recompute the five force MAEs. If the resulting row differs from 0.0508, 1.1734, 0.1597, 0.0039, 0.1425, then Table 2 is erroneous and all force-accuracy and transferability conclusions must be recomputed from corrected results. Also report whether GFM-Baseline-All's checkpoint is actually a single-head model or accidentally one of the dataset-specific models.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Sec. 5.1's Table 2 is internally inconsistent: the GFM-Baseline-All row (0.0508, 1.1734, 0.1597, 0.0039, 0.1425) is identical, to four decimals, to the Model-MPTrj row, and nearly identical to the Model-Alexandria row. A single model trained on all five datasets cannot have exactly the same force MAE on all five test sets as a model trained only on MPTrj. This suggests a copy/paste error or that the baseline force evaluation was computed with the wrong checkpoint. Because the paper's strongest claim—that MTL significantly improves force accuracy and transferability—relies directly on Table 2, this row invalidates the force-prediction comparisons until it is explained or corrected. The energy-per-atom alignment issue noted by the reader is also real: Sec. 4.1 says offsets were 'consistently aligned' but gives no method; if alignment is data-dependent, MTL heads can absorb dataset bias. Still, the table duplication is the more load-bearing concern because it is an internal inconsistency, not an external assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a model-parallelism strategy for multi-task learning (MTL) in graph foundation models for atomistic simulation. The method, implemented in HydraGNN, partitions dataset-specific output heads across GPUs while keeping the shared encoder synchronized via distributed data parallelism. The authors pre-train on five datasets totaling over 24 million structures (ANI1x, QM7-X, Transition1x, MPTrj, Alexandria), report energy and force MAE for seven model variants (Tables 1 and 2), and present weak- and strong-scaling experiments on Perlmutter, Aurora, and Frontier. The central claims are that MTL improves accuracy and transferability relative to single-dataset training and that the proposed 2D parallelization scales efficiently.","tokens_in":12574,"tokens_out":2085,"duration_ms":23741,"significance":"If the numerical results hold, the paper makes a useful engineering contribution: it demonstrates a practical way to scale MTL-based graph foundation model pre-training to many heterogeneous datasets, and it provides open-source artifacts and benchmarks across three major supercomputing platforms. The aggregation of 24 million structures from diverse organic and inorganic datasets is itself a notable resource. However, the accuracy and transferability claims rest on Tables 1 and 2, which contain an internal inconsistency in the force-prediction results, and on an energy-alignment procedure that is not described. These issues must be resolved before the scientific claims can be accepted. The scaling results are plausible and interesting, but they are not the main source of the paper's claimed scientific advance.","major_comments":[{"comment":"The GFM-Baseline-All row in Table 2 is identical, to four decimal places, to the Model-MPTrj row (0.0508, 1.1734, 0.1597, 0.0039, 0.1425), and the Model-Alexandria row differs only in the MPTrj column (0.0038 vs. 0.0039). A single model trained on all five datasets cannot produce exactly the same force MAE on all five test sets as a model trained only on MPTrj. This indicates a likely copy/paste error or evaluation with the wrong checkpoint. Because the paper's force-accuracy and transferability claims rely directly on this table, the force-prediction comparisons are invalidated until the table is corrected and the evaluation reproduced.","section":"Table 2"},{"comment":"The text states that 'we consistently aligned the energy per atom values across all the datasets' but provides no description of how the offsets were determined. If the per-dataset offsets are fitted to minimize training error, the MTL heads can absorb dataset-specific energy biases, which would confound the comparison between GFM-MTL-All and the single-dataset baselines in Table 1. The authors should specify the alignment method, state whether it uses only training data, and ideally show the sensitivity of the downstream MAE results to the alignment procedure.","section":"Sec. 4.1"},{"comment":"All MAE numbers in Tables 1 and 2 are reported without error bars or repeated-run statistics. Several differences that support the central claim are small, such as the MPTrj energy MAE of 0.0627 for GFM-MTL-All versus 0.0651 for Model-MPTrj, or the Alexandria force MAE of 0.0039 for GFM-MTL-All versus 0.0038 for Model-Alexandria. Without variance estimates or significance tests, the claim that MTL 'significantly improves' accuracy is not statistically supported. The authors should report mean and standard deviation over multiple seeds, or at least provide a clear justification for why single runs are sufficient.","section":"Sec. 5.1, Tables 1 and 2"}],"minor_comments":[{"comment":"The name 'NERSC-Pelrmutter' contains a typo and should be 'Perlmutter'.","section":"Sec. 5"},{"comment":"The conclusion contains 'NERSC=Perlmutter' and 'using using'; these should be fixed.","section":"Sec. 6"},{"comment":"The memory-scaling expressions 'Ps + (Ph * Nh)' and '(Ps + Ph)' assume that the number of processes equals the number of heads; this should be stated explicitly, and the notation Ps, Ph, Nh should be defined in one place.","section":"Sec. 4.3"},{"comment":"The scaling plots would benefit from a legend distinguishing MTL-base and MTL-par in all subplots, and from a description of how the plotted time is averaged over the three reported epochs.","section":"Fig. 4"},{"comment":"The phrase 'denominated as' is used where 'denoted as' is intended; this occurs several times in Sections 5.1 and 5.2.","section":"Sec. 3"}],"recommendation":"major_revision","confidential_remarks":"The internal inconsistency in Table 2 is the kind of issue that, if it reflects a copy/paste error, is fixable, but the authors must also provide the energy-alignment details and uncertainty quantification before the accuracy claims can be taken at face value. The scaling study is the strongest part of the paper; the scientific claims about MTL accuracy need more careful empirical support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate engineering contribution—distributing MTL heads across GPUs is a sensible idea and the 24M-structure pre-training is a real step up—but Table 2 has a row that is numerically identical to a single-dataset baseline, and that undermines the force-accuracy evidence until it's explained.\n\nWhat's genuinely new: the head-splitting scheme (memory per GPU drops from Ps + Nh*Ph to Ps + Ph) and the 2D integration with DDP. The scaling runs on Frontier, Perlmutter, and Aurora are useful, and the energy results in Table 1 support the MTL story: GFM-MTL-All is consistently the best or near-best on all five datasets, and the single-dataset baselines collapse out-of-distribution. The related work is fair and the writing is clear.\n\nThe soft spots, in order of importance. First, Table 2: the GFM-Baseline-All row is identical to Model-MPTrj (0.0508, 1.1734, 0.1597, 0.0039, 0.1425) and nearly identical to Model-Alexandria. A model trained on all five datasets cannot have exactly the same force MAE as a model trained only on MPTrj. This looks like a copy/paste error or the wrong checkpoint, and the paper's strongest claim—that MTL improves force accuracy—rests on this table. Fixing it is mandatory. Second, Section 4.1 says energy-per-atom values were 'consistently aligned' but never says how. If the offsets are fitted per dataset, the MTL heads can absorb dataset bias, and the transferability gains in Table 1 could be an artifact. Show the procedure. Third, there are no error bars or repeated seeds anywhere, so we can't tell if the differences are real. Fourth, the scaling results are mixed: MTL-base is sometimes faster on Perlmutter, and the strong-scaling advantage only appears at larger GPU counts; the paper should be more careful about claiming universal superiority.\n\nNone of these are fatal to the core idea, and the table error is probably an honest slip. But with the evidence as printed, the force-accuracy claim is not trustworthy.\n\nI'd send this to peer review with a request for a revised version that fixes Table 2, documents the alignment, and adds variance estimates. The engineering contribution is worth refereeing; the current manuscript just isn't ready to accept as-is.","headline":"A real scalability idea and the largest MTL pre-training run to date, but Table 2's identical rows and the undocumented energy alignment keep the force-accuracy claim from being trusted as printed.","tokens_in":12968,"tokens_out":2508,"would_cite":false,"duration_ms":24017,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multi-task learning with per-dataset output heads makes graph foundation models for atomistic data more accurate and transferable, and a new head-level parallelism scheme lets them scale to 24 million structures.","keywords":["graph neural networks","multi-task learning","model parallelism","distributed data parallelism","multi-fidelity data","atomistic modeling","graph foundation models","supercomputer scaling"],"falsifier":"Re-run GFM-MTL-All with the energy-per-atom alignment replaced by arbitrary per-dataset constants (or by offsets computed from a different reference total energy), and compare out-of-distribution MAE on the five test sets against Tables 1 and 2; if the MAE values stay essentially unchanged, the MTL advantage is an artifact of the alignment, while if they degrade sharply, the alignment is doing real work.","tokens_in":12022,"feed_emoji":"⚛️","tokens_out":11406,"duration_ms":110259,"temperature":0.7,"pith_summary":"Single-dataset pre-training fails to generalize across organic and inorganic chemistry, so a foundation model for atomistic data needs to learn from many sources at once. This paper argues that multi-task learning—one shared message-passing encoder with a separate output head for each dataset—is the right way to do that: pre-training on 24 million structures from five datasets with different DFT settings yields lower energy and force errors, and better transferability, than either dataset-specific models or one model trained on all data mixed together. The paper also introduces multi-task parallelism, which places each dataset head on its own GPU and synchronizes only the shared encoder, so memory per GPU no longer grows with the number of datasets. They report near-ideal strong scaling on up to 320 GPUs and weak scaling on up to 1,920 GPUs across three US supercomputers, suggesting that adding datasets to a foundation model can be as simple as adding GPUs for their heads.","feed_headline":"Multi-task training lifts accuracy across five atomistic datasets","feed_subtitle":"Per-dataset heads beat single-task baselines on 24M structures and scale to 1,920 GPUs","key_machinery":"The central object is the two-level hierarchical multi-task architecture in the authors' HydraGNN graph neural network. The first level of multi-task learning splits a shared message-passing encoder into one branch per dataset; the second level splits each branch into a head for energy per atom and a head for atomic forces. The load-bearing mechanism is multi-task parallelism: each process holds a copy of the shared encoder plus one complete dataset head, heads run forward and backward passes concurrently, and only the encoder gradients are averaged across processes during synchronization. This changes the memory footprint per GPU from $P_s + N_h P_h$ to $P_s + P_h$, so the number of datasets becomes an axis that scales independently of model depth and width. Process sub-groups carry out distributed data parallelism inside each head while a global group synchronizes the shared layers.","core_discovery":"The central claim is that multi-task learning should be the default pre-training scheme for graph foundation models on multi-source, multi-fidelity atomistic data, because it simultaneously gives high in-distribution accuracy and out-of-distribution transferability where single-dataset training and single-head all-data training fail. The authors report that a two-level MTL model—dataset-specific branches, each split into energy and force heads—achieves the best or near-best error on every one of the five datasets considered, for example cutting energy error on the inorganic MPTrj set from 0.4248 for the all-data baseline to 0.0627. They further claim that multi-task parallelism, integrated with distributed data parallelism through process sub-groups, makes this method scalable: memory per GPU drops from $P_s + N_h P_h$ to $P_s + P_h$, strong scaling stays near ideal up to 320 GPUs for large effective batches, and weak scaling holds up to 1,920 GPUs on the three supercomputers tested.","pith_inferences":["The reported accuracy gains depend on the unstated energy-per-atom alignment across the five datasets; a control experiment with randomized or permuted offsets would separate the effect of multi-task learning from the effect of that alignment.","Multi-task parallelism is not specific to graph neural networks: any multi-task architecture with independent per-task heads could use the same head-splitting scheme, so the method may transfer to other scientific foundation models.","The scaling comparison is against the same method without head-level parallelism; a comparison with tensor or pipeline parallelism under the same memory constraints would clarify when each strategy should be chosen.","A direct stress test is the paper's own future-work target of 359 million structures covering all natural elements: if the accuracy and scaling results persist at that size, the main practical obstacle to a universal atomistic foundation model is removed."],"forward_implications":["A single pre-trained GFM can serve both organic and inorganic downstream tasks, because the shared encoder learns common chemistry while each head specializes in its source dataset's data distribution.","Adding a new dataset to pre-training costs one additional GPU (or group of GPUs) for its head rather than a larger model or more memory per GPU, so the approach directly addresses the flood of new atomistic datasets.","The simultaneous prediction of energy per atom and atomic forces from the same shared representation means downstream users do not need separate models for the two quantities.","If the memory-scaling argument holds, the method becomes more attractive as the number of heads grows relative to the shared encoder, which is the regime typical of message-passing GNNs with many datasets.","Strong and weak scaling results imply that pre-training on hundreds of millions of structures—the paper's stated next step—is computationally reachable with current exascale systems."],"supporting_citations":[{"why":"Supplies the HydraGNN architecture, the hyperparameter configuration used for all experiments, and the earlier single-level MTL baseline.","marker":"[12]"},{"why":"The previous multi-task learner on about 4 million structures whose limited scale and lack of HPC scaling the paper extends.","marker":"[30]"},{"why":"Provides the ANI1x dataset of DFT energies and forces for organic molecules, one of the five training sources.","marker":"[23]"},{"why":"Provides the QM7-X dataset of organic molecular properties, the second training source.","marker":"[8]"},{"why":"Provides the Transition1x dataset of reactive DFT calculations, the third training source.","marker":"[19]"},{"why":"Provides the MPTrj dataset of inorganic materials DFT calculations, the fourth training source.","marker":"[10]"},{"why":"One of the Alexandria dataset references supplying the fifth training source of inorganic structures.","marker":"[16]"}],"fun_headline_variants":["Multi-task parallelism cuts energy error 85% on 24M structures","Per-dataset heads boost accuracy and scale on five atomistic sets","Multi-task training: better accuracy and transfer on 24M structures","Multi-task parallelism scales graph models to 1,920 GPUs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The energy-per-atom values across the five datasets were 'consistently aligned,' but the paper never describes how the alignment offsets were obtained; if those offsets are not physically correct, the per-dataset heads could simply absorb dataset-specific biases, making the reported transferability gains an artifact of the alignment rather than of multi-task learning.","fun_headline_variants_meta":{"raw":{"variants":["Multi-task parallelism cuts energy error 85% on 24M structures","Per-dataset heads boost accuracy and scale on five atomistic sets","Multi-task training: better accuracy and transfer on 24M structures","Multi-task parallelism scales graph models to 1,920 GPUs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000798,"raw_usage":{"total_tokens":3497,"prompt_tokens":920,"completion_tokens":2577,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":2501}},"tokens_in":536,"tokens_out":2577,"duration_ms":23003,"temperature":1.0,"reasoning_tokens":2501,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:18:24.194745+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run GFM-MTL-All with the energy-per-atom alignment replaced by arbitrary per-dataset constants (or by offsets computed from a different reference total energy), and compare out-of-distribution MAE on the five test sets against Tables 1 and 2; if the MAE values stay essentially unchanged, the MTL advantage is an artifact of the alignment, while if they degrade sharply, the alignment is doing real work.","supporting_citations":[{"cited_title":"Journal of Supercomputing81, Article 618 (Mar 2025), open Access; Published: 14 March 2025","cited_arxiv_id":null,"evidence_quote":"Supplies the HydraGNN architecture, the hyperparameter configuration used for all experiments, and the earlier single-level MTL baseline."},{"cited_title":"Scientific Data 9, 779 (2022)","cited_arxiv_id":null,"evidence_quote":"Provides the Transition1x dataset of reactive DFT calculations, the third training source."}],"review_version":1}