Pith. sign in

REVIEW 3 major objections 5 minor 23 references

Benchmarking large language models for materials synthesis: the case of atomic layer deposition

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read GPT-4o earns a passing but fragile grade on expert atomic layer deposition questions, with 36% of questions drawing at least one below-average score and five suspected hallucinations.

desk verdict A genuinely useful 70-question ALD benchmark with careful expert grading, but the headline correlations rest on internally inconsistent tables and clustered reviews; the stats need a redo before the main claims are reliable. read the letter →

arxiv 2412.10477 v1 pith:7HPPLW2G submitted 2024-12-13 cs.LG cond-mat.mtrl-scics.AI

classification cs.LGcond-mat.mtrl-scics.AI
keywords atomiclayerdepositionlargelanguagemodelsbenchmarkmaterialssynthesishallucinationdetectionexpertevaluationopen-endedquestionsGPT-4o
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces ALDbench, a benchmark of 70 open-ended questions about atomic layer deposition (ALD), a thin-film growth technique, with questions ranging from graduate level to state-of-the-art expert knowledge. Six ALD experts wrote the questions and seven experts graded the GPT-4o answers on four criteria: overall quality, specificity, relevance, and accuracy. The model earned a composite quality score of 3.7 out of 5, a passing grade, but 36% of questions received at least one below-average score and the authors identified at least five suspected hallucinations, mainly invented precursor chemistries. The paper also finds statistically significant correlations: harder questions get lower quality and relevance scores, and more specific questions get lower accuracy scores, while question difficulty and specificity are themselves nearly uncorrelated. The authors argue that multi-criteria, open-ended expert grading reveals failure modes that multiple-choice and NLP-style benchmarks miss.

What carries the argument

The central object is ALDbench, a hand-curated set of 70 open-ended questions about atomic layer deposition, each graded by multiple human ALD experts on question difficulty and specificity, and on the model's response quality, specificity, relevance, and accuracy using 1-5 Likert rubrics. The statistical engine is the 2x2 contingency table split at 'above average' (scores 4-5) versus 'at or below average' (scores 1-3), tested with the Fisher exact test to detect correlations between question attributes and response scores. The in-depth qualitative analysis of individual responses, looking for hallucinations and precision failures, completes the machinery.

What would settle it

Reanalyze the contingency tables with expert and question as random effects, for example using a mixed-effects logistic regression or a cluster-robust version of the Fisher exact test; if the associations between difficulty and quality, difficulty and relevance, and specificity and accuracy lose statistical significance, the paper's headline correlation claims would not survive.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that a state-of-the-art general LLM, GPT-4o, performs at a passing but uneven level on expert-level ALD knowledge: aggregate scores are above average across all four criteria, yet a substantial minority of questions draw low scores, and the model fabricates plausible-sounding but unreferenced chemistry, such as NF3 or SF6 as fluorine sources for MgF2 ALD and cobalt(II) nitrate as a Co3O4 precursor. The paper further claims that response quality and relevance drop as question difficulty rises, and accuracy drops as question specificity rises, based on Fisher exact tests of 2x2 contingency tables built from 236 expert reviews. It presents ALDbench itself as a reusable open-ended benchmark for materials synthesis that captures dimensions—relevance and specificity—that standard multiple-choice or NLP benchmarks do not.

Load-bearing premise

The statistical correlations rest on treating all 236 expert reviews as independent observations, but the reviews are clustered: the same seven experts graded many questions each, so the effective sample size is smaller and the reported p-values are likely too optimistic.

Editorial extensions

If this is right

  • GPT-4o's aggregate passing scores across all four criteria show that general-purpose LLMs already hold substantial declarative knowledge about a specialized synthesis field like ALD.
  • The statistically significant correlations mean that a user asking hard or highly specific ALD questions should expect lower-quality, less relevant, or less accurate answers than a user asking general ones.
  • Because question difficulty and question specificity are nearly uncorrelated (Pearson r = 0.12), benchmarks must grade both dimensions separately to reveal where LLMs fail.
  • Open-ended, expert-graded benchmarks can expose failure modes—such as invented precursor chemistries and imprecise quantitative ranges—that multiple-choice or NLP-style benchmarks would pass over.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the clustering of expert reviews is taken into account, the reported Fisher-exact p-values (0.033, 0.016, 0.007) are likely too small; a mixed-effects reanalysis could weaken or erase the claimed correlations.
  • A natural next experiment is to run ALDbench on the same model with access to the Atomic Limits ALD database or a retrieval tool; the paper's own hallucination analysis suggests that specificity and accuracy scores would rise.
  • The 'at least five' hallucination figure is probably a floor, since it depends on the reviewing experts' personal knowledge; automated checking of every proposed precursor pair against a database would give a reproducible rate.
  • The specificity-accuracy anticorrelation, if it generalizes, implies that LLM fluency in quantitative domains is partly a trade-off: precise numeric answers are sacrificed for plausible-sounding ranges, a testable hypothesis for other synthesis fields.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces ALDbench, a benchmark of 70 open-ended questions in atomic layer deposition (ALD), graded by seven domain experts who scored each question's difficulty and specificity and each GPT-4o response's overall quality, specificity, relevance, and accuracy on 1–5 Likert scales. The authors report a composite quality score of 3.7, with 36% of questions receiving at least one below-average score, at least five suspected hallucinations, and three statistically significant correlations from Fisher exact tests (quality–difficulty, relevance–difficulty, accuracy–specificity). The stated goal is to provide a domain-specific evaluation tool for LLMs in materials synthesis.

Significance. If the descriptive results stand, ALDbench is a valuable community resource: it is one of the few expert-reviewed, open-ended benchmarks in materials synthesis, and it deliberately includes questions at the frontier of expert knowledge. The multi-criteria rubric (quality, specificity, relevance, accuracy) and the qualitative hallucination analysis are useful contributions. The paper also ships the question set and responses in the Supporting Information, which supports reuse and independent verification. However, the paper's headline statistical claims are currently under-supported because of internal inconsistencies in the reported contingency tables and the use of a statistical test that assumes independence of clustered observations.

major comments (3)
  1. [§III B and Table VII] The first contingency table in Table VII (response quality vs. question difficulty) sums to 265 reviews (60+38+132+35), whereas §III B states that 236 independent reviews were gathered and every other contingency table in Tables VII–VIII sums to 235. This internal inconsistency means the reported p=0.033 for the quality–difficulty correlation cannot be reconstructed from the printed data. The authors should correct the table and recompute the p-value, and should also verify the marginal totals of all tables.
  2. [§III B and Tables VI–VIII] The Fisher exact tests treat each of the 236 reviews as an independent observation, but reviews are nested within questions (70 questions) and within the seven expert raters, who were free to review different subsets of questions. Ratings from the same expert or the same question are likely positively correlated, so the effective sample size is smaller than 236 and the reported p-values (0.033, 0.016, 0.007) are anticonservative. The authors should use an analysis that accounts for this clustering—for example, mixed-effects logistic regression with random intercepts for expert and question, or cluster-robust standard errors—or should aggregate to the question level and perform an appropriate test.
  3. [§III B and Table VI] The paper performs eight Fisher exact tests without any adjustment for multiple comparisons. Even if the first contingency-table issue is fixed, the quality–difficulty p=0.033 would not survive a Bonferroni correction (threshold 0.00625), and the relevance–difficulty p=0.016 would also fail. Only the accuracy–specificity p=0.007 would potentially remain significant, but its reliability is still subject to the clustering problem. The authors should either pre-specify a limited set of hypotheses, apply a correction, or present the results as exploratory rather than confirmatory.
minor comments (5)
  1. [§III B] The dichotomization of scores into above-average (4–5) and at-or-below-average (1–3) is arbitrary and discards ordinal information. The authors should justify this split or show that the conclusions are robust to alternative thresholds (e.g., median split per expert or the full ordinal scale).
  2. [General] No inter-rater reliability statistic (e.g., Cohen's kappa, weighted kappa, or intraclass correlation) is reported for the expert graders. Given the observed dispersion in Figures 1–3, such a measure would help the reader calibrate the reliability of the ratings.
  3. [References] Reference 11, 'Atomic limits ald database', is incomplete; the authors should provide a full citation or URL.
  4. [Text] There are several typographical errors, including 'In Figure 1 We show' (capital W) in §III A, and 'How does the temperature effect the growth' in Question 51 (should be 'affect').
  5. [Figures 1–2] The color-map representations of expert scores are difficult to read in grayscale print; a dot-plot or heatmap with numeric labels might be clearer.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an observational, human-graded benchmark study with no fitted parameter dressed as a prediction and no self-referential derivation chain.

full rationale

The paper's chain of reasoning is not a derivation: ALDbench is created from 70 human-written questions, GPT-4o responses are scored by seven co-author experts on four Likert criteria, and the headline findings are averaged scores and Fisher exact tests on the resulting contingency tables. Nothing is fitted and then re-predicted: the expert scores are the data, not an output of the model, and the correlations in Table VI are computed directly from those scores rather than from a model that was calibrated on them. The only self-citation (Ref. 7, a prior bibliometric study of ALD by one of the co-authors) is used as background for the field's applied relevance and is not load-bearing for any quantitative claim. Concerns that the Fisher tests treat clustered reviews as independent and that the printed contingency tables do not sum to the stated 236 reviews are validity/correctness issues about the reported p-values, not instances where a conclusion is identical to an input by construction. No circular step is therefore exhibited.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the validity of expert grading and on statistical independence of the 236 reviews. No new physical entities or fitted physical parameters are introduced; the only hand-chosen numbers are methodological thresholds and weighting schemes.

free parameters (2)
  • above-average threshold for review binarization = scores 4-5 above average; 1-3 at or below average
    Used to build 2x2 contingency tables for Fisher exact test in Section III B; this arbitrary choice affects the derived p-values.
  • equal expert weighting in composite scores = 1/N_expert per expert
    Composite scores average each expert's mean, ignoring differing review counts per expert; this choice affects the reported 3.7 composite score.
assumptions (4)
  • standard math Fisher exact test is valid for 2x2 contingency tables with independent observations.
    Used in Section III B to derive p-values; independence of observations is likely violated due to clustering.
  • domain assumption Expert Likert ratings are a valid measure of LLM response quality.
    No inter-rater reliability, calibration, or validation against objective criteria is reported in Section II.
  • domain assumption The 70 questions are representative of ALD knowledge needed by researchers.
    Questions were authored by six co-authors with expertise in ALD; the set is acknowledged as non-comprehensive in Section III A.
  • domain assumption Individual reviews are independent observations.
    Required for the Fisher exact test in Table VI; multiple reviews of the same question by the same experts violate this assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Benchmarking large language models for materials synthesis: the case of atomic layer deposition." pith.science (2026). https://pith.science/paper/7HPPLW2G

@misc{pith2026241210477,
  author       = {Pith},
  title        = {Pith review of: Benchmarking large language models for materials synthesis: the case of atomic layer deposition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7HPPLW2G}},
  note         = {Machine review of arXiv:2412.10477}
}
read the original abstract

In this work we introduce an open-ended question benchmark, ALDbench, to evaluate the performance of large language models (LLMs) in materials synthesis, and in particular in the field of atomic layer deposition, a thin film growth technique used in energy applications and microelectronics. Our benchmark comprises questions with a level of difficulty ranging from graduate level to domain expert current with the state of the art in the field. Human experts reviewed the questions along the criteria of difficulty and specificity, and the model responses along four different criteria: overall quality, specificity, relevance, and accuracy. We ran this benchmark on an instance of OpenAI's GPT-4o. The responses from the model received a composite quality score of 3.7 on a 1 to 5 scale, consistent with a passing grade. However, 36% of the questions received at least one below average score. An in-depth analysis of the responses identified at least five instances of suspected hallucination. Finally, we observed statistically significant correlations between the difficulty of the question and the quality of the response, the difficulty of the question and the relevance of the response, and the specificity of the question and the accuracy of the response as graded by the human experts. This emphasizes the need to evaluate LLMs across multiple criteria beyond difficulty or accuracy.

Figures

Figures reproduced from arXiv: 2412.10477 by the authors.

Figure 1
Figure 1. FIG. 1. Question grading in order of increasing average difficulty. The questions in the benchmark are well distributed across the 1 to 5 scale [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Question grading in order of increasing average specificity. Compared to the difficulty score, questions tend to be more clustered [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. A) Average difficulty score of the benchmark questions as [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: FIG. 5. Correlation between the average quality of the GPT4o re [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

23 extracted references · 9 canonical work pages

  1. [1]

    merlin.mbs aapmrev4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked

    FUNCTION id.bst "merlin.mbs aapmrev4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked" ENTRY address archive archivePrefix author bookaddress booktitle chapter collaboration doi edition editor eid eprint howpublished institution isbn issn journal key language month note number organization pages primaryClass publisher school SLACcitation series title translat...

  2. [2]

    merlin.mbs aipauth4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked

    FUNCTION id.bst "merlin.mbs aipauth4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked" ENTRY address archive archivePrefix author bookaddress booktitle chapter collaboration doi edition editor eid eprint howpublished institution isbn issn journal key language month note number organization pages primaryClass publisher school SLACcitation series title translat...

  3. [3]

    merlin.mbs aipnum4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked

    FUNCTION id.bst "merlin.mbs aipnum4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked" ENTRY address archive archivePrefix author bookaddress booktitle chapter collaboration doi edition editor eid eprint howpublished institution isbn issn journal key language month note number organization pages primaryClass publisher school SLACcitation series title translati...

  4. [4]

    Mirza , author N

    author author A. Mirza , author N. Alampara , author S. Kunchapu , author M. Ríos-García , author B. Emoekabu , author A. Krishnan , author T. Gupta , author M. Schilling-Wilhelmi , author M. Okereke , author A. Aneesh , author A. M. \ Elahi , author M. Asgari , author J. Eberhardt , author H. M. \ Elbeheiry , author M. V. \ Gil , author M. Greiner , auth...

  5. [5]

    author author A. D. \ White ,\ title title The future of chemistry is language , \ 10.1038/s41570-023-00502-0 journal journal Nature Reviews Chemistry \ volume 7 ,\ pages 457--458 ( year 2023 ) NoStop

  6. [6]

    author author K. M. \ Jablonka , author P. Schwaller , author A. Ortega-Guerrero , \ and\ author B. Smit ,\ title title Leveraging large language models for predictive chemistry , \ 10.1038/s42256-023-00788-1 journal journal Nature Machine Intelligence \ volume 6 ,\ pages 161--169 ( year 2024 ) NoStop

  7. [7]

    author author A. N. \ Rubungo , author K. Li , author J. Hattrick-Simpers , \ and\ author A. B. \ Dieng ,\ https://arxiv.org/abs/2411.00177 title LLM4Mat-Bench: Benchmarking large language models for materials property prediction , \ ( year 2024 ),\ http://arxiv.org/abs/2411.00177 arXiv:2411.00177 [cond-mat.mtrl-sci] NoStop

  8. [8]

    author author D. A. \ Boiko , author R. MacKnight , author B. Kline , \ and\ author G. Gomes ,\ title title Autonomous chemical research with large language models , \ 10.1038/s41586-023-06792-0 journal journal Nature \ volume 624 ,\ pages 570--578 ( year 2023 ) NoStop

Show all 23 references
  1. [9]

    author author S. M. \ George ,\ title title Atomic layer deposition: An overview , \ 10.1021/cr900056b journal journal Chemical Reviews \ volume 110 ,\ pages 111--131 ( year 2010 ) NoStop

  2. [10]

    Alvaro \ and\ author A

    author author E. Alvaro \ and\ author A. Yanguas-Gil ,\ title title Characterizing the field of atomic layer deposition: Authors, topics, and collaborations , \ 10.1371/journal.pone.0189137 journal journal PLOS ONE \ volume 13 ,\ pages 1--19 ( year 2018 ) NoStop

  3. [11]

    author author K. L. K. \ Lee , author C. Gonzales , author M. Nassar , author M. Spellings , author M. Galkin , \ and\ author S. Miret ,\ https://arxiv.org/abs/2309.05934 title MatSciML: A broad, multi-task benchmark for solid-state materials modeling , \ ( year 2023 ),\ http:...

  4. [12]

    Lee , author H

    author author Y. Lee , author H. Sun , author M. J. \ Young , \ and\ author S. M. \ George ,\ title title Atomic layer deposition of metal fluorides using HF--pyridine as the fluorine precursor , \ 10.1021/acs.chemmater.5b04360 journal journal Chemistry of Materials \ volume 2...

  5. [13]

    Pilvi , author T

    author author T. Pilvi , author T. Hatanpää , author E. Puukilainen , author K. Arstila , author M. Bischoff , author U. Kaiser , author N. Kaiser , author M. Leskelä , \ and\ author M. Ritala ,\ title title Study of a novel ALD process for depositing MgF2 thin films , \ 10.10...

  6. [14]

    https://doi.org/10.6100/alddatabase title Atomic limits ald database , \ NoStop

  7. [15]

    author author J. B. \ Kim , author D. R. \ Kwon , author K. Chakrabarti , author C. Lee , author K. Y. \ Oh , \ and\ author J. H. \ Lee ,\ title title Improvement in Al2O3 dielectric behavior by using ozone as an oxidant for the atomic layer deposition technique , \ 10.1063/1....

  8. [16]

    author author M. F. J. \ Vos , author H. C. M. \ Knoops , author R. A. \ Synowicki , author W. M. M. \ Kessels , \ and\ author A. J. M. \ Mackus ,\ title title Atomic layer deposition of aluminum fluoride using Al(CH3)3 and SF6 plasma , \ 10.1063/1.4998577 journal journal Appl...

  9. [17]

    Hornsveld , author W

    author author N. Hornsveld , author W. M. M. \ Kessels , author R. A. \ Synowicki , \ and\ author M. Creatore ,\ title title Atomic layer deposition of LiF using LiN(SiMe3)2 and SF6 plasma , \ 10.1039/D0CP05428C journal journal Phys. Chem. Chem. Phys. \ volume 23 ,\ pages 9304...

  10. [18]

    Kim , author D

    author author J. Kim , author D. Shim , author Y. Kim , \ and\ author H. Chae ,\ title title Atomic layer etching of Al2O3 with NF3 plasma fluorination and trimethylaluminum ligand exchange , \ 10.1116/6.0001616 journal journal Journal of Vacuum Science & Technology A \ volume...

  11. [19]

    author author A. M. Bran , author S. Cox , author O. Schilter , author C. Baldassari , author A. D. \ White , \ and\ author P. Schwaller ,\ title title Augmenting large language models with chemistry tools , \ 10.1038/s42256-024-00832-8 journal journal Nature Machine Intellige...

  12. [20]

    + command, where the argument is the citation key mentioned above. +

  13. [21]

    How to Grow

    + commands may be crafted by hand or, preferably, generated by using Bib . The AIP styles for REV 4 include Bib \ style files +aipnum.bst+ and +aipauth.bst+, appropriate for numbered and author-year bibliographies, respectively. REV 4 will automatically choose the style approp...

  14. [22]

    merlin.mbs apsrev4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked

    FUNCTION id.bst "merlin.mbs apsrev4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked" ENTRY address archive archivePrefix author bookaddress booktitle chapter collaboration doi edition editor eid eprint howpublished institution isbn issn journal key language month note number orga...

  15. [23]

    merlin.mbs apsrmp4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked

    FUNCTION id.bst "merlin.mbs apsrmp4-1.bst 2010-07-25 4.21a (PWD, AO, DPC) hacked" ENTRY address archive archivePrefix author bookaddress booktitle chapter collaboration doi edition editor eid eprint howpublished institution isbn issn journal key language month note number orga...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.