{"id":"5789690f-747a-4441-976d-c8ed6d794793","arxiv_id":"2411.15221","paper_version":2,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A community report describing 34 hackathon-built LLM applications for materials science and chemistry, with reflections on the event format and preliminary project results.","lead":"This paper reports the outcomes of a 2024 hackathon in which 34 teams built LLM-based tools for materials science and chemistry. It catalogs the projects and argues that such events show LLMs are becoming practical for rapid prototyping in these fields.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'significant improvements in LLM capabilities' claim in the abstract and conclusion is unsupported: no matched comparison with the 2023 hackathon is provided, so the trend is inferred only from model-version releases rather than from measured project outcomes.","rationale":"The reader's verdict of UNVERDICTED is appropriate for a community event report, and I found no evidence of fraud or internal inconsistency. The strongest claim in the paper is the year-over-year improvement in LLM capabilities. For that claim to hold, the paper would need to rule out confounds between the two events—differences in team composition, task difficulty, and model versions. It does not. Instead, the conclusion infers improvement from the release of new model versions, which is not a result of the hackathon. The absence of any matched benchmark makes the 'significant improvements' claim unfalsifiable in this paper. This is a real but genre-proportionate deficiency, so I do not move the verdict from UNVERDICTED. I partially agree with the reader's weakest assumption: the self-selectivity of teams is also a concern, but the sharper problem is that the paper's central comparative claim lacks any comparison. The proposed test—a matched-pair rerun of tasks common to both hackathons—would settle whether the claimed trend is real; if such tasks cannot be matched, the authors should soften the abstract to an anecdotal observation.","tokens_in":48437,"tokens_out":4501,"duration_ms":41324,"concrete_test":"Construct a matched-pair test using the two hackathon papers: pick an application area present in both (e.g., literature-to-structured-data extraction or property prediction from text), extract the task definition and dataset from the 2023 paper (ref [53]), and run the 2024-era models (GPT-4, Claude 3, Llama 3) and 2023-era models (GPT-3.5, Llama 2) on identical inputs with identical prompts and evaluation metrics. If 2024 models do not outperform 2023 models on the matched tasks, the 'significant improvements' claim is contradicted. If the two papers do not share any sufficiently specified task to allow such a match, the claim should be explicitly downgraded to anecdotal in the paper's conclusions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim with the strongest scientific content is stated in the abstract: 'the event highlighted significant improvements in LLM capabilities since the previous year's hackathon.' This is a comparative claim about a trend across two hackathon events, but the paper provides no matched comparison. Projects from 2024 (this paper) and 2023 (ref [53]) differ in teams, tasks, datasets, and evaluation protocols; no project here benchmarks current models against the models used last year. The Conclusion instead attributes the improvement to an external event: 'the performance across the diverse application space was improved simply via the release of new versions of Gemini, ChatGPT, Claude, Llama, and other models.' That is an inference from model-version releases, not a measurement from the reported work. Several 2024 projects actually report failure or non-significant results: LLMads describes hallucinated list data, the small-LLM concrete study says Phi-3 failed after five development cycles, G-Peer-T reports p=0.07, and LLMSpectrometry solved only 3 of 19 NMR spectra. These mixed outcomes make the blanket claim of 'significant improvements' especially fragile. If the improvement claim is removed, the paper still documents 34 prototypes, but its headline finding—the year-over-year trend—is unsupported. This is a load-bearing gap because the abstract promotes it as the main takeaway.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports on the second Large Language Model Hackathon for Applications in Materials Science and Chemistry, held in May 2024 with hybrid physical hubs and a global online hub. It describes 34 team submissions, of which 32 have written reports in the appendix, organizes them into seven application areas (property prediction, design, automation and interfaces, communication and education, data management, hypothesis generation, and knowledge extraction), and highlights exemplar projects in each area. The paper also discusses the event format and concludes that the event demonstrated the dual utility of LLMs as multipurpose predictive models and as rapid-prototyping platforms, and that LLM capabilities improved significantly since the previous year's hackathon. Individual appendix reports provide code links and, in several cases, candid statements of limitations.","tokens_in":48690,"tokens_out":4848,"duration_ms":49027,"significance":"The paper's main value is as an archival record of a large, open, community-driven prototyping event: it collects 32 project reports with public code repositories, spans a broad application taxonomy, and includes unusually honest reporting of failures and non-significant results. If its broader claims were supported, the paper would provide useful evidence that LLMs can be productively applied across a wide range of materials chemistry tasks within a 24-hour hackathon format. However, the headline claims of the abstract and conclusion go beyond what the evidence supports: no matched comparison with the previous hackathon is provided, and several reported projects show hallucinations, model failure, or non-significant results. These issues are local to the framing and can be corrected by softening the claims; the descriptive core of the paper remains valuable.","major_comments":[{"comment":"The abstract states that 'the event highlighted significant improvements in LLM capabilities since the previous year's hackathon', and the conclusion attributes this improvement to the release of new model versions. No matched comparison between the 2024 and 2023 hackathons is provided: the teams, tasks, datasets, and evaluation protocols differ, and no project benchmarks current models against the models used in the previous event. Moreover, several projects in the appendix report negative or mixed outcomes, including LLMSpectrometry solving only 3 of 19 NMR spectra (§5), the Phi-3 model failing after the fifth development cycle in the small-LLM concrete study (§8), hallucinated values in LLMads (§18), and a non-significant p=0.07 in G-Peer-T (§24). The improvement claim should either be removed or replaced by a carefully hedged statement that the breadth of prototypes is consistent with, but does not measure, year-over-year capability growth.","section":"Abstract; Conclusion; Appendix §§5, 8, 18, 24"},{"comment":"The manuscript reports '34 team submissions' but Table 1 and the appendix table of contents list only 32 projects. The abstract also says that each team submission is presented in a summary table and as brief papers in the appendix, which is inconsistent with the '32 submissions with a written description included here' formulation in the Event Overview. Please reconcile the total: either list the two additional submissions (even if they lack write-ups or code), or change the stated total to 32.","section":"Abstract; Event Overview; Table 1; Appendix TOC"},{"comment":"The claim that the event 'demonstrated the dual utility of LLMs' is stronger than the evidence warrants. The paper is a collection of self-selected, time-constrained prototypes, and several appendix reports explicitly note missing validation: Learning LOBSTERs states that no five-fold cross-validation was implemented (§1), NOMAD Query Reporter reports hallucinations on heterogeneous data (§19), and the small-LLM study reports a complete failure mode for Phi-3 (§8). The conclusion should be phrased in terms of demonstrated breadth of prototyping and potential utility, not demonstrated utility, unless the authors add a systematic validation criterion applied across projects.","section":"Conclusion; Abstract"}],"minor_comments":[{"comment":"The caption reads 'Overview of the tools developed by the various tools'; 'tools' should presumably be 'teams'.","section":"Table 1 caption"},{"comment":"The sentence 'all investigated models generated designs that outperformed the statistical baseline' is immediately followed by an exception for Phi-3 with design context, which failed after the fifth development cycle; please rephrase to state that all models outperformed the baseline in the first round, with the noted exception in later rounds.","section":"§8.4, Table 2"},{"comment":"In the NOMAD Query Reporter section, the in-text citations appear swapped: 'Query Reporter [1]' and 'NOMAD [2]' should refer to the GitHub repository and the NOMAD paper respectively, but the reference list assigns [1] to the NOMAD paper and [2] to the GitHub repository.","section":"§19, references"},{"comment":"The project name is spelled inconsistently as 'yeLLowhaMmer' in the overview and 'yeLLowhaMMer' in the appendix title; please standardize the spelling.","section":"Overview; Appendix §17"}],"recommendation":"major_revision","confidential_remarks":"This is a community event report rather than a methods paper, and its descriptive content is useful. The main risk is overclaiming in the abstract and conclusion; the unsupported year-over-year improvement claim and the count inconsistency are both fixable within the manuscript's scope. I see no concerns about authorship or citation practices beyond standard event-report conventions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful community catalogue, not a research result, and the one place where it reaches beyond its evidence is the abstract's claim of \"significant improvements\" over last year's hackathon. I'd send it to peer review as a report, but I'd ask the authors to cut or heavily qualify that trend claim.\n\nWhat's genuinely new: the 2024 edition of the hackathon report, with 34 projects organized into seven application areas, links to code for nearly all of them, and a decent summary of what was attempted. A few artifacts stand out as reusable: ChemQA is a new multimodal QA benchmark; the LK-99 Bayesian tracker is a clever demonstration of temporal evidence aggregation; yeLLowhaMMer shows a working multimodal agent for lab notebooks; LangSim is a clean interface between LLMs and ASE-style simulation workflows. The paper is also unusually honest for this genre: LLMads reports hallucinated values, the small-LLM concrete study says Phi-3 failed after five rounds, G-Peer-T reports p=0.07, and LLMSpectrometry solved only 3/19 NMR spectra. That candor makes the catalogue trustworthy.\n\nThe main soft spot is the year-over-year improvement claim. There is no matched comparison with the 2023 hackathon: teams, tasks, datasets, and protocols all differ, and no project benchmarks current models against last year's models. The conclusion attributes the improvement to external model releases, which is an inference, not a measurement. Several 2024 entries themselves report failures, so the blanket \"significant improvements\" reads as event enthusiasm rather than evidence. That said, this is a one-sentence overreach in an otherwise descriptive report, not a load-bearing flaw: if you delete the trend claim, the catalogue still stands.\n\nOther soft spots are minor and genre-typical: projects are self-selected and mostly unvalidated prototypes, and many apply established RAG, fine-tuning, and agent patterns to domain targets. The paper acknowledges most of this. I don't see any circularity problem; self-citations to the prior hackathon and related method papers are contextual.\n\nWho should read it: anyone wanting a quick map of what LLMs were tried for in materials and chemistry in 2024, or looking for starting points and code for a specific task. It is not a paper that resolves a scientific question.\n\nRecommendation: send it to a venue that publishes community reports. A serious referee can ask for a rewritten abstract and a limitations paragraph on selection bias and validation. It deserves referee time.","headline":"A useful, honest catalogue of 34 LLM prototypes in materials and chemistry; the abstract's 'significant improvements' trend claim is unsupported and should be qualified.","tokens_in":49970,"tokens_out":1999,"would_cite":true,"duration_ms":19251,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"34 teams produced working LLM prototypes across seven areas of materials science and chemistry, and the organizers read this as evidence that LLMs now serve as both general-purpose predictors and rapid-prototyping platforms.","keywords":["large language models","materials science","chemistry","hackathon","molecular property prediction","retrieval-augmented generation","scientific workflows","rapid prototyping"],"falsifier":"Inspect the linked repositories and rerun the key claims in a controlled setting: five-fold cross-validate the phonon peak predictor, re-run the XRD schema-filling pipeline on the same raw files and count hallucinated values, and test the novelty-scoring model on a larger abstract set. If the flagship prototypes do not reproduce or their error rates equal random baselines, the paper's evidence for expanded LLM capability collapses.","tokens_in":48272,"feed_emoji":"🧪","tokens_out":5202,"duration_ms":55941,"temperature":0.7,"pith_summary":"This report tries to establish that modern LLMs have become practical tools for materials science and chemistry, not just research curiosities. It summarizes 34 projects built within a 24-hour hackathon, covering molecular property prediction, molecular design, automation, education, data management, hypothesis evaluation, and literature mining. If the report's reading is right, the same class of models can serve as general-purpose prediction engines and as a fast prototyping layer for custom scientific tools. The paper also argues that the hybrid hackathon format, with physical hubs and an online community, supports durable new scientific collaborations.","feed_headline":"34 teams show LLMs remaking chemistry workflows in 24 hours","feed_subtitle":"A hackathon report finds the same models predict properties, run instruments, and mine the literature.","key_machinery":"The carrying mechanism is the time-boxed, open-ended hybrid challenge itself: 556 registered participants, 34 completed team submissions, each producing code and a short project report. The organizers use those reports as their dataset, group them into seven application areas, and compare the demonstrated capabilities against those of the previous year's event.","core_discovery":"The paper claims that second-generation LLMs now have a dual utility in materials and chemistry: they are multipurpose models that can be pointed at regression, generation, and extraction tasks, and they are platforms on which researchers can prototype custom applications in under a day. The evidence is the set of 34 hackathon submissions, each with code and a written report, which the organizers sorted into seven application areas. They observe that performance across this application space improved substantially since the previous year's hackathon, in many cases simply because newer model versions were released, and they take this as a sign that LLM use in these fields will keep expanding.","pith_inferences":["The report's emphasis on breadth over benchmarks suggests the fastest near-term gains will come in text-native data work, such as literature extraction, electronic lab notebooks, and report drafting, where success is judged by workflow completion rather than scientific accuracy.","Because many prototypes lack error bars, the honest reading is that LLMs widen what can be attempted, not yet what can be trusted; a reproducibility pass across the 34 repositories would separate stable tools from one-off demonstrations.","The low-data property-prediction results hint that context augmentation, such as bonding analysis and literature summaries, may matter more than model scale, a claim the paper supports only partially and that deserves direct comparison against classical machine-learning baselines.","If small locally hosted models continue to approach larger ones on design tasks, the practical ceiling may shift from model capability to prompt and workflow engineering, which would make data privacy less of an obstacle."],"forward_implications":["Low-data property prediction appears within reach: teams improved phonon peak and lithium-ion conductivity predictions by enriching composition inputs with bonding or literature context.","Design tasks can be bootstrapped with zero-shot LLM suggestions, and small locally hosted models beat random baselines in concrete formulation design, suggesting that private and data-sensitive laboratories could use them.","Natural-language interfaces for instruments and simulations are feasible: teams automated bulk-modulus calculations, microscope parameter estimation, and DFT setup, lowering the need for specialist operators.","Text-native workflows, such as knowledge-graph construction, schema filling, hypothesis evaluation, and question answering, are where LLMs most readily plug in, because these tasks align with the models' core strengths.","If the trend continues, newer model generations alone will broaden the range of research tasks that LLMs can handle without custom adaptation."],"supporting_citations":[{"why":"Establishes that LLMs can match or beat conventional machine learning on molecular property prediction in low-data settings, a premise for the property-prediction category.","marker":"[4]"},{"why":"Supplies further evidence that LLMs generalize to prediction tasks from limited data, cited by the organizers in framing that application area.","marker":"[5]"},{"why":"Recent work on LLMs for materials property prediction, used to support the claim that this area has advanced rapidly.","marker":"[6]"},{"why":"Defines retrieval-augmented generation, the technique that several teams use to ground molecular design and literature extraction in source documents.","marker":"[20]"},{"why":"Provides the agent-chaining layer that several submissions used to turn LLM prompts into executable workflows, underwriting the automation category.","marker":"[29]"},{"why":"Records the previous year's hackathon and its 14 examples, the baseline against which this paper argues LLM capability has significantly improved.","marker":"[53]"}],"fun_headline_variants":["LLM hackathon yields 34 projects spanning 7 chemistry tasks","34 teams show LLMs' dual role in materials and chemistry","Hackathon shows LLMs from property prediction to lab automation","LLMs: multipurpose models for chemistry in a day"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion rests on treating 34 self-selected, largely unvalidated team reports as representative evidence; many entries lack cross-validation, statistical significance, or quantitative checks, so if the reports flatter reality, the case for broad LLM utility is not made.","fun_headline_variants_meta":{"raw":{"variants":["LLM hackathon yields 34 projects spanning 7 chemistry tasks","34 teams show LLMs' dual role in materials and chemistry","Hackathon shows LLMs from property prediction to lab automation","LLMs: multipurpose models for chemistry in a day"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000924,"raw_usage":{"total_tokens":3937,"prompt_tokens":901,"completion_tokens":3036,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":2965}},"tokens_in":517,"tokens_out":3036,"duration_ms":17806,"temperature":1.0,"reasoning_tokens":2965,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:56:01.429260+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the linked repositories and rerun the key claims in a controlled setting: five-fold cross-validate the phonon peak predictor, re-run the XRD schema-filling pipeline on the same raw files and count hallucinated values, and test the novelty-scoring model on a larger abstract set. If the flagship prototypes do not reproduce or their error rates equal random baselines, the paper's evidence for expanded LLM capability collapses.","supporting_citations":[],"review_version":1}