{"id":"eb3ef774-5189-4305-bb33-d821a928e1e7","arxiv_id":"2505.12534","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"ChemPile is an open 75-billion-token, multimodal chemical dataset spanning education, papers, property tables, code, images, and reasoning traces, released for training chemical foundation models.","lead":"ChemPile is a new open dataset of 250 gigabytes, about 75 billion tokens, of chemistry-related text, images, code, and reasoning traces for training AI models. It is the largest openly released corpus of its kind, designed to give chemical AI models the same breadth of learning materials a human chemist would see.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 14.1B-token Paper subset rests on a chemistry classifier validated on only 150 entries from a different domain (FineWebMath) with F1≈0.77; if its precision on EuroPMC is materially lower, the 'curated chemical corpus' claim is overstated.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the Paper subset's BERT classifier is validated on 150 examples with F1≈0.77, and the datasheet concedes possible non-chemical articles. My read agrees and sharpens the point: the validation set is drawn from FineWebMath, not from EuroPMC, so the reported F1 is not even measured on the target distribution. The concern is about the central claim because 'curated chemical data' is a quality predicate, and the Paper subset is the largest human-text scientific component. The paper is otherwise internally consistent: token counts sum to 76.7B, the data release appears real, and the LIFT/mLIFT components are template-generated from manually curated tabular sources. The split inconsistency in Appendix K (the pseudo-code shows a random shuffle while Section 4.7 claims Murcko scaffold splitting) is a real secondary issue for benchmarking claims, but it does not bear on the scale or curation-quality claim as directly. The appropriate verdict remains CONDITIONAL: the concern is addressable with a precision audit, but until that audit is done, the quality claim for the Paper subset is not established.","tokens_in":26981,"tokens_out":6003,"duration_ms":62978,"concrete_test":"Download a random sample of 500 documents from the released ChemPile-Paper EuroPMC subset and have two independent chemistry-trained annotators classify each as 'chemistry research', 'chemistry-related but not research', or 'not chemistry'. Compute the classifier's precision against these labels, with a confidence interval. If the estimated precision falls below 0.70, or if more than 20% of sampled documents are 'not chemistry', the claim that ChemPile-Paper is curated chemical literature is not supported at the implied quality level.","verdict_should_be":"UNCHANGED","load_bearing_attack":"ChemPile's central claim is that it is the largest open curated chemical corpus. The largest human-written scientific component, ChemPile-Paper (14.1B of 76.7B tokens, roughly 18%), is selected by a BERT multilabel classifier described in Section 4.2 and Appendix N.3.1. The reported validation is F1≈0.77 on only 150 manually annotated entries, and those entries come from FineWebMath, not from EuroPMC, the actual target distribution. The classifier is then applied to EuroPMC biomedical literature, scoring the first five 512-token chunks per document. With no confidence interval and a small validation set, the true precision on EuroPMC could be substantially below the implied level. If precision is, say, 0.65 rather than 0.77, then roughly one in three Paper documents is not chemistry research; at the scale of 11.7M documents, that is a very large volume of off-topic text. The datasheet itself concedes: 'Some of the articles in the dataset might not include chemical research and only be related to chemistry.' Since the Paper subset is the largest natural-language scientific text component and is explicitly marketed as 'curated scientific literature filtered for chemical content,' this contamination risk directly bears on the quality half of the central claim. This is a load-bearing concern even though the total token count may still be correct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"ChemPile is an open, multimodal chemical corpus reported as 255 GB of compressed Parquet data (76.7B GPT-2 tokens, 260M documents) assembled from seven subsets: ChemPile-Education (LibreTexts textbooks, MIT OCW materials, YouTube lecture transcripts, US Olympiad problems), ChemPile-Paper (chemistry-filtered EuroPMC literature, ChemRxiv/BioRxiv/MedRxiv preprints, arXiv materials-science and physical-chemistry categories, and material safety data sheets), ChemPile-LIFT and ChemPile-mLIFT (template-generated language-interfaced tabular data with SMILES/SELFIES/InChI/IUPAC representations and, for mLIFT, molecular images), ChemPile-Code (keyword-filtered StarCoder and CodeParrot), ChemPile-Reasoning (Stack Exchange Q&A plus LLM-distilled spectral-elucidation traces), and ChemPile-Caption (100K image-caption pairs from LibreTexts). The paper claims that ChemPile is the largest open curated chemical corpus at a scale suited to foundation-model pretraining, with expert-reviewed curation, consistent HuggingFace interfaces, and leakage-controlled splits. The headline arithmetic is internally consistent (Table 1 sums to roughly 255 GB and 76.7B tokens). The principal weakness is the quality control of the chemistry-filtered Paper subset (14.1B tokens), which rests on a classifier validated on only 150 out-of-domain annotations; this is the load-bearing issue examined below.","tokens_in":27225,"tokens_out":16545,"duration_ms":159081,"significance":"If the quality claims hold, ChemPile is a substantial community resource: no openly released chemistry corpus at this scale (about 75B tokens) exists, and the comparison with ChemDFM's unreleased 34B-token corpus makes that gap concrete. The paper ships reproducible infrastructure: curation scripts on GitHub, a documented sampling engine with 1,636 expert-reviewed templates, OPSIN-validated SMILES-to-IUPAC conversion, RDKit-graph-based verification of distilled reasoning traces, and a consistent HuggingFace API. The internal accounting is consistent across Table 1 and the appendix, and the multiple-representation design for identical molecules is a real contribution for representation studies. The main uncertainty is the curation quality of the Paper subset, which is the largest component of naturally occurring scientific prose and is filtered by a modestly validated classifier; the significance of the whole corpus is conditional on substantiating that filter on the target distribution. The diversity evidence in Figure 2b is suggestive rather than quantitative.","major_comments":[{"comment":"The central 'curated chemical data' claim is load-bearing on the ChemPile-Paper subset (14.1B tokens, 11.7M documents), whose EuroPMC portion is selected by a BERT multilabel classifier trained on CAMEL data and validated on about 150 manually annotated entries from FineWebMath (F1 approximately 0.77). Three points make this validation insufficient for the claim. First, the validation set comes from a different distribution than the target corpus, so the reported F1 does not estimate precision on EuroPMC; with N = 150 the uncertainty is large (a 95% confidence interval on F1 = 0.77 spans roughly plus or minus 0.07), and F1 alone does not reveal the contamination rate, so precision and recall should be reported separately. Second, the manuscript does not state whether the ChemRxiv/BioRxiv/MedRxiv preprints that feed the same 14.1B-token subset pass through this or any chemistry filter; since BioRxiv and MedRxiv are broad biomedical servers, an unfiltered inclusion would be a large off-topic source. Third, the Appendix E datasheet itself concedes that 'some of the articles in the dataset might not include chemical research and only be related to chemistry.' The manuscript also reports '3.3 billion tokens' for the EuroPMC chemistry-filtered content while Table 1 lists 14.1B tokens for the whole Paper subset, so a per-source breakdown (documents and tokens for EuroPMC, each preprint server, arXiv, and MSDS) is needed to audit the total. I request: (i) classifier precision and recall on a held-out EuroPMC sample with confidence intervals; (ii) an explicit statement of the filtering applied to each Paper source; (iii) per-source size statistics; and (iv) release of the classifier and the 150-entry annotation set.","section":"§4.7, Appendix K, datasheets 'Data Splits'"},{"comment":"The split documentation is internally inconsistent and does not support the leakage-prevention claim. Every datasheet reports train/validation/test ratios of 0.9, 0.1, and 0.1, which sum to 1.1; taken literally this is not a valid partition. In the Appendix K pseudocode, with train_fraction = 0.9 and val_fraction = 0.1, the test set (all_molecules_list[train_size + val_size:end]) is empty, and the fallback random-assignment branch can never assign rows to the test class because train_fraction + val_fraction = 1.0. In addition, Section 4.7 states that SMILES-based datasets are split by RDKit Murcko scaffold, but the provided 'scaffold splitting' pseudocode performs a plain shuffle-and-cut of a molecule list with no scaffold computation anywhere in the algorithm. As written, the protocol is not reproducible, and a user cannot tell whether the shipped splits are scaffold-based (leakage-controlled, as claimed) or random. Please provide corrected pseudocode that matches the released implementation, state the actual split fractions, and describe how scaffolds are assigned across the global molecule list.","section":"§4.3, Appendix P"},{"comment":"Appendix P reports approximately 91% accuracy for the SMILES-to-IUPAC model, but the manuscript does not specify the acceptance criterion used in the OPSIN-based 'automatic verification' of generated IUPAC names, nor the fraction of entries that were accepted or discarded. If verification only checks that OPSIN can parse the generated name (syntactic validity), names that parse to a molecule different from the intended one will pass; if it checks round-trip equivalence with the input SMILES, the residual error rate is much lower. Since ChemPile-mLIFT is the largest subset by size (155 GB), the correctness of its IUPAC field materially affects the corpus-wide quality claim. Please report the verification criterion (syntactic parse versus canonical equivalence to the source SMILES), the acceptance rate, and the post-verification error rate on a held-out sample.","section":"§4.3, Appendix P"}],"minor_comments":[{"comment":"The diversity comparison does not specify which datasets are embedded, how many samples per dataset were used, or whether sampling was balanced; the conclusion that ChemPile 'spans a larger space' is qualitative, so please report the full dataset list, sample counts, and, ideally, a quantitative coverage or volume metric.","section":"§2.3, §3, Figure 2b"},{"comment":"The text describes ChemPile as released under a 'permissive license,' but the actual licenses include CC BY-NC-SA 4.0 (Education, LIFT, mLIFT, Caption) and CC BY-NC-ND 4.0 (Paper), which are non-commercial and, for Paper, no-derivatives; this wording is misleading for potential commercial pretraining use and should be corrected.","section":"§1, §3, Appendix J"},{"comment":"Token counts are computed with different tokenizers across corpora (GPT-2 tiktoken for ChemPile versus undisclosed tokenizers for ChemDFM and BioGPT), so a caveat that the comparison is tokenizer-dependent would keep the 'more than 50% larger' statement within its uncertainty.","section":"§3, Figure 2a"},{"comment":"The statement that 'no assessment was conducted regarding the model's adherence to instructed output formatting guidelines' should be reconciled with the claim that reasoning traces were parsed via the [START_REASONING]/[END_REASONING] tags; the appendix should also report the fraction of LLM generations discarded for failing parsing or correctness checks.","section":"§N.4"},{"comment":"The PaperScraper tool is cited to a drug-design review (Born and Manica, Current Medicinal Chemistry 2021); a direct citation of the PaperScraper repository or its accompanying paper would help readers locate the tool.","section":"§4.2, reference [77]"},{"comment":"The representation-embedding correlation analysis uses OpenAI's text-embedding-3-large, which is not reproducible and was not trained on chemical text; the conclusion that IUPAC embeddings are closest to established similarity measures is model-dependent and should be flagged as such.","section":"Appendix M"},{"comment":"The Known Limitations entry refers to 'the accuracy of the classifier used to select the code,' but Section 4.4 describes regular-expression-based keyword filtering; the terminology should be aligned across the paper and the datasheet.","section":"Appendix F (ChemPile-Code datasheet)"}],"recommendation":"major_revision","confidential_remarks":"To the editor: my recommendation rests on the Paper-subset validation gap and the split-documentation inconsistencies, both of which are fixable within revision; I would move toward rejection only if new evidence showed a collapse of classifier precision on EuroPMC or confirmed that BioRxiv/MedRxiv content was included with no topical filter. One citation-related item to verify during production: the CAMEL dataset cited as the classifier training source (20,000 examples per discipline) is, in its primary published form, an LLM agent-conversation corpus; the authors should be asked to point to the exact data artifact used for topic classification. Given that the paper's audience will rely on the filter, releasing the EuroPMC classifier and the annotation set as part of the data release should be encouraged."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, this is a real resource release, not vaporware: 250GB, 76.7B tokens, per-subset datasheets, HuggingFace hosting, and a sampling engine for the LIFT and mLIFT subsets. Second, the curation claim is weaker than the abstract implies, and the weak point is exactly where the reader's report puts it. ChemPile-Paper, the 14.1B-token subset, is filtered by a BERT classifier with F1≈0.77 measured on 150 manually annotated FineWebMath entries, not on EuroPMC, which is the actual target distribution. The datasheet itself concedes that some articles might only be related to chemistry rather than being chemical research. That is the soft spot to fix before this becomes canonical.\n\nThe genuinely new parts are worth crediting. No prior open corpus combines educational text, filtered papers, language-interfaced tabular data in multiple representations, code, reasoning traces, and image-caption pairs under one API. The 1,636 manually curated templates are a concrete contribution; the SMILES-to-IUPAC model, trained and checked against OPSIN, is reproducible; and the scaffold split across tabular datasets is a sensible attempt to prevent leakage. The internal accounting checks out: token counts sum correctly, sizes match, and the diversity analysis uses external embedding models. The curation process is documented to a degree that is unusual for this field.\n\nThe EuroPMC classifier concern is real and not manufactured. 150 validation samples from FineWebMath do not give a precision estimate on EuroPMC, and there is no confidence interval. With 11.7M documents, even a small drop in precision means a large volume of off-topic text. The fix is straightforward: annotate a sample from EuroPMC, report precision and recall with error bars, or describe that subset as “chemically related literature” rather than “curated chemical literature.” Also, two GitHub links in the text are blank placeholders; that is minor but sloppy and should be fixed. The LLM-generated reasoning traces are not fact-checked, though the paper does check SMILES validity; for a 2M-token subset this is minor, but it should be stated in the datasheet.\n\nThis is not a paper with a load-bearing flaw. The central claim—largest open curated chemical corpus—holds up if the Paper-subset validation is addressed or the claim is softened. Who is this for? Groups training chemical foundation models, benchmark builders, and anyone studying data mixing. It deserves a serious referee. I would send it to review with a request for the EuroPMC validation, the missing links, and a slightly more measured abstract.","headline":"ChemPile is a genuinely large, openly released chemical corpus that the field will likely use, but the quality of its largest paper subset rests on a classifier validated on 150 examples from a different domain; the paper should be published only after the authors either tighten that validation or soften the curation claim.","tokens_in":27891,"tokens_out":2170,"would_cite":true,"duration_ms":22617,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ChemPile gathers over 75 billion tokens of curated chemical data into an open corpus built for training chemical foundation models.","keywords":["chemical foundation models","large language models","dataset curation","multimodal learning","SMILES","SELFIES","IUPAC names","molecular representations"],"falsifier":"Draw a random sample of about 1,000 documents from the Paper subset and have two independent chemists label each as chemistry research or not; if the measured precision falls well short of the approximately 0.77 F1 reported, or if a large fraction of the corpus's biggest text subsets is clearly off-topic, the curated-chemical-corpus claim is contradicted.","tokens_in":26738,"feed_emoji":"🧪","tokens_out":9263,"duration_ms":89113,"temperature":0.7,"pith_summary":"The paper introduces ChemPile, an open dataset of about 255 GB and 76.7 billion tokens assembled for training and evaluating general-purpose foundation models in chemistry. It combines seven complementary subsets — educational texts, chemistry-filtered research papers, language-interfaced tabular data in text and image form, chemistry-related code, reasoning traces, and image-caption pairs — so that a model can learn chemistry the way a chemist does, from foundational concepts to specialized reasoning and multimodal problem solving. The authors claim this is the largest openly available curated chemical corpus and the only one at a scale suited to foundation-model pretraining. If the corpus is as clean and diverse as reported, it would let the community study data mixing, molecular representation choice, and scaling behavior in chemistry on reproducible, permissively licensed data with consistent train/validation/test splits.","feed_headline":"ChemPile: 75B curated tokens for chemical AI","feed_subtitle":"One open corpus unites papers, textbooks, code, reasoning traces, and molecular images for chemical foundation models.","key_machinery":"The load-bearing object is ChemPile itself, built as reproducible infrastructure rather than a single static file. Three mechanisms carry the argument: a transformer-based text classifier that filters a large literature corpus down to chemistry-related papers; a sampling engine that converts tabular chemistry datasets into natural-language questions by filling 1,636 expert-written templates with randomized synonyms, enumeration schemes, and multiple-choice options, then expands every entry into SMILES, SELFIES, InChI, IUPAC names, and images; and a split protocol that assigns every molecule to a global scaffold-based train/validation/test partition so the same molecular core never appears in both training and evaluation.","core_discovery":"ChemPile's central claim is that a single, openly released corpus can supply the volume, diversity, and quality that chemical foundation models have been missing. The dataset spans seven subsets built from very different sources: textbooks and lecture transcripts; research articles filtered from large abstract and full-text corpora; tabular chemistry datasets converted into natural language through 1,636 hand-written templates; those same tabular data expanded into multiple molecular representations and rendered molecular images; code filtered from large permissively licensed code corpora; community question-answer data; and synthetic reasoning traces for interpreting molecular spectra. Totaling about 255 GB, 76.7 billion tokens, and 260 million documents, the corpus is, the paper states, larger than the 34-billion-token corpus behind the largest previously reported chemical foundation model and orders of magnitude larger than released chemical instruction datasets. The paper also provides scaffold-based splits designed to keep the same molecular scaffold out of both training and test partitions, and it reports that embeddings of IUPAC names track molecular similarity more closely (r=0.722) than embeddings of SMILES strings (r=0.521), evidence that representation choice matters for how well the corpus can teach chemistry.","pith_inferences":["If the scale claim holds, chemical benchmarks should show scaling-law-like log-linear gains in downstream accuracy as ChemPile token counts grow; the paper does not itself present such scaling curves.","The reported IUPAC-embedding advantage suggests a single-representation ablation could show that IUPAC-heavy pretraining beats SMILES-only training for property prediction, but that model comparison is not in the paper.","Because much of the reasoning subset was generated by LLMs, models trained on ChemPile may inherit model-specific spectral-assignment errors; comparing downstream reasoning accuracy with and without the synthetic traces would separate distillation gains from distillation noise.","The template-sampling engine is domain-agnostic in design, so the same curation recipe could convert tabular data from other sciences into language-interfaced instruction data."],"forward_implications":["Chemical foundation models can be pretrained entirely on open data at a scale (roughly 76.7 billion tokens) that was previously available only to much smaller or closed chemical datasets.","Because the same molecules appear in SMILES, SELFIES, IUPAC, InChI, and rendered images, researchers can directly test which representation or combination transfers best to property prediction and inverse design.","The scaffold-based splits make benchmark results comparable across labs by preventing the same molecular scaffold from appearing in both training and test partitions.","The modular subsets support data-mixing studies, including how much code, reasoning-trace, or image-caption data improves chemical reasoning and multimodal understanding."],"supporting_citations":[{"why":"Supplies the source corpus of abstracts and full-text articles from which the Paper subset is filtered.","marker":"[74]"},{"why":"Provides the 20,000-examples-per-discipline training data used to build the chemical-content classifier.","marker":"[75]"},{"why":"Supplies the annotations used to estimate the classifier's F1 of approximately 0.77 on 150 cases.","marker":"[76]"},{"why":"Provides the tool used to collect and download preprint PDFs for the Paper subset.","marker":"[77]"},{"why":"Performs OCR-based text extraction from the preprint PDFs.","marker":"[78]"},{"why":"Source code corpus that is regex-filtered for chemistry-related code in the Code subset.","marker":"[81]"},{"why":"Underlying permissively licensed code corpus from which the StarCoder source was derived.","marker":"[82]"},{"why":"Defines the language-interfaced tabular format that the LIFT and mLIFT subsets instantiate with templates.","marker":"[25]"},{"why":"Supplies the scaffold-splitting convention used to build leakage-free splits across tabular subsets.","marker":"[39]"}],"fun_headline_variants":["ChemPile: 250GB open corpus for chemical AI training","ChemPile: 75B tokens spanning textbooks to spectral reasoning","ChemPile: 260M documents and 75B tokens for chemistry models","Open ChemPile: 250GB diverse data for chemical foundation models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire 'curated chemical corpus' claim rests on the assumption that the classifier used to select the paper subset is accurate enough, but it was validated on only about 150 manually labeled examples with an F1 near 0.77, so a modest drop in precision would admit substantial non-chemical text into the 14.1-billion-token Paper subset.","fun_headline_variants_meta":{"raw":{"variants":["ChemPile: 250GB open corpus for chemical AI training","ChemPile: 75B tokens spanning textbooks to spectral reasoning","ChemPile: 260M documents and 75B tokens for chemistry models","Open ChemPile: 250GB diverse data for chemical foundation models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00034,"raw_usage":{"total_tokens":1927,"prompt_tokens":1046,"completion_tokens":881,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":802}},"tokens_in":662,"tokens_out":881,"duration_ms":7604,"temperature":1.0,"reasoning_tokens":802,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:32:01.448499+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Draw a random sample of about 1,000 documents from the Paper subset and have two independent chemists label each as chemistry research or not; if the measured precision falls well short of the approximately 0.77 F1 reported, or if a large fraction of the corpus's biggest text subsets is clearly off-topic, the curated-chemical-corpus claim is contradicted.","supporting_citations":[{"cited_title":"Europe PMC in 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the source corpus of abstracts and full-text articles from which the Paper subset is filtered."},{"cited_title":"Trends in Deep Learning for Property-driven Drug Design","cited_arxiv_id":null,"evidence_quote":"Provides the tool used to collect and download preprint PDFs for the Paper subset."},{"cited_title":"StarCoder: may the source be with you!","cited_arxiv_id":null,"evidence_quote":"Source code corpus that is regex-filtered for chemistry-related code in the Code subset."},{"cited_title":"MoleculeNet: a benchmark for molecular machine learning","cited_arxiv_id":null,"evidence_quote":"Supplies the scaffold-splitting convention used to build leakage-free splits across tabular subsets."}],"review_version":1}