REVIEW 2 major objections 5 minor 59 references
On chemistry SMILES, BPE and Unigram-LM build near-disjoint vocabularies and cut molecules to different depths.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 03:41 UTC pith:AIZQ5DRQ
load-bearing objection Controlled, reproducible evidence that BPE and Unigram-LM stay near-disjoint on SMILES even with a fixed chemistry base; the algorithm is not a free default, though no LM is trained. the 2 major comments →
Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Holding the 165-token chemistry base, corpus, vocabulary size, and boundary policy fixed, BPE and Unigram-LM produce near-disjoint multi-glyph vocabularies (unweighted Jaccard at most 0.161, frequency-weighted at most 0.05) and Unigram-LM segments held-out molecules into 29–41 percent more tokens, with BPE a strict coarsening of Unigram-LM on 80–99 percent of molecules. The separation is stable across three corpus typologies, both boundary policies, and vocabulary sizes through an 8192-token anchor.
What carries the argument
The matched-condition grid (algorithm × corpus typology × vocabulary size × bracket boundary) over a shared 165-token OpenSMILES base, scored by membership overlap, fertility gap, and positional nestedness of cuts.
Load-bearing premise
The claim that the small-vocabulary regime is the practically relevant one rests on an NMT-derived rule of thumb for when token embeddings are learnable, whose transfer to chemistry is assumed rather than proven.
What would settle it
Train matched BPE and Unigram-LM models at the same small vocabulary on the same SMILES corpus and measure whether the tokenizer-level gaps in membership and fertility produce a measurable difference in held-out likelihood, generation validity, or property-prediction accuracy; if the algorithms remain interchangeable downstream, the modeling-decision claim weakens.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a controlled tokenizer-level comparison of BPE and Unigram-LM on chemistry SMILES. Holding fixed a complete 165-token OpenSMILES/Smirk base, it trains both algorithms across three corpus typologies (PubChem, ZINC-22, COCONUT), both bracket-boundary policies, and small target vocabularies (V∈{256,512,1024}, with 2048 and 8192 anchors), yielding 22 matched conditions. Across all of them the learned multi-glyph vocabularies are near-disjoint (unweighted Jaccard ≤0.161; frequency-weighted ≤0.05), Unigram-LM emits 29–41% more tokens per held-out molecule, and BPE’s parse is a strict coarsening of Unigram-LM’s on 80–99% of molecules. The separation is stable under hyperparameter sweeps, size-matched typology probes, cross-domain transfer, and non-canonical rewrites. No language models are trained; the claim is that the subword algorithm is a modeling decision rather than a free default.
Significance. If the measurements hold, the paper cleanly isolates an axis the chemistry-LM literature has largely treated as a free default. Prior chemistry head-to-heads either confounded algorithm with base/coverage or operated at large native V; this design removes those confounds and shows the natural-language BPE/Unigram structural contrast survives a tiny, valence-constrained alphabet. Strengths include the fixed complete base, byte-identical retrain checks, bootstrap CIs on held-out quantities, exhaustive sensitivity and interaction surfaces, size-matched typology probes, and full public release of tokenizers and per-condition measurements. That package is sufficient to reframe tokenizer choice as a design decision even without downstream LM results.
major comments (2)
- The central claim is framed as holding “where token embeddings are learnable,” certified by the transferred NMT bar F95%,100 (c100≥0.95 on learned pieces; §3.5, Appendix A.4, Table A14/A20). The Limitations section correctly flags the transfer as an assumption. The raw Jaccard and fertility numbers do not depend on the bar (it gates only rare-token diagnostics), but the interpretive step that the small-V grid is the practically decisive regime does. A short additional calibration—e.g., reporting clearance under a few alternative (p,n) pairs already recoverable from Table A20, or a brief note on how the bar would shift under a chemistry-specific frequency floor—would harden that framing without expanding scope.
- The paper asserts that the algorithm is a “modeling decision” while training no language models (§1, §5, Conclusion). Intrinsic separation need not imply a downstream gap (the manuscript cites Ali et al. 2024 on this point). The claim is carefully hedged as a precondition rather than a quality ranking, which is appropriate, but the Discussion should state more explicitly that the present evidence alone does not establish which arm is preferable for any particular chemical LM task; that remains the follow-on experiment the work motivates.
minor comments (5)
- Figure 1 caption and the nicotine/serotonin example would benefit from an explicit note that the illustrated gap (+5/+7) is larger than the held-out average (~one-third), so readers do not over-generalize from the figure.
- Table 2 packs seven scalars per condition; a short legend or footnote clarifying which Δ quantities are signed BPE−UL versus one-sided magnitudes would reduce lookup cost.
- The special-token budgeting offset (Unigram-LM carries six more content pieces at matched V; Appendix A.2) is correctly argued to shrink rather than inflate the fertility gap, but a one-sentence reminder in the main-text Methods would help readers who skip the appendix.
- A few long sentences in §4.1–§4.2 (membership and compatibility paragraphs) could be split for readability without changing content.
- The arXiv identifier in the header (2607.05691) and the Zenodo DOI are useful; ensuring both appear in the final Data availability statement will aid archival linking.
Circularity Check
Empirical tokenizer comparison with no circular derivation: overlap, fertility, and nestedness are measured on held-out data, not fitted then re-predicted.
full rationale
This paper is a controlled empirical comparison of two subword algorithms (BPE vs Unigram-LM) on chemistry SMILES. It trains tokenizers over a fixed external 165-token Smirk base, measures Jaccard overlap on multi-glyph pieces, fertility on held-out splits, boundary nestedness, absorption, and related diagnostics across corpora (PubChem, ZINC-22, COCONUT, REAL-Space), and reports that the algorithms do not converge. There is no derivation that claims a first-principles prediction from fitted parameters; the quantities are direct set and segmentation statistics (Eqs. 1–4 and positional cut agreement). Use of the Smirk base and a pinned fork is shared tooling, not a self-citation that forces the contrast. The NMT-derived F95%,100 learnability bar is an external assumption that frames the small-V regime; the paper itself flags the transfer as an assumption and states that the bar gates only rare-token diagnostics, not the headline overlap and fertility contrasts. No step reduces a claimed prediction to its inputs by construction. Score 0 is the correct honest finding.
Axiom & Free-Parameter Ledger
free parameters (4)
- Target vocabulary sizes V ∈ {256,512,1024} (+2048, 8192 anchors)
- F95%,100 learnability bar (c100 ≥ 0.95)
- Unigram-LM trainer defaults (max_piece_length=16, seed_size=1e6, n_sub_iterations=2, shrinking_factor=0.75)
- BPE min_frequency=2
axioms (6)
- domain assumption Smirk's 165-token OpenSMILES glyph base is complete for conformant SMILES and is a fair shared atomic level for both algorithms.
- domain assumption BPE (bottom-up greedy merge) and Unigram-LM (top-down likelihood prune) are the right algorithm pair; WordPiece is out of scope.
- domain assumption RDKit isomeric non-Kekulé canonicalization plus exact-string dedup defines the training distribution without distorting the algorithm contrast.
- ad hoc to paper The NMT-derived F95%,100 clearance bar transfers to chemistry token embeddings.
- ad hoc to paper Tokenizer-level membership, fertility, and nestedness differences are sufficient to call the subword algorithm a 'modeling decision' without training language models.
- standard math Standard Jaccard, fertility, imbalance D, and bootstrap-over-molecules statistics are appropriate effect-size measures for the contrast.
read the original abstract
Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding (BPE) from natural language with little scrutiny. In natural language, BPE's principal alternative, Unigram-LM, is known to build structurally different vocabularies. Whether that contrast survives in chemistry was open. We report a controlled comparison of BPE and Unigram-LM over a fixed 165-token chemistry base, at the small vocabulary sizes where token embeddings are learnable, across three corpus typologies (diverse, drug-like, natural-products) and both pre-tokenization boundary policies. The two do not converge. In all 22 matched conditions they build near-disjoint subword vocabularies: cross-algorithm Jaccard overlap on the learned pieces never exceeds 0.161, and at most 0.05 once weighted toward the high-frequency pieces a model updates most. Unigram-LM also segments held-out molecules into 29-41% more tokens; the arms largely agree on where to cut but not how deeply, so BPE's segmentation is a strict coarsening of Unigram-LM's on 80-99% of molecules. The separation holds across corpus, boundary, and vocabulary size, persisting even at eight times that scale. The subword algorithm is therefore a modeling decision, not a free default. The study trains no language models.
Figures
Reference graph
Works this paper leans on
-
[1]
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Schulze Buschhoff, Charvi Jain, Alexander E. Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suárez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, and Nicolas Flores-Herr. Token...
-
[2]
Stop taking tokenizers for granted: They are core design decisions in large language models
Sawsan Alqahtani, Mir Tafseer Nayeem, Md Tahmid Rahman Laskar, Tasnim Mohiuddin, and M Saiful Bari. Stop taking tokenizers for granted: They are core design decisions in large language models. InProceedings of the 2026 Conference of the European Chapter of the Association for Computational Linguistics, pages 8410–8432. Association for Computational Lingui...
-
[3]
Josep Arús-Pous, Simon V . Johansson, Oleksii Prykhodko, Esben Jannik Bjerrum, Christian Tyrchan, Jean-Louis Reymond, Hongming Chen, and Ola Engkvist. Randomized SMILES strings improve the quality of molecular generative models.Journal of Cheminformatics, 11(1):71, 2019. doi: 10.1186/s13321-019-0393-0
-
[4]
David Balcells and Bastian Bjerkem Skjelstad. tmQM dataset—quantum geometries and properties of 86k transition metal complexes.Journal of Chemical Information and Modeling, 60(12):6135–6146, 2020. doi: 10.1021/acs.jcim.0c01041
-
[5]
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. Byte pair encoding is suboptimal for language model pretraining. InFindings of the Association for Computational Linguistics: EMNLP 2020, pages 4617–4624, 2020. doi: 10.18653/v1/2020. findings-emnlp.414. URLhttps://aclanthology.org/2020.findings-emnlp.414
-
[6]
Gayane Chilingaryan, Hovhannes Tamoyan, Ani Tevosyan, Nelly Babayan, Lusine Khondkaryan, Karen Ham- bardzumyan, Zaven Navoyan, Hrant Khachatrian, and Armen Aghajanyan. BARTSmiles: Generative masked language models for molecular representations.Journal of Chemical Information and Modeling, 64(15):5832– 5843, 2024. doi: 10.1021/acs.jcim.4c00512
-
[7]
Chemberta: Large-scale self-supervised pretraining for molecular property prediction, 2020
Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. Chemberta: Large-scale self-supervised pretraining for molecular property prediction, 2020. 15 Where to cut, how deep: BPE vs Unigram-LMA PREPRINT
work page 2020
-
[8]
Two counterexamples to tokenization and the noiseless channel
Marco Cognetta, Vilém Zouhar, Sangwhan Moon, and Naoaki Okazaki. Two counterexamples to tokenization and the noiseless channel. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, pages 13674–13683, 2024. URL https://aclanthology.org/2024. lrec-main.1469
work page 2024
-
[9]
Investigating the effectiveness of BPE: The power of shorter sequences
Matthias Gallé. Investigating the effectiveness of BPE: The power of shorter sequences. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1375–1381. Association for Computational Linguistics,
work page 2019
-
[10]
URLhttps://aclanthology.org/D19-1141/
doi: 10.18653/v1/D19-1141. URLhttps://aclanthology.org/D19-1141/
-
[11]
Veronika Ganeeva, Kuzma Khrabrov, Artur Kadurin, and Elena Tutubalina. Measuring chemical LLM robustness to molecular representations: a SMILES variation-based framework.Journal of Cheminformatics, 17(1):164,
-
[12]
doi: 10.1186/s13321-025-01079-0
-
[13]
Finding the optimal vocabulary size for neural machine translation
Thamme Gowda and Jonathan May. Finding the optimal vocabulary size for neural machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3955–3964. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.findings-emnlp.352. URL https://aclanthology. org/2020.findings-emnlp.352/
-
[14]
Oleksandr O. Grygorenko, Dmytro S. Radchenko, Igor Dziuba, Alexander Chuprina, Kateryna E. Gubina, and Yurii S. Moroz. Generating multibillion chemical space of readily accessible screening compounds.iScience, 23 (11):101681, 2020. doi: 10.1016/j.isci.2020.101681
-
[15]
Hunter Heidenreich. Smirk fork for the vocabulary–tokenizer comparison study: shared glyph-id front-end with GpeTrainer (bpe) and a unigram-lm sibling trainer. GitHub release vtc-2026-05-24, 2026. URL https://github.com/hunter-heidenreich/smirk/releases/tag/vtc-2026-05-24 . Public, pinned fork of Smirk. The BPE and Unigram-LM arms share thecompute_alphabe...
work page 2026
-
[16]
Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
Sukjun Hwang et al. Dynamic chunking for end-to-end hierarchical sequence modeling, 2025. URL https: //arxiv.org/abs/2507.07955
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[17]
Craig A. James. OpenSMILES specification, version 1.0, 2016. URL http://opensmiles.org/opensmiles. html. Blue Obelisk project; dated 2016-05-15
work page 2016
-
[18]
Prathamesh Kalamkar, Ned Letcher, Meissane Chami, Sahger Lad, Shayan Mohanty, and Prasanna Pendse. The tokenization bottleneck: How vocabulary extension improves chemistry representation learning in pretrained language models, 2025
work page 2025
-
[19]
Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A. Shoemaker, Paul A. Thiessen, Bo Yu, Leonid Zaslavsky, Jian Zhang, and Evan E. Bolton. PubChem 2023 update.Nucleic Acids Research, 51(D1):D1373–D1380, 2023. doi: 10.1093/nar/gkac956
-
[20]
Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alán Aspuru-Guzik. Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation.Machine Learning: Science and Technology, 1(4):045024, 2020. doi: 10.1088/2632-2153/aba947
-
[21]
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. Subword regularization: Improving neural network translation models with multiple subword candidates. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66–75, Melbourne, Australia, 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1007. URLhttps://aclanth...
-
[22]
Taku Kudo and John Richardson. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, 2018. doi: 10.18653/v1/D18-2012. URL https://aclanthology.org/D18-2012
work page internal anchor Pith review doi:10.18653/v1/d18-2012 2018
-
[23]
Fishing for Magikarp: Automatically detecting under-trained tokens in large language models
Sander Land and Max Bartolo. Fishing for Magikarp: Automatically detecting under-trained tokens in large language models. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 11631–11646. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.emnlp-main.649. URLhttps://aclanthology.org/2024.emnlp-...
-
[24]
RDKit: Open-source cheminformatics, version 2026.03.1, 2026
Greg Landrum, Paolo Tosco, Brian Kelley, David Cosgrove, Riccardo Vianello, et al. RDKit: Open-source cheminformatics, version 2026.03.1, 2026. URL https://doi.org/10.5281/zenodo.19250388. Zenodo, release2026_03_1
-
[25]
Miguelangel Leon, Yuriy Perezhohin, Fernando Peres, Aleš Popoviˇc, and Mauro Castelli. Comparing SMILES and SELFIES tokenization for enhanced chemical language modeling.Scientific Reports, 14(1):25016, 2024. doi: 10.1038/s41598-024-76440-8. 16 Where to cut, how deep: BPE vs Unigram-LMA PREPRINT
-
[26]
Jianan Li, Keisuke Yanagisawa, Masahito Sugita, Takuya Fujie, Masahito Ohue, and Yutaka Akiyama. CycPeptM- PDB: A comprehensive database of membrane permeability of cyclic peptides.Journal of Chemical Information and Modeling, 63(7):2240–2250, 2023. doi: 10.1021/acs.jcim.2c01573
-
[27]
Xinhao Li and Denis Fourches. SMILES pair encoding: A data-driven substructure tokenization algorithm for deep learning.Journal of Chemical Information and Modeling, 61(4):1560–1569, 2021. doi: 10.1021/acs.jcim.0c01127
-
[28]
Haoran Lian, Yizhe Xiong, Jianwei Niu, Shasha Mo, Zhenpeng Su, Zijia Lin, Hui Chen, Peng Liu, Jungong Han, and Guiguang Ding. Scaffold-bpe: Enhancing byte pair encoding for large language models with simple and effective scaffold token removal, 2024. URLhttps://arxiv.org/abs/2404.17808
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[29]
SuperBPE: Space Travel for Language Models
Alisa Liu, Jonathan Hayase, Valentin Hofmann, Sewoong Oh, Noah A. Smith, and Yejin Choi. Superbpe: Space travel for language models. InProceedings of the Second Conference on Language Modeling (COLM), 2025. URLhttps://arxiv.org/abs/2503.13423
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[30]
HuggingFace’s tokenizers: Fast state-of-the-art tokenizers optimized for research and production
Anthony Moi and Nicolas Patry. HuggingFace’s tokenizers: Fast state-of-the-art tokenizers optimized for research and production. https://github.com/huggingface/tokenizers, 2023. URL https://github.com/ huggingface/tokenizers
work page 2023
-
[31]
Pit Neitemeier et al. Hierarchical autoregressive transformers: Combining byte- and word-level processing for robust, adaptable language models, 2025. URLhttps://arxiv.org/abs/2501.10322
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[32]
Noel M. O’Boyle. Towards a universal SMILES representation - a standard method to generate canonical SMILES based on the InChI.Journal of Cheminformatics, 4(1):22, 2012. doi: 10.1186/1758-2946-4-22
-
[33]
Noel M. O’Boyle and Andrew Dalke. DeepSMILES: An adaptation of SMILES for use in machine-learning of chemical structures.ChemRxiv, 2018. doi: 10.26434/chemrxiv.7097960.v1
-
[34]
Byte Latent Transformer: Patches Scale Better Than Tokens
Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, and Srinivasan Iyer. Byte latent transformer: Patches scale better than tokens, 2024. URLhttps://arxiv.org/abs/2412.09871
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[35]
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena V oita. BPE-dropout: Simple and effective subword regularization. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 1882–1892, 2020. doi: 10.18653/v1/2020.acl-main.170
-
[36]
Sridhar Radhakrishnan, Krish Mody, Arvind Venkatesh, and Ananth Venkatesh. Optimizing SMILES token sequences via trie-based refinement and transition graph filtering.Journal of Cheminformatics, 18(1):13, 2026. doi: 10.1186/s13321-025-01143-9
-
[37]
Bharath Ramsundar, Peter Eastman, Patrick Walters, Vijay Pande, Karl Leswing, and Zhenqin Wu.Deep Learning for the Life Sciences: Applying Deep Learning to Genomics, Microscopy, Drug Discovery, and More. O’Reilly Media, 2019. ISBN 978-1492039839
work page 2019
-
[38]
How Much is Enough? The Diminishing Returns of Tokenization Training Data
Varshini Reddy, Craig W. Schmidt, Yuval Pinter, and Chris Tanner. How much is enough? the diminishing returns of tokenization training data, 2025. URL https://arxiv.org/abs/2502.20273. ICML 2025 Tokenization Workshop (TokShop)
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[39]
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vuli´c, Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Pape...
-
[40]
ReactionT5: a large-scale pre-trained model towards application of limited reaction data
Tatsuya Sagawa and Ryosuke Kojima. ReactionT5: a large-scale pre-trained model towards application of limited reaction data.arXiv preprint, 2023. doi: 10.48550/arxiv.2311.06708
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2311.06708 2023
-
[41]
Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris C
Craig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris C. Tanner. Tokenization is more than compression. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 678–702. Association for Computational Linguistics, 2024. doi: 10.18653/ v1/2024.emnlp-main.40. URLhttps://acla...
work page 2024
-
[42]
Schmidt, Varshini Reddy, Chris Tanner, and Yuval Pinter
Craig W. Schmidt, Varshini Reddy, Chris Tanner, and Yuval Pinter. Boundless byte pair encoding: Breaking the pre-tokenization barrier. InConference on Language Modeling (COLM), 2025. URL https://arxiv.org/ abs/2504.00178
-
[43]
Japanese and Korean voice search
Mike Schuster and Kaisuke Nakajima. Japanese and Korean voice search. In2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5149–5152. IEEE, 2012. doi: 10.1109/ICASSP. 2012.6289079. 17 Where to cut, how deep: BPE vs Unigram-LMA PREPRINT
-
[44]
Philippe Schwaller, Théophile Gaudin, Dávid Lányi, Costas Bekas, and Teodoro Laino. “found in translation”: predicting outcomes of complex organic chemistry reactions using neural sequence-to-sequence models.Chemical Science, 9(28):6091–6098, 2018. doi: 10.1039/c8sc02339e
-
[45]
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics, pages 1715–1725,
-
[46]
URLhttps://aclanthology.org/P16-1162
doi: 10.18653/v1/P16-1162. URLhttps://aclanthology.org/P16-1162
-
[47]
Michael A. Skinnider. Invalid SMILES are beneficial rather than detrimental to chemical language models.Nature Machine Intelligence, 6(4):437–448, 2024. doi: 10.1038/s42256-024-00821-x
-
[48]
Maria Sorokina, Peter Merseburger, Kohulan Rajan, Mehmet Aziz Yirik, and Christoph Steinbeck. COCONUT online: Collection of open natural products database.Journal of Cheminformatics, 13(1):2, 2021. doi: 10.1186/ s13321-020-00478-9
work page 2021
-
[49]
Linguistic laws meet protein sequences: A comparative analysis of subword tokenization methods, 2024
Burak Suyunu, Enes Taylan, and Arzucan Özgür. Linguistic laws meet protein sequences: A comparative analysis of subword tokenization methods, 2024
work page 2024
-
[50]
Ülgen, Nilgün Karalı, and Arzucan Özgür
Asu Bü¸ sra Temizer, Gökçe Uludo˘gan, Rıza Özçelik, Taha Koulani, Elif Özkırımlı, Kutlu Ö. Ülgen, Nilgün Karalı, and Arzucan Özgür. Exploring data-driven chemical SMILES tokenization approaches to identify key protein-ligand binding moieties.Molecular Informatics, 43(3):e202300249, 2024. doi: 10.1002/minf.202300249
-
[51]
Benjamin I. Tingle, Khanh G. Tang, Mar Castanon, John J. Gutierrez, Munkhzul Khurelbaatar, Chinzorig Dandarchuluun, Yurii S. Moroz, and John J. Irwin. ZINC-22 – a free multi-billion-scale database of tangible compounds for ligand discovery.Journal of Chemical Information and Modeling, 63(4):1166–1176, 2023. doi: 10.1021/acs.jcim.2c01253
-
[52]
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: Open and efficient foundation language models, 2023. URL https: //arxiv.org/abs/2302.13971
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[53]
Ucak, Islambek Ashyrmamatov, and Juyong Lee
Umit V . Ucak, Islambek Ashyrmamatov, and Juyong Lee. Improving the quality of chemical language model outcomes with atom-in-SMILES tokenization.Journal of Cheminformatics, 15(1):55, 2023. doi: 10.1186/ s13321-023-00725-9
work page 2023
-
[54]
Alexius Wadell, Anoushka Bhutani, and Venkatasubramanian Viswanathan. Tokenization for molecular foundation models.Journal of Chemical Information and Modeling, 66(3):1384–1393, 2026. doi: 10.1021/acs.jcim.5c01856
-
[55]
Tokenization is sensitive to language variation
Anna Wegmann, Dong Nguyen, and David Jurgens. Tokenization is sensitive to language variation. InFindings of the Association for Computational Linguistics: ACL 2025, pages 10958–10983. Association for Computa- tional Linguistics, 2025. doi: 10.18653/v1/2025.findings-acl.572. URL https://aclanthology.org/2025. findings-acl.572/
-
[56]
SMILES, a chemical language and information system
David Weininger. SMILES, a chemical language and information system. 1. introduction to methodology and encoding rules.Journal of Chemical Information and Computer Sciences, 28(1):31–36, 1988. doi: 10.1021/ ci00057a005
work page 1988
-
[57]
Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S
Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay S. Pande. MoleculeNet: a benchmark for molecular machine learning.Chemical Science, 9 (2):513–530, 2018. doi: 10.1039/c7sc02664a
-
[58]
Barbara Zdrazil, Eloy Félix, Fiona Hunter, Emma J. Manners, James Blackshaw, Sybilla Corbett, Marleen De Veij, Harris Ioannidis, David Méndez, Juan F Mosquera, María Paula Magariños, Nicolas Bosc, Ricardo Arcila, Tevfik Kizilören, Anna Gaulton, A. Patrícia Bento, Melissa F. Adasme, Peter Monecke, Gregory A. Landrum, and Andrew R. Leach. The ChEMBL databas...
-
[59]
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. Tokenization and the noiseless channel. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5184–5207. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.acl-long.284. URLhttp...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.