Pith. sign in

REVIEW 2 major objections 5 minor 59 references

On chemistry SMILES, BPE and Unigram-LM build near-disjoint vocabularies and cut molecules to different depths.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 03:41 UTC pith:AIZQ5DRQ

load-bearing objection Controlled, reproducible evidence that BPE and Unigram-LM stay near-disjoint on SMILES even with a fixed chemistry base; the algorithm is not a free default, though no LM is trained. the 2 major comments →

arxiv 2607.05691 v1 pith:AIZQ5DRQ submitted 2026-07-06 cs.CL cs.LGq-bio.BM

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES

classification cs.CL cs.LGq-bio.BM
keywords SMILEStokenizationBPEUnigram-LMchemical language modelsvocabulary overlapfertilitysubword algorithms
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Chemical language models that read SMILES always start with a tokenizer, and the field has mostly inherited byte-pair encoding from natural language without testing whether that choice still matters. This paper holds the chemistry base fixed at 165 tokens and compares BPE against Unigram-LM at the small vocabulary sizes where embeddings can still be learned, across diverse, drug-like, and natural-product corpora and under both ways of treating bracketed atoms. In every matched condition the two algorithms build almost non-overlapping subword inventories and Unigram-LM emits roughly one-third more tokens per molecule. They largely agree on where to cut the string but not how deeply, so BPE's parse is usually a strict coarsening of Unigram-LM's. The algorithm is therefore a real modeling decision, not an interchangeable default; the study itself trains no language models and leaves the downstream quality comparison open.

Core claim

Holding the 165-token chemistry base, corpus, vocabulary size, and boundary policy fixed, BPE and Unigram-LM produce near-disjoint multi-glyph vocabularies (unweighted Jaccard at most 0.161, frequency-weighted at most 0.05) and Unigram-LM segments held-out molecules into 29–41 percent more tokens, with BPE a strict coarsening of Unigram-LM on 80–99 percent of molecules. The separation is stable across three corpus typologies, both boundary policies, and vocabulary sizes through an 8192-token anchor.

What carries the argument

The matched-condition grid (algorithm × corpus typology × vocabulary size × bracket boundary) over a shared 165-token OpenSMILES base, scored by membership overlap, fertility gap, and positional nestedness of cuts.

Load-bearing premise

The claim that the small-vocabulary regime is the practically relevant one rests on an NMT-derived rule of thumb for when token embeddings are learnable, whose transfer to chemistry is assumed rather than proven.

What would settle it

Train matched BPE and Unigram-LM models at the same small vocabulary on the same SMILES corpus and measure whether the tokenizer-level gaps in membership and fertility produce a measurable difference in held-out likelihood, generation validity, or property-prediction accuracy; if the algorithms remain interchangeable downstream, the modeling-decision claim weakens.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The manuscript reports a controlled tokenizer-level comparison of BPE and Unigram-LM on chemistry SMILES. Holding fixed a complete 165-token OpenSMILES/Smirk base, it trains both algorithms across three corpus typologies (PubChem, ZINC-22, COCONUT), both bracket-boundary policies, and small target vocabularies (V∈{256,512,1024}, with 2048 and 8192 anchors), yielding 22 matched conditions. Across all of them the learned multi-glyph vocabularies are near-disjoint (unweighted Jaccard ≤0.161; frequency-weighted ≤0.05), Unigram-LM emits 29–41% more tokens per held-out molecule, and BPE’s parse is a strict coarsening of Unigram-LM’s on 80–99% of molecules. The separation is stable under hyperparameter sweeps, size-matched typology probes, cross-domain transfer, and non-canonical rewrites. No language models are trained; the claim is that the subword algorithm is a modeling decision rather than a free default.

Significance. If the measurements hold, the paper cleanly isolates an axis the chemistry-LM literature has largely treated as a free default. Prior chemistry head-to-heads either confounded algorithm with base/coverage or operated at large native V; this design removes those confounds and shows the natural-language BPE/Unigram structural contrast survives a tiny, valence-constrained alphabet. Strengths include the fixed complete base, byte-identical retrain checks, bootstrap CIs on held-out quantities, exhaustive sensitivity and interaction surfaces, size-matched typology probes, and full public release of tokenizers and per-condition measurements. That package is sufficient to reframe tokenizer choice as a design decision even without downstream LM results.

major comments (2)
  1. The central claim is framed as holding “where token embeddings are learnable,” certified by the transferred NMT bar F95%,100 (c100≥0.95 on learned pieces; §3.5, Appendix A.4, Table A14/A20). The Limitations section correctly flags the transfer as an assumption. The raw Jaccard and fertility numbers do not depend on the bar (it gates only rare-token diagnostics), but the interpretive step that the small-V grid is the practically decisive regime does. A short additional calibration—e.g., reporting clearance under a few alternative (p,n) pairs already recoverable from Table A20, or a brief note on how the bar would shift under a chemistry-specific frequency floor—would harden that framing without expanding scope.
  2. The paper asserts that the algorithm is a “modeling decision” while training no language models (§1, §5, Conclusion). Intrinsic separation need not imply a downstream gap (the manuscript cites Ali et al. 2024 on this point). The claim is carefully hedged as a precondition rather than a quality ranking, which is appropriate, but the Discussion should state more explicitly that the present evidence alone does not establish which arm is preferable for any particular chemical LM task; that remains the follow-on experiment the work motivates.
minor comments (5)
  1. Figure 1 caption and the nicotine/serotonin example would benefit from an explicit note that the illustrated gap (+5/+7) is larger than the held-out average (~one-third), so readers do not over-generalize from the figure.
  2. Table 2 packs seven scalars per condition; a short legend or footnote clarifying which Δ quantities are signed BPE−UL versus one-sided magnitudes would reduce lookup cost.
  3. The special-token budgeting offset (Unigram-LM carries six more content pieces at matched V; Appendix A.2) is correctly argued to shrink rather than inflate the fertility gap, but a one-sentence reminder in the main-text Methods would help readers who skip the appendix.
  4. A few long sentences in §4.1–§4.2 (membership and compatibility paragraphs) could be split for readability without changing content.
  5. The arXiv identifier in the header (2607.05691) and the Zenodo DOI are useful; ensuring both appear in the final Data availability statement will aid archival linking.

Circularity Check

0 steps flagged

Empirical tokenizer comparison with no circular derivation: overlap, fertility, and nestedness are measured on held-out data, not fitted then re-predicted.

full rationale

This paper is a controlled empirical comparison of two subword algorithms (BPE vs Unigram-LM) on chemistry SMILES. It trains tokenizers over a fixed external 165-token Smirk base, measures Jaccard overlap on multi-glyph pieces, fertility on held-out splits, boundary nestedness, absorption, and related diagnostics across corpora (PubChem, ZINC-22, COCONUT, REAL-Space), and reports that the algorithms do not converge. There is no derivation that claims a first-principles prediction from fitted parameters; the quantities are direct set and segmentation statistics (Eqs. 1–4 and positional cut agreement). Use of the Smirk base and a pinned fork is shared tooling, not a self-citation that forces the contrast. The NMT-derived F95%,100 learnability bar is an external assumption that frames the small-V regime; the paper itself flags the transfer as an assumption and states that the bar gates only rare-token diagnostics, not the headline overlap and fertility contrasts. No step reduces a claimed prediction to its inputs by construction. Score 0 is the correct honest finding.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

The central claim is empirical and rests on standard tokenizer algorithms, a published chemistry base, public corpora, and a transferred NMT learnability threshold. No new physical entities are postulated. Free parameters are design choices (V grid, default trainer knobs, bar thresholds) rather than fitted constants that force the result. Domain assumptions about OpenSMILES completeness and RDKit canonicalization are explicit and largely stress-tested.

free parameters (4)
  • Target vocabulary sizes V ∈ {256,512,1024} (+2048, 8192 anchors)
    Chosen by hand to sit in the small-vocabulary regime where embeddings are argued to be learnable; the claim that divergence matters most here depends on this grid.
  • F95%,100 learnability bar (c100 ≥ 0.95)
    Threshold pair (p=0.95, n=100) transferred from Gowda & May NMT work; gates which conditions are certified and frames the 'learnable regime' narrative.
  • Unigram-LM trainer defaults (max_piece_length=16, seed_size=1e6, n_sub_iterations=2, shrinking_factor=0.75)
    Reference SentencePiece defaults held fixed for the headline grid; sensitivity shows shrinking factor moves the Unigram piece set, so the exact vocabulary is schedule-dependent.
  • BPE min_frequency=2
    Sennrich default; only free BPE training knob; swept but fixed for headline results.
axioms (6)
  • domain assumption Smirk's 165-token OpenSMILES glyph base is complete for conformant SMILES and is a fair shared atomic level for both algorithms.
    Installed as length-1 pieces for both arms (§3.3, Appendix B.1); conformance filter drops off-base molecules. Different bases would be separate experiments (Limitations).
  • domain assumption BPE (bottom-up greedy merge) and Unigram-LM (top-down likelihood prune) are the right algorithm pair; WordPiece is out of scope.
    Stated in §2.2; WordPiece noted as BPE-like and excluded.
  • domain assumption RDKit isomeric non-Kekulé canonicalization plus exact-string dedup defines the training distribution without distorting the algorithm contrast.
    §3.2, Appendix A.1; toolkit-dependent canonical SMILES is acknowledged; OpenBabel rewrite probe partially stress-tests write-stability.
  • ad hoc to paper The NMT-derived F95%,100 clearance bar transfers to chemistry token embeddings.
    Explicitly flagged as an assumption in Limitations; used to define the learnable regime and gate rare-token diagnostics (§3.5).
  • ad hoc to paper Tokenizer-level membership, fertility, and nestedness differences are sufficient to call the subword algorithm a 'modeling decision' without training language models.
    Core interpretive claim of abstract/conclusion; paper carefully withholds downstream quality claims but still elevates structural non-interchangeability to a modeling decision (§5).
  • standard math Standard Jaccard, fertility, imbalance D, and bootstrap-over-molecules statistics are appropriate effect-size measures for the contrast.
    §3.4–3.5; conventional tokenizer metrics carried from NLP literature.

pith-pipeline@v1.1.0-grok45 · 56979 in / 3940 out tokens · 47384 ms · 2026-07-11T03:41:29.486389+00:00 · methodology

0 comments
read the original abstract

Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding (BPE) from natural language with little scrutiny. In natural language, BPE's principal alternative, Unigram-LM, is known to build structurally different vocabularies. Whether that contrast survives in chemistry was open. We report a controlled comparison of BPE and Unigram-LM over a fixed 165-token chemistry base, at the small vocabulary sizes where token embeddings are learnable, across three corpus typologies (diverse, drug-like, natural-products) and both pre-tokenization boundary policies. The two do not converge. In all 22 matched conditions they build near-disjoint subword vocabularies: cross-algorithm Jaccard overlap on the learned pieces never exceeds 0.161, and at most 0.05 once weighted toward the high-frequency pieces a model updates most. Unigram-LM also segments held-out molecules into 29-41% more tokens; the arms largely agree on where to cut but not how deeply, so BPE's segmentation is a strict coarsening of Unigram-LM's on 80-99% of molecules. The separation holds across corpus, boundary, and vocabulary size, persisting even at eight times that scale. The subword algorithm is therefore a modeling decision, not a free default. The study trains no language models.

Figures

Figures reproduced from arXiv: 2607.05691 by Hunter Heidenreich.

Figure 1
Figure 1. Figure 1: Nicotine and serotonin under the two algorithms’ [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Cross-algorithm vocabulary overlap for 22 matched conditions, frequency-weighted (Jw, left) and unweighted (J, right) on a shared log axis; color and marker by corpus, filled NMB / open MB, with target V on the x-axis. Every condition stays near-disjoint, and Jw < J throughout. Structural variants in [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Segmentation and nesting on three held-out molecules under the matched PubChem [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Held-out fertility at V =1024 as a dumbbell chart: each row joins BPE (blue) to Unigram-LM (orange) on a shared tokens-per-molecule axis, so connector length is the absolute fertility gap, comparable across corpora. Unigram-LM lies to the right of BPE (more tokens, less compression) in every corpus and under both boundary policies (NMB, permeable MB), with the relative gap rel|∆f| (Eq. 2) annotated per row… view at source ↗
Figure 5
Figure 5. Figure 5: Scale does not dissolve the divergence, and could not be pushed much further if it did. Both panels track [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: BPE grows a vocabulary bottom-up; Unigram-LM prunes a large seed top-down. The [PITH_FULL_IMAGE:figures/full_fig_p019_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Per-knob sensitivity response curves on the fixed PubChem [PITH_FULL_IMAGE:figures/full_fig_p021_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: The four pairwise hyperparameter interactions, each a heatmap of the cross-arm vocabulary overlap [PITH_FULL_IMAGE:figures/full_fig_p021_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Algorithm×boundary interaction: the signed NMB−MB difference per contrast, plotted against V with corpus given by color and marker. The panels are, top to bottom, membership overlap J, the signed fertility gap fBPE − fUL (in tokens), and the signed imbalance gap DBPE − DUL; each carries a zero baseline off which the sign reads. No contrast’s outcome flips between the opaque-bracket (NMB) and permeable-brac… view at source ↗
Figure 10
Figure 10. Figure 10: The learned multi-glyph vocabularies are near-disjoint by composition, in every matched condition. Each bar [PITH_FULL_IMAGE:figures/full_fig_p025_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Glyph-length distribution of the BPE-only and Unigram-LM-only multi-glyph pieces (PubChem, ZINC [PITH_FULL_IMAGE:figures/full_fig_p027_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Unigram-LM piece-length distribution on PubChem ( [PITH_FULL_IMAGE:figures/full_fig_p027_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Adjacent base-glyph pairs within learned multi-glyph pieces (PubChem [PITH_FULL_IMAGE:figures/full_fig_p028_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Write-stability under non-canonical SMILES at [PITH_FULL_IMAGE:figures/full_fig_p033_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Structural-subword overlap for all 22 matched conditions, the appendix companion to [PITH_FULL_IMAGE:figures/full_fig_p035_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Cross-V trends of the three direct contrasts across the three multi-V corpora (the single-V REAL-Space anchor is omitted), the visual companion to [PITH_FULL_IMAGE:figures/full_fig_p036_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Token-distribution intrinsics at V =1024, one dumbbell per corpus×boundary joining BPE (blue) to Unigram￾LM (orange) on each metric’s own axis: imbalance D (divergence from uniform; left), normalized Shannon entropy η (center), and Rényi efficiency at α=2.5 (right). BPE is the more uniform arm in every condition and on all three metrics (lower D, higher η, higher R), a coherent within-family signature, wh… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

59 extracted references · 59 canonical work pages · 9 internal anchors

  1. [1]

    Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suárez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, and Nicolas Flores-Herr

    Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Schulze Buschhoff, Charvi Jain, Alexander E. Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suárez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, and Nicolas Flores-Herr. Token...

  2. [2]

    Stop taking tokenizers for granted: They are core design decisions in large language models

    Sawsan Alqahtani, Mir Tafseer Nayeem, Md Tahmid Rahman Laskar, Tasnim Mohiuddin, and M Saiful Bari. Stop taking tokenizers for granted: They are core design decisions in large language models. InProceedings of the 2026 Conference of the European Chapter of the Association for Computational Linguistics, pages 8410–8432. Association for Computational Lingui...

  3. [3]

    Johansson, Oleksii Prykhodko, Esben Jannik Bjerrum, Christian Tyrchan, Jean-Louis Reymond, Hongming Chen, and Ola Engkvist

    Josep Arús-Pous, Simon V . Johansson, Oleksii Prykhodko, Esben Jannik Bjerrum, Christian Tyrchan, Jean-Louis Reymond, Hongming Chen, and Ola Engkvist. Randomized SMILES strings improve the quality of molecular generative models.Journal of Cheminformatics, 11(1):71, 2019. doi: 10.1186/s13321-019-0393-0

  4. [4]

    tmQM dataset—quantum geometries and properties of 86k transition metal complexes.Journal of Chemical Information and Modeling, 60(12):6135–6146, 2020

    David Balcells and Bastian Bjerkem Skjelstad. tmQM dataset—quantum geometries and properties of 86k transition metal complexes.Journal of Chemical Information and Modeling, 60(12):6135–6146, 2020. doi: 10.1021/acs.jcim.0c01041

  5. [5]

    Byte pair encoding is suboptimal for language model pretraining

    Kaj Bostrom and Greg Durrett. Byte pair encoding is suboptimal for language model pretraining. InFindings of the Association for Computational Linguistics: EMNLP 2020, pages 4617–4624, 2020. doi: 10.18653/v1/2020. findings-emnlp.414. URLhttps://aclanthology.org/2020.findings-emnlp.414

  6. [6]

    BARTSmiles: Generative masked language models for molecular representations.Journal of Chemical Information and Modeling, 64(15):5832– 5843, 2024

    Gayane Chilingaryan, Hovhannes Tamoyan, Ani Tevosyan, Nelly Babayan, Lusine Khondkaryan, Karen Ham- bardzumyan, Zaven Navoyan, Hrant Khachatrian, and Armen Aghajanyan. BARTSmiles: Generative masked language models for molecular representations.Journal of Chemical Information and Modeling, 64(15):5832– 5843, 2024. doi: 10.1021/acs.jcim.4c00512

  7. [7]

    Chemberta: Large-scale self-supervised pretraining for molecular property prediction, 2020

    Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. Chemberta: Large-scale self-supervised pretraining for molecular property prediction, 2020. 15 Where to cut, how deep: BPE vs Unigram-LMA PREPRINT

  8. [8]

    Two counterexamples to tokenization and the noiseless channel

    Marco Cognetta, Vilém Zouhar, Sangwhan Moon, and Naoaki Okazaki. Two counterexamples to tokenization and the noiseless channel. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, pages 13674–13683, 2024. URL https://aclanthology.org/2024. lrec-main.1469

  9. [9]

    Investigating the effectiveness of BPE: The power of shorter sequences

    Matthias Gallé. Investigating the effectiveness of BPE: The power of shorter sequences. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1375–1381. Association for Computational Linguistics,

  10. [10]

    URLhttps://aclanthology.org/D19-1141/

    doi: 10.18653/v1/D19-1141. URLhttps://aclanthology.org/D19-1141/

  11. [11]

    Measuring chemical LLM robustness to molecular representations: a SMILES variation-based framework.Journal of Cheminformatics, 17(1):164,

    Veronika Ganeeva, Kuzma Khrabrov, Artur Kadurin, and Elena Tutubalina. Measuring chemical LLM robustness to molecular representations: a SMILES variation-based framework.Journal of Cheminformatics, 17(1):164,

  12. [12]

    doi: 10.1186/s13321-025-01079-0

  13. [13]

    Finding the optimal vocabulary size for neural machine translation

    Thamme Gowda and Jonathan May. Finding the optimal vocabulary size for neural machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3955–3964. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.findings-emnlp.352. URL https://aclanthology. org/2020.findings-emnlp.352/

  14. [14]

    Grygorenko, Dmytro S

    Oleksandr O. Grygorenko, Dmytro S. Radchenko, Igor Dziuba, Alexander Chuprina, Kateryna E. Gubina, and Yurii S. Moroz. Generating multibillion chemical space of readily accessible screening compounds.iScience, 23 (11):101681, 2020. doi: 10.1016/j.isci.2020.101681

  15. [15]

    Smirk fork for the vocabulary–tokenizer comparison study: shared glyph-id front-end with GpeTrainer (bpe) and a unigram-lm sibling trainer

    Hunter Heidenreich. Smirk fork for the vocabulary–tokenizer comparison study: shared glyph-id front-end with GpeTrainer (bpe) and a unigram-lm sibling trainer. GitHub release vtc-2026-05-24, 2026. URL https://github.com/hunter-heidenreich/smirk/releases/tag/vtc-2026-05-24 . Public, pinned fork of Smirk. The BPE and Unigram-LM arms share thecompute_alphabe...

  16. [16]

    Dynamic Chunking for End-to-End Hierarchical Sequence Modeling

    Sukjun Hwang et al. Dynamic chunking for end-to-end hierarchical sequence modeling, 2025. URL https: //arxiv.org/abs/2507.07955

  17. [17]

    Craig A. James. OpenSMILES specification, version 1.0, 2016. URL http://opensmiles.org/opensmiles. html. Blue Obelisk project; dated 2016-05-15

  18. [18]

    The tokenization bottleneck: How vocabulary extension improves chemistry representation learning in pretrained language models, 2025

    Prathamesh Kalamkar, Ned Letcher, Meissane Chami, Sahger Lad, Shayan Mohanty, and Prasanna Pendse. The tokenization bottleneck: How vocabulary extension improves chemistry representation learning in pretrained language models, 2025

  19. [19]

    Shoemaker, Paul A

    Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A. Shoemaker, Paul A. Thiessen, Bo Yu, Leonid Zaslavsky, Jian Zhang, and Evan E. Bolton. PubChem 2023 update.Nucleic Acids Research, 51(D1):D1373–D1380, 2023. doi: 10.1093/nar/gkac956

  20. [20]

    Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation.Machine Learning: Science and Technology, 1(4):045024, 2020

    Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alán Aspuru-Guzik. Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation.Machine Learning: Science and Technology, 1(4):045024, 2020. doi: 10.1088/2632-2153/aba947

  21. [21]

    Subword regularization: Improving neural network translation models with multiple subword candidates

    Taku Kudo. Subword regularization: Improving neural network translation models with multiple subword candidates. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66–75, Melbourne, Australia, 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1007. URLhttps://aclanth...

  22. [22]

    Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing

    Taku Kudo and John Richardson. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, 2018. doi: 10.18653/v1/D18-2012. URL https://aclanthology.org/D18-2012

  23. [23]

    Fishing for Magikarp: Automatically detecting under-trained tokens in large language models

    Sander Land and Max Bartolo. Fishing for Magikarp: Automatically detecting under-trained tokens in large language models. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 11631–11646. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.emnlp-main.649. URLhttps://aclanthology.org/2024.emnlp-...

  24. [24]

    RDKit: Open-source cheminformatics, version 2026.03.1, 2026

    Greg Landrum, Paolo Tosco, Brian Kelley, David Cosgrove, Riccardo Vianello, et al. RDKit: Open-source cheminformatics, version 2026.03.1, 2026. URL https://doi.org/10.5281/zenodo.19250388. Zenodo, release2026_03_1

  25. [25]

    Comparing SMILES and SELFIES tokenization for enhanced chemical language modeling.Scientific Reports, 14(1):25016, 2024

    Miguelangel Leon, Yuriy Perezhohin, Fernando Peres, Aleš Popoviˇc, and Mauro Castelli. Comparing SMILES and SELFIES tokenization for enhanced chemical language modeling.Scientific Reports, 14(1):25016, 2024. doi: 10.1038/s41598-024-76440-8. 16 Where to cut, how deep: BPE vs Unigram-LMA PREPRINT

  26. [26]

    CycPeptM- PDB: A comprehensive database of membrane permeability of cyclic peptides.Journal of Chemical Information and Modeling, 63(7):2240–2250, 2023

    Jianan Li, Keisuke Yanagisawa, Masahito Sugita, Takuya Fujie, Masahito Ohue, and Yutaka Akiyama. CycPeptM- PDB: A comprehensive database of membrane permeability of cyclic peptides.Journal of Chemical Information and Modeling, 63(7):2240–2250, 2023. doi: 10.1021/acs.jcim.2c01573

  27. [27]

    SMILES pair encoding: A data-driven substructure tokenization algorithm for deep learning.Journal of Chemical Information and Modeling, 61(4):1560–1569, 2021

    Xinhao Li and Denis Fourches. SMILES pair encoding: A data-driven substructure tokenization algorithm for deep learning.Journal of Chemical Information and Modeling, 61(4):1560–1569, 2021. doi: 10.1021/acs.jcim.0c01127

  28. [28]

    Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal

    Haoran Lian, Yizhe Xiong, Jianwei Niu, Shasha Mo, Zhenpeng Su, Zijia Lin, Hui Chen, Peng Liu, Jungong Han, and Guiguang Ding. Scaffold-bpe: Enhancing byte pair encoding for large language models with simple and effective scaffold token removal, 2024. URLhttps://arxiv.org/abs/2404.17808

  29. [29]

    SuperBPE: Space Travel for Language Models

    Alisa Liu, Jonathan Hayase, Valentin Hofmann, Sewoong Oh, Noah A. Smith, and Yejin Choi. Superbpe: Space travel for language models. InProceedings of the Second Conference on Language Modeling (COLM), 2025. URLhttps://arxiv.org/abs/2503.13423

  30. [30]

    HuggingFace’s tokenizers: Fast state-of-the-art tokenizers optimized for research and production

    Anthony Moi and Nicolas Patry. HuggingFace’s tokenizers: Fast state-of-the-art tokenizers optimized for research and production. https://github.com/huggingface/tokenizers, 2023. URL https://github.com/ huggingface/tokenizers

  31. [31]

    Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models

    Pit Neitemeier et al. Hierarchical autoregressive transformers: Combining byte- and word-level processing for robust, adaptable language models, 2025. URLhttps://arxiv.org/abs/2501.10322

  32. [32]

    Noel M. O’Boyle. Towards a universal SMILES representation - a standard method to generate canonical SMILES based on the InChI.Journal of Cheminformatics, 4(1):22, 2012. doi: 10.1186/1758-2946-4-22

  33. [33]

    O’Boyle and Andrew Dalke

    Noel M. O’Boyle and Andrew Dalke. DeepSMILES: An adaptation of SMILES for use in machine-learning of chemical structures.ChemRxiv, 2018. doi: 10.26434/chemrxiv.7097960.v1

  34. [34]

    Byte Latent Transformer: Patches Scale Better Than Tokens

    Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, and Srinivasan Iyer. Byte latent transformer: Patches scale better than tokens, 2024. URLhttps://arxiv.org/abs/2412.09871

  35. [35]

    BPE-dropout: Simple and effective subword regularization

    Ivan Provilkov, Dmitrii Emelianenko, and Elena V oita. BPE-dropout: Simple and effective subword regularization. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 1882–1892, 2020. doi: 10.18653/v1/2020.acl-main.170

  36. [36]

    Optimizing SMILES token sequences via trie-based refinement and transition graph filtering.Journal of Cheminformatics, 18(1):13, 2026

    Sridhar Radhakrishnan, Krish Mody, Arvind Venkatesh, and Ananth Venkatesh. Optimizing SMILES token sequences via trie-based refinement and transition graph filtering.Journal of Cheminformatics, 18(1):13, 2026. doi: 10.1186/s13321-025-01143-9

  37. [37]

    O’Reilly Media, 2019

    Bharath Ramsundar, Peter Eastman, Patrick Walters, Vijay Pande, Karl Leswing, and Zhenqin Wu.Deep Learning for the Life Sciences: Applying Deep Learning to Genomics, Microscopy, Drug Discovery, and More. O’Reilly Media, 2019. ISBN 978-1492039839

  38. [38]

    How Much is Enough? The Diminishing Returns of Tokenization Training Data

    Varshini Reddy, Craig W. Schmidt, Yuval Pinter, and Chris Tanner. How much is enough? the diminishing returns of tokenization training data, 2025. URL https://arxiv.org/abs/2502.20273. ICML 2025 Tokenization Workshop (TokShop)

  39. [39]

    How good is your tokenizer? on the monolingual performance of multilingual language models

    Phillip Rust, Jonas Pfeiffer, Ivan Vuli´c, Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Pape...

  40. [40]

    ReactionT5: a large-scale pre-trained model towards application of limited reaction data

    Tatsuya Sagawa and Ryosuke Kojima. ReactionT5: a large-scale pre-trained model towards application of limited reaction data.arXiv preprint, 2023. doi: 10.48550/arxiv.2311.06708

  41. [41]

    Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris C

    Craig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris C. Tanner. Tokenization is more than compression. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 678–702. Association for Computational Linguistics, 2024. doi: 10.18653/ v1/2024.emnlp-main.40. URLhttps://acla...

  42. [42]

    Schmidt, Varshini Reddy, Chris Tanner, and Yuval Pinter

    Craig W. Schmidt, Varshini Reddy, Chris Tanner, and Yuval Pinter. Boundless byte pair encoding: Breaking the pre-tokenization barrier. InConference on Language Modeling (COLM), 2025. URL https://arxiv.org/ abs/2504.00178

  43. [43]

    Japanese and Korean voice search

    Mike Schuster and Kaisuke Nakajima. Japanese and Korean voice search. In2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5149–5152. IEEE, 2012. doi: 10.1109/ICASSP. 2012.6289079. 17 Where to cut, how deep: BPE vs Unigram-LMA PREPRINT

  44. [44]

    found in translation

    Philippe Schwaller, Théophile Gaudin, Dávid Lányi, Costas Bekas, and Teodoro Laino. “found in translation”: predicting outcomes of complex organic chemistry reactions using neural sequence-to-sequence models.Chemical Science, 9(28):6091–6098, 2018. doi: 10.1039/c8sc02339e

  45. [45]

    Neural machine translation of rare words with subword units

    Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics, pages 1715–1725,

  46. [46]

    URLhttps://aclanthology.org/P16-1162

    doi: 10.18653/v1/P16-1162. URLhttps://aclanthology.org/P16-1162

  47. [47]

    Skinnider

    Michael A. Skinnider. Invalid SMILES are beneficial rather than detrimental to chemical language models.Nature Machine Intelligence, 6(4):437–448, 2024. doi: 10.1038/s42256-024-00821-x

  48. [48]

    COCONUT online: Collection of open natural products database.Journal of Cheminformatics, 13(1):2, 2021

    Maria Sorokina, Peter Merseburger, Kohulan Rajan, Mehmet Aziz Yirik, and Christoph Steinbeck. COCONUT online: Collection of open natural products database.Journal of Cheminformatics, 13(1):2, 2021. doi: 10.1186/ s13321-020-00478-9

  49. [49]

    Linguistic laws meet protein sequences: A comparative analysis of subword tokenization methods, 2024

    Burak Suyunu, Enes Taylan, and Arzucan Özgür. Linguistic laws meet protein sequences: A comparative analysis of subword tokenization methods, 2024

  50. [50]

    Ülgen, Nilgün Karalı, and Arzucan Özgür

    Asu Bü¸ sra Temizer, Gökçe Uludo˘gan, Rıza Özçelik, Taha Koulani, Elif Özkırımlı, Kutlu Ö. Ülgen, Nilgün Karalı, and Arzucan Özgür. Exploring data-driven chemical SMILES tokenization approaches to identify key protein-ligand binding moieties.Molecular Informatics, 43(3):e202300249, 2024. doi: 10.1002/minf.202300249

  51. [51]

    Tingle, Khanh G

    Benjamin I. Tingle, Khanh G. Tang, Mar Castanon, John J. Gutierrez, Munkhzul Khurelbaatar, Chinzorig Dandarchuluun, Yurii S. Moroz, and John J. Irwin. ZINC-22 – a free multi-billion-scale database of tangible compounds for ligand discovery.Journal of Chemical Information and Modeling, 63(4):1166–1176, 2023. doi: 10.1021/acs.jcim.2c01253

  52. [52]

    LLaMA: Open and Efficient Foundation Language Models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: Open and efficient foundation language models, 2023. URL https: //arxiv.org/abs/2302.13971

  53. [53]

    Ucak, Islambek Ashyrmamatov, and Juyong Lee

    Umit V . Ucak, Islambek Ashyrmamatov, and Juyong Lee. Improving the quality of chemical language model outcomes with atom-in-SMILES tokenization.Journal of Cheminformatics, 15(1):55, 2023. doi: 10.1186/ s13321-023-00725-9

  54. [54]

    Tokenization for molecular foundation models.Journal of Chemical Information and Modeling, 66(3):1384–1393, 2026

    Alexius Wadell, Anoushka Bhutani, and Venkatasubramanian Viswanathan. Tokenization for molecular foundation models.Journal of Chemical Information and Modeling, 66(3):1384–1393, 2026. doi: 10.1021/acs.jcim.5c01856

  55. [55]

    Tokenization is sensitive to language variation

    Anna Wegmann, Dong Nguyen, and David Jurgens. Tokenization is sensitive to language variation. InFindings of the Association for Computational Linguistics: ACL 2025, pages 10958–10983. Association for Computa- tional Linguistics, 2025. doi: 10.18653/v1/2025.findings-acl.572. URL https://aclanthology.org/2025. findings-acl.572/

  56. [56]

    SMILES, a chemical language and information system

    David Weininger. SMILES, a chemical language and information system. 1. introduction to methodology and encoding rules.Journal of Chemical Information and Computer Sciences, 28(1):31–36, 1988. doi: 10.1021/ ci00057a005

  57. [57]

    Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S

    Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay S. Pande. MoleculeNet: a benchmark for molecular machine learning.Chemical Science, 9 (2):513–530, 2018. doi: 10.1039/c7sc02664a

  58. [58]

    Barbara Zdrazil, Eloy Félix, Fiona Hunter, Emma J. Manners, James Blackshaw, Sybilla Corbett, Marleen De Veij, Harris Ioannidis, David Méndez, Juan F Mosquera, María Paula Magariños, Nicolas Bosc, Ricardo Arcila, Tevfik Kizilören, Anna Gaulton, A. Patrícia Bento, Melissa F. Adasme, Peter Monecke, Gregory A. Landrum, and Andrew R. Leach. The ChEMBL databas...

  59. [59]

    Tokenization and the noiseless channel

    Vilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. Tokenization and the noiseless channel. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5184–5207. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.acl-long.284. URLhttp...