Pith. sign in

REVIEW 4 major objections 4 minor 119 references

Beyond Text Compression: Evaluating Tokenizers Across Scales

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Small-model tokenizer rankings predict large-model performance on machine translation, where a 350M model with the right tokenizer matches a 2.7B model.

desk verdict A useful MT-specific empirical result, but the parameter/compute mismatch and overbroad claims need fixing before I'd trust it. read the letter →

arxiv 2506.03101 v1 pith:EB5FDIVG submitted 2025-06-03 cs.CL

classification cs.CL
keywords tokenizerevaluationscalinglawsmultilinguallanguagemodelsZipf'slawintrinsicmachinetranslationtokenizationsubword
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

To choose a tokenizer for a large language model, you usually have to train the model to find out, which is an expensive experiment. This paper claims you can instead train a small 350M-parameter model as a cheap stand-in, because tokenizer effects that matter at small scale also show up at larger scale. The claim is tested with six tokenizers at 350M and 2.7B parameters on English and multilingual tasks. The authors find that tokenizer choice barely matters for English tasks, but strongly and consistently affects machine translation, where a 350M model with a well-matched multilingual tokenizer can match or beat 2.7B models with English-centric tokenizers. They also propose intrinsic metrics based on how closely a tokenizer's frequency distribution follows Zipf's law, which predict translation performance better than text compression, and package them into a ranking framework.

What carries the argument

The load-bearing mechanism is the scale-consistency conjecture: if a tokenizer significantly affects model quality, its impact will manifest consistently across model scales. On the measurement side, the paper introduces four intrinsic metrics computed from the token-frequency distribution on task data, namely CARDINALITY (number of unique tokens), AUC (area under the log-log rank-frequency curve), SLOPE (the fitted power-law exponent), and POWER LAW (mean absolute deviation from a fitted power law), plus the standard COMPRESSION baseline. These feed a two-stage predictor: logistic regression or a linear-kernel SVM on pairwise metric differences, followed by a Bradley-Terry model that converts pairwise win probabilities into a global tokenizer ranking. The key identity at work is that a tokenizer whose token frequencies approximate a Zipfian power law is better aligned with natural-language statistics, and this alignment is what the metrics operationalize.

What would settle it

Retrain the same six tokenizers at a third scale, say 7B parameters, and recompute the tokenizer ranking on the same WMT21 translation pairs. If the 350M-to-2.7B ranking (Kendall's tau 0.87) does not persist at the larger jump, or if an English multiple-choice benchmark starts to show scale-consistent tokenizer effects beyond 2.7B, the scale-consistency conjecture and the task-specificity conclusion would need revision.

Watch

Extended reading notes

Core claim

The central claim is that tokenizer quality can be evaluated cheaply and reliably before large-scale training: rankings of tokenizers measured on 350M-parameter models predict rankings at 2.7B parameters with Kendall's tau of 0.87 for machine translation, at roughly 15% of the pretraining cost. The paper further claims that this scale-consistency does not hold for English-centric tasks, with multiple-choice benchmarks giving tau 0.33 and summarization giving -0.07, so the proxy method is task-specific. On the intrinsic side, the paper claims that deviation from a Zipfian power-law token-frequency distribution (POWER LAW) is the single most informative predictor of multilingual performance, that combining it with vocabulary cardinality and the slope of the rank-frequency curve in an SVM improves pairwise predictions, and that a Bradley-Terry model over those pairwise comparisons recovers the ground-truth tokenizer ranking across Czech, German, Russian, and Chinese. A direct corollary, stated in Section 4.3, is that a 350M model with the AYA23 tokenizer performs comparably to GPT-NEOX at 2.7B and surpasses GPT-2 at 2.7B, showing tokenizer choice can offset a roughly fivefold increase in parameters for translation.

Load-bearing premise

The entire proxy method rests on the assumption that tokenizer effects rank the same way at 350M and 2.7B parameters, an assumption the paper's own numbers confirm for machine translation but violate for English benchmarks and summarization.

Editorial extensions

If this is right

  • Pretraining a 350M-parameter proxy instead of a 2.7B model cuts the compute of an extrinsic tokenizer evaluation by roughly 85%, making systematic tokenizer comparison feasible before large runs.
  • For English-only modeling, tokenizer choice among the six evaluated has negligible downstream impact, so selection can focus on compression and inference speed rather than quality.
  • For multilingual generation, tokenizer selection is a first-order decision: the AYA23 tokenizer at 350M matches or beats English-centric tokenizers at 2.7B on translation, but larger vocabularies slow inference.
  • Intrinsic metrics based on Zipf's-law alignment, especially POWER LAW, predict multilingual performance better than the compression baseline, and combining CARDINALITY, POWER LAW, and SLOPE in a pairwise-then-Bradley-Terry framework yields reliable tokenizer rankings without training a model.
  • Tokenizers that overrepresent high-frequency tokens or rely on a long tail of rare units are penalized by the Zipfian metrics, giving a concrete design target for new multilingual tokenizers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If scale-consistency holds for translation but not for English benchmarks, a plausible generalization is that tokenizer impact is scale-consistent exactly when the task stresses vocabulary coverage for the target language; one testable extension is to apply the 350M-proxy protocol to code generation or biomedical text, where the authors expect the metric mix to change.
  • The Zipfian metrics are cheap enough to compute on unlabeled target-language text before any training, so a practical recipe suggested but not validated by the paper is to screen candidate tokenizers on a small sample of the target corpus and only pretrain proxies for the top few.
  • The paper's single-seed, single-pretraining-pipeline design leaves open how much of the measured ranking is tokenizer effect versus optimization noise; a direct follow-up would re-run the AYA23-versus-GPT-2 comparison with several seeds to bound the ranking's stability.
  • Because the AYA23 tokenizer's advantage shows up mainly on Chinese and other non-Latin scripts, the Zipfian metrics may be most informative for scripts where English-centric tokenizers fragment text heavily; testing on more languages beyond five would reveal whether the tau equals 0.87 result generalizes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper trains decoder-only transformer models at nominal 350M and 2.7B scales using six published tokenizers (four English-centric, two multilingual) on the FineWeb corpus, evaluates them on multiple-choice benchmarks, English summarization (X-SUM), and multilingual machine translation (four language pairs with English), and proposes four intrinsic metrics (CARDINALITY, AUC, SLOPE, POWER LAW) derived from Zipfian rank-frequency plots. It reports that tokenizer choice has little impact on English tasks, that a model using the AYA23 tokenizer outperforms larger English-tokenizer models in translation, and that the proposed metrics, alone or combined, can predict multilingual performance better than text compression. The paper closes with a two-stage framework (pairwise logistic/SVM plus Bradley-Terry aggregation) for intrinsic tokenizer ranking.

Significance. If the central claims were supported, the paper would offer a practical, low-cost way to select tokenizers for multilingual model development and new intrinsic diagnostics. The study is usefully broad: it spans six tokenizers, two scales, three task families, and four non-English languages across three scripts, with transparent pretraining and tuning details. However, the current evidence does not support the abstract's general claim of cross-scale predictability: scale consistency holds only for translation, the machine-translation results are confounded by vocabulary-size-induced changes in parameter count and compute, and the reported significance levels in the correlation analysis are not attainable with six tokenizers. The paper's strengths are its detailed reporting and its clear documentation of where tokenizer effects do and do not appear; these make the manuscript a useful starting point after substantial revision.

major comments (4)
  1. [Table 3] Table 3 reports Spearman correlations with n=6 tokenizers and marks several coefficients as significant at p<0.01 (e.g., COMPRESSION -0.59 for multiple-choice, CARDINALITY -0.79 for machine translation). With six data points, the exact two-tailed p-value for rho=-0.59 is approximately 0.22, and even rho=0.77 does not reach p<0.01; the critical rho for p<0.01 with n=6 is 1.0. The reported significance stars are therefore incorrect and the claim that the proposed metrics "correlate more strongly than text compression" lacks the stated statistical support. Please remove the asterisks, provide exact p-values (e.g., from permutation tests), or rephrase the correlation claims as descriptive effect sizes.
  2. [Section 4 and Tables 1/7] The claim in Section 4 that "the models vary only in their choice of tokenizer" is contradicted by Table 1 (and Table 7). The AYA23 350M model has 566M trainable parameters, 490 H100-hours, and 8.17 TFLOPs, whereas GPT-2 at the same nominal scale has 356M parameters, 210 H100-hours, and 5.58 TFLOPs. Thus the headline result that a "350M" AYA23 model surpasses a 2.7B GPT-2 model in translation conflates tokenizer choice with vocabulary size, embedding/softmax capacity, and training compute. The cross-scale tau=0.87 for machine translation in Table 3 may simply reflect that tokenizers with larger vocabularies are ordered the same way at both scales. To support the central tokenizer-selection claim, the study would need to control for parameter count or compute (e.g., by adjusting hidden size or training budget), or the claims must be substantially narrowed to "a model with a larger-vocabulary tokenizer and higher compute can outperform a smaller-vocabulary model with more parameters."
  3. [Section 3.1 and Table 3] The scale-consistency conjecture in Section 3.1 ("if a tokenizer significantly affects model quality, its impact will manifest consistently across different model scales") is not supported by the data: Table 3 shows across-scale Kendall's tau of 0.33 for multiple-choice benchmarks and -0.07 for summarization, with only machine translation reaching tau=0.87. The abstract's opening claim that "smaller models can accurately predict significant differences in tokenizer impact on larger models" is therefore overbroad. Either the paper should restrict its predictive claim to the translation scenario where consistency is actually observed, or it should provide a theoretical or empirical account of why consistency should hold selectively.
  4. [Abstract and Section 4.3] The abstract and Section 4.3 use the word "significant" without any variance estimate; each condition is trained with a single seed, so there is no confidence interval for the differences between tokenizers. The Limitations acknowledge this, but the abstract's "significant differences" and the comparison between AYA23 350M and GPT-2 2.7B are stated as if they were established effects. Please add uncertainty quantification or soften the wording to "reported differences."
minor comments (4)
  1. [Section 3.4] Section 3.4 states "we propose four new metrics"; CARDINALITY (number of unique tokens) is a standard measure in tokenizer analysis and is not new. Please rephrase to distinguish the genuinely new metrics from the existing one.
  2. [Figures 1 and 2] The legend in Figures 1 and 2 uses only color to distinguish the six tokenizers, making the curves difficult to separate in grayscale print. Consider adding markers or distinct line styles.
  3. [Table 6] Table 6 should state explicitly in the caption that Kendall's tau compares the Bradley-Terry predicted ranking with the ground-truth (MetricX-based) ranking, and should define the significance thresholds used for the asterisks.
  4. [Section 5.1] In Section 5.1, the F1 scores are reported as point estimates without confidence intervals or a chance-level baseline; given the small number of tokenizers and the correlated pairwise outcomes, please include a null baseline or a permutation interval to aid interpretation.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: scale-consistency and intrinsic-metric claims rest on independent empirical evaluations.

full rationale

The paper's chain of reasoning is empirical and self-contained rather than definitional. The central scale-consistency claim (Section 3.1 conjecture; Table 3 'Across scales 0.87' for MT) is a post-hoc rank correlation between independently trained 350M and 2.7B models; no parameter is fitted to the 2.7B performances and then reported as a prediction. The intrinsic metrics (CARDINALITY, AUC, SLOPE, POWER LAW) are defined a priori from Zipf's law and token-frequency statistics on task training data (Section 3.4, Figure 1), not from downstream MetricX or chrF scores, so their correlation with translation quality is an independent test of the stated hypothesis. The pairwise and Bradley-Terry frameworks (Sections 5.1-5.2) use leave-one-tokenizer-out and leave-one-language-out splits, which are legitimate supervised evaluations. The only self-citation (Rust et al. 2023 in Related Work, co-authored by J.F. Lotz) is descriptive and not load-bearing. The paper's Limitations section candidly states that larger-scale trends are unverified and that seed/hyperparameter sensitivity was not explored. A separate correctness concern, not a circularity, is the claim in Section 4 that 'the models vary only in their choice of tokenizer,' since Table 1/7 show AYA23 at the 350M tier uses 566M parameters, 490 H100 hours, and 8.17 TFLOPs versus GPT-2's 356M/210h/5.58 TFLOPs; this confounds tokenizer identity with vocabulary-driven capacity/compute, but it does not make any derived quantity equal to its input by construction.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central empirical claims rest on the scale-consistency conjecture and on the assumption that Zipfian alignment captures tokenizer quality for multilingual generation. The free parameters are methodological choices (rank cutoff, metric combination, hidden SVM hyperparameters) rather than physical constants. No new entities are introduced.

free parameters (3)
  • Upper rank cutoff for Zipf metrics = log(rank) <= 6
    Section 3.4 restricts AUC, POWER LAW, and SLOPE to tokens with log(rank)<=6, citing prior work on power-law minima. The cutoff influences all three metric values and correlations, and no sensitivity analysis is provided.
  • Intrinsic metric combination C+P+S = CARDINALITY, POWER LAW, SLOPE
    Section 5.1 compares combinations and reports the best one (C+P+S) on the same leave-one-tokenizer-out evaluations used to report F1, so the combination choice is selected on the evaluation data.
  • Model hyperparameters for SVM/logistic regression = cross-validated, not reported
    Section 5.1 states a cross-validated search for optimal hyperparameters, but the selected hyperparameters are not listed, leaving part of the framework unspecified and hard to reproduce exactly.
assumptions (3)
  • domain assumption Tokenizer effects are scale-consistent.
    Section 3.1 conjectures that tokenizer impact manifests consistently across model scales; the 350M-to-2.7B proxy method depends on this assumption.
  • domain assumption Zipf's law holds for natural language and token distributions should align with it.
    Section 3.4 hypothesizes that tokenizers whose token frequencies follow a Zipfian power law are better suited to generative modeling; this motivates all four intrinsic metrics.
  • standard math Power-law behavior only applies above a minimum rank.
    Section 3.4 truncates Zipf metrics to log(rank)<=6 based on Newman (2005) and Clauset et al. (2009).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Text Compression: Evaluating Tokenizers Across Scales." pith.science (2026). https://pith.science/paper/EB5FDIVG

@misc{pith2026250603101,
  author       = {Pith},
  title        = {Pith review of: Beyond Text Compression: Evaluating Tokenizers Across Scales},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EB5FDIVG}},
  note         = {Machine review of arXiv:2506.03101}
}
read the original abstract

The choice of tokenizer can profoundly impact language model performance, yet accessible and reliable evaluations of tokenizer quality remain an open challenge. Inspired by scaling consistency, we show that smaller models can accurately predict significant differences in tokenizer impact on larger models at a fraction of the compute cost. By systematically evaluating both English-centric and multilingual tokenizers, we find that tokenizer choice has negligible effects on tasks in English but results in consistent performance differences in multilingual settings. We propose new intrinsic tokenizer metrics inspired by Zipf's law that correlate more strongly with downstream performance than text compression when modeling unseen languages. By combining several metrics to capture multiple aspects of tokenizer behavior, we develop a reliable framework for intrinsic tokenizer evaluations. Our work offers a more efficient path to informed tokenizer selection in future language model development.

Figures

Figures reproduced from arXiv: 2506.03101 by the authors.

Figure 1
Figure 1. Token frequency plotted against frequency rank in log-log scale for English ( [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Token frequency plotted against frequency rank in log-log scale for Czech ( [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

119 extracted references · 7 canonical work pages

  1. [1]

    Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, and 110 others. 2024. https://doi.org/10.48550/arXiv.2404.14219 Phi-3 technical rep...

  2. [2]

    Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, and Yulia Tsvetkov. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.614 Do all languages cost the same? tokenization in the era of commercial language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9904--9923, ...

  3. [3]

    Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ond r ej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina Espa \ n a-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Kenneth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, ...

  4. [4]

    Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max L \"u bbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Buschhoff, Charvi Jain, Alexander Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, and 2 others. 2024. https://doi.org/10.18653/v1/2024.fin...

  5. [5]

    Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. https://doi.org/10.48550/arXiv.2311.16867 The falcon series of open language models . arXiv preprint

  6. [6]

    Norah Alzahrani, Hisham Alyahya, Yazeed Alnumay, Sultan AlRashed, Shaykhah Alsubaie, Yousef Almushayqih, Faisal Mirza, Nouf Alotaibi, Nora Al-Twairesh, Areeb Alowisheq, M Saiful Bari, and Haidar Khan. 2024. https://doi.org/10.18653/v1/2024.acl-long.744 When benchmarks are targets: Revealing the sensitivity of large language model leaderboards . In Proceed...

  7. [7]

    Reinald Kim Amplayo, Peter J Liu, Yao Zhao, and Shashi Narayan. 2023. https://openreview.net/forum?id=OIe3kpwl40D SMART : Sentences as basic units for text evaluation . In The Eleventh International Conference on Learning Representations

  8. [8]

    Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Kelly Marchisio, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Phil Blunsom, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker. 2024. https://doi.org/10.48550/arXiv.2404.14619 Aya 23: Open weight releases to further multilingua...

Show all 119 references
  1. [9]

    Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, and 3 ...

  2. [10]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan H...

  3. [11]

    Lo \"i c Barrault, Magdalena Biesialska, Ond r ej Bojar, Marta R. Costa-juss \`a , Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljube s i \'c , Christof Monz, Makoto Morishita, Ma...

  4. [12]

    Lo \"i c Barrault, Ond r ej Bojar, Marta R. Costa-juss \`a , Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias M \"u ller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. https://doi.org/10.1...

  5. [13]

    Lisa Beinborn and Yuval Pinter. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.272 Analyzing cognitive plausibility of subword tokenization . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4478--4486, Singapore. Association ...

  6. [14]

    Tamay Besiroglu, Ege Erdil, Matthew Barnett, and Josh You. 2024. https://doi.org/10.48550/arXiv.2404.10102 Chinchilla scaling: A replication attempt . arXiv preprint

  7. [15]

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, and 1 others. 2023. Pythia: A suite for analyzing large language models across training and scali...

  8. [16]

    Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, and 374 others

    BigScience Workshop , Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi,...

  9. [17]

    Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, and 1 others. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432--7439

  10. [18]

    Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. https://doi...

  11. [19]

    Ond r ej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz. 2018. https://doi.org/10.18653/v1/W18-6401 Findings of the 2018 conference on machine translation ( WMT 18) . In Proceedings of the Third Conference ...

  12. [20]

    Kaj Bostrom and Greg Durrett. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.414 Byte pair encoding is suboptimal for language model pretraining . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4617--4624, Online. Association for Computa...

  13. [21]

    Ralph Allan Bradley and Milton E Terry. 1952. Rank analysis of incomplete block designs i: The method of paired comparisons. Biometrika, 39(3/4):324--345

  14. [22]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...

  15. [23]

    Yekun Chai, Yewei Fang, Qiwei Peng, and Xuhong Li. 2024 a . https://doi.org/10.18653/v1/2024.findings-emnlp.86 Tokenization falling short: On subword robustness in large language models . In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 1582--159...

  16. [24]

    Yekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang, Yu Sun, and Hua Wu. 2024 b . https://doi.org/10.18653/v1/2024.emnlp-main.182 Autoregressive pre-training on pixels and texts . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 3...

  17. [25]

    Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. https://doi.org/10.48550/arXiv.1904.10509 Generating long sequences with sparse transformers . arXiv preprint

  18. [26]

    Leshem Choshen, Yang Zhang, and Jacob Andreas. 2024. https://doi.org/10.48550/arXiv.2410.11840 A hitchhiker's guide to scaling law estimation . arXiv preprint

  19. [27]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vino...

  20. [28]

    Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30

  21. [29]

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. https://doi.org/10.48550/arXiv.1803.05457 Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge . arXiv preprint

  22. [30]

    Aaron Clauset, Cosma Rohilla Shalizi, and M. E. J. Newman. 2009. https://doi.org/10.1137/070710111 Power-law distributions in empirical data . SIAM Review, 51(4):661--703

  23. [31]

    Marco Cognetta, Vil \'e m Zouhar, Sangwhan Moon, and Naoaki Okazaki. 2024. https://aclanthology.org/2024.lrec-main.1469/ Two counterexamples to tokenization and the noiseless channel . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Lang...

  24. [32]

    Gautier Dagan, Gabriel Synnaeve, and Baptiste Rozi\` e re. 2024. Getting the most out of your tokenizer for pre-training and domain adaptation. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org

  25. [33]

    Miguel Domingo, Mercedes Garc \' a-Mart \' nez, Alexandre Helle, Francisco Casacuberta, and Manuel Herranz. 2019. https://link.springer.com/chapter/10.1007/978-3-031-24337-0_38 How much does tokenization affect neural machine translation? In International Conference on Computa...

  26. [34]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zha...

  27. [35]

    Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023. https://doi.org/10.18653/v1/2023.wmt-1.51 Results of ...

  28. [36]

    Matthias Gall \'e . 2019. https://doi.org/10.18653/v1/D19-1141 Investigating the effectiveness of BPE : The power of shorter sequences . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natu...

  29. [37]

    Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...

  30. [38]

    Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein. 2024. https://openreview.net/forum?id=Y5inHAjMu0 Coercing LLM s to do and reveal (almost) anything . In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models

  31. [39]

    Daniela Gerz, Ivan Vuli \'c , Edoardo Maria Ponti, Roi Reichart, and Anna Korhonen. 2018. https://doi.org/10.18653/v1/D18-1029 On the relation between linguistic typology and (limitations of) multilingual language modeling . In Proceedings of the 2018 Conference on Empirical M...

  32. [40]

    Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty. 2024. https://doi.org/10.18653/v1/2024.findings-acl.134 Unpacking tokenization: Evaluating text compression and its correlation with model performance . In Findings of the Association for Comp...

  33. [41]

    Thamme Gowda and Jonathan May. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.352 Finding the optimal vocabulary size for neural machine translation . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3955--3964, Online. Association for Com...

  34. [42]

    Gregory Grefenstette. 1999. https://doi.org/10.1007/978-94-015-9273-4_9 Tokenization , pages 117--133. Springer Netherlands, Dordrecht

  35. [43]

    Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, and 2...

  36. [44]

    Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Y...

  37. [45]

    Jonathan Hayase, Alisa Liu, Yejin Choi, Sewoong Oh, and Noah A. Smith. 2024. https://openreview.net/forum?id=0SRg6Cwx3h Data mixture inference attack: BPE tokenizers reveal training data compositions . In ICML 2024 Workshop on Foundation Models in the Wild

  38. [46]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. https://openreview.net/forum?id=d7KBjmI3GmQ Measuring massive multitask language understanding . In International Conference on Learning Representations

  39. [47]

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Thomas Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Sim...

  40. [48]

    Valentin Hofmann, Janet Pierrehumbert, and Hinrich Sch \"u tze. 2021. https://doi.org/10.18653/v1/2021.acl-long.279 Superbizarre is not superb: Derivational morphology improves BERT `s interpretation of complex words . In Proceedings of the 59th Annual Meeting of the Associati...

  41. [49]

    Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert. 2022. https://doi.org/10.18653/v1/2022.acl-short.43 An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers . In Proceedings of the 60th Annual Meeting of the Associ...

  42. [50]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  43. [51]

    Juraj Juraska, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.wmt-1.63 M etric X -23: The G oogle submission to the WMT 2023 metrics shared task . In Proceedings of the Eighth Conference on Machin...

  44. [52]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. https://doi.org/10.48550/arXiv.2001.08361 Scaling laws for neural language models . arXiv preprint

  45. [53]

    Stav Klein and Reut Tsarfaty. 2020. https://doi.org/10.18653/v1/2020.sigmorphon-1.24 Getting the \# \# life out of living: How adequate are word-pieces for modelling complex morphology? In Proceedings of the 17th SIGMORPHON Workshop on Computational Research in Phonetics, Phon...

  46. [54]

    Tom Kocmi, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Nov \'a k, Ma...

  47. [55]

    Taku Kudo. 2018. https://doi.org/10.18653/v1/P18-1007 Subword regularization: Improving neural network translation models with multiple subword candidates . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), page...

  48. [56]

    Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. https://doi.org/10.18653/v1/D17-1082 RACE : Large-scale R e A ding comprehension dataset from examinations . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages...

  49. [57]

    Sander Land and Max Bartolo. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.649 Fishing for magikarp: Automatically detecting under-trained tokens in large language models . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 116...

  50. [58]

    Teven Le Scao, Thomas Wang, Daniel Hesslow, Stas Bekman, M Saiful Bari, Stella Biderman, Hady Elsahar, Niklas Muennighoff, Jason Phang, Ofir Press, Colin Raffel, Victor Sanh, Sheng Shen, Lintang Sutawika, Jaesung Tae, Zheng Xin Yong, Julien Launay, and Iz Beltagy. 2022. https:...

  51. [59]

    Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, and 40 o...

  52. [60]

    Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee. 2023. https://doi.org/10.48550/arXiv.2309.05463 Textbooks are all you need ii: phi-1.5 technical report . arXiv preprint

  53. [61]

    Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. https://doi.org/10.18653/v1/2022.acl-long.229 T ruthful QA : Measuring how models mimic human falsehoods . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pa...

  54. [62]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. https://doi.org/10.48550/arXiv.1907.11692 Roberta: A robustly optimized bert pretraining approach . arXiv preprint

  55. [63]

    Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, Richard Fan, Yi Gu, Victor Miller, Yonghao Zhuang, Guowei He, Haonan Li, Fajri Koto, Liping Tang, Nikhil Ranjan, and 8 others. 2024. https://ope...

  56. [64]

    Ilya Loshchilov and Frank Hutter. 2017. https://openreview.net/forum?id=Skq89Scxx SGDR : Stochastic gradient descent with warm restarts . In International Conference on Learning Representations

  57. [65]

    Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled weight decay regularization . In International Conference on Learning Representations

  58. [66]

    Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes

    Lovish Madaan, Aaditya K. Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes. 2024. https://doi.org/10.48550/arXiv.2406.10229 Quantifying variance in evaluation benchmarks . arXiv preprint

  59. [67]

    Guerreiro, Ricardo Rei, Duarte M

    Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, and André F. T. Martins. 2024. https://d...

  60. [68]

    Sachin Mehta, Mohammad Sekhavat, Qingqing Cao, Max Horton, Yanzi Jin, Frank Sun, Iman Mirzadeh, Mahyar Najibikohnehshahri, Dmitry Belenko, Peter Zatloukal, and Mohammad Rastegari. 2024. https://arxiv.org/abs/2404.14619 Openelm: An efficient language model family with open trai...

  61. [69]

    Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y

    Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoît Sagot, and Samson Tan. 2021. https://doi.org/10.48550/arXiv.2112.10508 Between words and characters: A brief history of open-vocabulary m...

  62. [70]

    Isabel Moreno-Sánchez, Francesc Font-Clos, and Álvaro Corral. 2016. https://doi.org/10.1371/journal.pone.0147073 Large-scale analysis of zipf’s law in english texts . PLOS ONE, 11(1):1--19

  63. [71]

    Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.154 C row S -pairs: A challenge dataset for measuring social biases in masked language models . In Proceedings of the 2020 Conference on Empirical Methods in Na...

  64. [72]

    Cohen, and Mirella Lapata

    Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. https://doi.org/10.18653/v1/D18-1206 Don`t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization . In Proceedings of the 2018 Conference on Empirical Methods in Natura...

  65. [73]

    MEJ Newman. 2005. https://doi.org/10.1080/00107510500052444 Power laws, pareto distributions and zipf's law . Contemporary Physics, 46(5):323--351

  66. [74]

    OpenAI . 2023. Gpt-4 technical report

  67. [75]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, a...

  68. [76]

    Guilherme Penedo, Hynek Kydl\' c ek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. https://proceedings.neurips.cc/paper_files/paper/2024/file/370df50ccfdf8bde18f8f9c2d9151bda-Paper-Datasets_and_Benchmarks_Track.pdf ...

  69. [77]

    Aleksandar Petrov, Emanuele La Malfa, Philip Torr, and Adel Bibi. 2023. https://openreview.net/forum?id=78yDLKi95p Language model tokenizers introduce unfairness between languages . In Thirty-seventh Conference on Neural Information Processing Systems

  70. [78]

    Piantadosi

    Steven T. Piantadosi. 2014. https://doi.org/10.3758/s13423-014-0585-6 Zipf's word frequency law in natural language: A critical review and future directions . Psychonomic Bulletin & Review, 21(5):1112--1130

  71. [79]

    John C. Platt. 1999. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. In Advances in Large Margin Classifiers, pages 61--74. MIT Press

  72. [80]

    Maja Popovi \'c . 2015. https://doi.org/10.18653/v1/W15-3049 chr F : character n-gram F -score for automatic MT evaluation . In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392--395, Lisbon, Portugal. Association for Computational Linguistics

  73. [81]

    Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020. https://doi.org/10.18653/v1/2020.acl-main.170 BPE -dropout: Simple and effective subword regularization . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1882--1892, O...

  74. [82]

    Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf Language models are unsupervised multitask learners

  75. [83]

    Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott

    Phillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott. 2023. https://openreview.net/forum?id=FkSp8VW8RjH Language modelling with pixels . In The Eleventh International Conference on Learning Representations

  76. [84]

    Phillip Rust, Jonas Pfeiffer, Ivan Vuli \'c , Sebastian Ruder, and Iryna Gurevych. 2021. https://doi.org/10.18653/v1/2021.acl-long.243 How good is your tokenizer? on the monolingual performance of multilingual language models . In Proceedings of the 59th Annual Meeting of the ...

  77. [85]

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99--106

  78. [86]

    Elizabeth Salesky, David Etter, and Matt Post. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.576 Robust open-vocabulary translation from visual text representations . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7235--725...

  79. [87]

    Jonne Saleva and Constantine Lignos. 2023. https://doi.org/10.18653/v1/2023.insights-1.7 What changes when you randomly choose BPE merge operations? not much. In Proceedings of the Fourth Workshop on Insights from Negative Results in NLP, pages 59--66, Dubrovnik, Croatia. Asso...

  80. [88]

    Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.40 Tokenization is more than compression . In Proceedings of the 2024 Conference on Empirical Methods in Natural Languag...

  81. [89]

    Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. https://doi.org/10.18653/v1/2020.acl-main.704 BLEURT : Learning robust metrics for text generation . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881--7892, Online. Ass...

  82. [90]

    Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. https://doi.org/10.18653/v1/P16-1162 Neural machine translation of rare words with subword units . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages ...

  83. [91]

    Claude E. Shannon. 1948. https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=6773024 A mathematical theory of communication . The Bell System Technical Journal, 27(3):379--423

  84. [92]

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020. https://doi.org/10.48550/arXiv.1909.08053 Megatron-lm: Training multi-billion parameter language models using model parallelism . arXiv preprint

  85. [93]

    Felix Stollenwerk. 2023. https://doi.org/10.48550/arXiv.2304.14780 Training and evaluation of a multilingual tokenizer for gpt-sw3 . arXiv preprint

  86. [94]

    Jimin Sun, Patrick Fernandes, Xinyi Wang, and Graham Neubig. 2023. https://doi.org/10.18653/v1/2023.findings-eacl.128 A multi-dimensional evaluation of tokenizer-free multilingual pretrained models . In Findings of the Association for Computational Linguistics: EACL 2023, page...

  87. [95]

    Yintao Tai, Xiyang Liao, Alessandro Suglia, and Antonio Vergari. 2024. https://doi.org/10.18653/v1/2024.findings-acl.874 PIXAR : Auto-regressive language modeling in pixel space . In Findings of the Association for Computational Linguistics: ACL 2024, pages 14673--14695, Bangk...

  88. [96]

    Sho Takase, Shun Kiyono, Sosuke Kobayashi, and Jun Suzuki. 2024. https://doi.org/10.48550/arXiv.2312.16903 Spike no more: Stabilizing the pre-training of large language models . arXiv preprint

  89. [97]

    Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, and Ngai Wong. 2024. https://openreview.net/forum?id=sKCKPr8cRL Scaling laws with vocabulary: Larger models deserve larger vocabularies . In The Thirty-eighth Annual Conference on Neural In...

  90. [98]

    Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler

    Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. 2023. https://openreview.net/forum?id=6ruVLB727MC UL 2: Unifying language learning paradigms . ...

  91. [99]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 a . https://arxiv.org/abs/2302.13971 LLa...

  92. [100]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  93. [101]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html Attention is all you need . In Advances in Neural I...

  94. [102]

    Shibo Wang and Pankaj Kanwar. 2019. https://cloud.google.com/blog/products/ai-machine-learning/bfloat16-the-secret-to-high-performance-on-cloud-tpus Bfloat16: The secret to high performance on cloud tpus . Blog Post

  95. [103]

    Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.72 Frequency effects on syntactic rule learning in transformers . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 932--948...

  96. [104]

    Jason Wei, Najoung Kim, Yi Tay, and Quoc Le. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.963 Inverse scaling can become U -shaped . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 15580--15591, Singapore. Association for C...

  97. [105]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. https://openreview.net/forum?id=yzkSU5zd...

  98. [106]

    Haoran Xu, Young Jin Kim, Amr Sharaf, and Hany Hassan Awadalla. 2024 a . https://openreview.net/forum?id=farT6XXntP A paradigm shift in machine translation: Boosting translation performance of large language models . In The Twelfth International Conference on Learning Representations

  99. [107]

    Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim. 2024 b . https://openreview.net/forum?id=51iwkioZpn Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation . In F...

  100. [108]

    Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022. https://doi.org/10.1162/tacl_a_00461 B y T 5: Towards a token-free future with pre-trained byte-to-byte models . Transactions of the Association for Computa...

  101. [109]

    Shaked Yehezkel and Yuval Pinter. 2023. https://doi.org/10.18653/v1/2023.eacl-main.45 Incorporating context into subword vocabularies . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 623--635, Dubrovnik, Cr...

  102. [110]

    Zheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Winata, Stella Biderman, Edward Raff, Dragomir Radev, and Vassilina Nikoulina. 2023. https://doi....

  103. [111]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4...

  104. [112]

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...

  105. [113]

    Jun Zhao, Zhihao Zhang, Luhui Gao, Qi Zhang, Tao Gui, and Xuanjing Huang. 2024. https://doi.org/10.48550/arXiv.2401.01055 Llama beyond english: An empirical study on language capability transfer . arXiv preprint

  106. [114]

    George K. Zipf. 1949. Human Behaviour and the Principle of Least Effort. Addison-Wesley

  107. [115]

    George Kingsley Zipf. 1935. The Psychobiology of Language. Houghton-Mifflin, New York, NY, USA

  108. [116]

    Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, and Xide Xia. 2024. https://doi.org/10.48550/arXiv.2412.10360 Apollo: An exploration of video understanding in lar...

  109. [117]

    Vil \'e m Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023. https://doi.org/10.18653/v1/2023.acl-long.284 Tokenization and the noiseless channel . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (...

  110. [118]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  111. [119]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.