Pith. sign in

REVIEW 1 cited by

Training and Evaluation of a Multilingual Tokenizer for GPT-SW3

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.14780 v1 pith:SWLAK2OW submitted 2023-04-28 cs.CL cs.AI

classification cs.CLcs.AI
keywords tokenizergpt-sw3multilingualadditionalgorithmanalyzedatadetailed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper provides a detailed discussion of the multilingual tokenizer used for GPT-SW3. It was trained on the Nordic Pile using the SentencePiece library and the BPE algorithm. We outline the tokenizer's most important features and share details on its learned vocabulary. In addition, we systematically analyze the properties and evaluate the performance of the tokenizer with regard to the different languages present in the data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Text Compression: Evaluating Tokenizers Across Scales

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Tokenizer choice matters mostly for multilingual tasks, and 350M-parameter models can predict 2.7B model ranking on translation but not on English benchmarks.

Pith tools