REVIEW 2 cited by
Bridging the Gap for Tokenizer-Free Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Purely character-based language models (LMs) have been lagging in quality on large scale datasets, and current state-of-the-art LMs rely on word tokenization. It has been assumed that injecting the prior knowledge of a tokenizer into the model is essential to achieving competitive results. In this paper, we show that contrary to this conventional wisdom, tokenizer-free LMs with sufficient capacity can achieve competitive performance on a large scale dataset. We train a vanilla transformer network with 40 self-attention layers on the One Billion Word (lm1b) benchmark and achieve a new state of the art for tokenizer-free LMs, pushing these models to be on par with their word-based counterparts.
Forward citations
Cited by 2 Pith papers
-
Byte Latent Transformer: Patches Scale Better Than Tokens
A byte-level transformer that dynamically groups bytes into entropy-based patches matches token-based LLM performance at 8B scale and opens a new patch-size scaling axis for fixed inference cost.
-
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
A hierarchical byte-to-word-to-byte transformer matches subword-tokenizer LLMs at 1B-7B scale while being more robust to input corruption and faster to adapt to new languages.
Discussion (0). Continue with ORCID to comment.