Pith. sign in

Paper Citation Record · LEDGER

Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2112.10508.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.10508 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:30:40.453767Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

106
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7a1c99ac-44f9-4d62-887c-8c9e45f07472 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:11.075762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:f4827cccb5ea5d09f3c2879fc5b0c0cadcd9e9e583f304a707f52fd884108637

Observation f105b10b-2fd9-4768-bee4-7a87a9141599 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 278

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:11.646731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:22258a41ba60b7abb72032e518c9fbdb00f66d4eae6a11e6c04a8f42eabcd102

Observation c204e047-8388-4312-9431-0d41cc1ddd0b · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:19:46.729797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:d16d4045c2856b882bee7b756055d2ff5a2ff1959ed4f8d26e65980af4a11ef8

Observation d570476c-b086-43ce-89e5-2fdbafa6d6fa · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.607604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:1af02a1d0051a3eeed3b140ac56e7d790d62acdf37d2e2391df9d3adb81c38aa

Observation 4bd88b63-8ba7-4bb3-a279-175228a7692a · inbound

Jamba: A Hybrid Transformer-Mamba Language Model cites this paper.

Jamba: A Hybrid Transformer-Mamba Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:11:27.200041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T14:11:27.156350Z digest=sha256:438072487129b465288cab2eb94baf6fba7ee67c858541c8948723011c4deb71

Observation 565770a0-fe2f-430e-bb0b-eedcd7a0a046 · inbound

Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers cites this paper.

Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 288

Resolution
unresolved
no resolver link, observed 2026-08-09T16:30:40.453767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:30:40.453767Z digest=sha256:61c8ef0f587f411d082216ac1ec4e9bc9197df3f94f894b4e20bc89e0543b36a

Observation 4f013cee-8e99-4990-b05f-77b42f49ad0b · inbound

Comparative analysis of subword tokenization approaches for Indian languages cites this paper.

Comparative analysis of subword tokenization approaches for Indian languages Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:43.616119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:43.616119Z digest=sha256:bcf0540d9cb7c5c26b8cf74667d61b40e3c1c21cf2b444ed5cf569b33bef5811

Observation 1f1860a8-eae6-4cc5-b178-70030171147b · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:11.036474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:11.036474Z digest=sha256:993a23060ded33275b82b28cbb01a9d5675c9235f482beb7ad6da65327835a1b

Observation bde9fb3c-eb54-42c1-b85b-04165b01106a · inbound

AI Agents for Conversational Patient Triage: Preliminary Simulation-Based Evaluation with Real-World EHR Data cites this paper.

AI Agents for Conversational Patient Triage: Preliminary Simulation-Based Evaluation with Real-World EHR Data Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:38.211148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:38.211148Z digest=sha256:636f79e57d1ceeeeb9799150312b31b15db5892699d4025b19bd93c523944f3c

Observation f176c6cb-890a-4371-b7b6-e53edebd78bf · inbound

Bit-level BPE: Below the byte boundary cites this paper.

Bit-level BPE: Below the byte boundary Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:47.776405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:47.776405Z digest=sha256:ffb9dca01c9d4681ed302e5c8260e68a46b4d2afa428d82e9027357706a9709c

Observation c4f9ba5f-e3f0-4321-a224-fd5e3daa375c · inbound

Is There a Case for Conversation Optimized Tokenizers in Large Language Models? cites this paper.

Is There a Case for Conversation Optimized Tokenizers in Large Language Models? Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:19:15.192201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:19:15.192201Z digest=sha256:f4ef80fdab5bea0c92e27007f255af412c51001af3a05a01851f0e85330b459d

Observation 30ff154b-ccfc-425f-815b-7b6da12380b0 · inbound

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language cites this paper.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.446794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.446794Z digest=sha256:ac80f34d71da18dd723127391d8f8db99e17145684ef7490a7291232a02aaaae

Observation a3c0260e-6aa5-4a87-8872-85a002c12540 · inbound

Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models cites this paper.

Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T22:43:23.055609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:43:23.055609Z digest=sha256:8a9c6aac4657c6b89886d41b935ee9fdc2f0180fa763cc991a32bff625569049

Observation 23e77333-f6c1-4321-9e02-8bce1b0b9835 · inbound

FlowletFormer: Network Behavioral Semantic Aware Pre-training Model for Traffic Classification cites this paper.

FlowletFormer: Network Behavioral Semantic Aware Pre-training Model for Traffic Classification Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:26:02.073354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:26:02.073354Z digest=sha256:bb39e510f6721e0539afffdd7e687e13aebc00cb852ab9b74b2f64caf295cfe9

Observation c7abf7d9-1473-425e-b7cf-f7e24b68acfd · inbound

Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model cites this paper.

Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T13:26:15.129449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:26:15.129449Z digest=sha256:10be85feeb033f9c00f98819092dd8b64880f0c35e1b633e74d29180fdd2c3f3

Observation abb38b1c-7949-4d7b-94ac-cfc58ebfb653 · inbound

Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning cites this paper.

Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:01:20.305515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:01:20.305515Z digest=sha256:c5299d66758b105e0f7346a26a4928285f2a0ce23d2a8b2b6cb821fcc3a71340

Observation aeb96fa4-aa26-40e8-9ff8-adabf20151cb · inbound

Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali cites this paper.

Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:49:36.205280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T00:48:32.432546Z digest=sha256:22a5647ccf25786d2eaf35da076bbbdd68d97bfab24326a1e8b1e609c5a83ecc

Observation a2515f8e-e534-46a0-8c5e-42044d551339 · inbound

BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources cites this paper.

BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:01:02.017768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:22:34.046014Z digest=sha256:697ffc972140650c43c70d10eb7ab1827a9c1265e93bf4066d4c012949a0be05

Observation b47d1164-f5aa-4c99-91d0-46eeade1d351 · inbound

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages cites this paper.

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:38:21.835952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T14:33:36.100966Z digest=sha256:d1b5fa922ebe1e8b88d0f680aabbcea2fdec61457a4f89f132d259efb09576f1

Observation 93d585ba-5dcb-4b0c-8628-a416be556035 · inbound

Translating Signals to Languages for sEMG-Based Activity Recognition cites this paper.

Translating Signals to Languages for sEMG-Based Activity Recognition Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:44:42.857609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T07:41:26.887209Z digest=sha256:6f3928749dd0b7844eaa91149b3c35190205893a8919bf41b91b586dcb909003

Observation 533da98c-0bd8-4f52-a6a3-7915e716954c · inbound

Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models cites this paper.

Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.399554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:50:25.019889Z digest=sha256:8a069bf1ac61d52f1b75bf0b06356c511abe8c99649e443c4883ebd14ffaaaab

Observation 85e7a8d4-9ca4-40a2-9acf-a8dfb0cd6e7d · inbound

The price of incrementality in k-center clustering cites this paper.

The price of incrementality in k-center clustering Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:28.310361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:41:41.502995Z digest=sha256:3b96d77d1b776d3d8361616a479cb2d2d7e6f88627ff401df5df8e95cab71b76

Observation cab02254-afc5-4f27-95c4-0ee3846bfd24 · inbound

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet cites this paper.

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:39:34.660865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T16:53:25.556222Z digest=sha256:c10763679df0d8e7a35c91e5fd42a9a0d5e09757ce4cf927e93aea81aea5bcda

Observation 60b98f86-83b7-4010-af89-a83726b84097 · inbound

MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment cites this paper.

MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:51.297220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T04:59:05.057552Z digest=sha256:c3f2da128db12ceaa5bcf04d8a6731a2f4a59d78409b116d7b7725bbb3019d4c