Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2112.10508.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:43.616119Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
106
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 7a1c99ac-44f9-4d62-887c-8c9e45f07472 · inbound
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f105b10b-2fd9-4768-bee4-7a87a9141599 · inbound
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 278
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c204e047-8388-4312-9431-0d41cc1ddd0b · inbound
BloombergGPT: A Large Language Model for Finance Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d570476c-b086-43ce-89e5-2fdbafa6d6fa · inbound
A Comprehensive Overview of Large Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bd88b63-8ba7-4bb3-a279-175228a7692a · inbound
Jamba: A Hybrid Transformer-Mamba Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f013cee-8e99-4990-b05f-77b42f49ad0b · inbound
Comparative analysis of subword tokenization approaches for Indian languages Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f1860a8-eae6-4cc5-b178-70030171147b · inbound
Beyond Text Compression: Evaluating Tokenizers Across Scales Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde9fb3c-eb54-42c1-b85b-04165b01106a · inbound
AI Agents for Conversational Patient Triage: Preliminary Simulation-Based Evaluation with Real-World EHR Data Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f176c6cb-890a-4371-b7b6-e53edebd78bf · inbound
Bit-level BPE: Below the byte boundary Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4f9ba5f-e3f0-4321-a224-fd5e3daa375c · inbound
Is There a Case for Conversation Optimized Tokenizers in Large Language Models? Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ff154b-ccfc-425f-815b-7b6da12380b0 · inbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3c0260e-6aa5-4a87-8872-85a002c12540 · inbound
Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e77333-f6c1-4321-9e02-8bce1b0b9835 · inbound
FlowletFormer: Network Behavioral Semantic Aware Pre-training Model for Traffic Classification Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7abf7d9-1473-425e-b7cf-f7e24b68acfd · inbound
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb38b1c-7949-4d7b-94ac-cfc58ebfb653 · inbound
Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeb96fa4-aa26-40e8-9ff8-adabf20151cb · inbound
Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2515f8e-e534-46a0-8c5e-42044d551339 · inbound
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b47d1164-f5aa-4c99-91d0-46eeade1d351 · inbound
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93d585ba-5dcb-4b0c-8628-a416be556035 · inbound
Translating Signals to Languages for sEMG-Based Activity Recognition Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 533da98c-0bd8-4f52-a6a3-7915e716954c · inbound
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85e7a8d4-9ca4-40a2-9acf-a8dfb0cd6e7d · inbound
The price of incrementality in k-center clustering Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cab02254-afc5-4f27-95c4-0ee3846bfd24 · inbound
Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60b98f86-83b7-4010-af89-a83726b84097 · inbound
MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.