Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T23:50:17.987906Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.15232.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T23:50:17.987906Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cf5b3712-064f-4da3-af85-7839bd8ff29d · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 388f9141-6904-4258-aa13-eb0a526551ad · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04aa6d4-4a75-48c0-8333-d70e8c144124 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Getting the most out of your tokenizer for pre-training and domain adaptation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ac9973-c0d4-46de-a51f-ccbd3f3e8d61 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Token distillation: Attention-aware input em- beddings for new tokens.arXiv preprint arXiv:2505.20133,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8d0d90-5de9-4746-96ec-4f0e3608f8da · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Sailor: Open Language Models for South-East Asia
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04464007-c7da-4791-9430-4eb37113238b · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0900b006-84b7-4846-a024-4497586a2d37 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1948917c-3699-49d6-85c4-5440c92a8391 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938266b8-643d-4ffd-92d1-7230a5756e4d · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cab3315e-dfe5-43a9-830d-37610feac956 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c09525-8a6c-4248-b807-1a8bab64d848 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Measuring Mathematical Problem Solving With the MATH Dataset
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3293b783-0028-4a41-b754-8b7453a8ef18 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd963db7-8803-49aa-b137-ec61ab57d8d8 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e2d6db-0ba2-426c-8812-52262f7417e0 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs TokAlign: Efficient Vocabulary Adaptation via Token Alignment
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106f1fc8-3e33-4744-a3a2-939966ea043e · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs doi: 10.18653/v1/2024.acl-long
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ab653b-ccb1-46df-93d9-de795a0531f5 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Let's Verify Step by Step
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a553e361-61d8-44dc-b874-88787803b766 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs LFM2 technical report.arXiv preprint arXiv:2511.23404,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5acf4e6d-2eaa-4133-858a-412334be0a1d · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b05f5d2-02e8-4df7-9b44-3ff484999259 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Nandini Mundra, Aditya Nanda Kishore Khandavally, Raj Dabre, Ratish Puduppully, Anoop Kunchukut- tan, and Mitesh M
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892cae98-8bbc-446a-9d92-c75d44097134 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd508f2c-7658-47dd-b894-9c75674811c2 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs AdaptiVocab: Enhancing LLM Efficiency in Focused Domains through Lightweight Vocabulary Adaptation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91b7392c-684f-4381-84d1-860805078905 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a12afaeb-34ce-473a-91f9-9ad1b79e3171 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9508b108-dc2a-4c07-9d6d-88c92ac8f8c9 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Yamshchikov, and Mark Fishel
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b58d0ec-ee68-4219-8ae6-7c166a46e79f · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40fb040e-78e9-4687-8003-0f6170f7b558 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b73e10c2-9ce7-4601-b604-06344521055a · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6679158-4ec1-4fb8-ada4-164eca059df1 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs On-Device Language Models: A Comprehensive Review
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 099d2ab8-56c7-492f-8cad-0347b17c91ea · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs An empirical study on cross-lingual vocab- ulary adaptation for efficient language model inference
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250276d5-f2ef-4e92-b32b-46eb7cd802b6 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98690ae9-4327-4c8d-9400-abb96731d278 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs arXiv:2406.11477
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ff7ab0b-4cec-4b2e-9778-d73b7da54a1d · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Instruction-Following Evaluation for Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23905fa8-91f9-46ec-8343-c9477964fadd · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs •Math (3 benchmarks).GSM8K (Cobbe et al., 2021), MATH500 (the 500-problem test subset of MATH (Hendrycks et al.,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93dc6229-83cb-4524-8f91-561589b41875 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs (2024)), and GSM-Plus (Li et al., 2024)
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5069654b-6246-4b0b-8d39-45a06169958b · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs •Multilingual (2 benchmarks).MMMLU (OpenAI,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff637248-d043-442d-89d3-9bf1d2562fa7 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a12a49f-824c-44c2-9ad7-50cfbcafca57 · outbound
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566d3c34-cc63-4f5f-b2ea-a9eae8368a52 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4719273d-bfbc-449d-908e-435d491b0555 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c3bb9c1-032c-4457-8e07-62f919a85f61 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7869ccf8-b4d6-47e5-b8f2-960106574478 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Do all languages cost the same? Tokenization in the era of commercial language models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e07e81-f001-498f-b758-7bdd1108e15b · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs Sailor: Open language models for South-East Asia
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda33b7d-680d-4834-b29b-413bed8bdb87 · outbound
In-Place Tokenizer Expansion for Pre-trained LLMs doi: 10.18653/v1/2026.findings-eacl.341
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.