Pith. sign in

Paper Citation Record · LEDGER

In-Place Tokenizer Expansion for Pre-trained LLMs

As of 14 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.15232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15232 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:50:17.987906Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf5b3712-064f-4da3-af85-7839bd8ff29d · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

In-Place Tokenizer Expansion for Pre-trained LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:14.205257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:14.205257Z digest=sha256:427b8070a389ca0e49580925a4e52ee6868420557229c0c5318c98b685f58c83

Observation 388f9141-6904-4258-aa13-eb0a526551ad · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

In-Place Tokenizer Expansion for Pre-trained LLMs Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:14.519304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:14.519304Z digest=sha256:2531fba01f5b543381aa5787332e042527b62031d5a8dea41cd512702e8a1c43

Observation d04aa6d4-4a75-48c0-8333-d70e8c144124 · outbound

This paper cites Getting the most out of your tokenizer for pre-training and domain adaptation.

In-Place Tokenizer Expansion for Pre-trained LLMs Getting the most out of your tokenizer for pre-training and domain adaptation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:14.734727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:14.734727Z digest=sha256:f2e42e76925b1178b6758037b276ac3a0a759084bca962eba41d885d37f44c58

Observation a9ac9973-c0d4-46de-a51f-ccbd3f3e8d61 · outbound

This paper cites Token distillation: Attention-aware input em- beddings for new tokens.arXiv preprint arXiv:2505.20133,.

In-Place Tokenizer Expansion for Pre-trained LLMs Token distillation: Attention-aware input em- beddings for new tokens.arXiv preprint arXiv:2505.20133,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:14.831833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:14.831833Z digest=sha256:73c880b5bb0dfbd342b4410e5543e8bd67429aa91632de9677cfbc530fe9af83

Observation 1b8d0d90-5de9-4746-96ec-4f0e3608f8da · outbound

This paper cites Sailor: Open Language Models for South-East Asia.

In-Place Tokenizer Expansion for Pre-trained LLMs Sailor: Open Language Models for South-East Asia

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.036700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.036700Z digest=sha256:9c78934df08ae49dd4680f8eb09f6a775d1bc3cbb4e5022f9d1e367449d44c2a

Observation 04464007-c7da-4791-9430-4eb37113238b · outbound

This paper cites Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities.

In-Place Tokenizer Expansion for Pre-trained LLMs Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.109835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.109835Z digest=sha256:ae8caeeeb7fc88b88f28f3bce0670c2cb6d0f80ba45ac712c07ad5ba3bcfed3b

Observation 0900b006-84b7-4846-a024-4497586a2d37 · outbound

This paper cites Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit.

In-Place Tokenizer Expansion for Pre-trained LLMs Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.197912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.197912Z digest=sha256:2ac2d557f6bb40265485554bb4fca02a3c532db78f3b8329802be60abdfb8c02

Observation 1948917c-3699-49d6-85c4-5440c92a8391 · outbound

This paper cites The Llama 3 Herd of Models.

In-Place Tokenizer Expansion for Pre-trained LLMs The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.276274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.276274Z digest=sha256:cef9d5403098809e8b0341979123a8e9ef3376cc11fa7a41e146c466f7a0586d

Observation 938266b8-643d-4ffd-92d1-7230a5756e4d · outbound

This paper cites ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model.

In-Place Tokenizer Expansion for Pre-trained LLMs ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.357730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.357730Z digest=sha256:db0103a3df80a73c2a6fe1cb3abdd9ea78275596eff302a049fbaaad5a192894

Observation cab3315e-dfe5-43a9-830d-37610feac956 · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

In-Place Tokenizer Expansion for Pre-trained LLMs Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.426936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.426936Z digest=sha256:1c36fcf410044a35c3582dde1ea4af205afaa322df0afcee3751f0312b365f1d

Observation 94c09525-8a6c-4248-b807-1a8bab64d848 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

In-Place Tokenizer Expansion for Pre-trained LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.518068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.518068Z digest=sha256:739e5e23f281484acaddf0c9b5d1d1114ad3f5f3dfdf7800ce40f0e6fc1d005c

Observation 3293b783-0028-4a41-b754-8b7453a8ef18 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

In-Place Tokenizer Expansion for Pre-trained LLMs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.611343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.611343Z digest=sha256:22353ef16e2efed3a995428f792945d6d526bb28e46394f3ce4cef00512f0427

Observation fd963db7-8803-49aa-b137-ec61ab57d8d8 · outbound

This paper cites Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language Models.

In-Place Tokenizer Expansion for Pre-trained LLMs Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.667569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.667569Z digest=sha256:363214033bd341ace10a5e7be0674f7c7c6c9cbd687de53f18bdfc70c47c6044

Observation c3e2d6db-0ba2-426c-8812-52262f7417e0 · outbound

This paper cites TokAlign: Efficient Vocabulary Adaptation via Token Alignment.

In-Place Tokenizer Expansion for Pre-trained LLMs TokAlign: Efficient Vocabulary Adaptation via Token Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.766936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.766936Z digest=sha256:8bd8ea04eb7bb19443c8e41523a191822e2bb54955021363e0b562d17ecd7397

Observation 106f1fc8-3e33-4744-a3a2-939966ea043e · outbound

This paper cites doi: 10.18653/v1/2024.acl-long.

In-Place Tokenizer Expansion for Pre-trained LLMs doi: 10.18653/v1/2024.acl-long

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.828568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.828568Z digest=sha256:672975bbc65b31eb544480dccf7eb78b5b6f06a473d0e55ab609eb09fc55b3eb

Observation 61ab653b-ccb1-46df-93d9-de795a0531f5 · outbound

This paper cites Let's Verify Step by Step.

In-Place Tokenizer Expansion for Pre-trained LLMs Let's Verify Step by Step

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.916953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.916953Z digest=sha256:ec0f4e69552e7285fab4df16cc5c15e32c731fc8ebf79f00d96ac418d0dee5b0

Observation a553e361-61d8-44dc-b874-88787803b766 · outbound

This paper cites LFM2 technical report.arXiv preprint arXiv:2511.23404,.

In-Place Tokenizer Expansion for Pre-trained LLMs LFM2 technical report.arXiv preprint arXiv:2511.23404,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.013219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.013219Z digest=sha256:e2626248228e4d9f30bf6c5346764dfe61915ee9227a56d0843353705ef9f9fc

Observation 5acf4e6d-2eaa-4133-858a-412334be0a1d · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

In-Place Tokenizer Expansion for Pre-trained LLMs Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.078214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.078214Z digest=sha256:0ed2af25acd2445206d7943141216004cdcb20da8c84d1843213516a9df987ea

Observation 1b05f5d2-02e8-4df7-9b44-3ff484999259 · outbound

This paper cites Nandini Mundra, Aditya Nanda Kishore Khandavally, Raj Dabre, Ratish Puduppully, Anoop Kunchukut- tan, and Mitesh M.

In-Place Tokenizer Expansion for Pre-trained LLMs Nandini Mundra, Aditya Nanda Kishore Khandavally, Raj Dabre, Ratish Puduppully, Anoop Kunchukut- tan, and Mitesh M

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.200437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.200437Z digest=sha256:9f6b8c7841846d825292d863879a9f24c8857f6ee8b15709e8f9d3b188a71e52

Observation 892cae98-8bbc-446a-9d92-c75d44097134 · outbound

This paper cites An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models.

In-Place Tokenizer Expansion for Pre-trained LLMs An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.259225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.259225Z digest=sha256:8797f0fb341e6c8b379d72a63dce81d7792d0b7ae9bc9dca513cb5576db9dc35

Observation cd508f2c-7658-47dd-b894-9c75674811c2 · outbound

This paper cites AdaptiVocab: Enhancing LLM Efficiency in Focused Domains through Lightweight Vocabulary Adaptation.

In-Place Tokenizer Expansion for Pre-trained LLMs AdaptiVocab: Enhancing LLM Efficiency in Focused Domains through Lightweight Vocabulary Adaptation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.366515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.366515Z digest=sha256:9c4868c21296ed06d428599a58740ff9bb3da14d2f31022b3093ca178d00a22e

Observation 91b7392c-684f-4381-84d1-860805078905 · outbound

This paper cites Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning.

In-Place Tokenizer Expansion for Pre-trained LLMs Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.551996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.551996Z digest=sha256:76a65b17eba27d9151eee6b5be85a05b8310c0221d5a7b105026b808d90dfe64

Observation a12afaeb-34ce-473a-91f9-9ad1b79e3171 · outbound

This paper cites FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language.

In-Place Tokenizer Expansion for Pre-trained LLMs FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.706422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.706422Z digest=sha256:5e109823d66bad77d4882333ca98a69337fc40bec3c7c8ccde55d52b1bed8c5d

Observation 9508b108-dc2a-4c07-9d6d-88c92ac8f8c9 · outbound

This paper cites Yamshchikov, and Mark Fishel.

In-Place Tokenizer Expansion for Pre-trained LLMs Yamshchikov, and Mark Fishel

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.828605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.828605Z digest=sha256:1d285a621fcc700a60bb7c4f2d901baa59aeed1a948451aec2aae172474ccf84

Observation 1b58d0ec-ee68-4219-8ae6-7c166a46e79f · outbound

This paper cites Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP.

In-Place Tokenizer Expansion for Pre-trained LLMs Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.052920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.052920Z digest=sha256:39baa3774b60dd3bdce7053aac714743725ca9ded6fa4f25a988b9c078f11f2c

Observation 40fb040e-78e9-4687-8003-0f6170f7b558 · outbound

This paper cites Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation.

In-Place Tokenizer Expansion for Pre-trained LLMs Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.232027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.232027Z digest=sha256:c10a344c0da44c30ce9c01f046c56319053c5b50c8b8bec7ba73bcc950e39631

Observation b73e10c2-9ce7-4601-b604-06344521055a · outbound

This paper cites Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model.

In-Place Tokenizer Expansion for Pre-trained LLMs Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.306885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.306885Z digest=sha256:c657e233166274400071a0b86dec363ee4a85f0fe38e0b70c0eb9cb584b283be

Observation d6679158-4ec1-4fb8-ada4-164eca059df1 · outbound

This paper cites On-Device Language Models: A Comprehensive Review.

In-Place Tokenizer Expansion for Pre-trained LLMs On-Device Language Models: A Comprehensive Review

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.391198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.391198Z digest=sha256:2b643da97d7b81c51cf76c3bebcebc72d67614a3a0d86a21fe548695feb60114

Observation 099d2ab8-56c7-492f-8cad-0347b17c91ea · outbound

This paper cites An empirical study on cross-lingual vocab- ulary adaptation for efficient language model inference.

In-Place Tokenizer Expansion for Pre-trained LLMs An empirical study on cross-lingual vocab- ulary adaptation for efficient language model inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.453883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.453883Z digest=sha256:c5d4066faa5239805d214c58cc7c02811acc877c754e927696a0fcdd9022759f

Observation 250276d5-f2ef-4e92-b32b-46eb7cd802b6 · outbound

This paper cites An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference.

In-Place Tokenizer Expansion for Pre-trained LLMs An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.541365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.541365Z digest=sha256:f63db4c2cb39b8cda7450d1f26a5b23ab06d3d97f9079f1cc62a9b724bcc4f16

Observation 98690ae9-4327-4c8d-9400-abb96731d278 · outbound

This paper cites arXiv:2406.11477.

In-Place Tokenizer Expansion for Pre-trained LLMs arXiv:2406.11477

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.601724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.601724Z digest=sha256:5a4998b41a0932e64715b53de71cd82df0b110d37b1abd9505385445b61baa54

Observation 2ff7ab0b-4cec-4b2e-9778-d73b7da54a1d · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

In-Place Tokenizer Expansion for Pre-trained LLMs Instruction-Following Evaluation for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.660388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.660388Z digest=sha256:f5a273a6435cc5d4ef7fa1cf9e4caa25b203e494cf1c050640509a0b8e5fe139

Observation 23905fa8-91f9-46ec-8343-c9477964fadd · outbound

This paper cites •Math (3 benchmarks).GSM8K (Cobbe et al., 2021), MATH500 (the 500-problem test subset of MATH (Hendrycks et al.,.

In-Place Tokenizer Expansion for Pre-trained LLMs •Math (3 benchmarks).GSM8K (Cobbe et al., 2021), MATH500 (the 500-problem test subset of MATH (Hendrycks et al.,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.708762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.708762Z digest=sha256:7d1a327ae45c01334ee92f623f8324ea2d253eaf315f35c406ca46019c8e1718

Observation 93dc6229-83cb-4524-8f91-561589b41875 · outbound

This paper cites (2024)), and GSM-Plus (Li et al., 2024).

In-Place Tokenizer Expansion for Pre-trained LLMs (2024)), and GSM-Plus (Li et al., 2024)

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.762279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.762279Z digest=sha256:774545dc3875faddd0ee1f3273c9f7aaf86739566bbb8a984b3e04a83910db93

Observation 5069654b-6246-4b0b-8d39-45a06169958b · outbound

This paper cites •Multilingual (2 benchmarks).MMMLU (OpenAI,.

In-Place Tokenizer Expansion for Pre-trained LLMs •Multilingual (2 benchmarks).MMMLU (OpenAI,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.846638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.846638Z digest=sha256:ca215db70be881722ee7ae8d65dfb8f288017bb20ac81af88b752aef3bec240c

Observation ff637248-d043-442d-89d3-9bf1d2562fa7 · outbound

This paper cites an unresolved cited work.

In-Place Tokenizer Expansion for Pre-trained LLMs Unresolved cited work

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-01T23:50:17.925297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.925297Z digest=sha256:51a0133f9bd9e046294f5886ee53daf8cf3d9c17c380a7d43906e3d002e46d68

Observation 9a12a49f-824c-44c2-9ad7-50cfbcafca57 · outbound

This paper cites Zero-shot.

In-Place Tokenizer Expansion for Pre-trained LLMs Zero-shot

Reference 43

Resolution
malformed identifier
no resolver link, observed 2026-08-01T23:50:17.987906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.987906Z digest=sha256:6a6ea9934340223333409d902a2cb50cfa5557ec4dddf79ff96a04cd39d7c922

Observation 566d3c34-cc63-4f5f-b2ea-a9eae8368a52 · outbound

This paper cites Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning.

In-Place Tokenizer Expansion for Pre-trained LLMs Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.166923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.166923Z digest=sha256:3ef3b13fc3bdd7f562dfa5f1c151b2af66552039dda9a15902752ad8630d7049

Observation 4719273d-bfbc-449d-908e-435d491b0555 · outbound

This paper cites Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca.

In-Place Tokenizer Expansion for Pre-trained LLMs Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:14.638495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:14.638495Z digest=sha256:714d67ed3b0f1c42fe132fa31fdbe4c0296b8faf72bb71b946ebfde04ea10055

Observation 8c3bb9c1-032c-4457-8e07-62f919a85f61 · outbound

This paper cites Tower: An Open Multilingual Large Language Model for Translation-Related Tasks.

In-Place Tokenizer Expansion for Pre-trained LLMs Tower: An Open Multilingual Large Language Model for Translation-Related Tasks

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:14.394538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:14.394538Z digest=sha256:a46e5b4f787a7601119f91bc49a21a90df90be1d6e49568e63bbfb618ac6796c

Observation 7869ccf8-b4d6-47e5-b8f2-960106574478 · outbound

This paper cites Do all languages cost the same? Tokenization in the era of commercial language models.

In-Place Tokenizer Expansion for Pre-trained LLMs Do all languages cost the same? Tokenization in the era of commercial language models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:14.276234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:14.276234Z digest=sha256:d139b07636d2a6a1838f62aeb293895d783f6b71f1cb300155bd76c0e169eb44

Observation 13e07e81-f001-498f-b758-7bdd1108e15b · outbound

This paper cites Sailor: Open language models for South-East Asia.

In-Place Tokenizer Expansion for Pre-trained LLMs Sailor: Open language models for South-East Asia

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:14.948905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:14.948905Z digest=sha256:a5ee2e1c24eee55dd7f65cb2a2260104c3106959ca77cd0a8fe89aa50c94db5d

Observation cda33b7d-680d-4834-b29b-413bed8bdb87 · outbound

This paper cites doi: 10.18653/v1/2026.findings-eacl.341.

In-Place Tokenizer Expansion for Pre-trained LLMs doi: 10.18653/v1/2026.findings-eacl.341

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.948239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.948239Z digest=sha256:f39a876f55c23b24fe78195452cf497dde5960af123d880f1380bf581ed072f9

Pith citing papers

No inbound Pith citation observations are available.