Pith. sign in

Paper Citation Record · LEDGER

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 2 inbound Pith citation observations for arXiv:2412.06926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06926 v5

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:25:22.152544Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:55:55.460756Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:39:34.656602Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51720b38-9869-4158-ae8f-f9bb8beb6575 · outbound

This paper cites Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.056655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.056655Z digest=sha256:5bfef6e20998965b41200be4b25c7af1683192264024f2dab93d6ad9b8665654

Observation 17cf01e3-9573-4395-a4c5-70f20dfccd92 · outbound

This paper cites Language Models are Few-Shot Learners.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Language Models are Few-Shot Learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.060360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.060360Z digest=sha256:1b00fd0d489b3185a62f62c001ba2d459ff66598ccd33d03aeb132034a1aebb9

Observation 6043def0-88ca-44b4-bcd1-e126b0660993 · outbound

This paper cites Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:25:22.480560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.063543Z digest=sha256:095e0145207592c631af50d933f846620890c8a3aebe8836524f77cd76d8216c

Observation 8545b9be-fb32-46ea-8015-d5721aaa00b7 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.066887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.066887Z digest=sha256:971d685a59fb01dbb566a375b335d28b744cf6a322835c55b90f8e5797ec2209

Observation 14a6a898-b4ef-4761-b192-cd8d6c3ac452 · outbound

This paper cites Getting the most out of your tokenizer for pre-training and domain adaptation.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Getting the most out of your tokenizer for pre-training and domain adaptation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.070092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.070092Z digest=sha256:068780306e8592851a9d447ab19a15c5ff953d6406c5bbe76e37cf9d947c60e6

Observation 3e7ce535-6458-4f8c-8b6a-af2d7a9614df · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.073299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.073299Z digest=sha256:1e96442d7921760e9f0cda16649264cc9933ee0b298ffbdd945f99b97efb68e7

Observation 887604cf-d2a5-43f5-a6ff-bdd32e6b1c6d · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:25:22.472945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.076728Z digest=sha256:ec37f109091a600afe58f8fbb0c09e90f1009de79b7587e7b9045cfe7e7c66c0

Observation d9a96d13-12fd-4b5b-a2d4-0017a7478305 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.079698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.079698Z digest=sha256:aa3ef95cbb9e989a8c1f25c93b939709ffa77911fdbfbc3b487e7e2454f3f25d

Observation c68e6069-a91a-4f68-bdc4-909124d1e8eb · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.082230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.082230Z digest=sha256:1e685a4ce049b610667bc6452199a9cca75b884c8b7f2f8a240cff02ae30f7df

Observation da4b618c-e179-4e97-9fc2-39c8889b919b · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.084667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.084667Z digest=sha256:fb9fa1939fb8a07b4d8dcaa2968f8ef4ff83fdac3b522e89871f7ea83a3f7323

Observation d7e50150-09ab-48db-a0b1-8b271f46c416 · outbound

This paper cites The Llama 3 Herd of Models.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.087056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.087056Z digest=sha256:2323c7d5824b107cf909016f182e8c022cb62d41e8e9378616e2fd5f7ebfc9dd

Observation 795d1b2d-ea62-4a05-b114-c34f2442cbbc · outbound

This paper cites Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.090169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.090169Z digest=sha256:99e13b1d5d00f872216369c5ceaee4e5903a35576aa568b33b520dc4a09a888e

Observation 18ee15c1-4447-45a4-9d49-d709d12cb944 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:25:22.465279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.092780Z digest=sha256:91d84959a41ccec1ff923dedc5c35a20ce01fc89d2a9e4e5597c43717a0e5d23

Observation ed137759-f8af-484b-84ca-ed6ab4f02134 · outbound

This paper cites Low-resource Languages: A Review of Past Work and Future Challenges.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Low-resource Languages: A Review of Past Work and Future Challenges

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.095097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.095097Z digest=sha256:00ca51a6ced0842abfaa05585435f3e21c8f2bcf643847ba3911f16012a6d416

Observation 28e134b2-db04-4550-8691-f0a7fa367d64 · outbound

This paper cites Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:25:22.456538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.097769Z digest=sha256:da0875c2300bb8858f383a4609282d18070061b9044431efc1989306df043de3

Observation f7158363-8079-4502-b296-89e553b3ccb9 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 16

Resolution
verified exact
doi, observed 2026-08-11T19:25:22.191554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.100121Z digest=sha256:0eb084af6f6c6c52def21caa21a9f6e13349a924e7ef5efe984b725148366e35

Observation b8c417db-9ab5-422c-8fb2-e32e59599c64 · outbound

This paper cites A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.102716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.102716Z digest=sha256:12f9faace212ee279f9fb482b4717a30b26eaed86bffa89b310cb9fbd32e95a0

Observation c37c32ba-e990-455c-960b-bb511e0f5a95 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:25:22.447015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.105358Z digest=sha256:38c82951bc9d92922a13a68f9306c0e9d793d9bbcff2ca840b2014cfa79bf878

Observation fa61c109-b017-491e-af8c-6168aea6e493 · outbound

This paper cites GPT-4 Technical Report.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.107648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.107648Z digest=sha256:95f28acca23b4219f78535b41c645bc7d7a997e20e7fc020511af07a5b06ea05

Observation f25817aa-4ee6-4b60-b2b6-49c50c437dc3 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.110282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.110282Z digest=sha256:1d44c98b1d8a03b83a9e5d3b1d3aadaf8816617e09f535f603fd0840f1c718c0

Observation 8f3d0465-aa73-4646-aace-d3067162310b · outbound

This paper cites Language Model Tokenizers Introduce Unfairness Between Languages.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Language Model Tokenizers Introduce Unfairness Between Languages

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.112975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.112975Z digest=sha256:959c33439539c47949823ebb727e5870e94684822d19fd5c78c9544544bc096e

Observation 679ce153-a688-4453-b7c4-f7c1063e7ae3 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:25:22.439237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.115428Z digest=sha256:f1054a2273e5e1f3d4ddd4fd2f9977ec4bf413620a9c70cf80930206974be712

Observation c878b7df-5fb4-4505-b0fb-5617c76bb5f9 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.117661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.117661Z digest=sha256:22b13be333b84581dfcb5f4a7d7d4cf912b5f67d62813ab5d4bffb9ceef8c4da

Observation 33fdf193-adb9-4028-b21b-50f1956cecad · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.120049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.120049Z digest=sha256:ff1630059f1ca60a6d8eac35cd0008ee3b54319e4ac4773c66d10699330e9f77

Observation 0c3fc772-0f05-4fce-86cc-faade0aab335 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.122282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.122282Z digest=sha256:13b36e58b9ab1b80978b4d7ecb06dbf442f3f356794c3e91f9bc32e1a4644634

Observation afd4083e-b090-46fe-8f22-d2ab3ccb1d34 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:25:22.431371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.125122Z digest=sha256:045d5a4ef8ffb02913a0b39856c14f0f31ed23fbd6e31fa2e60bed6b199e1bbd

Observation 032d1d6d-e205-49f8-a007-6722a45a082b · outbound

This paper cites Tokenization Is More Than Compression.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Tokenization Is More Than Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.128101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.128101Z digest=sha256:1cd893d473450d3cd1828f990e2755bc19facf4857541e906594a94f5ca82b6a

Observation a58f0b5d-74f9-4135-9e9a-3fc4309e54bf · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:25:22.423545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.131350Z digest=sha256:3c65bc00241edeb7115bfbe8c1bcf432a4601dbba4522d9feccfe4580465615b

Observation d8506af5-1699-4158-907c-e3d8d2866223 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:25:22.415568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.134149Z digest=sha256:5b7240cc127c6b8aa9db1da65bb71dc8cc35dc546c3018bf890496fe9633d7d2

Observation 1ad9fc31-0882-4ba0-8f4c-311a49ae3cc3 · outbound

This paper cites Greed is All You Need: An Evaluation of Tokenizer Inference Methods.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Greed is All You Need: An Evaluation of Tokenizer Inference Methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.136962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.136962Z digest=sha256:ef6d86eaeefabbbb8bb91fd7206a9ec73fb8cbd2cda520c802c9443f7443d016

Observation 82ac2e93-56b0-485e-a80c-a8017d75d31e · outbound

This paper cites Egalitarian Language Representation in Language Models: It All Begins with Tokenizers.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Egalitarian Language Representation in Language Models: It All Begins with Tokenizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.140105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.140105Z digest=sha256:cca3e59987fa02362de1ed82969d456e7669f206b0a76be2019536e26f2bc867

Observation ccb193a0-747a-4e40-8aed-aef7dcf40493 · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.142947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.142947Z digest=sha256:dba6ccfa2b1d872e0e61943b14ff93caeded30ec72c201d2af2fc2bd0a03a4bd

Observation ff1acca1-06b5-43db-a08b-42d9dc54bc1e · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.145180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.145180Z digest=sha256:8593e044eda2b021704ed1af0f7884f2f64038e3b8a99d528fdfd086fe4cc1b2

Observation e1ce7be9-2059-42e0-a329-203afb4b2c6d · outbound

This paper cites an unresolved cited work.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:25:22.397573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T19:25:22.147457Z digest=sha256:19a81971e47bb04204415fa5346d6baffb097da99bc5f433d88ecf8cb3258f36

Observation ea23ced9-9282-4f0e-8776-7c6ebd66f0ca · outbound

This paper cites online" 'onlinestring :=.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models online" 'onlinestring :=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.149689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.149689Z digest=sha256:85e5ce0e3de72935715e3d30c7d725e0e7466feef3a920faf15df35fdb799421

Observation df4a8c44-be81-498c-a11f-39a5ff5ef582 · outbound

This paper cites write newline.

When Every Token Counts: Optimal Segmentation for Low-Resource Language Models write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:25:22.152544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:25:22.152544Z digest=sha256:d84523c53b1432023961591bfd11d91140f8a1975cecc91d3c4d3f35a6e3fa12

Pith citing papers

Observation 472eac41-6e95-4b58-9343-b1c85a10bd84 · inbound

Tokenization Matters: Improving Zero-Shot NER for Indic Languages cites this paper.

Tokenization Matters: Improving Zero-Shot NER for Indic Languages When Every Token Counts: Optimal Segmentation for Low-Resource Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:55:55.460756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:55:55.460756Z digest=sha256:7bfa8cd9b558a245f13f5cf1414e340ec89c77332e0a732011062a0c23485fb4

Observation 7adebcde-5119-423a-a6da-c2bca8560b34 · inbound

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet cites this paper.

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet When Every Token Counts: Optimal Segmentation for Low-Resource Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:39:34.657942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T16:53:25.556222Z digest=sha256:73852dba1e24ca1e222fa22f0f54272fca2d3ce3e48ef651961c755c89092f5d