Pith. sign in

Paper Citation Record · LEDGER

Language Model Tokenizers Introduce Unfairness Between Languages

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2305.15425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.15425 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:32:25.383880Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

29
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5493db6b-1d07-41e4-a038-85a36e15f45b · inbound

Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models cites this paper.

Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models Language Model Tokenizers Introduce Unfairness Between Languages

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:25.383880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:25.383880Z digest=sha256:1e43e9655a15a2ca63d06c2358ab2beaf2ebf0b7b47a032abd55036268f447ea

Observation 1121a55b-fb62-46b6-b0d0-c893765cad1f · inbound

Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts cites this paper.

Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts Language Model Tokenizers Introduce Unfairness Between Languages

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:00.875465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:00.875465Z digest=sha256:69c66e66d5553e5bb0b1b21004fc8a6fce4474a229a2389e40e103086149fc25

Observation 2108986b-cfda-40b9-b2b3-3d827f213f76 · inbound

Causal Estimation of Tokenisation Bias cites this paper.

Causal Estimation of Tokenisation Bias Language Model Tokenizers Introduce Unfairness Between Languages

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:43.589887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:15:43.589887Z digest=sha256:8031988a5fbba5d0c594845d3006d8e30413b29ad0b81f8d0aeaf7fda2d6b3d6

Observation a4448f31-a4c7-4396-90ac-f0d963686242 · inbound

Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala cites this paper.

Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala Language Model Tokenizers Introduce Unfairness Between Languages

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:37:53.013482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:36:27.421985Z digest=sha256:d6620b5af214cc9c5a73d84db14e1045e5cf3c7d31597eeaa415d6af950d4311

Observation 487f9c6b-8106-4cad-9807-69ef24b82406 · inbound

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization cites this paper.

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization Language Model Tokenizers Introduce Unfairness Between Languages

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.781850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:21:56.826488Z digest=sha256:b7c432e80178d9f18f41227421aa4656f5735859b971fa385cb3f0e34ddeb4c7

Observation f4f19e13-e5cb-451f-860c-91646d30f46c · inbound

Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study cites this paper.

Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study Language Model Tokenizers Introduce Unfairness Between Languages

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:09:02.558978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T21:08:44.929762Z digest=sha256:0c35559dc60faf4f0df8173a874aadfdaa4655c99389c6693ad4b561e1e916ef

Observation 54ff702d-6fd9-47d9-af6c-f107095d4e8e · inbound

LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations cites this paper.

LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations Language Model Tokenizers Introduce Unfairness Between Languages

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T16:03:33.014552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T15:59:56.231210Z digest=sha256:c0e0790c83ff8c986eaffc8f7f430328f8129cd981f8e958d4921bbda8bf39db

Observation fb7e9804-4133-4a84-a9e0-8ffa26f9554d · inbound

Tokenization with Split Trees cites this paper.

Tokenization with Split Trees Language Model Tokenizers Introduce Unfairness Between Languages

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.838187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:48:29.391178Z digest=sha256:9b422384478e37dfecaa2aa72be7ea1f4de4ef7038db940119e116a3fad5d1cc

Observation 899c837d-d4cd-46af-94a5-8b9f213d26de · inbound

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty cites this paper.

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty Language Model Tokenizers Introduce Unfairness Between Languages

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:14:40.909415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:08:33.172423Z digest=sha256:cf393c07902108a1a6e1d6ad0aa309d46db77920ccce0ed696ec53667d02f12b

Observation b2321e79-237a-41f1-a0db-c5d3bd003ee3 · inbound

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base cites this paper.

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base Language Model Tokenizers Introduce Unfairness Between Languages

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:13:14.768133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:12:16.235649Z digest=sha256:b3aec4fddda40e86a2a445f36444f9823449e2e387b90284099110de108f3f4e

Observation 69bdf272-e63f-4d5d-9a5d-062dab6f10de · inbound

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators cites this paper.

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators Language Model Tokenizers Introduce Unfairness Between Languages

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:01:38.447490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:01:38.447490Z digest=sha256:4bd3dc05d1b9542f42e30b2793cc2ee6eeac37849bd20345ee991be3c90b3949

Observation 8afba8f7-16bf-4d88-83a8-5fab1838c950 · inbound

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis cites this paper.

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis Language Model Tokenizers Introduce Unfairness Between Languages

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T23:49:11.023254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:49:11.023254Z digest=sha256:ec46c2bc7235f0c67974d90de1fbc40f48da56bdd0564aea875d179783035c0f