Pith. sign in

Paper Citation Record · LEDGER

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages

As of 14 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2411.12240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12240 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:54:44.629337Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact4
  • verified fuzzy8
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b307a898-43ce-4a72-826a-247b91890e07 · outbound

This paper cites Future applications of generative large language models: A data-driven case study on ChatGPT,.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Future applications of generative large language models: A data-driven case study on ChatGPT,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:54:45.365443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.457136Z digest=sha256:479ad59fadf0527e130370c17562af0c5885c7380f9bb5921596e471166b961c

Observation bdc5874f-b701-48c3-8d8f-0f4157f2eb87 · outbound

This paper cites A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.463460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.463460Z digest=sha256:ed44f245d22a865a963eadaf7399806435cc9dd5fce90b71ef4ec5c79874558a

Observation 9e196a95-df44-4810-95b1-b367abf80d81 · outbound

This paper cites Performance Evaluation of Tokenizers in Large Language Models for the Assamese Language.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Performance Evaluation of Tokenizers in Large Language Models for the Assamese Language

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.469236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.469236Z digest=sha256:2a849505bed266c4550db51d3e043b0b2c4d77105e65b39781b40887a4d1136f

Observation c4ce1d40-672f-40b4-919f-9c96686dd475 · outbound

This paper cites Problematic Tokens: Tokenizer Bias in Large Language Models.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Problematic Tokens: Tokenizer Bias in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.475699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.475699Z digest=sha256:9a5883b9ba80a54ef898fdf20012b198f0e54250761d7572b1dbdbb33c9fd37c

Observation 93670e64-474c-4aa9-abfc-dccfb31b51fb · outbound

This paper cites Fast WordPiece Tokenization.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Fast WordPiece Tokenization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.482728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.482728Z digest=sha256:0748dd2a3752923744f154c137d4d897b06e7cd036ba38916a14a1fb3ab2d618

Observation 4d168711-a754-48e4-9afa-1be96606319e · outbound

This paper cites Better Than Whitespace: Information Retrieval for Languages without Custom Tokenizers.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Better Than Whitespace: Information Retrieval for Languages without Custom Tokenizers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:54:45.116212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.488420Z digest=sha256:05d5b79ec1f2571d87659f79522ed52abfea46c28baae43a40a044ba4ccedad9

Observation 9e05e673-3fa1-45a6-82c4-98086381ac5a · outbound

This paper cites Theoretical Analysis of Byte-Pair Encoding.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Theoretical Analysis of Byte-Pair Encoding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.495748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.495748Z digest=sha256:c8f522ca14078c2999f383822dbf6ef54f7a602dc8e4cc8c2a34c23ed9f225a2

Observation f4da99aa-bb51-4302-bbf6-8a9d2fbd2d83 · outbound

This paper cites A Formal Perspective on Byte-Pair Encoding.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages A Formal Perspective on Byte-Pair Encoding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:54:45.056557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.507691Z digest=sha256:0b10562929e95162503eae4e049462f514975ada9ea9eaeefdce730534127bb2

Observation 9f134283-c4fb-4a41-b185-28cd5f01f74f · outbound

This paper cites IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.514121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.514121Z digest=sha256:0379e1fca65a2e4630905c5cb29c924b22ee4b189b5aee0f65d8826a49824837

Observation 65a09826-9ad3-4417-be83-2a23f7364573 · outbound

This paper cites EU Tokenizer Performance,.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages EU Tokenizer Performance,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:54:45.348250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.520706Z digest=sha256:a58a78de1c2b7f8693af940ab2f56c8412909af38b7f66ee5d057ace60eff701

Observation c7feb563-ae2d-4fe0-89a8-13e459f578f1 · outbound

This paper cites Tokenizer performance on EU languages,.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Tokenizer performance on EU languages,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:54:45.330806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.525997Z digest=sha256:edcdc576e2f4611f16bfee799d89afb4fc809641743693484c32e044916b4122

Observation d1276e2f-baa4-4a15-be6f-dd0ab55f55bd · outbound

This paper cites Getting the most out of your tokenizer for pre-training and domain adaptation.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Getting the most out of your tokenizer for pre-training and domain adaptation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.532574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.532574Z digest=sha256:805bd5b54abda005290c7612e48b5d79dc6f665eb1c722014d3b6e59ee9d497a

Observation 57728926-8cd4-431f-ab3d-e4fd7e58c836 · outbound

This paper cites Exploring the New Frontier of AI: OpenAI’s GPT-4-O for Indic Languages,.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Exploring the New Frontier of AI: OpenAI’s GPT-4-O for Indic Languages,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:54:45.314664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.537517Z digest=sha256:9269fe329dd064dea15166d6b4772673294e2d8f58b4254f87e3d407d326ffa0

Observation 24fdd2f3-b6eb-4552-8d77-401b8ef443c8 · outbound

This paper cites GPT-4 Technical Report.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.548238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.548238Z digest=sha256:3df4a8eefd2e4ea33e7f254e373c9db1ec1a6cab92d199764b0ca32b7488123e

Observation ddd4f0dd-65bb-4b9b-9d5d-5a545d785210 · outbound

This paper cites SUTRA: Scalable Multilingual Language Model Architecture.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages SUTRA: Scalable Multilingual Language Model Architecture

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.553032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.553032Z digest=sha256:f52bb2fe57c534dcbdf835c14e074f36bf72150b8abb5081ba47cb14430d56d7

Observation 3c1138af-fa54-42d5-b563-8d454189ddcf · outbound

This paper cites Distributionally Robust Receive Combining.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Distributionally Robust Receive Combining

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.559532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.559532Z digest=sha256:093fbc30c6e9d27a65483b6c27ab1a1b19a825f3536ab96bfc4483e37a12b250

Observation d55eef09-3f1e-4fea-8c52-51212f13dca0 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.564594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.564594Z digest=sha256:149543dfd9cff23b9417729aa4529fc806297cfc427745cb768c1a65664d896b

Observation f1efb486-06c9-4ee2-831c-ee9a07d27476 · outbound

This paper cites Large Language Models: A Survey.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Large Language Models: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.569735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.569735Z digest=sha256:1771ef1e690e56508d3847b7268c0c512ed0d943535b558d039a2119127fc860

Observation 9c002b24-a011-407e-9aae-565566611598 · outbound

This paper cites A Comprehensive Overview of Large Language Models.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages A Comprehensive Overview of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.576196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.576196Z digest=sha256:6aad0f6f1954d1acf62a97367900bf8272f29d3f38e3372c8ca24bcd6e71f72a

Observation 19316873-8bf3-4147-a305-6d729d8c49b7 · outbound

This paper cites Multilingual Tokenization Efficiency in Large Language Models: A Study on Indian Languages,.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Multilingual Tokenization Efficiency in Large Language Models: A Study on Indian Languages,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:54:45.298041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.581199Z digest=sha256:9c52c2aee078cf3dbc7ae1dd2c223768c28c74c3c5a8073ba5d2f3d341e9b8ec

Observation 792b8595-885a-47bd-9ac3-52acfb58d6c0 · outbound

This paper cites Impact of Tokenization on Language Models: An Analysis for Turkish,.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Impact of Tokenization on Language Models: An Analysis for Turkish,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.587223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.587223Z digest=sha256:9cce889544839f23379578e003b4df6b4816f1bfc78a1fb3f097dca421ed579c

Observation 0b9737d6-8c6e-4504-b487-9124b746303e · outbound

This paper cites Generative Models For Indic Languages: Evaluating Content Generation Capabilities.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Generative Models For Indic Languages: Evaluating Content Generation Capabilities

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:54:45.280154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.592173Z digest=sha256:0de270e3f546589f521d4afaedc090ff333747c87fc121c26e418beef4f8d5ba

Observation 33426183-9ddc-409a-b3d2-061d52ecd62e · outbound

This paper cites IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-08-12T17:54:44.597231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.597231Z digest=sha256:fb5a173a58fc8ba4a4d850cc33bb78b944ceab253ddb60345a3b9110b7906856

Observation 498a5256-af69-4339-9302-d4c3f9b80264 · outbound

This paper cites INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.602678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.602678Z digest=sha256:a66dcbc09c50cd05a678b961cb6bf00628a947c9c59dac2a6f71076da78a439b

Observation e7b7e122-125e-44e2-b945-98b1d830029c · outbound

This paper cites IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.608824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.608824Z digest=sha256:d1eed36d2cc9532c1396d7f652a17c175bec0c1797339f370fc7a18f88f5d214

Observation 7693b119-1852-4f86-ad18-c313674a8d44 · outbound

This paper cites Evaluating Various Tokenizers for Arabic Text Classification.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Evaluating Various Tokenizers for Arabic Text Classification

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:54:44.732025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.613732Z digest=sha256:f88d95121e5fb09ba60143fbb9ec247e96918083b88151c344e09526b6780b22

Observation 7b91a00b-0822-40a1-8bc1-7fd2e9e8305a · outbound

This paper cites Eighth Schedule,.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Eighth Schedule,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:54:45.237388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.619631Z digest=sha256:cfdd615358f91611b3f3e97b7ed28ebfa2118e3606db844805d396b1e3a3714a

Observation c6ebf67c-6188-4f0e-bc60-68bd0e310338 · outbound

This paper cites PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T17:54:44.624855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:54:44.624855Z digest=sha256:834c686cc90640a6ce103f6612de7740799b69b2ffaa955d590b57251337be38

Observation e5248096-f133-4a6b-ae11-dea79b005ed2 · outbound

This paper cites জীৱনৰ পিৰসেৰ মািহত হাৱােটা বানীয়।.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages জীৱনৰ পিৰসেৰ মািহত হাৱােটা বানীয়।

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:54:45.218843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.629337Z digest=sha256:117c046de48c82e03439d3633e79221683dc780dde18b81e27e3ad876a883a6f

Observation 19088322-b37b-499c-aa0c-1b9c568e0d2e · outbound

This paper cites Available: https://techcommunity.microsoft.com/blog/azure-ai-services-blog/ exploring-the-new-frontier-of-ai-openais-gpt-4-o-for-indic-languages/4142383.

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages Available: https://techcommunity.microsoft.com/blog/azure-ai-services-blog/ exploring-the-new-frontier-of-ai-openais-gpt-4-o-for-indic-languages/4142383

Reference 2024

Resolution
verified exact
raw_fallback, observed 2026-08-12T17:54:44.989557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:54:44.543187Z digest=sha256:acf95449b5480642819c876d9847d9e25333c997d00df167095a60b9989fc9ce

Pith citing papers

No inbound Pith citation observations are available.