Pith. sign in

Paper Citation Record · LEDGER

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 2 inbound Pith citation observations for arXiv:2502.07057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07057 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:58:58.597059Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:55:55.418419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:19.438760Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8435d93-5282-4c6e-976b-0d38f7f10bbc · outbound

This paper cites Tokenization Is More Than Compression.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenization Is More Than Compression

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.892514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.514018Z digest=sha256:dd643328fad6cef4aade0c85b0df6eb7aeb1e3b424a4e2ebf66ab0c4ef48afd7

Observation d374156d-3e5b-47e0-86a6-10cce3c7338e · outbound

This paper cites an unresolved cited work.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:58:58.881244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.518772Z digest=sha256:35bf5c8039d3837a32b1aa97b21f20e2ff76ebf7620c8a76f9e0b011f0f53dc5

Observation ddbab64a-6d8f-43eb-bb42-51dfb39d4943 · outbound

This paper cites How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in japanese.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in japanese

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.869782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.522901Z digest=sha256:54628676ea30ef7ed8549d9f138317107f2221d5d33ecd3d16a1503789cc4f83

Observation 4da0aa48-8e4f-490b-9630-431a09e54865 · outbound

This paper cites Critical tokenization and its properties.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Critical tokenization and its properties

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.857711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.526306Z digest=sha256:aeb9979cc38d04cd34c01efcfdc6bcf09342e42b0ba82bd03c558cdb05033ad2

Observation 5f10b29d-b30a-42b0-bc2f-ae8cdf34b639 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.530907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.530907Z digest=sha256:3895b909567b860927bed8aa18ee476865dcc0fbefa7ec4893047118411f81bb

Observation 5926943e-bc6e-4ab8-afe4-06ef4863c5d4 · outbound

This paper cites Chemotactic motility-induced phase separation.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Chemotactic motility-induced phase separation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.535300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.535300Z digest=sha256:009e3d7f2585e457674ee912990fa6d8d56d495bc8453c6516c3b9d3ca457e6d

Observation 2f903b61-5838-4090-a979-bd945a6b9c11 · outbound

This paper cites Formalizing BPE Tokenization.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Formalizing BPE Tokenization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.539958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.539958Z digest=sha256:278aee2f1787179400c5416dedda41fec220f71ead565fa487da4e8e86f73fd0

Observation af1abc08-fbc1-4fbf-aa29-71647036cd41 · outbound

This paper cites The Technical User’s Introduction to LLM Tokenization.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark The Technical User’s Introduction to LLM Tokenization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.846367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.544214Z digest=sha256:9f7617b9d61ebcfee25b001fab4ba7e91183ba29bb9ad9ec81a0c352b0cc4423

Observation 3335b46e-21c7-45a8-9e96-c7d68c67edff · outbound

This paper cites A New Algorithm for Data Compression, 1994.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark A New Algorithm for Data Compression, 1994

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.834784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.548336Z digest=sha256:6808d38956ba73da4e3c8973a7c7a2316b87579e8631951ccffbb42d3c8b5b39

Observation 4a6da132-4bf4-4a48-881b-c155a60e938c · outbound

This paper cites github.com/riotu-lab/aranizer, December 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark github.com/riotu-lab/aranizer, December 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.822901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.551338Z digest=sha256:356b57ad8d0515f3b5ca13f78fe1caa6d2602f27153eba2202b91cb67ae6ad79

Observation 06daf0fe-fe4f-4285-82de-016854153ab7 · outbound

This paper cites So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer, December 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer, December 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.810188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.554589Z digest=sha256:17dd495e42247bf8cb2d997b123a829e30b0f0bc118e2b5724bca1d4c0c99a91

Observation 1e9a9452-5499-4bf5-905d-83e47c234d98 · outbound

This paper cites Arabic Tokenizers Leaderboard - a Hugging Face Space by MohamedRashad.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Arabic Tokenizers Leaderboard - a Hugging Face Space by MohamedRashad

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.793754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.558056Z digest=sha256:09701e637d89a0dd6a63b759fafd3bd23616402137d15ce29c0bb33a275564d7

Observation fb2149e2-4a93-417c-994c-ca0684d2ef07 · outbound

This paper cites NbAiLab/tokenizer-benchmark, November 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark NbAiLab/tokenizer-benchmark, November 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.782283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.561881Z digest=sha256:e37dc20d9c95bc64647f79ef2a223bcc3d02bafc2fe71c8af122bf5b20a5eb19

Observation 099c238d-30de-48db-8698-9fb6448049b4 · outbound

This paper cites Tokenizing on scale.Preprocessing large text corpora on the lexical and sentence level.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenizing on scale.Preprocessing large text corpora on the lexical and sentence level

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.768447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.565542Z digest=sha256:7d0d7a842abd3242c52bc6ac4c57c8e929702ff48eeb92d814ab20e0a8062ab6

Observation 128c76ba-c58e-4c16-bb51-c43d74aab12c · outbound

This paper cites Analysis of Subword Tokenization Approaches for Turkish Language.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Analysis of Subword Tokenization Approaches for Turkish Language

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.754936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.568407Z digest=sha256:0e7c6b87cce13533a689100ff24ce17cdbc12801f09302c470e981ca949da0f1

Observation 89aab7e3-7877-4b3d-9daf-13b67e6a345f · outbound

This paper cites EuroLLM: Multilingual Language Models for Europe.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark EuroLLM: Multilingual Language Models for Europe

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.571329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.571329Z digest=sha256:d1dc184b3a0af52cfa7fe287159a767c431eda3c66c3b4ae84490d4d8984b225

Observation 2a8a8885-e358-415a-b3dd-2f010c29951e · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models, June.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models, June

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.742478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.575667Z digest=sha256:90b0fe074910975f198ab1963a1d7deb248ccbbbb6f71e7477db6b2e9f728024

Observation da76be89-48bc-459d-93d1-f78fcf1932b6 · outbound

This paper cites Not All Tokens Are What You Need for Pretraining.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Not All Tokens Are What You Need for Pretraining

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.730830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.583279Z digest=sha256:0e6e7da278b0beb4ef391f588cd3cdb9c7f332ce7817b511fdbda05841d8cf5e

Observation 2cfa602c-f3d2-4591-929f-7b3963a6b6ef · outbound

This paper cites Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.586474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.586474Z digest=sha256:65f6efbe1bbb9f60cff37b624c19c27b486717d59b1eea6a091189e16242cf81

Observation a9daa4db-dbf0-47ca-b0c4-818c15cd2366 · outbound

This paper cites ITU Turkish NLP Web Service.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ITU Turkish NLP Web Service

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.718416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.590224Z digest=sha256:d6902eb0502cf5cbe82b4af6ba7692693f97b80325af5f35a2ee1ea952faa4dd

Observation 617aa5f9-e6b1-4558-b0ff-741e6ce74e0b · outbound

This paper cites ahmetax/kalbur, October 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ahmetax/kalbur, October 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.706260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.593375Z digest=sha256:b151965078718deb83cc139b9982a0a9672f9d62ac77feb07148d133c4916823

Observation 10db0fcf-e5b1-4566-a060-6c33370a1828 · outbound

This paper cites Ali Bayram.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Ali Bayram

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.694284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T13:58:58.597059Z digest=sha256:867a0a5f421b2802c21bcdc1e10fe55c0535212f9fee1de41896b8a6bc0d66fc

Observation 87669eec-12e7-4ce2-b555-855d4fbc350a · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.579433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.579433Z digest=sha256:2a91efc9a1405c7174700e68f44f57e6b397012dfd8b6348bafcc36231e57880

Pith citing papers

Observation a092e4ed-44e1-46e1-b0ae-cb18ff71aaa6 · inbound

Tokenization Matters: Improving Zero-Shot NER for Indic Languages cites this paper.

Tokenization Matters: Improving Zero-Shot NER for Indic Languages Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:55:55.418419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:55:55.418419Z digest=sha256:364edb4bae07308539b4c19e3b7aab34b1a9d7a20f359f336ac3025a59b4c6b4

Observation 9c59362a-d240-4fc7-96b8-55705f8931c1 · inbound

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish cites this paper.

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.441232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T20:53:57.646472Z digest=sha256:24b905542d20a3e9c9dd9c22dd91122b47483e93a22f92262b825bd287d8efc0