Pith. sign in

Paper Citation Record · LEDGER

Comparative analysis of subword tokenization approaches for Indian languages

As of 17 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2505.16868.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16868 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:45.919746Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb1ca4a8-9723-42c3-a201-cce50e45c2b7 · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:53.598816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:40.732995Z digest=sha256:fc25bc33da360f8f210796970ab53440b2170e41f041c7b752c7b280378dba04

Observation 0156cc44-19ec-44c7-940b-1662bb841aaf · outbound

This paper cites An awkward disparity between BLEU/RIBES scores and human judgements in machine translation.

Comparative analysis of subword tokenization approaches for Indian languages An awkward disparity between BLEU/RIBES scores and human judgements in machine translation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:53.378184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:40.813720Z digest=sha256:433d4946988ef94caf6b439c063250ca469fe6a57edba0263b7e3148257893ef

Observation c4150d97-819e-4845-9aaf-13846d41016b · outbound

This paper cites METEOR: An automatic metric for MT evaluation with improved correlation with human judgments.

Comparative analysis of subword tokenization approaches for Indian languages METEOR: An automatic metric for MT evaluation with improved correlation with human judgments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:53.114813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:40.920925Z digest=sha256:0ebf33fdfb3a96606a9c507cf9f55bfdf552e82e155199a5ed3fa2f2fb729201

Observation bb282389-3dcd-4731-9c10-a63a2318367b · outbound

This paper cites A study of translation edit rate with targeted human annotation.

Comparative analysis of subword tokenization approaches for Indian languages A study of translation edit rate with targeted human annotation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:52.849174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:41.000492Z digest=sha256:6816cb8f439b55a7a873f1677b27ec35243bfa328c2ef714bc6fc08ebcac91aa

Observation 585289e0-bb35-4ec9-81ae-36141b49a028 · outbound

This paper cites chrF: character n -gram F-score for automatic MT evaluation.

Comparative analysis of subword tokenization approaches for Indian languages chrF: character n -gram F-score for automatic MT evaluation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:52.584017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:41.108401Z digest=sha256:cd777f09220c4ed96db49c499f6e7a01a80fe7613f97619c76a388ab51614374

Observation e6ea1fe1-48fa-4e29-9ab5-29f925f3497f · outbound

This paper cites COMET: A Neural Framework for MT Evaluation.

Comparative analysis of subword tokenization approaches for Indian languages COMET: A Neural Framework for MT Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:41.199975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:41.199975Z digest=sha256:4fffc1320eab864be3190ee0f722d1a391b52c5606757965cdf5746584ae29f2

Observation 1ebcd9e6-1d42-4f66-ad66-bf03338a7f47 · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:52.396246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:41.284404Z digest=sha256:8397ced6b49afd7201d44bae3719f63764c525952f12f4abdd6f4f55e3a11c9d

Observation 58ae7b42-df44-452a-afe3-5a15d57f399f · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Comparative analysis of subword tokenization approaches for Indian languages No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:41.422461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:41.422461Z digest=sha256:48f1c5371daea86f6529a2e18280cd4d27fc85eb32be88ea66fa8d495e18cf20

Observation d5bd4a5b-0163-4bf4-8052-2839630a70dc · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Comparative analysis of subword tokenization approaches for Indian languages Neural Machine Translation of Rare Words with Subword Units

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:41.529877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:41.529877Z digest=sha256:e5bb95314ae8dcd585274578dde58bdd7070967952a99851bd2e8d3ea399b63f

Observation 1f70bd03-c8ea-47fa-9067-88fc3f01d8a4 · outbound

This paper cites K., & Patra, B.

Comparative analysis of subword tokenization approaches for Indian languages K., & Patra, B

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:52.194389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:41.618807Z digest=sha256:e1e7c770881ad33632a4773860d8072b5eda64aa2cfb2060752d9466cca54a16

Observation 4ef9b725-680d-4f7a-a34d-beb3d2db41b6 · outbound

This paper cites B., Panda, D., Mishra, T.

Comparative analysis of subword tokenization approaches for Indian languages B., Panda, D., Mishra, T

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:51.927502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:41.707611Z digest=sha256:02570480f836c90d259677095d775fa19d7622e824e628b9dcadb8885d1778b3

Observation 551d8edb-cb64-4d4c-b150-1f71133c0996 · outbound

This paper cites B., Biradar, A., Mishra, T.

Comparative analysis of subword tokenization approaches for Indian languages B., Biradar, A., Mishra, T

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:51.655445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:41.803355Z digest=sha256:9f585730624da7da8378721b293cec8e0363d8cddec5a3f0a2841008ff68696b

Observation 1afd84be-9fb9-4b5d-9950-761ae9f6e73a · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Comparative analysis of subword tokenization approaches for Indian languages Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:41.907403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:41.907403Z digest=sha256:5b5637e4e889bf4b16d58d70f450e48c0da45baca4b2d5dffae06e99d1eac402

Observation 49cb7729-eb5b-446c-831a-5860104c1ea3 · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:51.339239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:41.987662Z digest=sha256:57f3a64d7cb3c8e165776e6e31d2aa067421db9b6bdb1726d55a35a8bd47907a

Observation 9dc1623c-3682-40d0-a077-31310bb4604f · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:51.076132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:42.083468Z digest=sha256:20ded99d486e47388b5d7b3728b0f7b05168d421d75c719aa57243b231008415

Observation 69713da9-d7fc-4d45-9cc8-78893e6d380e · outbound

This paper cites Statistical machine translation.

Comparative analysis of subword tokenization approaches for Indian languages Statistical machine translation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:50.872502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:42.166468Z digest=sha256:535813caddcadf25e10ec9ef46e6e7aa85c2f7b0cec2505871582e4e23d39a08

Observation 3bc8e54e-839f-4145-94cc-aecdce8c0b3f · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Comparative analysis of subword tokenization approaches for Indian languages Neural Machine Translation by Jointly Learning to Align and Translate

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:42.262089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:42.262089Z digest=sha256:8df69592a0a3944b109ab20bdf2e6d457322232b486f61c3b30cc448dcb8b817

Observation d9debab4-e691-4772-88d1-ae86f1796e02 · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:50.602641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:42.366538Z digest=sha256:3baad27ce0450891d554060304688b852c2450dfc400f76192282b1361538954

Observation 1606030e-d832-4f04-ac59-f19e4594868e · outbound

This paper cites Massively Multilingual Neural Machine Translation.

Comparative analysis of subword tokenization approaches for Indian languages Massively Multilingual Neural Machine Translation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:42.479881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:42.479881Z digest=sha256:41ecfdaa01dc8fadaed3fbde29f1587f20aee89d0358b55c96c18b64e4fb6ad9

Observation f6626084-c8ea-48df-9bdd-a08cb8a77682 · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:50.349360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:42.572397Z digest=sha256:236341a54b23138017d155ebe25f0e753199fb5435450c8b225f2d7268510420

Observation b1c26973-80a4-4afe-bd67-e6ecc1701507 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Comparative analysis of subword tokenization approaches for Indian languages SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:42.668489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:42.668489Z digest=sha256:0c393ed1010e837fbe54d086d7f1eb03227b6d0221021972e9a6773590222fa9

Observation ba047e57-4187-44eb-8d9a-ebeefb80b38f · outbound

This paper cites Machine Translation Approaches and Survey for Indian Languages.

Comparative analysis of subword tokenization approaches for Indian languages Machine Translation Approaches and Survey for Indian Languages

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:56:46.539097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:42.801331Z digest=sha256:214a6a83f19d9a217c37ace344ffef18d8a1f75030fe6f71246409d5a02e3477

Observation d6f7ee31-1969-458a-9ec1-855f6315827b · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:50.100829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:42.917746Z digest=sha256:bdb181fc41d8310694860e045c462419e1c85a0ab5da757415f3044682b036e5

Observation 9fcba593-8f1d-4005-ac65-798d9aa798f5 · outbound

This paper cites Morphology: Indian languages and European languages.

Comparative analysis of subword tokenization approaches for Indian languages Morphology: Indian languages and European languages

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:49.850961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:43.049181Z digest=sha256:22f03b77de13480dac2fc16aebacdbb08007c53b9e3682ae2f22a130faf20381

Observation 17b62402-c801-424e-a09b-fefe4788404e · outbound

This paper cites Fast WordPiece Tokenization.

Comparative analysis of subword tokenization approaches for Indian languages Fast WordPiece Tokenization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:43.166127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:43.166127Z digest=sha256:2c5c5cf8ed3a820eab33c4cd0390037157ba67b3f1ad2ac6f252ab9ee5b552bc

Observation 2edf3d60-c5b2-4672-90bd-c34226b1aa9b · outbound

This paper cites NLTK documentation.

Comparative analysis of subword tokenization approaches for Indian languages NLTK documentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:49.659129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:43.332058Z digest=sha256:d900d377beaafbff8b8d0be1b68fc78a71d5c83af5a4bc1e8d5cd767d7d20a64

Observation 24fd864b-8ebd-421d-bc1b-768dec29dbe0 · outbound

This paper cites D., Tetreault, J., & Stent, A.

Comparative analysis of subword tokenization approaches for Indian languages D., Tetreault, J., & Stent, A

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:49.424184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:43.492695Z digest=sha256:597f51a1943ff45f4fd128b7d248851d80def8ce2db0055bf2c29b627e3b72b9

Observation 4f013cee-8e99-4990-b05f-77b42f49ad0b · outbound

This paper cites Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP.

Comparative analysis of subword tokenization approaches for Indian languages Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:43.616119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:43.616119Z digest=sha256:03bae166187c4e3fd0c1f1c3de094ccadc5f47ee302ae692d6e5418fe9e70af7

Observation 30e5eb63-3725-4b29-b9a2-43d52bfbc586 · outbound

This paper cites Character-based Neural Machine Translation.

Comparative analysis of subword tokenization approaches for Indian languages Character-based Neural Machine Translation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:43.749854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:43.749854Z digest=sha256:75c50939afb309918b5a28855abb3e9f893780a073daafb66574e15f844cddcb

Observation 7a1c699a-9b42-4830-9f31-bba3923be6ca · outbound

This paper cites An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks.

Comparative analysis of subword tokenization approaches for Indian languages An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:56:46.375650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:43.850716Z digest=sha256:cefb83abea7cdbf50f07c02d4e3b1bc09cb76335de44f9b8ce39e283ffe4fc0e

Observation 299d4fb3-4312-4b58-bcb0-72b75f872f79 · outbound

This paper cites BPE-Dropout: Simple and Effective Subword Regularization.

Comparative analysis of subword tokenization approaches for Indian languages BPE-Dropout: Simple and Effective Subword Regularization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:44.126108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:44.126108Z digest=sha256:7d5c46198bc72565cb5acbff45c3f97f9b0ea8314429bef3403217cf2db931e4

Observation 2354ad57-c2db-4543-93a7-5f74cd55ecca · outbound

This paper cites Byte Pair Encoding is Suboptimal for Language Model Pretraining.

Comparative analysis of subword tokenization approaches for Indian languages Byte Pair Encoding is Suboptimal for Language Model Pretraining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:44.277726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:44.277726Z digest=sha256:2f8a358c66cc6bf078a6fcb4f349b92a4020d5cf9827c1aa53880f5bec4622d0

Observation 13c3f9df-408b-40e9-a3e0-c357cbb5ea39 · outbound

This paper cites Using Integrated Gradients and Constituency Parse Trees to explain Linguistic Acceptability learnt by BERT.

Comparative analysis of subword tokenization approaches for Indian languages Using Integrated Gradients and Constituency Parse Trees to explain Linguistic Acceptability learnt by BERT

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:56:46.136522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:44.384852Z digest=sha256:c2b5c3b391fc45b9ebcb1b83b958b4496213993980f7aeaabb1ae8fd0719e32e

Observation 8aa51c85-b120-41dd-bf41-a894350007cd · outbound

This paper cites HAN: hierarchical association network for computing semantic relatedness.

Comparative analysis of subword tokenization approaches for Indian languages HAN: hierarchical association network for computing semantic relatedness

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:49.182855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:44.559715Z digest=sha256:18dedf7661170f202a027771667a8c77c8bb6c1ffd57cf80d18e8124f917b7a5

Observation afe603d7-9c9b-432d-bd6d-c39ba4577c88 · outbound

This paper cites Meaningless yet meaningful: Morphology grounded subword- level NMT.

Comparative analysis of subword tokenization approaches for Indian languages Meaningless yet meaningful: Morphology grounded subword- level NMT

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:48.952378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:44.716253Z digest=sha256:8c2089ae4ed45992f2bc5f7b010a3106e34d5eadc59300a356430fac75e3edce

Observation 73ad5378-3fce-4a90-b788-46f1acddc5e2 · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:48.737339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:44.805572Z digest=sha256:56ffbcec4c62cf323b0f98076041fc1a6d3dbafedb8a82571010b7b73d7b1cf3

Observation 6b3376c9-85a7-4743-917c-d9eeb2bbc44e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Comparative analysis of subword tokenization approaches for Indian languages BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:44.948945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:44.948945Z digest=sha256:1cc82518ceab950122465feb72af54264e1ea6fd325258f5b4c77c17627f57ea

Observation fa1e4781-9d82-4b9a-bd61-c3bb47ad0cab · outbound

This paper cites AI as the next GPT: a Political -Economy Perspective.

Comparative analysis of subword tokenization approaches for Indian languages AI as the next GPT: a Political -Economy Perspective

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:48.544357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:45.049589Z digest=sha256:91716f3bfbd86a3f09dacbdbfcef52afd219faf829e6fcaaeff9fe7022e0d959

Observation 5da2293b-5724-472f-a723-ab5bc8ff582d · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:48.346280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:45.152504Z digest=sha256:fdcf337dfcd26515d605bc4153f8e786aa6c37c7a9f716c67781a6084ce88707

Observation f6f2050a-1697-458d-b091-17fafffa61ad · outbound

This paper cites fairseq: A Fast, Extensible Toolkit for Sequence Modeling.

Comparative analysis of subword tokenization approaches for Indian languages fairseq: A Fast, Extensible Toolkit for Sequence Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:45.244129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:45.244129Z digest=sha256:624e83ed45dc7faeaf3d95f0ce528ffa145857d0b226f108e9f14d7b1147a2e0

Observation e8ee23cc-5c85-42ad-8818-fc4235c0b593 · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:48.099408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:45.360960Z digest=sha256:37a06f0310378291d252eefd6685dd1617d7873444f0fd31c1695c53156b3716

Observation 6914086e-43ac-4199-9df7-1636f4fdc64b · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:47.821428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:45.437018Z digest=sha256:0a1756c74e4656b945aa351041abe41570d53dad5614ec53fc5330b92412cacc

Observation 717053e7-a5b4-48d7-9fce-a8e9a5f59f03 · outbound

This paper cites Moses-Statistical Machine Translation System.

Comparative analysis of subword tokenization approaches for Indian languages Moses-Statistical Machine Translation System

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:47.538867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:45.548671Z digest=sha256:beedbe77f0814aa1b952c174e458ed6f2f13cefa0cea048fd9799a6416c9f030

Observation 1bf3a788-d361-49cf-be01-6b13ca1437e3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Comparative analysis of subword tokenization approaches for Indian languages Adam: A Method for Stochastic Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:45.645407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:45.645407Z digest=sha256:245b0eeb541829b8ef90f974e3778937c8a904767e41477393d91947801d573f

Observation 68133417-a204-4aac-a1cf-e6726cb1a396 · outbound

This paper cites Post, A call for clarity in reporting BLEU scores, in Proceedings of the Third Conference on Machine Translation: Research Papers.

Comparative analysis of subword tokenization approaches for Indian languages Post, A call for clarity in reporting BLEU scores, in Proceedings of the Third Conference on Machine Translation: Research Papers

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:47.359452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:45.739641Z digest=sha256:641eca999400cf41344f582a56f71b2c0987c3dd203d47d8a1b6751fdbd1ce03

Observation a80c834e-7b5d-484e-8ecf-b16e383ec87a · outbound

This paper cites an unresolved cited work.

Comparative analysis of subword tokenization approaches for Indian languages Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:56:47.089618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:45.830297Z digest=sha256:3a40a13213f16f1b7fb65d908986babda06cf0ccb919e41981d281efed85d2cc

Observation fff5327f-622b-4c4d-8999-9b0fc0d45680 · outbound

This paper cites B., Choudhury, S., Mishra, T.

Comparative analysis of subword tokenization approaches for Indian languages B., Choudhury, S., Mishra, T

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:46.811586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:56:45.919746Z digest=sha256:457442866a504a6bd2ba3f414ea364e69afc5bbb287c6274274f877939792a36

Pith citing papers

No inbound Pith citation observations are available.