Pith. sign in

Paper Citation Record · LEDGER

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

As of 9 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2508.01006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01006 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T06:00:10.327109Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 548a52aa-2253-4475-81d1-0927df9c063e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.281920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.281920Z digest=sha256:e75a43c88b42ca4158c1b8dd247c7f54eaf574720692ae54cac54dd9920ea386

Observation 47537837-aa1d-4b6d-8a6b-7c73a663c655 · outbound

This paper cites MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.296880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.296880Z digest=sha256:9af1cd1db8f4e0868cc4d410031168ff490b125fea7ef23935ba2242593dfc64

Observation f6ac2c71-d32d-4fac-8f5a-8aa11ec5954a · outbound

This paper cites Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T06:00:10.374786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T06:00:10.300890Z digest=sha256:85f72a8b5ccf6b6a5432a810624116dda45fe4b88fb416990a63c3397ec2e941

Observation 9b21b02a-1d0d-4cfb-b3bb-a103a75aaca7 · outbound

This paper cites In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:17.008156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T06:00:10.304908Z digest=sha256:8b67d138ffa3f1331694741cf2d681c63e543978ac0d42c41278dd9d952cbe17

Observation 325c05ce-bb73-4f46-91eb-d7f90fe50223 · outbound

This paper cites In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4606–4634, Abu Dhabi, United Arab Emirates.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4606–4634, Abu Dhabi, United Arab Emirates

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.984655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T06:00:10.316058Z digest=sha256:6768921dae45888c83fc5f8aa0aaa446d79fda0a44d97d8d4867e178d22a6996

Observation 658c8df3-eb7a-4f0c-90ec-0c39dc780deb · outbound

This paper cites https: //huggingface.co/large-traversaal/ Alif-1.0-8B-Instruct.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu https: //huggingface.co/large-traversaal/ Alif-1.0-8B-Instruct

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.973213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T06:00:10.319635Z digest=sha256:8aad0de85f905a2629c1852e5237e80c389bd1f704c84e41a4a7c9b16fc3cc93

Observation d6c2eb2c-9940-4d45-9194-ce4a9ebbffa2 · outbound

This paper cites In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.962267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T06:00:10.323430Z digest=sha256:4b0d411a8770607ce9eae0a6ba71e10c919d3254548e8501da4af244288bbc97

Observation 1a95b647-5a3a-400a-b0b8-921d650a2b73 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.285644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.285644Z digest=sha256:e38554b40ec9bffff68490805402deada155032aba8a1144cad5412f76f6384a

Observation 833c7619-d4b4-4f6c-8b1c-ef0c6e9dde59 · outbound

This paper cites an unresolved cited work.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-06T06:00:16.950764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T06:00:10.327109Z digest=sha256:9b744cca2d73dde440a150490be3f97c527334dc0f1fdf35e61897e9c0c9ae61

Observation 903cf07f-df61-470c-a178-bc27745dfb5a · outbound

This paper cites In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, Online and Punta Cana, Dominican Republic.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, Online and Punta Cana, Dominican Republic

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:17.020156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T06:00:10.293383Z digest=sha256:593b4cd8d4c6eb5eecb4ec56252d1e71be79ae09e73f37e3612ca82907ae10e2

Observation 62d44409-ffa4-4a70-acb5-b3d8dcc8b068 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Crosslingual Generalization through Multitask Finetuning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.308640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.308640Z digest=sha256:b0b0fba4076b7e5b4f8647afc19169c7ad08edb43685d1730c38c7ee6068d991

Observation 12c83f7b-ce1b-4839-b665-57a31b8abe16 · outbound

This paper cites In Findings of the Association for Computational Linguistics: EACL 2023 , pages 1581–1594, Dubrovnik, Croatia.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu In Findings of the Association for Computational Linguistics: EACL 2023 , pages 1581–1594, Dubrovnik, Croatia

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:00:16.995724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T06:00:10.312630Z digest=sha256:98688e5afee1cc5c4a895d1de71ea19f0937e589afcd52cfead8ad51f4b6e175

Observation c3f58f95-83ac-4dde-b071-2c9f97a73365 · outbound

This paper cites The Llama 3 Herd of Models.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.289493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.289493Z digest=sha256:54b0b9837d7f8a661ef3d375ec159f8261da376d087e23701dc06b04291487db

Observation b6b58527-6c82-42f4-836d-554e91e84c54 · outbound

This paper cites Preprint, arXiv:2506.13487.

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu Preprint, arXiv:2506.13487

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T06:00:10.277918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:00:10.277918Z digest=sha256:0d874ae0d19502761d44eb0a68b81babff1729ab63924e776dd43522b21b0345

Pith citing papers

No inbound Pith citation observations are available.